Chapter 06 · Decisions
Part 2 · The foundations7 of 35
Why our AI agent runs outside the Next.js app
An agent turn takes fifteen seconds to two minutes; a page render takes milliseconds. Why we split the assistant into its own Python service, and what that boundary has cost us.
The Margin3 min read
The app is TypeScript. The assistant is Python, in a separate service, on its own hostname, behind an internal key. Every time I explain that, someone asks why it is not just a route in the Next app. It would be fewer moving parts, one deploy and one language.
What a model call actually does to a web app#
A request that renders a page finishes in tens of milliseconds. A request that asks a model to do something takes fifteen seconds if it goes well and two minutes if the task is real, because a turn is a loop rather than a single call. Read, think, call a tool, read the result, think again.
Two kinds of request
Tens of milliseconds
15 seconds
2 minutes
Those two things do not belong in the same process. You can put them there, but everything you tune for one is wrong for the other. Connection pools sized for fast requests, memory limits set for rendering, timeouts that make sense for a page load: put a two-minute agent loop through them and the app falls over while someone is loading their dashboard, which is a lot worse than a slow assistant.
Separating them means the assistant can be slow, or busy, or briefly broken, and the product stays a product.
What the runtime is actually holding#
The deeper reason is that the agent runtime does a different KIND of work, and mixing it in hides that. It:
- holds conversation state across turns;
- decides which tools exist for this particular user, on this tier, in this workspace;
- passes every prompt through a guard before it reaches a model;
- meters what a turn cost and writes that down;
- reads which model provider to use from a setting, so swapping one is a config change rather than a release.
Written as routes inside the web app, those become scattered middleware nobody can see the shape of. Written as a service, they are the service: one place you can point at and say, this is what happens when someone asks our product a question.
The bill for the boundary#
It is not free, and I would rather be honest about the parts that have hurt.
Two languages means two of everything: schemas, validation, test harnesses, CI. Change the shape of a card and you change it in TypeScript and again in Python, and there is no compiler that will tell you when you have only done one.
The wire between them is a place data goes wrong. We shipped a bug where JSON was encoded twice, once by our own call site and once by the database driver, and the data landed as a string containing a string. Everything "worked". Nothing was readable. That entire class of failure does not exist inside one process.
And the service is reachable from the internet, which is a sentence with real consequences. Two routes once shipped with no authentication on them, one of which let an outside caller hand work to our agents. They were found and closed, and the lesson stuck: a service on its own hostname does not inherit your app's session. Every route has to say out loud that it requires the internal key, and forgetting is silent until someone else notices.
Why it stays#
Because the alternative fails in a worse way.
An assistant embedded in the web app can only ever be reached from the web app. Ours is reached from the web app, the phone app, a gateway that lets another agent hand it work, and background jobs that run with nobody watching. That is only possible because the runtime was never a feature of one frontend.
And because it can be replaced. The model provider is a setting. The conversation loop is a file. A better model arrives every few months, and the change stays inside a service whose entire job is being the part of the system that changes. We have swapped the default model more than once this year and the web app did not notice.