Field notes

Chapter 04 · Decisions

Part 2 · The foundations5 of 35

MCP tools in practice: our assistant calls the same tools

Instead of a chat panel with a private back door, every capability in The Margin is an MCP tool, and Margin Intelligence, outside agents and tests all call the same ones.

The Margin4 min read

By the time we got to the AI work, the product already existed: boards, notes, habits, expenses, sharing and a family suite, all running offline on a real local database. I think that order is the reason the assistant turned out to be worth anything.

The shape almost everyone builds#

The default way to add AI to an existing product is to bolt a chat panel onto the side and wire it to a private back door. The panel gets its own endpoints, its own queries, its own slightly different idea of what a card is. It can do maybe fifteen things, chosen by whoever built the panel.

It demos beautifully and it rots immediately, for one structural reason: the assistant's path through your system is a different path from the one your UI takes. Two implementations of "create a card" drift apart the moment either one changes. The panel starts lying by omission, because nobody updated its private copy of reality when a field was added last week.

I did not want to maintain two products.

The front door#

So the assistant does not get a back door. It gets the front door, and the front door is a tool surface that anything can call.

Every capability in the product is exposed as a named tool with a schema. The count moves whenever the product does, because those are the same list. Creating a card, splitting an expense, logging a habit, drawing on a whiteboard, saving a recipe from a link, putting its ingredients on the shopping list, reading a note by pages when it is too long to read at once. The built-in assistant calls those tools. So can anything else that speaks the same protocol, which in practice means any AI client the user already trusts.

That reframing does more work than it looks like.

  • The assistant is a user of the product. It has the same permissions model, the same entitlement gates and the same workspace scoping as a person at a keyboard. When it tries something it should not be allowed to do, the code that refuses a human refuses it too. There is no privileged path to forget to secure.
  • The surface is testable. A tool with a schema can be called deterministically, without a model in the loop, which turns "does the AI work" from a vibe into a test suite.
  • The product stops being the only place you use it. Someone who would rather drive their workspace from an editor or a terminal comes in through the same door, and we built nothing extra for them.

Handing over the whole job#

The part of this I did not plan for is that the door goes both ways.

An outside agent can authenticate and call the tools one at a time, which is the obvious use and the one we designed for. It can also hand a whole task to one of our own specialists, the planner or the researcher or the one that triages an inbox, and collect the result when it is finished. That is closer to hiring a contractor than to operating a machine.

We got that second mode almost for free, and only because the specialists were already users of the same tool surface. Had they been functions buried inside a chat panel, there would have been nothing to hand a job to.

The law that keeps it honest#

There is a rule here that has bitten us more than once, so it is now written down in the repository as a law.

Every capability must ship in both tool servers and in the runtime's own allowlist. For a long time even that was not enough. The allowlist only decided which names the runtime would refuse, and making a tool callable by our own assistant took a hand-written wrapper that nobody had listed as a step. Miss any piece and the tool exists, passes review, is documented, and cannot be called. Nothing fails. The assistant, asked to use it, tells the user with total confidence that the product does not have that feature.

That exact thing happened with whiteboards. The drawing tools were implemented, registered on one server, permitted by the allowlist, and available to custom agents. The one agent every user actually talks to had never been given them. Asked to lay out a diagram, it replied that there was no whiteboard capability in this workspace, which was true of its own toolset and false of the product it was speaking for.

When we finally counted

200+

named tools, one list shared by the assistant, outside agents and the tests

74

allowlisted tools with no wrapper, so our own assistant could not call them

The fix was to make the allowlist do the registering. Every name on it now becomes a typed tool the assistant can call, with its schema generated from the same models the tool server uses, and a test fails if the two ever disagree. Adding a capability to the list is now the whole job. The list has passed 300 names since then, and Margin Intelligence can call all of them except a handful we hold back on purpose, each one named in a single list.

What it cost#

The cost is discipline. A capability is done when it is reachable from every surface that should have it, which is later than when it works. That is more steps per feature, permanently.

What it saved us is a second product. There is one implementation of creating a card, so the assistant and the interface have nothing to disagree about.

I did not expect how much this changed what I ask it. When it can reach everything, "add a task" gets boring fast, and the requests drift toward the ones that cross modules, like whether this month's grocery spend has anything to do with the meal plan nobody followed.