Field notes

Chapter 23 · Decisions

A game mind that doesn't burn tokens: two tiers of character AI

The break room's characters play and talk with a local mind that runs every frame for free, and ask Margin Intelligence for help only outside play, through a shared cache, under hard caps.

The Margin5 min read

Most of the time, when something happens in a break room game, the character watching it says nothing. Every event goes to the character's memory and mood, and only now and then does it come back as a line. How often depends on how much of a show-off that character is, and each one has a cooldown. A quiet character can go a whole match without a word.

It is also the cheapest possible answer to the founder's last request for the games. He wanted the opponents and the cast to be "kind of intelligent, so it's not just a dumb, same programmatic way of playing each time", powered by Margin Intelligence, the assistant built into the rest of The Margin, and without waste.

Those pull in opposite directions. A game draws sixty frames a second. A model call is slow by that clock, costs money every time, and sometimes never answers. So before anything else we wrote down one rule with no exceptions: no model call sits on the frame loop, or on any path that decides a move while you are playing.

That leaves two tiers.

The part that runs every frame#

The first tier is ordinary code shipped inside the game, running on the device with no network, and it costs nothing to call sixty times a second.

A character facing a decision lists what it could do and scores each option with weighted considerations, each bent through a response curve, then picks among the best few with weighted odds, so the strongest option usually wins but not every time. The same ball twice can get two different answers. Each member of the cast carries a small personality vector (aggression, patience, risk, showmanship, kindness, precision) that leans those scores, so two characters facing the same board make different calls.

Mistakes are modeled the way people make them. A character reacts a little late, or its aim scatters, or it jumps early, and all of that widens under pressure and late in a long rally. There is no random blunder roll anywhere. A miss happens when the small errors add up to more than the shot forgave.

The game also learns you. It keeps named features like reach.left or trouble.deep, faded on a 21-day half-life and capped at 48, stored with your game progress so the laptop and the phone both know them. When both have been playing, their two copies merge. In Volley, difficulty drifts to keep things close, but only inside a narrow band around the level you picked, and the daily court and two-player matches don't drift at all. Characters remember the last few things that happened and carry a mood, and the mood is what their face shows.

You can hear that learning in The Week, once. Jot watches which days you fill last, and when the player model is confident that one day stays clearly emptier than the rest, Jot says so at the start of an Endless or Zen round: "You tend to keep Fridays clear. I've noticed." The whole remark is worked out on the device from the player model, with no model call anywhere near it. If the model isn't sure yet, Jot keeps quiet, and there is a test for that silence as well as for the line.

Table for Six has the most characters and none of them are the cast. Each guest gets a temperament read from who they are, so a chatty guest shows off, a shy one hardly speaks, a child fidgets and an elder is content to wait. Each carries its own mood and memory, down to who it has sat beside tonight. None of that reaches the rules. A guest's face still says exactly whether their wishes hold, because that face is the clue you are playing by.

The part that asks Margin Intelligence#

The second tier is Margin Intelligence itself, reached through the same service the assistant runs on, and it is allowed three jobs, all of them before or beside play.

Before a match it may be asked once for a plan: which tactics to lean on against this kind of player, and which of their habits are worth testing. The plan comes back in the game's own vocabulary and the first tier reads it as weight adjustments. The request gives up after a second and a half, and the match never waits for it. Slow, offline, capped or switched off, the match starts unplanned, with no spinner. A plan that arrives after the first serve gets ignored.

Dialogue is generated in small batches for each game and character, cached on the server and shared by everyone who plays. On the device a line is picked from that cache in the same instant as the event. Characters the cache hasn't reached yet fall back to lines bundled with the app.

And the model can write flavor: story beats for Loose Ends and words for Paper Plane's Letter run. Whatever comes back is checked against a schema on the server and again on the device, and a Letter run uses a word only if the run can actually spell it. The levels themselves still come from the generators and solvers in the previous chapter.

Where the tokens stop#

The server runs four checks before any call. There is a kill switch, and off means off at once. Then the shared cache, keyed by the kind of player (game, mode, character, difficulty and a coarse bucket of the player model), so one answer serves everybody who lands in the same bucket. A stale entry keeps being served while it refreshes. Then the caps, one per person per day and one for everybody per day, claimed atomically before the call. If the database can't say whether there is budget left, the answer is no. Last, the call itself is deduplicated, so a crowd opening the same game at once shares one request, and a call still running when the wait expires finishes anyway and fills the cache for the next person.

Game calls don't touch anyone's AI action allowance. They stop when Margin's daily spend ceiling is reached, keep their own ledger of calls and tokens, run at a low reasoning setting, and have a Games tab on the admin observability page.

That ledger showed the cache doing most of the work within days. Dialogue requests in the past week came back from the cache almost every time. The caps matter more than we expected, because every character asks for its own lines, so a long session can reach the per-person daily cap. When it does, the bundled lines fill in and nobody sees a gap.

Plans are the awkward part. When we compared models on production, a fresh plan took about eight seconds at the median, and the game stops waiting after one and a half. So every plan a match has used so far came from the cache, filled by a request that finished after its own match had already started. The first tier plays a decent match on its own, and if cached plans turn out to feel no different from no plan, the plan request is the first thing we'd switch off.