Don't pay a frontier model to be a shell script.
A Continuum Engine session starts by paying for 300 KB of instructions, and on 26 September its sessions ran restarts and store queries by hand that code could run for no tokens at all. Glue Becomes Code retires each step on the cheapest rung that can do it: code, an MCP tool, a local model measured on the job, and a paid model only with confirmation.
Standing order
Every session pays for 300 KB before it reads a file.
A Claude Code session in Continuum Engine starts by loading .claude/CLAUDE.md,
everything it imports with @, and every file in .claude/rules/
without a paths: key. On 28 September that was 13 files and 300,198 bytes:
the rules, the project map, the local model protocol and every model's settings, the
system as built, why it exists, the queue of what's next, and the reference repos on
disk. At about four bytes a token that is 75,000 tokens, in the context window on
every turn of every session.
Prompt caching keeps the rereading cheap. Claude Code writes the set to a one-hour
cache on the first request, at twice the input price, and each later turn reads it back
at a twentieth of the input price on Opus 5.5. Caching gives none of the room back: the
tokens are in the window on every turn, they bring the next compaction closer, and the
project's CLAUDE.md is loaded again after it.
The set is paid for on purpose. CLAUDE.md says what it buys: "You should
never need to go looking for what _ref/ holds, what is built, or what to
work on next." The model settings are always on because "a measurement taken under
different settings is a different measurement, and the drift is invisible." Each file is
held to one test, whether every session needs it or only the sessions that touch its
area, and only the first kind stays always on.
The always-on set, file by file
| File | Holds | Bytes |
|---|---|---|
.claude/CLAUDE.md | identity, commands, how sessions work | 6,862 |
docs/RULES.md | the coding rules and the Definition of Done | 27,576 |
docs/STRUCTURE.md | every project, and where it sits | 31,710 |
docs/models/LOCAL_LLM.md | the protocol for any local model | 39,512 |
docs/models/MODEL_RUNTIME.md | every model role's load and inference values | 32,936 |
docs/arch/current/README.md | the system as built, derived from the code | 21,765 |
docs/arch/FINDINGS.md | where the code stands, and the queue | 66,251 |
docs/arch/PURPOSE.md | why Continuum exists | 23,489 |
docs/REFS.md | the reference repos on disk | 47,832 |
four rules without paths: | the @ lines that import the last four | 2,265 |
| Total | 13 files | 300,198 |
Five more rules carry paths: and load only when a session touches a
matching file: the journal's conventions, the milestone architecture, GitHub
tracking, the harness itself, and what to do with a project's README.
Stuck on you
Glue is a manual step that works only while a session runs.
The engine's rules call it "any manual step an agent runs to keep Continuum working."
On 26 September Claude stopped and started the AppHost by process id through WMI,
queried Postgres, Redis and Temporal through docker exec, and messaged
another session about build windows and process ids. The next morning LM Studio had
restarted overnight with nothing loaded, and the profile went back on by hand. Each
step cost tokens every time, and none of them happened while no session was running.
Tristan
"Less tokens for manual labour."
Tristan
"If you end up using glue to harness something, please make it code, or if not deterministic, pick a model for it to run against locally. Obviously, accuracy matters for the LLM and we will continue learning lessons as we build this way."
The rule that came out of it, Glue Becomes Code, retires each step on the cheapest rung that can do it:
- Plain code or a Terminal verb, for anything deterministic.
- An MCP tool, which answers in one compact reply where
docker exec … psqlreturned pages. - A local model, for routine judgement, once a battery has measured it on that job.
- A paid model, only with the person's confirmation, under a cap of $1 a day across every paid provider.
The build list written that day estimates that about half its steps were rung-one work
done by a model. The WMI went first: continuum stack start|stop|status,
proven the same day. Hand-run polls of Temporal's HTTP API became
continuum ops queues, and on 28 September the hand-typed terminates became
continuum ops terminate, proven live after the day's reset. The test for each one is
whether it still works with no session running.
# 27 September: typed into a session, every time
docker exec continuum-temporal-<suffix> sh -c 'temporal workflow terminate --address 127.0.0.1:7233 --query "…"'
# 28 September: a Terminal verb, and no session needed
continuum ops terminate <type> --reason "…"
Small change
Every token tactic is a piece of glue retired.
- Rules scoped to paths. Five of the engine's nine rules load only when a session touches their files. This site's 67 KB animation guide went behind one on 28 September, loading when a session opens an animation script, component or skill, and its always-on set fell from 161 KB to 94 KB.
- A README on touch. A hook hands a session a project's README the first time it reads a file in that project, once per session.
STRUCTURE.mdwas rewritten on 22 September, when its file tree had reached 45,898 characters under one heading, 70 times the median section, and each project's detail moved into the README that arrives with its code. - Hooks that say nothing until there's something to act on. A hook that asks at each prompt whether anything is waiting prints nothing when nothing is, so it costs no tokens while things are quiet.
- The catch-up at session start, under a budget. The last session's open questions arrive with the new one, under 8,000 bytes, for reasons the next entry tells.
- Compaction rescue. Before a compaction a hook saves the person's messages the journal hasn't recorded, and hands them back after it.
- Subagents by tier. On 23 September the first architecture survey launched 13 agents at once, each costing about 250,000 to 450,000 tokens. From then on, at most four ran at a time, with Sonnet reading and Opus judging, and from 27 September one or two.
- No polling. A session waits for the notification, since every look is a turn.
- No command that stops for approval. On 28 September a delete on a path built from a variable sat on an approval prompt for 40 minutes while Tristan was away from the desk. Runs now write to a new folder instead of deleting to make room.
Tristan, 27 September
"It's imperative, we are running out of spend, to reduce the costs."
Tristan, 28 September
"How could you have let me hang?"
Home grown
Judgement goes to a local model once a battery has measured it.
Summarising long output, classifying a line, triaging a failure and writing a first
draft go to the RTX 5090 once a model has been measured on that exact job. Anything not
yet measured stays on the rung above. The engine's LOCAL_LLM.md holds the
rules every integration has hit:
- Send every sampler, every request. An omitted one is the runtime's chat default, and LM Studio's repeat penalty of 1.1 punishes JSON, which repeats by construction.
- Read the request on the wire. On 22 September three samplers were missing from the wire while the configuration looked right. On 28 September the same client library turned a schema's
maxLengthinto description text, removing the bound that stops a model looping inside a string. - Check the response's
model. LM Studio answers a request for a model that isn't loaded with one that is. On 24 September that recorded 46 claims as the extractor's, and the extractor hadn't answered any of them. - A big model loads alone. Loaded beside the chat and embedding models, the extractor warmed up at 4.1 tokens a second, with its layers on the CPU.
The receptionist that routes Continuum's messages is on the same path: a routing battery decides which local model makes the choice.
Old hands
Google's SRE book calls it toil, and caps it at half an engineer's time.
Its chapter on eliminating toil defines it as work "that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows," and keeps it "below 50% of each SRE's time." Glue is toil done by a model: agent toil. It is paid in tokens, and it comes back every session, since a session starts with nothing but what was written down.
Mitchell Hashimoto's "Engineer the Harness" (February 2026) says what to do about a mistake: "you take the time to engineer a solution such that the agent never makes that mistake again." Every's compound engineering (Dan Shipper and Kieran Klaassen, December 2025) expects "each feature to make the next feature easier to build." Toyota's jidoka is a line that stops itself when something is wrong, and the lease gate is one.
What this work adds is the ladder, whose third rung is a local model measured on the job before it is trusted; a hook that notices the agent's own glue while the session is still running; and the test that a retired step works with no session running.
The kit: the glue detector, the context budget, the local model
The same practices as hooks that run by themselves: Node with no dependencies, each
one silent on any error and done in about 50 ms. They live in
agent-tools/hooks/harness, and one command merges them into a project's
existing .claude/settings.json.
node agent-tools/hooks/harness/install.mjs path/to/project --with-claude-md
cd path/to/project
# The runtime's own model ids, then measure one before trusting it
node .claude/hooks/harness/local-llm.mjs --list
node .claude/hooks/harness/local-llm.mjs --model <id> --battery
# Rung 3: a first read of long output, off the paid model
npm test 2>&1 | node .claude/hooks/harness/local-llm.mjs --model <id> --system "Name the failing test, and nothing else."
glue-detector.mjs(Stop, UserPromptSubmit) counts the session's shell commands by shape: an inline script is one step whatever it says, and paths, strings, ids and numbers are replaced. On the third repeat it appends a row toGLUE.mdand tells Claude in one line: "Glue ×3 this session:curl -s -m 3 http://localhost:<n>/v1/models→ rung 2, MCP tool or CLI verb". Reading and editing files,git, builds and tests don't count: they're the work. A new session hears only of the steps two sessions repeated, never the whole file, which grows by a row per session per step.context-budget.mjs(SessionStart) resolves the always-on set as Claude Code does, everyCLAUDE.md, its imports to four hops and every rule withoutpaths:, and says one line only when it is over budget, naming the three largest files.local-llm.mjspipes stdin to an OpenAI-compatible endpoint with every sampler set, saves the request body, refuses a reply from a model it didn't ask for, and ships a ten-item battery to measure a model before trusting it.claude-block.mdis a 15-lineCLAUDE.mdblock with only what the hooks can't enforce: the ladder, the context budget's question, sharing the machine, and spend.
The defaults, all in one place (lib/common.mjs), so a project's
harness.json holds only what it changes:
| Setting | Default |
|---|---|
| A step becomes glue | on its 3rd repeat in a session |
| A new session hears of it | once 2 sessions repeated it, 5 steps at most |
| Context budget | 25,000 tokens, estimated at 4 bytes a token |
| A hook's output | 600 to 6,000 characters by hook, never over 9,000 |
| Local model | LM Studio at localhost:1234, no model until you name one; temperature 0.2, top_p 0.95, top_k 40, min_p 0.05, repeat penalty 1.0, 1,024 tokens, seed 42 |
Next: Sessions that share a machine take leases, and The codebase remembers itself.