From the build

Dev Entries

Lessons from building Continuum, each one taken from the real work and the journal that recorded it. Two to five minutes apiece, and most of them learned the interesting way.

Entries

2026-09-28 · 5 min read · Continuum Engine

Don't pay a frontier model to be a shell script.

A Continuum Engine session started by paying for 300 KB of instructions, and on 26 September its sessions ran restarts and store queries by hand that code could run for no tokens at all. Glue Becomes Code retires each step on the cheapest rung that can do it: code, an MCP tool, a local model measured on the job, and a paid model only with confirmation.

  • m2
  • model-routing
  • tool-access
2026-09-28 · 5 min read · Continuum Engine

Sessions that share a machine take leases.

Two Claude Code sessions on one machine share its stack, its GPU, its working tree and each other's processes. Continuum's sessions take leases on what they hold, keep their own state where no other session writes, and hand off through a journal whose catch-up reached each new session as a 2,000-character preview for nine days, until the hook was given a budget.

  • m2
  • deployable
2026-09-28 · 4 min read · Continuum Engine

The codebase remembers itself.

Continuum ingests its own documentation into its own memory, and Claude Code reaches it with two MCP tools. On 28 September the manifest seeded 208 entries from 128 sources, and Lightning extracted 1,557 facts from them in 19 minutes. A session pays for that knowledge only when it asks.

  • m2
  • tool-access
  • knowledge-extraction
  • one-truth
2026-09-28 · 4 min read · Continuum Engine

CLM-8B only answers multiple choice.

VentureBeat reported that Stanford and NVIDIA's open CLM-8B runs agent decisions up to 9× faster than Jev. Read from its source, it is a frozen Qwen3-8B encoder with two small heads that score a list of options, which fits the choices Continuum's receptionist makes on every message and nothing that has to write or abstain.

  • m2
  • model-routing
  • tool-access
2026-09-26 · 4 min read · Continuum Engine

The only timeout that fired was the HTTP client's.

A 32,000-token job failed at 100 seconds as a direct call and finished in 192 through a Temporal workflow. Claude's first write-up then blamed Temporal for an arm that failed on an HTTP client's timeout, though none of Temporal's own limits had fired.

  • m2
  • durable-workflows
  • honest-measurement
2026-09-26 · 4 min read · Continuum Engine

vLLM wouldn't repeat itself.

At seed 42, one request at a time, LM Studio's Lightning gave one answer 200 times in 200, and vLLM's answered one prompt at seven different lengths. Tristan was sure a setting was missing. vLLM has one, and in 0.29.0 it refuses every model Continuum runs.

  • m2
  • honest-measurement
  • provenance
2026-09-26 · 4 min read · Continuum Engine

The quality numbers were measuring markdown.

The first durable-extraction pass rejected every fact in LM Studio's smoke test for evidence not found in the section. The evidence was there, behind bold, backticks and blockquote markers, and so was the evidence behind 588 of vLLM's 824 rejections.

  • m2
  • knowledge-extraction
  • honest-measurement
2026-09-24 · 5 min read · Continuum Engine

The only part of a fact a machine can check is the quote.

Continuum pulls facts out of its own documents with a model. Three days of runs, a judge that had never seen the rules, and a design session about gravity worked out what it can actually check.

  • m2
  • knowledge-extraction
  • provenance
2026-02-24 · 4 min read · Continuum Engine

Most of the reader's saving was the benchmark's.

The first week of Continuum was a Redis stream writer, a reader, and a BenchmarkDotNet project to hold them to account. The reader's allocations fell by up to 43%, and most of that was the benchmark.

  • m0
  • streaming
  • async-inference