From the build

Dev Entries

Lessons from building Continuum, each one taken from the real work and the journal that recorded it. Two to five minutes apiece, and most of them learned the interesting way.

Entries

2026-09-28 · 4 min read · Continuum Engine

Extraction already wrote the question each fact answers.

Tristan called it pre-cognition: questions written ahead of time for every fact, so a question can match a question. The question-first extraction prompt was already producing them, Continuum embedded each one beside its statement, and its retrieval half turned out to have a name: HyPE, hypothetical prompt embeddings.

  • m2
  • knowledge-extraction
  • recall
2026-09-28 · 5 min read · Continuum Engine

Don't pay a frontier model to be a shell script.

A Continuum Engine session started by paying for 300 KB of instructions, and on 26 September its sessions ran restarts and store queries by hand that code could run for no tokens at all. Glue Becomes Code retires each step on the cheapest rung that can do it: code, an MCP tool, a local model measured on the job, and a paid model only with confirmation.

  • m2
  • model-routing
  • tool-access
2026-09-28 · 5 min read · Continuum Engine

Sessions that share a machine take leases.

Two Claude Code sessions on one machine share its stack, its GPU, its working tree and each other's processes. Continuum's sessions take leases on what they hold, keep their own state where no other session writes, and hand off through a journal whose catch-up reached each new session as a 2,000-character preview for nine days, until the hook was given a budget.

  • m2
  • deployable
2026-09-28 · 4 min read · Continuum Engine

The codebase remembers itself.

Continuum ingests its own documentation into its own memory, and Claude Code reaches it with two MCP tools. On 28 September the manifest seeded 208 entries from 128 sources, and Lightning extracted 1,557 facts from them in 19 minutes. A session pays for that knowledge only when it asks.

  • m2
  • tool-access
  • knowledge-extraction
  • one-truth
2026-09-28 · 4 min read · Continuum Engine

CLM-8B only answers multiple choice.

VentureBeat reported that Stanford and NVIDIA's open CLM-8B runs agent decisions up to 9× faster than Jev. Read from its source, it is a frozen Qwen3-8B encoder with two small heads that score a list of options, which fits the choices Continuum's receptionist makes on every message and nothing that has to write or abstain.

  • m2
  • model-routing
  • tool-access
2026-09-26 · 4 min read · Continuum Engine

The only timeout that fired was the HTTP client's.

A 32,000-token job failed at 100 seconds as a direct call and finished in 192 through a Temporal workflow. Claude's first write-up then blamed Temporal for an arm that failed on an HTTP client's timeout, though none of Temporal's own limits had fired.

  • m2
  • durable-workflows
  • honest-measurement
2026-09-26 · 4 min read · Continuum Engine

vLLM wouldn't repeat itself.

At seed 42, one request at a time, LM Studio's Lightning gave one answer 200 times in 200, and vLLM's answered one prompt at seven different lengths. Tristan was sure a setting was missing. vLLM has one, and in 0.29.0 it refuses every model Continuum runs.

  • m2
  • honest-measurement
  • provenance
2026-09-26 · 4 min read · Continuum Engine

The quality numbers were measuring markdown.

The first durable-extraction pass rejected every fact in LM Studio's smoke test for evidence not found in the section. The evidence was there, behind bold, backticks and blockquote markers, and so was the evidence behind 588 of vLLM's 824 rejections.

  • m2
  • knowledge-extraction
  • honest-measurement
2026-09-24 · 4 min read · Continuum Engine

A fact turned out to need an observer.

Continuum began with chunks and embeddings and moved to extracting facts. Defining a fact took the primary sources, a thought experiment about Alice, and an observer, and left a memory where truth depends on whose view it is.

  • m2
  • knowledge-extraction
  • provenance
2026-02-24 · 4 min read · Continuum Engine

Most of the reader's saving was the benchmark's.

The first week of Continuum was a Redis stream writer, a reader, and a BenchmarkDotNet project to hold them to account. The reader's allocations fell by up to 43%, and most of that was the benchmark.

  • m0
  • streaming
  • async-inference