From the build
Dev Entries
Lessons from building Continuum, each one taken from the real work and the journal that recorded it. Two to five minutes apiece, and most of them learned the interesting way.
Entries
The only timeout that fired was the HTTP client's.
A 32,000-token job failed at 100 seconds as a direct call and finished in 192 through a Temporal workflow. Claude's first write-up then blamed Temporal for an arm that failed on an HTTP client's timeout, though none of Temporal's own limits had fired.
vLLM wouldn't repeat itself.
At seed 42, one request at a time, LM Studio's Lightning gave one answer 200 times in 200, and vLLM's answered one prompt at seven different lengths. Tristan was sure a setting was missing. vLLM has one, and in 0.29.0 it refuses every model Continuum runs.
The quality numbers were measuring markdown.
The first durable-extraction pass rejected every fact in LM Studio's smoke test for evidence not found in the section. The evidence was there, behind bold, backticks and blockquote markers, and so was the evidence behind 588 of vLLM's 824 rejections.
The only part of a fact a machine can check is the quote.
Continuum pulls facts out of its own documents with a model. Three days of runs, a judge that had never seen the rules, and a design session about gravity worked out what it can actually check.
Most of the reader's saving was the benchmark's.
The first week of Continuum was a Redis stream writer, a reader, and a BenchmarkDotNet project to hold them to account. The reader's allocations fell by up to 43%, and most of that was the benchmark.
No entry matches that search, so try fewer words or clear a tag.