Recall finds what would come to mind, and Continuum says so when nothing does.
A question is put to Continuum's memory by its words, by its meaning, by both fused, and by ACT-R activation, four methods under test side by side. The model answers from what comes back and says the memory doesn't hold the answer when nothing does, and the review pipeline puts the same questions to every method.
UPDATED 2026-09-28 · VERSION 0.1
Recall
A search index returns its nearest matches whether or not they answer the question.
Continuum's memory recalls: it ranks what would come to mind, and when nothing in it
answers, the answer says so. Every MCP client, Claude Code, Copilot and Cursor among them,
every custom integration over the HTTP context API and ICortexClient, and
every person, through the Dashboard and the Discord bot, asks the same memory.
How best to find what it holds is being measured now. Four methods run side by side, and the review pipeline puts the same questions to each of them.
Nothing to declare
When the memory holds nothing that answers, Continuum says so.
A model answers anything it's asked, from whatever its training left behind. Continuum puts the recalled facts in front of it with one rule: answer from these alone, and say so when they don't answer. The review pipeline's prompt puts it exactly: "If they don't answer the question, reply exactly: Memory doesn't hold the answer."
On 28 September the first experiments asked, among their questions, who won the 1998 FIFA World Cup, which nothing in the memory knows. The closed-book baseline, the same model with no memory, answered France. Every method backed by the memory, by words, by meaning, fused or by activation, came back with the memory not holding the answer. One question is a single reading, and the probe sets grow from here.
Activation adds a line of its own. Under it, a fact is retrieved only above the
retrieval threshold, τ, and a fact below it drops into
belowThreshold, returned beside the answer, so "why wasn't this recalled?"
always has a reply.
Horses for courses
Each method finds a different fact first, so every one of them is under test.
| Method | Finds a fact by | Suits a question that |
|---|---|---|
| Lexical | Its words: Postgres full-text search over each fact's question, statement and subject, weighted in that order, the question's words OR'd together | Turns on one rare word |
| RAG | Its meaning: the question's embedding against each fact's statement in Qdrant, and against the questions the extractor wrote for it, each fact ranked by the nearer of the two | Is phrased nothing like the fact |
| Hybrid | Both: the 50 best facts by words and the 50 nearest statements and 50 nearest questions by meaning, pooled, and ordered by reciprocal rank fusion, 1 / (60 + rank) summed over the two lists a fact appears in | Needs a word match and a meaning match at once |
| Activation | The same pool, ordered by ACT-R activation and cut at the retrieval threshold | Should reach what a mind would reach first, and nothing when nothing clears |
| Graph, in the playground | The three nearest facts, then up to two hops out along the fact edges in FalkorDB | Is answered by a fact connected to the one that matches |
One question from 28 September shows why no single method is enough. Asked "What does the Gateway never do?", meaning returned six general facts about the Gateway, and not the answer. The words found it, because the fact that answers is one of the few that says "never". Meaning matches a paraphrase, and words match the one word that decides it. Hybrid holds both, and activation orders what they find by what the memory has behind it.
A caller doesn't choose: POST /v1/facts/recall, the MCP tool
context_recall, the Terminal's continuum context recall and a
chat turn all recall with the ranking the experiments are running, and the playground
and the experiments name a method per run, so two can be compared on the same question.
POST /v1/retrieval/playground
{ "query": "What does the Gateway never do?",
"retrievers": ["LexicalFacts", "DenseFacts", "Recall", "Activation"], "topK": 10 }
HTTP/1.1 200 OK
{ "results": [
{ "retriever": "Activation", "note": "…",
"items": [ { "rank": 1, "content": "…", "score": …,
"metadata": { "similarity": "…", "base_level": "…", "spreading": "…",
"ceiling": "…", "final": "…", "ceiling_applied": "no" } } ],
"belowThreshold": [ … ] },
… ],
"traceId": "…" }Top of mind
Activation ranks a fact by its history and its connections, and caps it at its evidence.
Activation is ACT-R's equation for how easy a memory is to reach right now, computed for
every candidate at the moment of the question and never written back:
A = min(B + Σ W·S + ε, τ + band). The
activation series tells the science; here is
what each term does in Continuum.
- Base level,
B.ln(β + Σ t^−d)over the ages of a fact's references, withd = 0.5, the decay Anderson and Schooler measured and flashcard schedulers found again. References older than the window are counted by Petrov's tail, which is ACT-R's own optimised learning. The floor,β = 0.002, was fitted in the evaluation harness, and it is what keeps a faded fact reachable by a strong enough cue. Until answers cite facts, a fact's one reference is its creation. - Spreading,
Σ W·S. What is in play lends strength to what it connects to. In fact recall the question is the one cue, and lends by meaning, scaled up toS_max = 4. In the design, the graph lends too: a declared connection lendsS_max − ln(fan + 1), so a vague cue with many connections lends less than a precise one, and it halves with every hop, which is Soar's multi-hop design where ACT-R spreads a single step. - The ceiling,
τ + band. However often a fact is recalled, it can rise no higher than its evidence: the kind of author who asserted it (a person 4, an agent 2, the system 1.5), plus 1 if it is grounded, 1 for each further independent author up to 2, and up to 2 either way for its record of being right. It is a cap, never an added weight, so repetition can't buy authority. - The threshold,
τ = −4.0. A fact is retrieved above it and not below, unless the validity gate has already excluded it. τ is the value the experiments are calibrating.
Every fact comes back with its score taken apart: similarity, base level, spreading, the raw sum, the ceiling, the final score, whether the ceiling applied, whether it is contested, and, under hybrid, its rank in each list and its fused score. The playground shows the lot, and lists the facts below the threshold under "why these weren't retrieved".
Join the dots
Recall finds the facts, and the answer still has to blend them.
A chat turn recalls up to eight facts and hands them to the model in one system message: "Facts recalled from this person's memory in Continuum for their message, strongest first. Use the ones that answer it, say so when none do, and state nothing as fact that isn't here or in the conversation." Then the statements, numbered. What comes back is grounded, and it reads like a pick from a list: the one statement that answers, repeated, where an answer should put several facts together into what they mean.
The model has four jobs with recalled facts: pick the ones that answer, combine several into one reply, qualify what they leave open, and refuse when none answer. Refusing works, and combining is the problem the experiments are for. Tristan, on 27 September: "If we are supplying questions, facts, and the user's question... What exactly is the LLM's purpose?"
The experiments come at it from three sides. From the memory's side, a fact is an index
and its section is the payload, so the grounded harness,
continuum model-quality grounded, can hand the model the same facts five
ways: none, plain, structured, scored, or with their source sections attached. From the
graph's side, the design resolves facts that say the same thing into one, so a question
reaches the facts around its answer as well as the answer. From the answer's side, each
variant in the review pipeline can pin its own chat configuration, so a prompt or a model
is a variant as a retriever is.
Trial and error
The review pipeline puts the same questions to every method, and a person rates the answers.
An experiment is questions, one per line, with ! marking the ones the memory
can't answer, crossed with variants: a method each, and one with none, the closed-book
baseline. Up to 50 questions and 12 variants make an experiment, each pairing a trial,
and the experiment runs as a Temporal workflow on the evaluation queue.
A trial retrieves with its method, assembles the prompt from what came back, each entry
cut at 2,000 characters, and asks the production chat model for the answer. Then it
checks the answer: that it finished and isn't only reasoning, that every citation like
[2] points at an entry it was given, and whether it abstained. Each step is
traced, and the Dashboard lays a run out as a tree on a timeline. A person rates each
answer good or bad beside the scoreboard.
- A balanced objective. Never fit a parameter on a metric that a weaker setting improves for free. Fitted on recall alone, the threshold falls until it stops firing, and the number improves while the gate dies, so abstention is scored beside recall.
- The exam isn't in the corpus. The journal, the benchmarks and the experiments' own documents are kept out of what is ingested, in code, since the document that would contaminate a run is always written after any list of them.
- No quiet substitutes. A trial fails if the runtime served a different model than the one asked for, or if its method can't run.
The next stages are designed: a model as judge beside the person, questions that need several facts to answer, rewriting a question before it is searched, and the grounded probe set as the experiments' standing dataset.
For the record
ClickHouse keeps what happened to each fact, and never what it said.
Base level decays over a fact's history of use, and that history is a record the other
stores can't rebuild. The event plane is ClickHouse's context_events: an
append-only receipt per thing that happened to a fact, Used,
Corroborated, Rejected, Vindicated,
Refuted or Traversed, with no content column, kept 180 days.
In the design, a Used receipt is written when an answer cites a fact, and
activation reads the windows of receipts as each fact's references.
Two kinds are left out on purpose. Created is read from the fact itself.
Presented must never be recorded, because being shown a fact is not
evidence of using it: the first reinforcement loop recorded every fact recall returned,
strengthened whatever the retriever already favoured, and was removed a day after it was
built.
Why it's built this way
Every candidate is scored once, by a theory that can be checked.
Candidates are pooled, then fused.
The first design was a cascade: Qdrant first, then graph expansion, then scoring over what they found. Recall was bounded by whatever Qdrant returned first, and a fact that meaning missed could never be scored at all, which is exactly how an assistant forgets the rule it was given. So candidates come from every source at once, pooled, fused by rank and scored once. The pool is two lists, not three: with statements and questions as two dense lists beside one of words, meaning outvoted the words and lost "never".
Activation has to beat hybrid to earn its place.
A learned ranker needs labelled relevance data, and there was none. ACT-R gives a prior fitted against human memory over four decades, with every term named, which is what makes "why wasn't this recalled?" answerable. It still has to win: what activation adds to the fused ranking is measured on the same questions, and it stays for what it adds.
A score is computed when it's asked for.
Activation depends on now: the decay term is a function of the time since each reference. A score written back to the row is stale the moment after the write, and it would alter a fact from a model that may later be found wrong, which the context store's first law forbids. Computed at query time, a corrected constant rescores the whole memory on the next question, with no migration.
- A receipt per reference. A
reference_countcolumn is a ratchet with no memory of when: old use never fades, and the decay term stops meaning anything. Base level needs each reference's time. Receipts are a columnar workload, append-only, never updated and read in time ranges, so they went to ClickHouse, over PostgreSQL, KurrentDB, Marten, Kafka and the time-series databases. - Each mechanism to its architecture. The papers were read again from the PDFs in September, and two corrections were structural. Petrov's hybrid, which the specification had treated as new, is ACT-R's optimised learning, and multi-hop spreading, credited to ACT-R for months, is Soar's.
Lab notebook
Recall was specified and fitted offline, then put to live questions in a week.
| When | What landed |
|---|---|
| 20 Sep | Candidates pooled from every source and fused, in place of the cascade |
| 21 Sep | The event plane on ClickHouse; the papers read again; the base-level floor fitted in the harness; the reinforcement loop removed |
| 27 Sep | Fact recall through Cortex, for the MCP tool, the Terminal and every chat turn |
| 28 Sep | The playground, traces and experiments; the lexical arm over facts; the first experiments, where every method backed by the memory declined the question it couldn't answer |
| 28 Sep | Questions embedded beside each fact's statement, and hybrid and activation added as variants to the review pipeline |