Extraction already wrote the question each fact answers.
Tristan called it pre-cognition: questions written ahead of time for every fact, so a question can match a question. The question-first extraction prompt was already producing them, Continuum embedded each one beside its statement, and its retrieval half turned out to have a name: HyPE, hypothetical prompt embeddings.
Ask around
Facts are stored as statements and asked for as questions.
With extraction working, Tristan hit the question every store hits after ingest.
Tristan
"I don't know how to make our data store of facts and document/sections actually useful and retrievable after ingestion."
He had two ideas. The first was background processing that evaluates the search results themselves, sub-conscious tuning. The second was the one he called pre-cognition.
Tristan
"…we have a set of facts, pre-cognition would like creating cached questions one could ask about those facts in advance, storing them for retrieval/caching purposes."
A fact is stored as a statement, and a reader asks for it with a question. Writing the question ahead of time closes that gap: a question then matches a question. Claude cited doc2query for the technique, from memory, and noted it hadn't been read.
Ask first
The extractor was already writing the question first.
The same day, a study of eight extraction prompts over four rounds had settled on one built around a question. For each sentence, the model writes the question someone would ask that the sentence answers, then the complete answer as a standalone sentence. On ten held-out sections it produced no rejected quotes against ten for the prompt it replaced, and no fragments against three. The journal's verdict: question-first framing "moved the output most". It kept a rule's reason with the rule, and it kept "is meant to" on intent.
The question came free with each fact. That is the pre-computed question Tristan had just proposed, already sitting in the output, and the decision that followed was to keep it in the fact's interpretation, the versioned layer from the last entry, so recall could match questions against questions.
A question can also carry a false premise. "How does retrieving context score candidates?" turned a plan into a present fact, because the answer had to be written as if the question were true.
Two of a kind
A question is a second point beside its statement.
On 28 September the projector began embedding each interpretation's question with its
statement, as a second point in the same Qdrant collection: a derived id, a
field: "question" in the payload, written in the same request. No new
collection and no change to the ledger. The existing 1,587 facts were re-projected once
by hand, statements and questions, in 34 seconds.
Three's a crowd
Three equal lists let the questions outvote the words.
The first design fused three lists at equal weight: words, statements and questions. It fixed one probe. "Which host is the only one allowed to call a model itself?" moved its answer from second, tied, to first. It broke another. "What does the Gateway never do?" lost its "never" answers, because the stored question "What does the Gateway do?" sits next to "…never do?" in embedding space, and two dense lists outvoted the one lexical list.
The kept design has one dense list. Each fact ranks by the nearer of its statement and its question, so meaning and words get one vote each. "The Gateway never calls a model." now ranks first for the "never" question, fourth in the lexical list and fifth in the dense one. The morning's dense-only recall had it nowhere in its top six. Four questions are not a probe set, and whether the questions help or hurt negation is for a probe set to say.
Great minds
Pre-cognition's retrieval half had a name, and Tristan hadn't been taught it.
That night Tristan asked whether pre-cognition already existed under another name.
Tristan
"You also forgot to tell me what pre-cognition was, vs. HyDE I got it wrong, but then you didn't bring up 'Hypothetical Prompt Embeddings'. I told you have to teach me if it already exists."
Tristan
"I was feeling all good yesterday having come up with that all on my own."
The retrieval half is HyPE, hypothetical prompt embeddings (Vake, Vičič and Tošić, 2026). Claude had filed his idea in the glossary under doc2query, the lexical ancestor, and never taught him the name or the paper. HyPE moves HyDE's generation from query time to indexing: several questions written ahead for each chunk, each embedded and pointing back to it, so a query is matched question to question with no model call at query time. The journal read the abstract and not the paper.
Nothing about pre-cognition trains an embedding. The questions are plain text, and the embedder Continuum already runs turns each one into a vector that goes into the store beside its statement's, which is HyPE with one question. The question also stays as text that outlives any one index, and the probe set below is made of it. Tristan's version goes further than the paper in two more ways: several questions per fact, and questions written in the background while the card is idle. The fix on Claude's side was a memory that says to teach the name, the paper and the difference in the same reply, never only to map it in a doc.
Held out
A question can be a key or a probe, and never both.
His first idea and his second turned out to be one loop. A question's gold answer is the fact it was written from, so scoring recall needs no judge: the store asks itself questions whose answers it knows and measures what comes back. The written questions are the probe set, Tristan's point that the second idea feeds the first.
The trap is a question that is both a retrieval key and a probe, which measures the store finding its own index. So the questions are split, most kept as keys and some held out as probes, with unanswerable questions to test abstention. A fact that no sensible question can be written for is probably noise, which gives the status-line facts a measured way out.