Contents
Dev Entries clm_receptionist

CLM-8B only answers multiple choice.

VentureBeat reported that Stanford and NVIDIA's open CLM-8B runs agent decisions up to 9× faster than Jev. Read from its source, it is a frozen Qwen3-8B encoder with two small heads that score a list of options, which fits the choices Continuum's receptionist makes on every message and nothing that has to write or abstain.

2026-09-28· 4 min read · Continuum Engine

Tick the box

CLM-8B scores options against a state, and never writes a word.

VentureBeat reported that Stanford and NVIDIA's open CLM-8B caches reusable agent actions and runs up to 9× faster than Jev. Tristan asked whether Continuum could use either. Claude answered from CLM's README, its model card and its source (heads.py, embedder.py, schema.py) rather than the article.

CLM calls itself a "System One" model. Base Qwen3-8B, the chat model rather than Qwen3-Embedding, runs frozen as an encoder under vLLM's pooling runner: last-token pooling, 4,096 dimensions, 2,048 tokens of input. Two 20M-parameter MLP heads, 75 MB between them, map the state and each candidate down to 512 dimensions. A candidate's score is a scaled cosine, and a softmax over the question's candidates is the answer. It asks three kinds of question, Noul (yes or no), Choice and Score, under Apache-2.0. The heads are locked to that encoder and its pooling.

# One Choice, in CLM's shape. The encoder is frozen Qwen3-8B,
# and each head maps its 4,096 dimensions down to 512.
state  = state_head(encode(message))
routes = [cand_head(encode(r)) for r in ROUTES]  # embedded once
logits = exp(logit_scale) * cosine(state, routes)
probs  = softmax(logits)  # sums to 1 over ROUTES

Cache and carry

Its speed comes from embedding the options once.

The caching in the headline is ordinary embedding reuse. A fixed set of actions is embedded once, so each call embeds only the new state. The authors give 28 ms for a new state on an RTX 4090, and under 1 ms for one it has seen before.

TypeSafe's Jev (15 September) takes the same shape of request. It runs only as a hosted service, in early access, billed per input token. Continuum is local-first, and Jev as a dependency would send every routed message to a third party, so it was ruled out on that alone.

Front desk

The receptionist makes a choice from a short list on every message.

Every plain message that reaches Continuum on a channel goes to the receptionist: Discord today, and each further channel through an adapter of its own. Its job, as designed on 26 September, is to answer the message itself, hand it to the live Claude Code session, start or resume a headless one, or ask what was meant.

Tristan

"Nano can even respond, it's I guess what an agent would do, make a determination."

Tristan

"I guess we will need a receptionist agent role."

So receptionist is a capability in the model registry, beside chat, extraction and judge, and its model and settings are a row in model_configs. Today it is whichever acceptable model is already loaded, the Nemotron 3 Nano 4B first, at 2.8 GB beside the embedder. In a two-message probe the Nano answered one itself and handed the other to a Claude session with the right tool call.

Tristan

"It also needs to have the same set of tools/mcp access that Claude has access to and the context appropriately."

The receptionist gets Continuum's MCP surface, the 23 tools Claude uses, and every tool call runs as the person who sent the message, with their roles. A 4B model reading a DM can never do more than its sender could, however it is talked to.

The route is a Choice over four fixed candidates, which is exactly CLM's shape, and so is picking a tool: 23 tools are a fixed list, embedded once. CLM can't answer or fill in a tool's arguments, so it would assist the Nano rather than replace it. It would pick the route and narrow the tools to the few that fit the message, and the Nano would write the answer or the call, reading a few tool schemas instead of all 23. The messaging build list runs every job on the cheapest tier that can do it, plain code first and the local model for routine judgement, and a scorer that judges without generating sits between the two.

None of the above

A softmax always picks one, so CLM stays out of recall.

CLM's probabilities are relative to the candidates it is given, and they always sum to one. A route list carries its own way out, the "ask" candidate, so a routing Choice can decline. Recall has no such candidate. Its candidates are the store's entries, and when none of them is relevant the softmax still ranks one first. Deciding that nothing clears the bar is the retrieval threshold's job, and activation's score breaks down term by term where CLM's is one opaque number.

Nor does it make any permission decision, a use its authors disclaim themselves. Continuum never lets a model's judgement be the permission: the tool call runs as its sender whatever the model decides.

Pay your way

CLM joins the routing battery, where it has to pay for 16.4 GB.

The encoder is 16.4 GB in BF16, on a card where the Nano takes 2.8 GB. Quantising it, or serving it through llama.cpp's --pooling last, would give embeddings the heads were never trained on, and the decisions would drift by an amount nobody has measured.

The authors' own material has warning signs. Their verifier results were measured with heads fine-tuned for the task, and the released checkpoint isn't one of them. In their quickstart, "Is this urgent?" scores 0.41 for a customer who was charged twice and can't get through, which anyone who has been that customer would score higher.

Decision

The routing battery planned for the receptionist, with ambiguous requests and ones that try to talk it into escalating, gains CLM as a third arm beside the Nano and Lightning. It is adopted only if it beats them on accuracy and on correct abstention by enough to pay for its VRAM. If it does, the heads are ported to C# (two MLPs, a cosine and a softmax) behind the existing embedding pipeline, so no Python runs in production.

← All Dev Entries