Continuum Engine

Part Two

Forgetting Is a Forecast

Every time your memory lets something go, it is placing a bet about tomorrow - using odds it has been quietly collecting your entire life.

That's Fit To Print

Two researchers stopped studying people and started studying newspapers.

In 1991, John Anderson and Lauren Schooler did something slightly odd for memory researchers. They went and looked at the world instead.

They counted how often words show up in newspaper headlines, how often a parent hears one particular child's name, and how often a given email address gets used. For each one they asked a single question:

If I needed this t days ago, what are the odds I need it today?

Every source gave back the same shape of answer: steep at first, then a long shallow tail that never quite reaches zero. That shape turned out to be a near-perfect match for the shape of human forgetting.

Place a bet on your inbox

Think of somebody you emailed yesterday, and somebody you last emailed a year ago. Which of them is more likely to turn up in your inbox today?

You answered without doing any maths, and you would win that bet most days. The year-old contact is a long shot, but still in the race: people do resurface. That is the long shallow tail.

Your memory is forecasting, and it is well calibrated to the actual statistics of the world you live in.

Picture a bookmaker in the back of your head, pricing every memory you own at odds of being needed tomorrow. Calibrated means the prices are honest: the things it offers at evens really do come back about half the time.

The forgetting curve is that bookmaker's price list, drawn as a graph.

Anderson & Schooler, Reflections of the environment in memory, Psychological Science, 1991. Everything in this series rests on it.

In the swimming pool from part one, depth is these odds. A memory priced at short odds bobs near the surface; a long shot sits down in the dark. Activation measures odds, and odds are stored on a compressed scale.

A Warning About Scale

These numbers are lying to you, politely.

Odds get very large and very small very quickly, so they are stored compressed, the same trick we use for earthquakes.

A magnitude 7 earthquake is about thirty times worse than a magnitude 6: picture thirty magnitude-6 quakes going off under the same city at once. The numbers look small and friendly and evenly spaced, and they are quietly enormous.

On a scale like that, each step up the dial multiplies whatever was there before. It is called a logarithmic scale, and activation is one. A "+2 bump" is a 7.4× swing in the odds - enough to turn a memory the bookmaker had at eight to one against into something close to a coin flip.

Every instinct you have about scores - let's weight that a bit higher - is calibrated for a scale this is not.

Zero does not mean nothing

An activation of 0 means evens - a coin flip. It does not mean "no evidence." So the cut-off for what counts as retrievable has to be measured, never assumed. Set it to zero because zero feels like the natural boundary and your system will quietly return everything, or nothing, forever.

You cannot just bolt on a similarity score

Search engines score "how alike are these two things" from -1 to 1. Activation scores "what are the odds this is needed" on an open-ended scale. Adding one to the other is adding metres to kilograms.

It is the most common mistake in this field. We made it too, and it cost us months.

The Equation

Every time you use a memory, you ring a bell.

The bell starts loud and fades at a known rate. It never quite stops - it just gets quieter and quieter until it is lost under everything else in the room.

Every time you have ever used that memory, you rang the bell again. Some of those rings were years ago and are almost silent. One might have been this morning and is still humming.

How loud the memory is right now is the sum of every bell you ever struck for this memory, each at whatever volume it has faded to.

That total is the number. ACT-R calls it base-level activation, and calls the fading of each bell base-level decay. Written down, the whole idea is one line:

B = ln( Σ  age^-0.5 )
         ▲   ▲
         │   └── every past use, faded by how long ago it was
         └────── add them all up

That exponent of 0.5 means each use fades as one over the square root of its age. In bell terms: wait four times as long and the bell is half as loud; wait a hundred times as long and it is a tenth as loud.

A use this long ago is still worth
1 second1.0
100 seconds0.1
1 hour0.017
1 day0.0034
1 week0.0013
1 month0.00061
1 year0.00018

A memory you used a year ago is worth 0.00018. A memory you glanced at an hour ago is worth 0.017 - which means it would take ninety-four of those year-old uses to match that one glance. Stretch it to a use one second ago and you need five and a half thousand. Fill a concert hall, hand everyone a bell, have them all ring it once a year ago, and today the whole hall is as loud as a single bell struck a second ago.

Two problems, one equation

Take a song you played ten times last month and one you heard once on the radio this morning. Notice which one is humming in your head right now.

Ten uses a month ago add up to 0.0061. One use an hour ago is 0.017. The single recent use wins, by nearly three to one.

Nobody wrote a rule saying "recent beats frequent." The fading does it. Every system that carries a separate recency score and a separate popularity score - and then employs somebody to tune the balance between them - is solving a problem this equation does not have.

Nothing ever needs deleting

That year-old use is not pruned, archived or special-cased. There is no cleanup job. It simply stops mattering, smoothly, without anyone having to choose a cut-off date. In the pool, it just sinks a little further every day.

Compare the obvious alternative - a counter you tick up on every use. A counter has no memory of when. Fifty uses last year and fifty this morning give exactly the same number, which is a spectacular way to be wrong about what somebody currently cares about.

The ln wrapped around the outside of that line is a unit conversion.

It is the natural logarithm, the same squashing trick as the Richter scale. It turns "total loudness" into the bookmaker's "odds this is needed" - and it is the reason everything else in the model can simply add on top instead of multiplying.

Curve Ball

The most famous graph in memory research is the wrong shape.

In 1885 a German psychologist called Hermann Ebbinghaus decided to study forgetting by memorising thousands of meaningless syllables and testing himself at intervals for months, with no other participants. He was the experiment.

Out of that came the forgetting curve, the most reproduced graph in the psychology of memory. Ebbinghaus fitted his data with a formula built on the logarithm of elapsed time, but the curve that got reproduced ever since is a simpler exponential.

Almost every AI memory system built in the last three years also uses an exponential, mostly for the same reason a student reaches for one: it is the curve everybody has seen. Very few of them checked the alternative.

The alternative is the power law, the curve the bells follow. The two look nearly identical for about a day, then diverge.

Picture two coastlines. On one, the land runs flat and then stops at a cliff, and whatever walks off the edge is gone. On the other, a very long beach slopes gently into the sea, and you can wade out a long way before your feet leave the sand. Exponential decay is the cliff. The power law is the beach.

To line them up fairly, both curves below start at full strength at one day old. The exponential is given a half-life - the time it takes a memory to lose half its strength - of one day.

Age Power law Exponential
1 day1.001.00
1 week0.380.016
1 month0.180.0000000019
1 year0.052~0

Both curves set to 1.00 at one day; the exponential halves every day.

At one month the power law still holds a fifth of its strength, and the exponential is nearly nine orders of magnitude down - about five hundred million times weaker than it was at one day. Build a memory on the exponential and anything older than a fortnight is unreachable rather than faint.

Walk out onto the beach

Try to recall a conversation you had about a month ago. Who was there, roughly what it was about.

Most people get something back: faint and patchy, but there. On the cliff, that memory would be at about one part in five hundred million, somewhere far below the edge. Your own head just voted for the beach.

The best evidence comes from flashcards

There is an algorithm called FSRS that schedules revision in modern flashcard apps - the ones students use to cram vocabulary before an exam. It is fitted against tens of millions of real recall events, each one a person looking at a card and either remembering or not. Flip one card every two seconds, without stopping to sleep, and ten million of them takes you the better part of a year.

Its forgetting curve uses an exponent of −0.5: the same value as ACT-R's decay, which rests on the newspaper headlines Anderson analysed in 1991. FSRS arrived at it in 2023, three decades later, by different people, using a completely different method, on a dataset thousands of times larger.

FSRS did not start there. Earlier versions used an exponential. They switched, because the power law fit the data better.

That formula also contains a stray 1 + whose only job is to stop the curve exploding at the instant a card is first seen. Continuum's version of the ACT-R equation needed a similar constant at the old end: a small floor, β, inside the log, which stops an unused memory fading towards minus infinity. ACT-R's own base-level constant sits outside the log and does not do that job.

In the pool, β is the bottom. Without it the pool has no floor, and a memory nobody touches keeps sinking forever, which is a strange property for a swimming pool to have.

Next

Part 3 · What brings it back

So far everything depends on history alone. But you do not remember in a vacuum - what is already in your head changes what surfaces next. One consequence is that knowing more about something can make it harder to remember.

What Brings a Memory Back →