Carebun Interactive/Memory Memory labCheat sheet

Part V · Practice

Chapter 23

The memory lab

A whole memory layer that runs in your browser, and the main calculators of the book in one place. Learn memory by moving its numbers.

25 min read · interactive

On this page
  1. How to use the lab
  2. Part one: a memory layer in your browser
  3. Part two: seven experiments
  4. More instruments
  5. Common questions
  6. Carry this
  7. Check yourself

In this chapter, we will learn by doing. First we run a complete memory layer inside this page: we tell it things, ask it things and let months pass. Then we walk through seven experiments, each with one question, one prediction to make and one instrument to test it. Nothing here needs an account, a key or a network call.

How to use the lab

Predict first, then move the slider. Each experiment below asks a question. Write your guess down before you touch anything. A wrong guess that you then correct stays in your head far longer than a right answer that you only read.

Part one: a memory layer in your browser

The simulator below has a secretary, a box and a librarian. It is small enough to run in a web page, and it follows the same design as the real library.

PieceIn the simulatorIn hanumemAI
The secretaryA handful of rules that look for patterns such as "I live in ..."An LLM that reads the whole session
MeaningShared words plus a small table of related wordsA neural embedding with 2,560 numbers
Exact wordsA BM25-style scoreBM25 from SQLite's full-text index
NamesCapitalised words in the questionEntities written by the extractor
Categories, write policy, dedup, closing a card, two dates, top k, decay, eraseThe sameThe same

A guided tour in eight steps

  1. Press "Play the whole script". Eight messages are sent as Maya. The log shows the newest line on top and keeps the last 12 lines, so read it from the bottom up. Count the cards in the box: 8 open and 1 closed. Three sentences got no card. Which ones, and why?
  2. Find the closed card. "Maya lives in London" is struck through. It has an end date. It was not deleted.
  3. Find what was refused. The wifi password produced a refusal. Look at the log line: it names the kind of information and never the secret.
  4. Ask "What should I cook for dinner?" The question does not contain the words "vegetarian" or "allergy". Look at the column called meaning. That is how the cards were found.
  5. Ask "Gift ideas for Sam?" Now look at the column called names. The birthday card has a 1 there, because it contains the name Sam. Set the weight of names to 0 and ask again. The score of the card falls by 0.10, which is the weight 0.1 times the 1. It still wins, because the meaning and the words also match.
  6. Ask "Where did I live before?" The closed card may now compete, because the question is about the past. Its line ends with "outdated as of".
  7. Press "Let 30 days pass" three times, then "Run forget()". The dentist visit expires. The allergy does not. Events fade. Facts and rules stay.
  8. Press delete_user. The box is empty. One call, everything gone, vectors included.

Part two: seven experiments

Experiment 1: the bill for having no memory

Question: Maya has used the assistant for 100 sessions. How many times cheaper is one message with memory than one message with her whole history?

Predict: ten times? A hundred? More?

What to notice: set the sessions to 100. The answer is 2,000 times. Then raise the facts injected from 3 to 30. Memory is still 200 times cheaper. The saving does not depend on injecting almost nothing. It depends on the read cost being flat.

Experiment 2: size a product in one minute

Question: a product has 5 million users and 10 percent of them are active each day. How many searches per second must the memory store serve at peak?

Predict: tens, hundreds or thousands? Then set the first two sliders and leave the others at the class values.

daily active     5,000,000 x 10%      =   500,000
sessions a day   500,000 x 2          = 1,000,000
messages a day   1,000,000 x 6        = 6,000,000
reads            6,000,000 / 86,400   = about 69 per second
at peak          69 x 3               = about 208 per second

What to notice: reads follow messages, writes follow sessions, and storage follows registered users. Double the daily active share and the storage does not move at all.

Experiment 3: three dials

Question: which single signal, used alone, gives the worst answer to a gift question: relevance, recency or importance?

What to notice: to use one signal alone, set the other two weights to 0. With recency alone, a note about printer ink wins and the birthday is last. With importance alone, the card about the children (importance 7) beats the card that says Sam loves vinyl records (importance 6). No single signal is enough, which is why the score is a sum.

Experiment 4: how many cards on the desk?

Question: for the dinner question, what is the smallest k that brings both the diet and the allergy to the desk?

First pick the question "What should I cook for dinner tonight?" in the instrument. Then move k. The scores for this question are illustrative, as the instrument says.

What to notice: with k = 3 the allergy stays behind, in fourth place with 0.352 against 0.357. With k = 4 it arrives, for 8 more tokens. The class design uses a top 3 of about 75 tokens. hanumemAI uses k = 30 inside a token budget, about 1,470 tokens on LoCoMo. Top 3 is a dial, not a law, and the safe setting depends on what a miss costs.

Experiment 5: a box that never forgets

Question: after two years, how many cards does one active user have if nothing is ever merged or expired?

cards per month    2 sessions x 3 facts x 30 days = 180
after 24 months    180 x 24                       = 4,320

What to notice: 4,320 cards, against the 200 the class sized for. The instrument counts months of 30 days. With years of 365 days the same sum gives 4,380, the number in Chapter 12. Storage is the small problem. The large problem is that several cards may now say where Maya lives, and the librarian may pick a stale one first.

Experiment 6: shrink the vectors

Question: 200 million memories with 768 numbers each. How many gigabytes do the vectors need, and how many after storing each number in one byte instead of four?

float32, 4 bytes    200,000,000 x 768 x 4 B = 614.4 GB
int8, 1 byte        200,000,000 x 768 x 1 B = 153.6 GB, shown as 154 GB

What to notice: 614 GB becomes 154 GB. The text and payload, 50 GB, cannot be compressed this way. At 4 bits per number, recall falls by 1.6 points in our measurement, from 1.000 to 0.984. At 2 bits it falls by 6, to 0.940. Recall here means how many of the true 10 nearest cards the smaller vectors still find.

Experiment 7: the daily bill

Question: which line of the memory bill is the largest with a small assistant model? And with a frontier model?

extraction          $126.00
update decisions    $127.80
write path          126.00 + 127.80 = $253.80 of $312

What to notice: with a small assistant the write path is the bill: extraction and update decisions are $253.80 of $312. With a frontier assistant the 75 injected tokens become the largest line, $540. Memory costs 5.8 percent of resending history with the small assistant and 0.76 percent with the frontier one.

More instruments

Eight more instruments from earlier chapters are here as well, for reference.

The day the desk overflows

One long session and its rolling summary

How fast recency fades

Who waits for the secretary

What is on the desk

One line of code between two lives

Ask the box about any day

A score and its price

Common questions

Are the numbers in the lab real?

Each instrument says where its numbers come from. Calculators derive everything from the sliders and print the arithmetic. Tables marked measured come from the class experiments or from the hanumemAI experiment log. Scores marked illustrative were made up in the style of the measured ones, to give you more questions to play with.

Why does the simulator sometimes miss a fact that I typed?

Its secretary is a short list of patterns. It understands "I live in Paris" and not "Paris has been home since spring". A real extractor is an LLM, which reads both. The simulator is honest about the rest of the system, and deliberately weak in this one place so that it can run in a page.

Does anything I type leave my browser?

No. The lab makes no network calls. The box lives in the page's memory and disappears when you reload.

Can I run the real thing?

Yes. Chapter 22 has a program of about sixty lines and its real output. The ten class experiments run on a laptop with Ollama, a program that runs models locally. Some of them also need a Qdrant container. The chapters that use an experiment name it and show the command.

Carry this

  • Memory's read cost is flat. That single property produces the 200 times and the 2,000 times.
  • Reads follow messages, writes follow sessions, storage follows registered users.
  • No single signal ranks well alone. Recency alone retrieves trivia.
  • k is a dial. Set it by what a missed card costs, and cap it with a token budget.
  • Forgetting is precision. Closing a card keeps the past answerable. Deleting it does not.

Check yourself

1. In the simulator, play the script and ask "Any good bakeries near me?". Which card must be on the desk, and which card must not be?

Answer

The Paris card, "Maya moved from London to Paris", must be there. The card "Maya lives in London" must not, because it is closed. A design without closing would hold both as equally true, and the librarian could hand over London first.

2. Use the scale sheet. A product has 10 million registered users, 5 percent daily active, 3 sessions a day and 8 messages per session. What is the read rate?

Answer
daily active     10,000,000 x 5%    =    500,000
sessions a day   500,000 x 3        =  1,500,000
messages a day   1,500,000 x 8      = 12,000,000
reads            12,000,000 / 86,400 = about 139 per second

3. In experiment 7, why does switching the assistant from a small model to a frontier model change only one line of the memory bill?

Answer

Only the injected memory tokens are read by the assistant model. Extraction and update decisions run on the small write-path model, embeddings on the embedding model, and the store is storage. So only the injected-tokens line is priced at the assistant's rate: $27 becomes $540.