Part V · Practice
Chapter 23
The memory lab
A whole memory layer that runs in your browser, and the main calculators of the book in one place. Learn memory by moving its numbers.
On this page
In this chapter, we will learn by doing. First we run a complete memory layer inside this page: we tell it things, ask it things and let months pass. Then we walk through seven experiments, each with one question, one prediction to make and one instrument to test it. Nothing here needs an account, a key or a network call.
How to use the lab
Predict first, then move the slider. Each experiment below asks a question. Write your guess down before you touch anything. A wrong guess that you then correct stays in your head far longer than a right answer that you only read.
Part one: a memory layer in your browser
The simulator below has a secretary, a box and a librarian. It is small enough to run in a web page, and it follows the same design as the real library.
| Piece | In the simulator | In hanumemAI |
|---|---|---|
| The secretary | A handful of rules that look for patterns such as "I live in ..." | An LLM that reads the whole session |
| Meaning | Shared words plus a small table of related words | A neural embedding with 2,560 numbers |
| Exact words | A BM25-style score | BM25 from SQLite's full-text index |
| Names | Capitalised words in the question | Entities written by the extractor |
| Categories, write policy, dedup, closing a card, two dates, top k, decay, erase | The same | The same |
A guided tour in eight steps
- Press "Play the whole script". Eight messages are sent as Maya. The log shows the newest line on top and keeps the last 12 lines, so read it from the bottom up. Count the cards in the box: 8 open and 1 closed. Three sentences got no card. Which ones, and why?
- Find the closed card. "Maya lives in London" is struck through. It has an end date. It was not deleted.
- Find what was refused. The wifi password produced a refusal. Look at the log line: it names the kind of information and never the secret.
- Ask "What should I cook for dinner?" The question does not contain the words "vegetarian" or "allergy". Look at the column called meaning. That is how the cards were found.
- Ask "Gift ideas for Sam?" Now look at the column called names. The birthday card has a 1 there, because it contains the name Sam. Set the weight of names to 0 and ask again. The score of the card falls by 0.10, which is the weight 0.1 times the 1. It still wins, because the meaning and the words also match.
- Ask "Where did I live before?" The closed card may now compete, because the question is about the past. Its line ends with "outdated as of".
- Press "Let 30 days pass" three times, then "Run forget()". The dentist visit expires. The allergy does not. Events fade. Facts and rules stay.
- Press delete_user. The box is empty. One call, everything gone, vectors included.
Part two: seven experiments
Experiment 1: the bill for having no memory
Question: Maya has used the assistant for 100 sessions. How many times cheaper is one message with memory than one message with her whole history?
Predict: ten times? A hundred? More?
What to notice: set the sessions to 100. The answer is 2,000 times. Then raise the facts injected from 3 to 30. Memory is still 200 times cheaper. The saving does not depend on injecting almost nothing. It depends on the read cost being flat.
Experiment 2: size a product in one minute
Question: a product has 5 million users and 10 percent of them are active each day. How many searches per second must the memory store serve at peak?
Predict: tens, hundreds or thousands? Then set the first two sliders and leave the others at the class values.
daily active 5,000,000 x 10% = 500,000 sessions a day 500,000 x 2 = 1,000,000 messages a day 1,000,000 x 6 = 6,000,000 reads 6,000,000 / 86,400 = about 69 per second at peak 69 x 3 = about 208 per second
What to notice: reads follow messages, writes follow sessions, and storage follows registered users. Double the daily active share and the storage does not move at all.
Experiment 3: three dials
Question: which single signal, used alone, gives the worst answer to a gift question: relevance, recency or importance?
What to notice: to use one signal alone, set the other two weights to 0. With recency alone, a note about printer ink wins and the birthday is last. With importance alone, the card about the children (importance 7) beats the card that says Sam loves vinyl records (importance 6). No single signal is enough, which is why the score is a sum.
Experiment 4: how many cards on the desk?
Question: for the dinner question, what is the smallest k that brings both the diet and the allergy to the desk?
First pick the question "What should I cook for dinner tonight?" in the instrument. Then move k. The scores for this question are illustrative, as the instrument says.
What to notice: with k = 3 the allergy stays behind, in fourth place with 0.352 against 0.357. With k = 4 it arrives, for 8 more tokens. The class design uses a top 3 of about 75 tokens. hanumemAI uses k = 30 inside a token budget, about 1,470 tokens on LoCoMo. Top 3 is a dial, not a law, and the safe setting depends on what a miss costs.
Experiment 5: a box that never forgets
Question: after two years, how many cards does one active user have if nothing is ever merged or expired?
cards per month 2 sessions x 3 facts x 30 days = 180 after 24 months 180 x 24 = 4,320
What to notice: 4,320 cards, against the 200 the class sized for. The instrument counts months of 30 days. With years of 365 days the same sum gives 4,380, the number in Chapter 12. Storage is the small problem. The large problem is that several cards may now say where Maya lives, and the librarian may pick a stale one first.
Experiment 6: shrink the vectors
Question: 200 million memories with 768 numbers each. How many gigabytes do the vectors need, and how many after storing each number in one byte instead of four?
float32, 4 bytes 200,000,000 x 768 x 4 B = 614.4 GB int8, 1 byte 200,000,000 x 768 x 1 B = 153.6 GB, shown as 154 GB
What to notice: 614 GB becomes 154 GB. The text and payload, 50 GB, cannot be compressed this way. At 4 bits per number, recall falls by 1.6 points in our measurement, from 1.000 to 0.984. At 2 bits it falls by 6, to 0.940. Recall here means how many of the true 10 nearest cards the smaller vectors still find.
Experiment 7: the daily bill
Question: which line of the memory bill is the largest with a small assistant model? And with a frontier model?
extraction $126.00 update decisions $127.80 write path 126.00 + 127.80 = $253.80 of $312
What to notice: with a small assistant the write path is the bill: extraction and update decisions are $253.80 of $312. With a frontier assistant the 75 injected tokens become the largest line, $540. Memory costs 5.8 percent of resending history with the small assistant and 0.76 percent with the frontier one.
More instruments
Eight more instruments from earlier chapters are here as well, for reference.
The day the desk overflows
One long session and its rolling summary
How fast recency fades
Who waits for the secretary
What is on the desk
One line of code between two lives
Ask the box about any day
A score and its price
Common questions
Are the numbers in the lab real?
Each instrument says where its numbers come from. Calculators derive everything from the sliders and print the arithmetic. Tables marked measured come from the class experiments or from the hanumemAI experiment log. Scores marked illustrative were made up in the style of the measured ones, to give you more questions to play with.
Why does the simulator sometimes miss a fact that I typed?
Its secretary is a short list of patterns. It understands "I live in Paris" and not "Paris has been home since spring". A real extractor is an LLM, which reads both. The simulator is honest about the rest of the system, and deliberately weak in this one place so that it can run in a page.
Does anything I type leave my browser?
No. The lab makes no network calls. The box lives in the page's memory and disappears when you reload.
Can I run the real thing?
Yes. Chapter 22 has a program of about sixty lines and its real output. The ten class experiments run on a laptop with Ollama, a program that runs models locally. Some of them also need a Qdrant container. The chapters that use an experiment name it and show the command.
Carry this
- Memory's read cost is flat. That single property produces the 200 times and the 2,000 times.
- Reads follow messages, writes follow sessions, storage follows registered users.
- No single signal ranks well alone. Recency alone retrieves trivia.
- k is a dial. Set it by what a missed card costs, and cap it with a token budget.
- Forgetting is precision. Closing a card keeps the past answerable. Deleting it does not.
Check yourself
1. In the simulator, play the script and ask "Any good bakeries near me?". Which card must be on the desk, and which card must not be?
Answer
The Paris card, "Maya moved from London to Paris", must be there. The card "Maya lives in London" must not, because it is closed. A design without closing would hold both as equally true, and the librarian could hand over London first.
2. Use the scale sheet. A product has 10 million registered users, 5 percent daily active, 3 sessions a day and 8 messages per session. What is the read rate?
Answer
daily active 10,000,000 x 5% = 500,000 sessions a day 500,000 x 3 = 1,500,000 messages a day 1,500,000 x 8 = 12,000,000 reads 12,000,000 / 86,400 = about 139 per second
3. In experiment 7, why does switching the assistant from a small model to a frontier model change only one line of the memory bill?
Answer
Only the injected memory tokens are read by the assistant model. Extraction and update decisions run on the small write-path model, embeddings on the embedding model, and the store is storage. So only the injected-tokens line is priced at the assistant's rate: $27 becomes $540.