Carebun Interactive/Memory Memory labCheat sheet

Part V · Practice

Chapter 25

Whiteboard it cold

The whole book on a few pages: one paragraph, one picture, the pocket numbers, twelve sentences, a 45-minute interview script and the ten mistakes to avoid.

14 min read

On this page
  1. The one-paragraph answer
  2. The picture
  3. The pocket numbers
  4. Twelve sentences to carry
  5. The 45-minute interview script
  6. Ten mistakes that fail the interview
  7. Carry this
  8. Check yourself
  9. Back to Thursday

In this chapter, we will learn to give the whole design from memory, with a pen and an empty whiteboard. Everything here was taught earlier. This is the sheet to read the night before, and the script to follow on the day.

The one-paragraph answer

If you may say only one paragraph, say this one.

An LLM is stateless, so memory is a layer around it. After each session an extractor model reads the transcript once, in the background. It writes a few short, self-contained, dated facts into a store, where every row carries a user_id. Before each answer we embed the question and run a search filtered to that one user. Only the top few facts go into the prompt: about 75 tokens, flat, however much is stored. The read path is synchronous and stays under 50 ms. The write path is asynchronous and may take seconds. When a fact changes we close the old row and point it at the new one, so the past stays answerable. Events fade after about 90 days unused. Secrets are refused at write time, and the user can see, edit and delete everything. We measure it in two layers, and we never quote a score without its models, its judge and its token cost.

The picture

One diagram holds both paths. Draw the read path on top with solid arrows and the write path below with dashed ones. Write the speed next to each.

READ PATH   synchronous, about 28 per second, under 50 ms, no LLM call

  user message
       |
       v
  orchestrator ----> session store (this chat's turns, about 1 ms)
       |
       +-----------> query embedder (10 to 30 ms)
       |                    |
       |                    v
       |             memory store, filter: user_id = maya (1 to 10 ms)
       |                    |
       |                    v   top 3 facts, about 75 tokens
       +-----------> prompt = system + memories + session turns + message
                            |
                            v
                     assistant LLM ----> first token at 200 to 500 ms, then a stream


WRITE PATH  asynchronous, about 4.6 per second, seconds are fine

  session ends  - - >  extraction queue  - - >  extractor LLM (temperature 0, strict JSON)
                                                    |   1,500 tokens in, about 3 facts out
                                                    v
                                         decide: new, duplicate, or replaces an old fact
                                                    |
                                                    v
                                         memory store: insert the new row,
                                         close the old row (valid_to, superseded_by)

The pocket numbers

The pocket numbers are the few figures that size the whole system, and the traffic figures all come from five inputs. Learn the inputs and the derivations, not the results. Then a different product is only a different multiplication.

the five inputs
registered users                 1,000,000
daily active users                 200,000
sessions per user per day                2
turns per session                        6
tokens per session                   1,500

Storage needs a few design choices on top: 200 memories per user, 768 numbers per vector, about 250 bytes of text and payload per memory, and 3 replicas. Traffic and cost need two more: a peak of 3 times the average, and the prices of the models.

NumberValueHow it is derivedChapter
Registered users1,000,000given5
Daily active users200,000given5
Sessions per day400,000200,000 x 2 per user5
Messages per day2,400,000400,000 x 6 turns5
Tokens per session1,500given; 1,500 / 6 = about 250 per turn5
Transcript per user per year1,095,000 tokens2 x 365 = 730 sessions, x 1,5001
Facts per sessionabout 3, at about 25 tokensextractor average6
Injected memoryabout 75 tokenstop 3 x 258
History against memory, session 10200 times15,000 / 752
History against memory, session 1002,000 times150,000 / 752
Readsabout 28 per second, peak about 832,400,000 / 86,400, then x 35
Writesabout 4.6 per second, peak about 14400,000 / 86,400, then x 35
Read budgetunder 50 msembed 10 to 30 + search 1 to 103
First token200 to 500 msthe assistant model3
One vector3,072 B768 x 4 bytes7
One memoryabout 3.3 KB (3,322 B)3,072 + about 250 of text and payload7
Memories200,000,0001,000,000 x 200 per user5
Raw storeabout 664 GB200,000,000 x 3,322 B16
Vectors after INT8154 GB200,000,000 x 768 B16
One full copy204 GB154 + 50 of text and payload16
On disk612 GB204 x 3 replicas16
Memory layer per day, small assistantabout $312126 + 127.80 + 1.56 + 27 + 3016
Memory layer per day, frontier assistantabout $825126 + 127.80 + 1.56 + 540 + 3016
Resending 10 sessions per day$5,400 or $108,0002,400,000 messages x 15,000 tokens = 36 billion tokens, x $0.15 or $3 per million16
Recency0.995 ^ hourshours since last access9
Event expiryabout 90 daysafter last access12
the two checks worth doing in your head

memory against history   $312 / $5,400     = 5.8 percent    (small assistant)
                         $825 / $108,000   = 0.76 percent   (frontier assistant)

vectors in the store     200,000,000 x 3,072 B = 614 GB
                         614 GB / 664 GB   = 92 percent of the raw size

Twelve sentences to carry

One sentence for each big idea, in the order of the book.

  1. An LLM is a frozen file plus a desk, so memory is a layer around the model, not inside it.
  2. History grows with every session and memory stays flat: about 75 tokens today and in ten years.
  3. The loop is recall, answer, remember: read urgent, write lazy.
  4. Short-term memory is the context window. Long-term memory is the box, with facts, events and rules.
  5. The secretary keeps about 5 percent, and a card must stand on its own, with real names and real dates.
  6. One store for all users, and the user_id filter on every query is the privacy wall.
  7. Similarity finds the topic. Recency and importance decide what matters. Meaning, words and names search together.
  8. The top k is a dial, and the budget is counted in tokens, not rows.
  9. Never tear up a card: close it, date it, and point it at its replacement.
  10. Forgetting is precision: dedup, supersede, decay, erase.
  11. Memory fails quietly, so test in two layers and label every wrong answer.
  12. A score is a claim until it names its answer model, its judge, its question count and its tokens.

The 45-minute interview script

The script follows eight sections, always in the same order. The minutes are a guide. Speak the numbers aloud as you write them, and draw before you explain.

MinutesSectionWhat to sayWhat to draw
0 to 61. Requirements and scopeTell Maya's story in two lines. Kill the three fixes: resend history, bigger window, fine-tune per user. State five functional requirements and the targets: under 50 ms, about 75 tokens, write off the critical path. Say what is out of scope.The five inputs, then reads 28 per second, writes 4.6 per second, 664 GB.
6 to 122. High-level architectureRecall, answer, remember. Two paths with two speeds. Walk one request from message to first token, then to the background extraction.The picture above, with the milliseconds.
12 to 183. Data and knowledge pipelineOne row, field by field. What the extractor keeps and drops. What an embedding is. One collection filtered by user. Profile against collection.One row with its fields, and 768 x 4 + 250 = 3,322 B.
18 to 234. Models and promptsThree roles: the assistant is hot, the extractor is warm, the embedder is married to the store. Name the prompts: extraction, injection, update decision, answer.The injection prompt in five lines, with dated facts.
23 to 305. OrchestrationThe rolling summary for long sessions. Scoring with relevance, recency and importance. Consolidation: new, duplicate or replacement. Supersede instead of overwrite. Decay by category.London closed, Paris open, with the pointer between them.
30 to 356. GuardrailsThe filter is the boundary. The write policy. The four user controls. The right to be forgotten reaches vectors, caches and logs. Poisoning and its four defences.RIGHT and WRONG: the query with and without the filter.
35 to 407. Scaling and costStateless pods, a session store, a sharded memory store with three replicas, a queue with workers. Quantize the vectors. Read the daily bill line by line.664 GB to 204 GB to 612 GB, then $312 against $5,400.
40 to 458. Monitoring and evaluationEight dials. Two layers of golden sets. The three public benchmarks and what each one asks. The loop from user signals to a replayed golden set. End with what you would build first.The loop: signals, triage, failing layer, fix, replay, ship.

Ten mistakes that fail the interview

Each mistake below breaks a rule that an earlier chapter paid for.

  1. Forgetting the user_id filter. Nothing crashes. Another user's life appears in the results. (Chapter 15)
  2. Overwriting instead of superseding. "Where did I use to live?" can never be answered again. (Chapter 11)
  3. Putting extraction on the critical path. The user waits for two model calls where one would do. (Chapter 13)
  4. Quoting a benchmark score without its judge and token cost. The judge alone is worth about 5 points on identical answers. (Chapter 18)
  5. Proposing to resend all history, or to wait for a bigger window. The bill grows with every session. (Chapter 2)
  6. Storing everything. Stale and trivial cards take the places of the right ones. (Chapter 12)
  7. Ranking by similarity alone. It cannot tell a birthday from printer ink. (Chapter 9)
  8. Extracting from tool output or web pages. One injected sentence becomes a permanent implant. (Chapter 15)
  9. Sizing storage on active users. Traffic follows active users. Storage follows registered ones. (Chapter 5)
  10. Having no plan to test it. Memory fails quietly, and without a golden set nobody notices. (Chapter 17)

Carry this

  • One paragraph, one picture, five inputs. Everything else can be derived from them.
  • Eight sections in a fixed order: requirements, architecture, data, models and prompts, orchestration, guardrails, scale and cost, monitoring and evaluation.
  • Read urgent, write lazy. Filter by user. Supersede, never overwrite.
  • Memory costs 5.8 percent of resending history with a small assistant, and 0.76 percent with a frontier one.
  • A number without its derivation, or a score without its judge, is a rumour.

Check yourself

1. Without looking: derive the reads per second and the writes per second from the five inputs.

Answer
sessions per day    200,000 x 2         =   400,000
messages per day    400,000 x 6         = 2,400,000
reads               2,400,000 / 86,400  = about 28 per second, peak x 3 = about 83
writes                400,000 / 86,400  = about 4.6 per second, peak x 3 = about 14

Reads follow messages. Writes follow sessions, because extraction runs once per session.

2. Derive the size of the store, raw and on disk.

Answer
one memory     768 x 4 + about 250          =       3,322 B
memories       1,000,000 x 200              = 200,000,000
raw            200,000,000 x 3,322 B        = about 664 GB
INT8 vectors   200,000,000 x 768 B          = 154 GB
one copy       154 + 50                     = 204 GB
three replicas 204 x 3                      = 612 GB

3. Maya moved from London to Paris. Name what happens to the London row, and two questions that only this design can answer.

Answer

The London row is kept. It receives an end date and a pointer to the Paris row. Only this design can answer "where did I live before Paris?" and "when did I move?". An overwrite answers neither.

4. A vendor slide says "94 percent on LongMemEval". Which four things do you ask for before you believe it?

Answer

The answer model, the judge, the number of questions, and the context tokens per question. A sample can flatter by ten points: our 100-question sample scored 91 and the full 500 scored 81.0. A different judge can move identical answers by about five.

5. The assistant gives a wrong answer about the user. Name the five labels, and what each one tells you to fix.

Answer

NOT STORED: the fact never became a card, so look at extraction. NOT RETRIEVED: the card exists but did not reach the prompt, so look at retrieval. IGNORED: the card was on the desk and the answer missed it, so look at the answer prompt or the reader. WRONG DATE: a problem of time. JUDGE: the grader is wrong, so note it and do not chase it.

Back to Thursday

The book began with a Thursday. Maya asked for a snack for her kids' school trip, and an assistant that had said "Noted!" on Monday suggested peanut butter energy balls.

Now run that Thursday again. On Monday night, after the reply was sent, the secretary read the chat and wrote one card: Maya has a serious peanut allergy. On Thursday the librarian found it in a few milliseconds, filtered to Maya's box alone, and placed it on the desk with two others. The model was the same model. It was still frozen, and it still forgot everything when the chat ended. It suggested cheese and fruit.

That is all memory is: a few true cards, kept outside the model, and the right ones handed over at the right moment. You can now draw it, size it, price it, test it and defend it. Close the book and try the whiteboard cold.