Part I · Novice
Chapter 5
Requirements and the scale math
What the memory layer must do, how fast, and how big it gets. Every number is derived from five inputs, so you can derive them again for any system.
On this page
In this chapter, we will learn how to turn a vague wish, "the assistant should remember its users", into a list of requirements and a page of numbers. We will derive every number from five inputs. The goal is not to memorise the results. The goal is to be able to produce them again, for any system, on a whiteboard.
The task
The class sets the task in one sentence: design the memory layer that makes an assistant remember its users across sessions. It comes with five inputs.
registered users 1,000,000 daily active users 200,000 sessions per user per day 2 turns per session 6 (6 user messages, 6 replies) tokens per session 1,500 so 1,500 / 6 = about 250 per turn
Daily active users are the people who use the product on a given day. Only one in five registered users shows up each day. This difference matters: traffic comes from active users, and storage comes from all users.
Functional requirements: what it must do
A functional requirement is a thing the system must be able to do. The class notes list five.
| Requirement | In Maya's words | Where we build it |
|---|---|---|
| Remember durable memories across sessions: facts, events, rules | "You know I am vegetarian." | Chapter 6 |
| Use them at the right moment, without being asked | "I should not have to remind you." | Chapters 8 and 9 |
| Update when life changes | "I moved to Paris" must beat "lives in London". | Chapters 10 and 11 |
| Never leak one user's memories to another | "Tom must never see my allergy." | Chapter 15 |
| The user can see, edit and delete everything | "Show me what you know. Forget that." | Chapter 15 |
Notice the second one. Maya does not say "use your memory". She asks for a snack, and the allergy must arrive by itself. A memory that has to be asked for is a search box, not a memory.
Non-functional targets: how well it must do it
A non-functional target is a number that says how fast, how small or how reliable the system must be. A target without a reason is a guess, so each one below carries its reason.
| Requirement | Target | Why this number |
|---|---|---|
| Read-path overhead | under 50 ms | invisible next to the model's 200 to 500 ms time to first token |
| Injected memory budget | about 75 tokens | the top 3 facts, flat no matter how much is stored |
| Write path | off the critical path | the user never waits for it |
| Memory freshness | by the next session | background extraction lands within seconds |
| Delete profile | within days, weeks or months | as the applicable rule requires |
The critical path is the chain of steps the user waits for. Everything on it must be fast. Everything off it may be slow.
Look at how the 50 ms was chosen. Nobody asked "how fast can a database be?". The question was "what is the largest delay the user will not notice?". The model already takes 200 to 500 ms before its first token. Fifty more is at most a quarter of the fastest case and a tenth of the slowest.
50 / 200 = 25 percent of the fastest first token 50 / 500 = 10 percent of the slowest first token
Scale math, step by step
Scale math is the arithmetic that turns the five inputs into traffic and storage. We do it in two halves: how often the system is called, and how much it must hold.
Traffic: how often
daily active users 200,000 sessions per day 200,000 x 2 = 400,000 user messages per day 400,000 x 6 = 2,400,000 seconds in a day 24 x 60 x 60 = 86,400 read path (once per message) 2,400,000 / 86,400 = about 28 per second write path (once per session) 400,000 / 86,400 = about 4.6 per second
Two things to see here.
First, reads and writes are counted in different units. The librarian works once per message. The secretary works once per session. A session has 6 messages, so there are 6 times more reads than writes.
28 / 4.6 = about 6
Second, traffic is not spread evenly over the day. People chat in the evening more than at four in the morning. Peak is the busiest moment, and the class sizes it at three times the average.
read path peak 27.8 x 3 = 83.3 about 83 per second write path peak 4.63 x 3 = 13.9 about 14 per second
We multiply the unrounded rates, 27.8 and 4.63. Rounding first would give 28 x 3 = 84.
We build for the peak, not for the average. A system sized for 28 reads per second fails every evening.
Storage: how much
Start with one card and count its bytes. A byte (B) is the unit of storage: one letter of plain text is about one byte. A kilobyte (KB) is 1,000 bytes. A gigabyte (GB) is 1,000,000,000 bytes.
one memory text about 25 tokens x about 4 letters = about 100 B vector 768 numbers x 4 bytes each = 3,072 B payload ids, dates, category, source = about 150 B total 100 + 3,072 + 150 = 3,322 B about 3.3 KB
The vector is the list of numbers that lets the librarian search by meaning. Chapter 7 explains it. For now, notice its weight: 3,072 of 3,322 bytes. The sentence itself is tiny. The numbers that describe it are thirty times heavier.
3,072 / 3,322 = 92 percent of a memory is its vector
Now multiply up.
memories per user about 200
total memories 1,000,000 x 200 = 200,000,000
raw size 200,000,000 x 3,322 B = 664,400,000,000 B
= about 664 GB
Storage uses registered users, all one million of them. A user who did not visit today still has cards in the box.
Change the inputs below. Watch which results move when you change the active users, and which move when you change the vector size. The widget adds the text and the payload into one figure of 250 B (100 + 150).
Where does 200 memories per user come from?
It is a design assumption, and an honest book should say so. An active user produces about 3 candidate facts per session.
candidates per day 2 sessions x 3 = 6 candidates per year 6 x 365 = 2,190 candidates in 2 years 2,190 x 2 = 4,380 4,380 / 200 = about 22 times more than we sized for
These are candidates, not cards. Most of them repeat something already known or replace an older card. A store that keeps every candidate is 22 times larger and much less accurate. Keeping the box at 200 cards is the work of consolidation (merging a new fact into the cards already there) and forgetting, in Chapters 10 and 12.
Scope: what is in and what is out
Scope is the line around the problem we agree to solve. Drawing it early keeps a design discussion from wandering.
| In scope | Out of scope |
|---|---|
| short-term and long-term memory | retrieval over documents (RAG) |
| fact extraction and the asynchronous write path | memory shared between several agents |
| embedding storage in a vector database | tools and planning |
| retrieval at inference time | |
| consolidation and conflict resolution | |
| forgetting and decay | |
| privacy of stored memories |
Inference is the moment the model produces an answer. "Retrieval at inference time" means the librarian works while the user waits.
Try it yourself: size a different product
No computer is needed for this one. Take a smaller product and run the same steps on paper. A study assistant has 50,000 registered students, 20,000 active each day, 3 sessions a day and 8 turns per session. Each student has about 100 memories.
sessions per day 20,000 x 3 = 60,000 messages per day 60,000 x 8 = 480,000 reads per second 480,000 / 86,400 = about 5.6 peak x 3 = about 17 writes per second 60,000 / 86,400 = about 0.7 peak x 3 = about 2 total memories 50,000 x 100 = 5,000,000 raw size 5,000,000 x 3,322 B = 16,610,000,000 B = about 16.6 GB
Seventeen reads a second and 16.6 GB fit on one modest server. The same method that gave 664 GB for the big product says this one needs no cluster at all. (A cluster is a group of machines working as one.) That is what scale math is for: it tells us how much machinery not to build.
Common questions
Why is the peak three times the average? Who decided that?
It is a rule of thumb for consumer products, where usage gathers in the waking hours of each region. A real system measures its own peak. In a design discussion, state the factor you chose and move on. The reasoning matters more than the exact factor.
Why does a number take 4 bytes?
Each number in the vector is stored as a 32-bit float, and 32 bits are 4 bytes. It can be stored more coarsely. With 1 byte per number, the vector shrinks from 3,072 to 768 bytes with almost no loss in search quality. Chapter 16 does this and cuts 664 GB to 204 GB.
Is 664 GB a lot?
It is too much for one laptop and ordinary for a production database. It is large enough that we must think about splitting the data over machines and keeping copies. It is small next to the transcripts themselves, which the memory layer does not keep.
Why are reads counted per message and writes per session?
Every message needs the right cards before its answer, so recall runs each time. The secretary works better on a whole conversation, so extraction runs once at the end. That choice divides the write load by six.
Why is "delete profile" allowed to take weeks?
A good design stops using the memory at once. Removing every copy takes longer, because copies can also sit in backups, caches and logs. Laws and company rules set the final deadline, which is why the class notes say "as per rule".
What if my users write much longer messages?
Then tokens per session goes up, and with it the cost of the write path, because the secretary reads the whole transcript. The read path barely changes, since it injects a fixed budget. The benchmark called BEAM in Chapter 18 has assistant turns that run to 500 words, and it taught us to cap what is injected in tokens, not in rows.
Carry this
- Five inputs produce everything: 1,000,000 users, 200,000 active, 2 sessions, 6 turns, 1,500 tokens.
- Reads: 2,400,000 messages / 86,400 seconds = about 28 per second, peak about 83. Writes: 400,000 sessions / 86,400 = about 4.6 per second, peak about 14.
- One memory is about 3.3 KB, and 92 percent of it is the vector.
- 200,000,000 memories x 3,322 B = about 664 GB.
- Traffic follows active users. Storage follows registered users.
Check yourself
1. Daily active users double to 400,000. Registered users stay at 1,000,000. What happens to reads per second and to the raw size of the store?
Answer
messages per day 400,000 x 2 x 6 = 4,800,000 reads per second 4,800,000 / 86,400 = about 56 (doubled) raw size 1,000,000 x 200 x 3,322 B = about 664 GB (unchanged)
Traffic follows active users. Storage follows registered users.
2. We switch to an embedding with 1,536 numbers per vector. How big is one memory, and how big is the store?
Answer
vector 1,536 x 4 = 6,144 B one memory 100 + 6,144 + 150 = 6,394 B store 200,000,000 x 6,394 B = about 1,279 GB
Doubling the vector almost doubles the store, because the vector is nearly all of a memory.
3. Why is the read target 50 ms and not 5 ms or 500 ms?
Answer
The model's first token takes 200 to 500 ms. Fifty milliseconds is small enough to go unnoticed beside it, and large enough to fit an embedding call of 10 to 30 ms and a search of 1 to 10 ms. Five would be impossible with an embedding call. Five hundred would double the wait.