Carebun Interactive/Memory Memory labCheat sheet

Part I · Novice

Chapter 5

Requirements and the scale math

What the memory layer must do, how fast, and how big it gets. Every number is derived from five inputs, so you can derive them again for any system.

10 min read · interactive

On this page
  1. The task
  2. Functional requirements: what it must do
  3. Non-functional targets: how well it must do it
  4. Scale math, step by step
  5. Scope: what is in and what is out
  6. Try it yourself: size a different product
  7. Common questions
  8. Carry this
  9. Check yourself

In this chapter, we will learn how to turn a vague wish, "the assistant should remember its users", into a list of requirements and a page of numbers. We will derive every number from five inputs. The goal is not to memorise the results. The goal is to be able to produce them again, for any system, on a whiteboard.

The task

The class sets the task in one sentence: design the memory layer that makes an assistant remember its users across sessions. It comes with five inputs.

registered users                 1,000,000
daily active users                 200,000
sessions per user per day                2
turns per session                        6   (6 user messages, 6 replies)
tokens per session                   1,500   so 1,500 / 6 = about 250 per turn

Daily active users are the people who use the product on a given day. Only one in five registered users shows up each day. This difference matters: traffic comes from active users, and storage comes from all users.

Functional requirements: what it must do

A functional requirement is a thing the system must be able to do. The class notes list five.

RequirementIn Maya's wordsWhere we build it
Remember durable memories across sessions: facts, events, rules"You know I am vegetarian."Chapter 6
Use them at the right moment, without being asked"I should not have to remind you."Chapters 8 and 9
Update when life changes"I moved to Paris" must beat "lives in London". Chapters 10 and 11
Never leak one user's memories to another"Tom must never see my allergy."Chapter 15
The user can see, edit and delete everything"Show me what you know. Forget that."Chapter 15

Notice the second one. Maya does not say "use your memory". She asks for a snack, and the allergy must arrive by itself. A memory that has to be asked for is a search box, not a memory.

Non-functional targets: how well it must do it

A non-functional target is a number that says how fast, how small or how reliable the system must be. A target without a reason is a guess, so each one below carries its reason.

RequirementTargetWhy this number
Read-path overheadunder 50 msinvisible next to the model's 200 to 500 ms time to first token
Injected memory budgetabout 75 tokensthe top 3 facts, flat no matter how much is stored
Write pathoff the critical paththe user never waits for it
Memory freshnessby the next sessionbackground extraction lands within seconds
Delete profilewithin days, weeks or monthsas the applicable rule requires

The critical path is the chain of steps the user waits for. Everything on it must be fast. Everything off it may be slow.

Look at how the 50 ms was chosen. Nobody asked "how fast can a database be?". The question was "what is the largest delay the user will not notice?". The model already takes 200 to 500 ms before its first token. Fifty more is at most a quarter of the fastest case and a tenth of the slowest.

50 / 200  =  25 percent of the fastest first token
50 / 500  =  10 percent of the slowest first token

Scale math, step by step

Scale math is the arithmetic that turns the five inputs into traffic and storage. We do it in two halves: how often the system is called, and how much it must hold.

Traffic: how often

daily active users                                     200,000
sessions per day         200,000 x 2              =    400,000
user messages per day    400,000 x 6              =  2,400,000

seconds in a day         24 x 60 x 60             =     86,400

read path   (once per message)   2,400,000 / 86,400  =  about 28 per second
write path  (once per session)     400,000 / 86,400  =  about 4.6 per second

Two things to see here.

First, reads and writes are counted in different units. The librarian works once per message. The secretary works once per session. A session has 6 messages, so there are 6 times more reads than writes.

28 / 4.6  =  about 6

Second, traffic is not spread evenly over the day. People chat in the evening more than at four in the morning. Peak is the busiest moment, and the class sizes it at three times the average.

read path peak     27.8 x 3   =  83.3    about 83 per second
write path peak    4.63 x 3   =  13.9    about 14 per second

We multiply the unrounded rates, 27.8 and 4.63. Rounding first would give 28 x 3 = 84.

We build for the peak, not for the average. A system sized for 28 reads per second fails every evening.

Storage: how much

Start with one card and count its bytes. A byte (B) is the unit of storage: one letter of plain text is about one byte. A kilobyte (KB) is 1,000 bytes. A gigabyte (GB) is 1,000,000,000 bytes.

one memory
  text       about 25 tokens x about 4 letters     =  about 100 B
  vector     768 numbers x 4 bytes each            =      3,072 B
  payload    ids, dates, category, source          =  about 150 B
  total      100 + 3,072 + 150                     =      3,322 B   about 3.3 KB

The vector is the list of numbers that lets the librarian search by meaning. Chapter 7 explains it. For now, notice its weight: 3,072 of 3,322 bytes. The sentence itself is tiny. The numbers that describe it are thirty times heavier.

3,072 / 3,322  =  92 percent of a memory is its vector

Now multiply up.

memories per user                                   about 200
total memories    1,000,000 x 200           =     200,000,000
raw size          200,000,000 x 3,322 B     =     664,400,000,000 B
                                            =     about 664 GB

Storage uses registered users, all one million of them. A user who did not visit today still has cards in the box.

Change the inputs below. Watch which results move when you change the active users, and which move when you change the vector size. The widget adds the text and the payload into one figure of 250 B (100 + 150).

Where does 200 memories per user come from?

It is a design assumption, and an honest book should say so. An active user produces about 3 candidate facts per session.

candidates per day     2 sessions x 3       =      6
candidates per year    6 x 365              =  2,190
candidates in 2 years  2,190 x 2            =  4,380

4,380 / 200  =  about 22 times more than we sized for

These are candidates, not cards. Most of them repeat something already known or replace an older card. A store that keeps every candidate is 22 times larger and much less accurate. Keeping the box at 200 cards is the work of consolidation (merging a new fact into the cards already there) and forgetting, in Chapters 10 and 12.

Scope: what is in and what is out

Scope is the line around the problem we agree to solve. Drawing it early keeps a design discussion from wandering.

In scopeOut of scope
short-term and long-term memoryretrieval over documents (RAG)
fact extraction and the asynchronous write pathmemory shared between several agents
embedding storage in a vector databasetools and planning
retrieval at inference time
consolidation and conflict resolution
forgetting and decay
privacy of stored memories

Inference is the moment the model produces an answer. "Retrieval at inference time" means the librarian works while the user waits.

Try it yourself: size a different product

No computer is needed for this one. Take a smaller product and run the same steps on paper. A study assistant has 50,000 registered students, 20,000 active each day, 3 sessions a day and 8 turns per session. Each student has about 100 memories.

sessions per day     20,000 x 3             =      60,000
messages per day     60,000 x 8             =     480,000

reads per second     480,000 / 86,400       =  about 5.6     peak x 3 = about 17
writes per second     60,000 / 86,400       =  about 0.7     peak x 3 = about 2

total memories       50,000 x 100           =   5,000,000
raw size             5,000,000 x 3,322 B    =  16,610,000,000 B  =  about 16.6 GB

Seventeen reads a second and 16.6 GB fit on one modest server. The same method that gave 664 GB for the big product says this one needs no cluster at all. (A cluster is a group of machines working as one.) That is what scale math is for: it tells us how much machinery not to build.

Common questions

Why is the peak three times the average? Who decided that?

It is a rule of thumb for consumer products, where usage gathers in the waking hours of each region. A real system measures its own peak. In a design discussion, state the factor you chose and move on. The reasoning matters more than the exact factor.

Why does a number take 4 bytes?

Each number in the vector is stored as a 32-bit float, and 32 bits are 4 bytes. It can be stored more coarsely. With 1 byte per number, the vector shrinks from 3,072 to 768 bytes with almost no loss in search quality. Chapter 16 does this and cuts 664 GB to 204 GB.

Is 664 GB a lot?

It is too much for one laptop and ordinary for a production database. It is large enough that we must think about splitting the data over machines and keeping copies. It is small next to the transcripts themselves, which the memory layer does not keep.

Why are reads counted per message and writes per session?

Every message needs the right cards before its answer, so recall runs each time. The secretary works better on a whole conversation, so extraction runs once at the end. That choice divides the write load by six.

Why is "delete profile" allowed to take weeks?

A good design stops using the memory at once. Removing every copy takes longer, because copies can also sit in backups, caches and logs. Laws and company rules set the final deadline, which is why the class notes say "as per rule".

What if my users write much longer messages?

Then tokens per session goes up, and with it the cost of the write path, because the secretary reads the whole transcript. The read path barely changes, since it injects a fixed budget. The benchmark called BEAM in Chapter 18 has assistant turns that run to 500 words, and it taught us to cap what is injected in tokens, not in rows.

Carry this

  • Five inputs produce everything: 1,000,000 users, 200,000 active, 2 sessions, 6 turns, 1,500 tokens.
  • Reads: 2,400,000 messages / 86,400 seconds = about 28 per second, peak about 83. Writes: 400,000 sessions / 86,400 = about 4.6 per second, peak about 14.
  • One memory is about 3.3 KB, and 92 percent of it is the vector.
  • 200,000,000 memories x 3,322 B = about 664 GB.
  • Traffic follows active users. Storage follows registered users.

Check yourself

1. Daily active users double to 400,000. Registered users stay at 1,000,000. What happens to reads per second and to the raw size of the store?

Answer
messages per day    400,000 x 2 x 6        =  4,800,000
reads per second    4,800,000 / 86,400     =  about 56     (doubled)
raw size            1,000,000 x 200 x 3,322 B  =  about 664 GB  (unchanged)

Traffic follows active users. Storage follows registered users.

2. We switch to an embedding with 1,536 numbers per vector. How big is one memory, and how big is the store?

Answer
vector        1,536 x 4              =  6,144 B
one memory    100 + 6,144 + 150      =  6,394 B
store         200,000,000 x 6,394 B  =  about 1,279 GB

Doubling the vector almost doubles the store, because the vector is nearly all of a memory.

3. Why is the read target 50 ms and not 5 ms or 500 ms?

Answer

The model's first token takes 200 to 500 ms. Fifty milliseconds is small enough to go unnoticed beside it, and large enough to fit an embedding call of 10 to 30 ms and a search of 1 to 10 ms. Five would be impossible with an embedding call. Five hundred would double the wait.