Part I · Novice
Chapter 3
The card box: recall, answer, remember
One loop sits under every memory system: recall before the answer, remember after it. We meet the secretary, the librarian and the millisecond budget.
On this page
In this chapter, we will learn what a memory layer is made of. We will meet the card box that we use for the rest of the book, the two people who run it, and the loop that every memory system follows: recall, answer, remember. Then we will follow one message through the system and count its milliseconds.
What memory is
A memory layer is a system that turns conversations into short facts, stores them outside the model, and places the relevant few into the next prompt.
This is the design that Chapter 2 pushed us into. Here it is with its numbers.
session transcript write path extractor memory read path next session's
about 1,500 tokens ---- (background) ---> LLM ---> store ---- about 75 ----> prompt ---> assistant
0 to N facts tokens LLM
about 3 on average,
about 25 tokens each
Three names in that picture stay with us.
- The extractor is the secretary. It is a small LLM whose only job is to read a transcript and write facts.
- The memory store is the box. It is an ordinary database.
- Retrieval is the librarian's work: finding the right cards for one question.
Notice how much the secretary throws away. 1,500 tokens go in. About 3 facts of 25 tokens come out, which is 75 tokens. The secretary keeps one twentieth of what was said.
in 1,500 tokens out 3 x 25 = 75 tokens kept 75 / 1,500 = 5 percent
The loop: recall, answer, remember
Every memory system runs the same loop around the model: recall before the answer, remember after it.
user message ---> RECALL ---------> LLM ---------> ANSWER
(read memory) .
^ .
| v
memory store <. . . . . . . . . . REMEMBER
(write memory)
solid arrows: the user is waiting dotted arrows: nobody is waiting
- Recall. The librarian pulls the cards that match the message.
- Answer. The model reads the message and the cards, and replies.
- Remember. The secretary reads the conversation and writes new cards for next time.
The model itself never changes. It is the same frozen file from Chapter 1. Memory is a layer around the model. ChatGPT's memory, the open-source library Mem0 and the system HydraDB all run this loop. They differ in how the secretary writes and how the librarian searches. We compare them in Chapter 19.
Two paths with two speeds
The read path is everything that happens between the user's message and the answer. The write path is everything that happens after the answer to update memory.
These two paths have opposite needs, and the whole architecture follows from that.
| Read path (recall) | Write path (remember) | |
|---|---|---|
| Who is waiting | the user | nobody |
| How it runs | synchronous: inside the request | asynchronous: in the background |
| How often | once per message, about 28 per second | once per session, about 4.6 per second |
| Time allowed | under 50 milliseconds | seconds are fine |
| Extra LLM call for memory | no, only the answer itself | yes, the extractor |
Synchronous means the user waits until the step is finished. Asynchronous means the step is handed to a background worker and the user moves on. The two rates come from the workload, and Chapter 5 derives them.
The rule is short: read urgent, write lazy. Anything slow or expensive belongs on the write path, where no one is watching the clock.
The architecture
An orchestrator is the piece of the app that receives the message and calls everything else in the right order. It holds no memory itself. Here is what it talks to.
READ PATH: synchronous, about 28 per second, every millisecond visible
user message ---> orchestrator ---> session store (this chat's turns)
|
+---------> query embedder ---> memory store
filter: user_id
|
top 3, about 75 tokens
v
session turns ---------------> prompt assembly ---> assistant LLM ---> stream to user
WRITE PATH: asynchronous, about 4.6 per second, seconds are fine
after the response is sent . . > extraction queue ---> extractor LLM ---> update decision ---> memory store
about 3 facts ADD / UPDATE /
DELETE / NOOP
A few new words, each in one line.
- The session store holds the turns of the chat that is happening now. It is the short-term side, and Chapter 4 covers it.
- The embedder turns a sentence into a list of numbers so that sentences with similar meaning can be found. Chapter 7 explains it.
- The filter on user_id makes sure the librarian only opens Maya's box, never Tom's.
- The queue is a waiting line for background work.
- The update decision chooses, for each new fact, whether to add it, change an old card, remove an old card or do nothing. Chapter 10 is about that choice.
One request, millisecond by millisecond
A latency budget is the time each step is allowed to take. Here is the life of one message, with the budgets from the class notes.
| Step | What happens | Time |
|---|---|---|
| 1 | The message travels from the app to the orchestrator | 1 to 5 ms |
| 2 | The orchestrator fetches this session's turns | about 1 ms |
| 3 | The question is embedded | 10 to 30 ms |
| 4 | The memory store is searched, filtered to this user | 1 to 10 ms |
| 5 | The prompt is assembled: system text, memories, history, message | under 1 ms |
| 6 | The assistant LLM produces its first token | 200 to 500 ms |
| 7 | The answer streams to the user | 50 or more tokens per second |
| 8 | After the session: the transcript is put on the queue | background |
| 9 | The extractor writes facts and decides the operations | about 1 s per call |
| 10 | The facts are written to the store | about 10 ms per write |
The system text in step 5 is the app's fixed instruction to the model, such as "You are Maya's personal assistant".
A millisecond (ms) is one thousandth of a second. Add up what memory costs the waiting user. It is steps 3 and 4.
best case 10 + 1 = 11 ms
worst case 30 + 10 = 40 ms
target under 50 ms
the model's first token 200 to 500 ms
memory's share, worst case against a 500 ms first token
40 / (40 + 500) = about 7 percent of the wait
Time to first token is the wait before the first piece of the answer appears. Next to it, 40 milliseconds of recall cannot be felt. That is why the target is 50 milliseconds: recall must be invisible beside the model.
Step through the request below. Watch which steps the user waits for, and which begin only after the answer has left. The widget's clock counts every step before the model, steps 1 to 5, so it shows a slightly larger range than our sum of steps 3 and 4.
best case 1 + 1 + 10 + 1 + 0 = 13 ms worst case 5 + 1 + 30 + 10 + 1 = 47 ms still under 50 ms
Try it yourself: the read path, timed
The class experiment retrieval-at-inference runs the whole read path on a laptop and times
each step. It stores 8 facts about Maya and asks one question. Here are the parts of its output that matter
to us. The cosine score beside each fact says how close its meaning is to the question.
Higher is closer, and Chapter 7 shows how it is computed.
question: "Any ideas for a weekend activity with the family?"
selected, top 3 by cosine score (filter: user_id = maya):
0.436 Maya has two kids, aged 6 and 9.
0.382 Maya is training for the Berlin marathon in September.
0.373 Maya's partner Sam has a birthday on June 18.
step measured
embed the question 44.7 ms
vector search, top 3 of 8 32.1 ms
LLM call with 3 memories 2,266.0 ms
Retrieval adds 44.7 + 32.1 = 76.8 ms before a 2,266 ms LLM call:
about 3.4% of the wait, and about 95 extra prompt tokens.
recall 44.7 + 32.1 = 76.8 ms
share 76.8 / 2,266 = about 3.4 percent
tokens 3 dated facts 71
fixed wrapper text 24
total 71 + 24 = 95 extra prompt tokens
The laptop run takes 76.8 ms, over the 50 ms target. On the laptop, each step is a separate web request, and the store had only just started. A production system keeps the embedder and the store running and ready, which is how the class arrives at its budget. Even at 76.8 ms, recall is about 3.4 percent of the wait.
The three facts cost 71 tokens, close to our budget of 75. The other 24 tokens are the fixed lines around them, such as "You are Maya's personal assistant". They do not grow with the box.
ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/retrieval-at-inference
pip install -r requirements.txt
python main.py
Common questions
Why does the read path not use an LLM to pick the memories?
Because an LLM call takes hundreds of milliseconds and the read path has fifty. Recall uses arithmetic on numbers, which takes a few milliseconds. The expensive thinking was already done by the secretary at write time.
Why write once per session and not after every message?
A whole session gives the secretary more context. "She moved there for the job" only makes sense next to the earlier lines. One call per session is also cheaper: 400,000 calls a day and not 2,400,000.
What if the background write fails?
The transcript waits on the queue, so the work can be tried again. The user's answer was already sent and is not affected. The only cost of a delay is that a new fact reaches the box a little later.
Is the memory store a special kind of database?
It is an ordinary database that can also search by meaning. Products such as Qdrant and pgvector do this. Our own library uses SQLite, a database that lives in a single file. Chapter 7 opens the box and counts the bytes.
Who decides what is worth a card?
The extractor LLM, guided by a written instruction called a prompt. It keeps durable facts and drops small talk, jokes and one-off requests. Chapter 6 shows the prompt and a real run.
Does recall run on every message, even "thanks"?
In the simple design, yes. It costs about 40 milliseconds and 75 tokens, so skipping it saves little. Some systems skip recall for very short messages. It is an optimisation, not a requirement.
Carry this
- The box is the store, the secretary is the extractor, the librarian is retrieval.
- The loop is recall, answer, remember. The model in the middle never changes.
- Read urgent, write lazy: recall is synchronous and under 50 ms, remembering is asynchronous and may take seconds.
- Recall costs 11 to 40 ms (embed 10 to 30, search 1 to 10) beside a first token of 200 to 500 ms.
- The secretary keeps about 5 percent: 1,500 tokens in, about 75 out.
Check yourself
1. A teammate proposes running the extractor before the answer, so that memory is always up to date. The extractor takes about 1 second. What happens to the user's wait?
Answer
It grows by about 1 second on every message, from a first token at 200 to 500 ms to one at 1,200 to 1,500 ms. And nothing is gained, because facts from this session are already on the desk in the session's turns.
2. Embedding takes 25 ms and search takes 8 ms. The model's first token arrives after 300 ms. What share of the wait is recall?
Answer
recall 25 + 8 = 33 ms total 33 + 300 = 333 ms share 33 / 333 = about 10 percent
3. Which steps of the loop touch the memory store, and which of them does the user wait for?
Answer
Recall reads the store and remember writes to it. The user waits only for recall. Remember runs in the background after the answer is sent.