Part II · Apprentice
Chapter 8
The librarian: retrieval at inference
Follow one question through the read path: embed, search, pick the top few, inject, answer. All of it in the time the cursor blinks.
On this page
In this chapter, we will learn how the right memories reach the model at the moment a question is asked. We will follow one question through every step, time each step, count the tokens it adds, and find out how many memories are worth handing over.
What is the read path?
The read path is everything that happens between the user's message arriving and the model starting to answer. It is also called retrieval at inference time. Inference is the word for a model producing an answer.
"weekend activity embed filtered search top 3 inject generate
with the family?" ---> ~20 ms ---> user_id = maya ---> of N ---> ~75 tokens ---> the answer
~5 ms stored
Five steps. Let us take them in order.
- Embed the question. The embedding model turns the question into a vector, its place on the map of meaning from Chapter 7.
- Filtered search. The store compares that vector with Maya's memories only, and scores each one.
- Top k. We keep the best few. The letter k is simply the name for "how many". The class uses k = 3.
- Inject. The chosen facts are written into the prompt, in front of the question.
- Generate. The assistant model answers, with the facts on its desk.
The important thing is where the path starts. It starts from the question, not from the store. We never ask "what do we know about Maya?" We ask "what do we know that helps with this?"
Choosing, and leaving behind
Retrieval is selection: most memories must stay behind.
Maya has eight memories in the experiment. She asks for a weekend activity with the family. Here are all eight scores.
selected, top 3 by cosine score (filter: user_id = maya):
0.436 Maya has two kids, aged 6 and 9.
0.382 Maya is training for the Berlin marathon in September.
0.373 Maya's partner Sam has a birthday on June 18.
stayed behind, 5 of 8, never enter the prompt:
0.341 Maya works as a data engineer at a fintech company.
0.335 Maya lives in Paris.
0.314 Maya is vegetarian.
0.297 Maya prefers short, direct answers.
0.293 Maya has a serious peanut allergy.
The peanut allergy came last. Is that a mistake? For this question, no. The allergy is critical for a snack question and has nothing to do with a weekend outing. Selection means the allergy waits in the box until a question needs it.
But look closer and two things should bother us. "Lives in Paris" stayed behind, and a weekend activity surely depends on the city. And the gaps are small: 0.373 got in, 0.341 did not. Try other questions below, and move k. Watch which memory stays behind each time, and what each extra memory costs in tokens. The widget counts the bare sentences without their dates, so its totals are lower than the ones we measure later in this chapter.
The injection prompt
The injection prompt is the short text that carries the chosen memories to the model. Here is the one from the class notes.
SYSTEM: You are Maya's personal assistant. Known facts about the user (most relevant first, with dates): - Maya has two kids, aged 6 and 9. (since 2026-01) - Maya is training for the Berlin marathon. (since 2026-03) - Maya's partner Sam has a birthday on June 18. (since 2026-02) Personalize the answer using these facts where relevant. Do not recite these facts back unless asked. USER: Any ideas for a weekend activity with the family?
Every line has a reason.
| Line | Why it is there |
|---|---|
| "Known facts about the user" | tells the model these are trusted background, not part of the question |
| "most relevant first" | models pay most attention to what comes first |
| "with dates" | lets the model judge how fresh a fact is |
| "where relevant" | permission to ignore a fact that does not fit |
| "Do not recite these facts" | see below |
Why "do not recite"?
Without that line, a model tends to show off what it knows. Maya asks for a weekend idea and hears: "As a mother of two children aged 6 and 9, who is training for the Berlin marathon, and whose partner Sam has a birthday on June 18, you might enjoy..."
That is unpleasant. It feels like being watched. A friend who knows you has kids suggests the zoo. She does not announce that she knows you have kids. Good memory is felt in the quality of the answer. It is not read out loud.
Try it yourself: the read path, end to end
The class experiment retrieval-at-inference runs all five steps and times each one.
ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/retrieval-at-inference
pip install -r requirements.txt
python main.py
First the answer with the three memories injected, then the same question with none.
(b) Generate WITH the 3 retrieved memories as the system prompt
| How about planning a fun day at the local amusement park? Your kids will love it,
| and you can enjoy some quality time together while Sam could surprise you both by
| joining in on the fun!
(c) Generate with NO memories, same question, for contrast
| How about planning a nature scavenger hunt in your local park or nearby woods? It's
| a fun way to get everyone moving and exploring together while keeping an eye out for
| specific items on a custom list you create. Alternatively, you could organize a DIY
| craft day at home, where each family member can work on their own project using
| materials you gather from around the house or a local store.
Same model, same question: the only difference is the 3 retrieved facts.
The second answer is fine for anybody. The first is for Maya: it knows there are kids, and it knows Sam. Then the experiment shows where the time and the tokens went.
step measured
embed the question 44.7 ms
vector search, top 3 of 8 32.1 ms
LLM call with 3 memories 2,266.0 ms
LLM call with no memories 4,350.0 ms
prompt tokens with memories 141
prompt tokens without 46
memory overhead 141 - 46 = 95 tokens
the 3 dated facts 71
the fixed wrapper 24 (persona line plus two instructions)
The 50 millisecond budget
The read path may add at most 50 milliseconds, because the user is waiting. A millisecond, ms, is one thousandth of a second.
Why 50? Because of what comes after. An assistant model needs 200 to 500 ms before its first word appears. This is called the time to first token. A delay of 50 ms in front of that is not something a person can feel.
class budget embed the question 10 to 30 ms filtered search 1 to 10 ms read path total 11 to 40 ms under the 50 ms target model's first token 200 to 500 ms measured on a laptop embed + search 44.7 + 32.1 = 76.8 ms LLM call 2,266 ms share of the wait 76.8 / 2,266 = about 3.4 percent
The laptop run landed over the target, at 76.8 ms. Both steps were network round trips to local programs. One of them was cold, which means it had just started and had nothing loaded yet. A production system keeps the index loaded, or warm, and places the embedding service next to the store. Even at 76.8 ms, retrieval was about 3.4 percent of the wait.
The flat cost
The read cost does not grow with the size of the store.
memory overhead in the experiment 95 tokens the 3 dated facts 71 tokens (class budget: about 75) the fixed wrapper 24 tokens (never grows) store holds 8 memories prompt carries 3 about 95 tokens store holds 800 memories prompt carries 3 about 95 tokens
This is the promise from Chapter 1 kept. Resent history grows with every session. Injected memory stays at about 95 tokens whether Maya has used the assistant for a week or for ten years.
Top 3 is a dial, not a law
The number k trades two things against each other: the chance that the needed memory is included, and the number of tokens the model must read.
The first is called recall: of the memories that hold the answer, what share did we retrieve? The word has two uses in this book. In Chapter 3 "recall" was a step of the loop. Here it is a measure. With k = 3 and a simple question, recall is good. With a hard question that needs four facts from four different months, three slots cannot be enough.
We measured this curve on the LoCoMo development set, the 385 test questions introduced in Chapter 6. Its questions are much harder than Maya's. The test is free to run, because it only checks whether the rows that hold the answer came back. That share is called evidence recall. No model is called (runs 0005-free-k15, 0005-free-k50 and 0005-free-k100). A row here is either a card, which is an extracted fact, or a raw turn, which is one original message. So k = 30 rows is not 30 cards.
| k | Evidence recall | Context tokens per question | Tokens for each extra point |
|---|---|---|---|
| 15 | 83.0 % | 875 | |
| 30 | 87.6 % | 1,489 | about 133 |
| 50 | 90.5 % | 2,324 | about 288 |
| 100 | 93.3 % | 4,431 | about 753 |
from k = 15 to k = 30 (1,489 - 875) / (87.6 - 83.0) = 614 / 4.6 = about 133 tokens per point from k = 30 to k = 50 (2,324 - 1,489) / (90.5 - 87.6) = 835 / 2.9 = about 288 tokens per point from k = 50 to k = 100 (4,431 - 2,324) / (93.3 - 90.5) = 2,107 / 2.8 = about 753 tokens per point
Each extra point of recall costs more than the one before. That is a curve of diminishing returns. But the real surprise came when we let the model answer.
So which is right, 3 or 30? Both. The class design serves a personal assistant with short facts and simple questions, and 3 is a sound start. A benchmark with multi-step questions over long conversations needs 30. The right k is measured, not guessed.
Budget in tokens, not in rows
A read budget is a cap on how many tokens of memory may enter the prompt.
Counting rows assumes all rows are the same size. They are not. On the BEAM benchmark the assistant's messages run to about 500 words each, and 30 rows came to 12,380 tokens per question.
saving 12,380 - 4,227 = 8,153 tokens per question share 8,153 / 12,380 = about 66 percent
A librarian who hands over fewer, better pages helps more than one who hands over the shelf.
Common questions
What if no memory is relevant to the question?
The search still returns its top k, because something is always nearest. That is why the prompt says "where relevant". Some systems also set a minimum score and inject nothing below it. The risk is a threshold that is too strict, which silently drops the one fact that mattered.
Does retrieval run on every message, or once per session?
On every message. Each message is a new question, and a new question needs different cards. In the class baseline that is 2.4 million retrievals a day, about 28 per second.
Why did the peanut allergy rank last for a family outing? Is that not dangerous?
For that question the allergy is not needed. The danger is real when food is involved and similarity alone misses it. That is the reason for importance scores, which let a critical fact outrank a merely similar one. Chapter 9 builds them.
Why not let the model search the memory itself, as a tool?
Some agents do. It costs an extra model call, often a second or more, and the model must first realise that it should search. It cannot know that it should ask about an allergy it has never heard of. Automatic retrieval in front of every answer is faster and catches what the model does not know to look for.
Should the most relevant memory be first or last in the prompt?
The class prompt puts the most relevant first. In hanumemAI we tried ordering by relevance against ordering by date, and the score was identical (experiment 0011). We kept date order, because a timeline is easier for people to read when they debug.
Is a big context window a reason to raise k to 200?
It makes it possible, not wise. Our measurements show no gain in accuracy from extra rows while the bill rises. Mem0 reports about 7,000 tokens per question for its top 200 memories. hanumemAI reaches its LoCoMo scores with about 1,470.
Carry this
- The read path has five steps: embed, filtered search, top k, inject, generate. It starts from the question.
- The budget is under 50 ms, invisible next to the model's 200 to 500 ms to first token.
- Three facts cost about 95 tokens, 71 of facts and 24 of wrapper, whether the store holds 8 memories or 800.
- k is a dial. Recall rises with k, but each point costs more tokens, and more rows did not raise accuracy.
- Cap the read in tokens, not in rows. On BEAM, 66 percent fewer tokens cost nothing in score.
Check yourself
1. A system embeds in 25 ms and searches in 8 ms. The model's first token arrives after 300 ms. What share of the wait before the first word is the read path?
Answer
read path 25 + 8 = 33 ms total wait 33 + 300 = 333 ms share 33 / 333 = about 10 percent
It is inside the 50 ms budget, and no user would notice it.
2. Moving from k = 30 to k = 50 raised recall by 2.9 points. Why might the answers still get worse?
Answer
Recall only says the needed row was somewhere in the block. The model still has to find it and use it. Twenty extra rows are mostly distraction, and a small model reads a long block less carefully. In our run accuracy moved from 91.17 to 90.91, no better, while tokens rose by 67 percent.
3. Why does the injection prompt say "Do not recite these facts back unless asked"?
Answer
Without it the model lists what it knows about the user before answering. That wastes words and feels like surveillance. Memory should improve the answer quietly.