Carebun Interactive/Memory Memory labCheat sheet

Part II · Apprentice

Chapter 8

The librarian: retrieval at inference

Follow one question through the read path: embed, search, pick the top few, inject, answer. All of it in the time the cursor blinks.

11 min read · interactive

On this page
  1. What is the read path?
  2. Choosing, and leaving behind
  3. The injection prompt
  4. Try it yourself: the read path, end to end
  5. The 50 millisecond budget
  6. The flat cost
  7. Top 3 is a dial, not a law
  8. Common questions
  9. Carry this
  10. Check yourself

In this chapter, we will learn how the right memories reach the model at the moment a question is asked. We will follow one question through every step, time each step, count the tokens it adds, and find out how many memories are worth handing over.

What is the read path?

The read path is everything that happens between the user's message arriving and the model starting to answer. It is also called retrieval at inference time. Inference is the word for a model producing an answer.

"weekend activity         embed          filtered search        top 3          inject         generate
 with the family?"  --->  ~20 ms   --->  user_id = maya   --->  of N     --->  ~75 tokens ---> the answer
                                         ~5 ms                  stored

Five steps. Let us take them in order.

  1. Embed the question. The embedding model turns the question into a vector, its place on the map of meaning from Chapter 7.
  2. Filtered search. The store compares that vector with Maya's memories only, and scores each one.
  3. Top k. We keep the best few. The letter k is simply the name for "how many". The class uses k = 3.
  4. Inject. The chosen facts are written into the prompt, in front of the question.
  5. Generate. The assistant model answers, with the facts on its desk.

The important thing is where the path starts. It starts from the question, not from the store. We never ask "what do we know about Maya?" We ask "what do we know that helps with this?"

Choosing, and leaving behind

Retrieval is selection: most memories must stay behind.

Maya has eight memories in the experiment. She asks for a weekend activity with the family. Here are all eight scores.

    selected, top 3 by cosine score (filter: user_id = maya):
        0.436  Maya has two kids, aged 6 and 9.
        0.382  Maya is training for the Berlin marathon in September.
        0.373  Maya's partner Sam has a birthday on June 18.

    stayed behind, 5 of 8, never enter the prompt:
        0.341  Maya works as a data engineer at a fintech company.
        0.335  Maya lives in Paris.
        0.314  Maya is vegetarian.
        0.297  Maya prefers short, direct answers.
        0.293  Maya has a serious peanut allergy.

The peanut allergy came last. Is that a mistake? For this question, no. The allergy is critical for a snack question and has nothing to do with a weekend outing. Selection means the allergy waits in the box until a question needs it.

But look closer and two things should bother us. "Lives in Paris" stayed behind, and a weekend activity surely depends on the city. And the gaps are small: 0.373 got in, 0.341 did not. Try other questions below, and move k. Watch which memory stays behind each time, and what each extra memory costs in tokens. The widget counts the bare sentences without their dates, so its totals are lower than the ones we measure later in this chapter.

The injection prompt

The injection prompt is the short text that carries the chosen memories to the model. Here is the one from the class notes.

SYSTEM:
You are Maya's personal assistant.
Known facts about the user (most relevant first, with dates):
- Maya has two kids, aged 6 and 9.            (since 2026-01)
- Maya is training for the Berlin marathon.    (since 2026-03)
- Maya's partner Sam has a birthday on June 18. (since 2026-02)
Personalize the answer using these facts where relevant.
Do not recite these facts back unless asked.

USER:
Any ideas for a weekend activity with the family?

Every line has a reason.

LineWhy it is there
"Known facts about the user"tells the model these are trusted background, not part of the question
"most relevant first"models pay most attention to what comes first
"with dates"lets the model judge how fresh a fact is
"where relevant"permission to ignore a fact that does not fit
"Do not recite these facts"see below

Why "do not recite"?

Without that line, a model tends to show off what it knows. Maya asks for a weekend idea and hears: "As a mother of two children aged 6 and 9, who is training for the Berlin marathon, and whose partner Sam has a birthday on June 18, you might enjoy..."

That is unpleasant. It feels like being watched. A friend who knows you has kids suggests the zoo. She does not announce that she knows you have kids. Good memory is felt in the quality of the answer. It is not read out loud.

Try it yourself: the read path, end to end

The class experiment retrieval-at-inference runs all five steps and times each one.

ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/retrieval-at-inference
pip install -r requirements.txt
python main.py

First the answer with the three memories injected, then the same question with none.

(b) Generate WITH the 3 retrieved memories as the system prompt

    | How about planning a fun day at the local amusement park? Your kids will love it,
    | and you can enjoy some quality time together while Sam could surprise you both by
    | joining in on the fun!

(c) Generate with NO memories, same question, for contrast

    | How about planning a nature scavenger hunt in your local park or nearby woods? It's
    | a fun way to get everyone moving and exploring together while keeping an eye out for
    | specific items on a custom list you create. Alternatively, you could organize a DIY
    | craft day at home, where each family member can work on their own project using
    | materials you gather from around the house or a local store.

    Same model, same question: the only difference is the 3 retrieved facts.

The second answer is fine for anybody. The first is for Maya: it knows there are kids, and it knows Sam. Then the experiment shows where the time and the tokens went.

    step                          measured
    embed the question                44.7 ms
    vector search, top 3 of 8         32.1 ms
    LLM call with 3 memories       2,266.0 ms
    LLM call with no memories      4,350.0 ms

    prompt tokens with memories        141
    prompt tokens without               46
    memory overhead               141 - 46 = 95 tokens
      the 3 dated facts                 71
      the fixed wrapper                 24  (persona line plus two instructions)

The 50 millisecond budget

The read path may add at most 50 milliseconds, because the user is waiting. A millisecond, ms, is one thousandth of a second.

Why 50? Because of what comes after. An assistant model needs 200 to 500 ms before its first word appears. This is called the time to first token. A delay of 50 ms in front of that is not something a person can feel.

class budget
  embed the question       10 to 30 ms
  filtered search           1 to 10 ms
  read path total          11 to 40 ms      under the 50 ms target

  model's first token     200 to 500 ms

measured on a laptop
  embed + search           44.7 + 32.1   =  76.8 ms
  LLM call                                  2,266 ms
  share of the wait        76.8 / 2,266  =  about 3.4 percent

The laptop run landed over the target, at 76.8 ms. Both steps were network round trips to local programs. One of them was cold, which means it had just started and had nothing loaded yet. A production system keeps the index loaded, or warm, and places the embedding service next to the store. Even at 76.8 ms, retrieval was about 3.4 percent of the wait.

The flat cost

The read cost does not grow with the size of the store.

memory overhead in the experiment      95 tokens
  the 3 dated facts                    71 tokens     (class budget: about 75)
  the fixed wrapper                    24 tokens     (never grows)

store holds     8 memories     prompt carries 3     about 95 tokens
store holds   800 memories     prompt carries 3     about 95 tokens

This is the promise from Chapter 1 kept. Resent history grows with every session. Injected memory stays at about 95 tokens whether Maya has used the assistant for a week or for ten years.

Top 3 is a dial, not a law

The number k trades two things against each other: the chance that the needed memory is included, and the number of tokens the model must read.

The first is called recall: of the memories that hold the answer, what share did we retrieve? The word has two uses in this book. In Chapter 3 "recall" was a step of the loop. Here it is a measure. With k = 3 and a simple question, recall is good. With a hard question that needs four facts from four different months, three slots cannot be enough.

We measured this curve on the LoCoMo development set, the 385 test questions introduced in Chapter 6. Its questions are much harder than Maya's. The test is free to run, because it only checks whether the rows that hold the answer came back. That share is called evidence recall. No model is called (runs 0005-free-k15, 0005-free-k50 and 0005-free-k100). A row here is either a card, which is an extracted fact, or a raw turn, which is one original message. So k = 30 rows is not 30 cards.

kEvidence recallContext tokens per questionTokens for each extra point
1583.0 %875
3087.6 %1,489about 133
5090.5 %2,324about 288
10093.3 %4,431about 753
from k = 15 to k = 30     (1,489 -   875) / (87.6 - 83.0)  =   614 / 4.6  =  about 133 tokens per point
from k = 30 to k = 50     (2,324 - 1,489) / (90.5 - 87.6)  =   835 / 2.9  =  about 288 tokens per point
from k = 50 to k = 100    (4,431 - 2,324) / (93.3 - 90.5)  = 2,107 / 2.8  =  about 753 tokens per point

Each extra point of recall costs more than the one before. That is a curve of diminishing returns. But the real surprise came when we let the model answer.

So which is right, 3 or 30? Both. The class design serves a personal assistant with short facts and simple questions, and 3 is a sound start. A benchmark with multi-step questions over long conversations needs 30. The right k is measured, not guessed.

Budget in tokens, not in rows

A read budget is a cap on how many tokens of memory may enter the prompt.

Counting rows assumes all rows are the same size. They are not. On the BEAM benchmark the assistant's messages run to about 500 words each, and 30 rows came to 12,380 tokens per question.

saving     12,380 - 4,227    =  8,153 tokens per question
share       8,153 / 12,380   =  about 66 percent

A librarian who hands over fewer, better pages helps more than one who hands over the shelf.

Common questions

What if no memory is relevant to the question?

The search still returns its top k, because something is always nearest. That is why the prompt says "where relevant". Some systems also set a minimum score and inject nothing below it. The risk is a threshold that is too strict, which silently drops the one fact that mattered.

Does retrieval run on every message, or once per session?

On every message. Each message is a new question, and a new question needs different cards. In the class baseline that is 2.4 million retrievals a day, about 28 per second.

Why did the peanut allergy rank last for a family outing? Is that not dangerous?

For that question the allergy is not needed. The danger is real when food is involved and similarity alone misses it. That is the reason for importance scores, which let a critical fact outrank a merely similar one. Chapter 9 builds them.

Why not let the model search the memory itself, as a tool?

Some agents do. It costs an extra model call, often a second or more, and the model must first realise that it should search. It cannot know that it should ask about an allergy it has never heard of. Automatic retrieval in front of every answer is faster and catches what the model does not know to look for.

Should the most relevant memory be first or last in the prompt?

The class prompt puts the most relevant first. In hanumemAI we tried ordering by relevance against ordering by date, and the score was identical (experiment 0011). We kept date order, because a timeline is easier for people to read when they debug.

Is a big context window a reason to raise k to 200?

It makes it possible, not wise. Our measurements show no gain in accuracy from extra rows while the bill rises. Mem0 reports about 7,000 tokens per question for its top 200 memories. hanumemAI reaches its LoCoMo scores with about 1,470.

Carry this

  • The read path has five steps: embed, filtered search, top k, inject, generate. It starts from the question.
  • The budget is under 50 ms, invisible next to the model's 200 to 500 ms to first token.
  • Three facts cost about 95 tokens, 71 of facts and 24 of wrapper, whether the store holds 8 memories or 800.
  • k is a dial. Recall rises with k, but each point costs more tokens, and more rows did not raise accuracy.
  • Cap the read in tokens, not in rows. On BEAM, 66 percent fewer tokens cost nothing in score.

Check yourself

1. A system embeds in 25 ms and searches in 8 ms. The model's first token arrives after 300 ms. What share of the wait before the first word is the read path?

Answer
read path       25 + 8        =  33 ms
total wait      33 + 300      = 333 ms
share           33 / 333      = about 10 percent

It is inside the 50 ms budget, and no user would notice it.

2. Moving from k = 30 to k = 50 raised recall by 2.9 points. Why might the answers still get worse?

Answer

Recall only says the needed row was somewhere in the block. The model still has to find it and use it. Twenty extra rows are mostly distraction, and a small model reads a long block less carefully. In our run accuracy moved from 91.17 to 90.91, no better, while tokens rose by 67 percent.

3. Why does the injection prompt say "Do not recite these facts back unless asked"?

Answer

Without it the model lists what it knows about the user before answering. That wastes words and feels like surveillance. Memory should improve the answer quietly.