Carebun Interactive/Memory Memory labCheat sheet

Part I · Novice

Chapter 3

The card box: recall, answer, remember

One loop sits under every memory system: recall before the answer, remember after it. We meet the secretary, the librarian and the millisecond budget.

9 min read · interactive

On this page
  1. What memory is
  2. The loop: recall, answer, remember
  3. Two paths with two speeds
  4. The architecture
  5. One request, millisecond by millisecond
  6. Try it yourself: the read path, timed
  7. Common questions
  8. Carry this
  9. Check yourself

In this chapter, we will learn what a memory layer is made of. We will meet the card box that we use for the rest of the book, the two people who run it, and the loop that every memory system follows: recall, answer, remember. Then we will follow one message through the system and count its milliseconds.

What memory is

A memory layer is a system that turns conversations into short facts, stores them outside the model, and places the relevant few into the next prompt.

This is the design that Chapter 2 pushed us into. Here it is with its numbers.

session transcript        write path         extractor        memory        read path        next session's
about 1,500 tokens  ---- (background) --->   LLM       --->   store   ---- about 75 ---->   prompt  --->  assistant
                                             0 to N facts                 tokens                            LLM
                                             about 3 on average,
                                             about 25 tokens each

Three names in that picture stay with us.

  • The extractor is the secretary. It is a small LLM whose only job is to read a transcript and write facts.
  • The memory store is the box. It is an ordinary database.
  • Retrieval is the librarian's work: finding the right cards for one question.

Notice how much the secretary throws away. 1,500 tokens go in. About 3 facts of 25 tokens come out, which is 75 tokens. The secretary keeps one twentieth of what was said.

in      1,500 tokens
out     3 x 25   =  75 tokens
kept    75 / 1,500  =  5 percent

The loop: recall, answer, remember

Every memory system runs the same loop around the model: recall before the answer, remember after it.

user message ---> RECALL ---------> LLM ---------> ANSWER
                  (read memory)                       .
                     ^                                .
                     |                                v
               memory store  <. . . . . . . . . .  REMEMBER
                                                   (write memory)

solid arrows: the user is waiting        dotted arrows: nobody is waiting
  1. Recall. The librarian pulls the cards that match the message.
  2. Answer. The model reads the message and the cards, and replies.
  3. Remember. The secretary reads the conversation and writes new cards for next time.

The model itself never changes. It is the same frozen file from Chapter 1. Memory is a layer around the model. ChatGPT's memory, the open-source library Mem0 and the system HydraDB all run this loop. They differ in how the secretary writes and how the librarian searches. We compare them in Chapter 19.

Two paths with two speeds

The read path is everything that happens between the user's message and the answer. The write path is everything that happens after the answer to update memory.

These two paths have opposite needs, and the whole architecture follows from that.

Read path (recall)Write path (remember)
Who is waitingthe usernobody
How it runssynchronous: inside the requestasynchronous: in the background
How oftenonce per message, about 28 per secondonce per session, about 4.6 per second
Time allowedunder 50 millisecondsseconds are fine
Extra LLM call for memoryno, only the answer itselfyes, the extractor

Synchronous means the user waits until the step is finished. Asynchronous means the step is handed to a background worker and the user moves on. The two rates come from the workload, and Chapter 5 derives them.

The rule is short: read urgent, write lazy. Anything slow or expensive belongs on the write path, where no one is watching the clock.

The architecture

An orchestrator is the piece of the app that receives the message and calls everything else in the right order. It holds no memory itself. Here is what it talks to.

READ PATH: synchronous, about 28 per second, every millisecond visible

  user message ---> orchestrator ---> session store (this chat's turns)
                         |
                         +---------> query embedder ---> memory store
                                                         filter: user_id
                                                              |
                                                     top 3, about 75 tokens
                                                              v
                     session turns ---------------> prompt assembly ---> assistant LLM ---> stream to user


WRITE PATH: asynchronous, about 4.6 per second, seconds are fine

  after the response is sent . . > extraction queue ---> extractor LLM ---> update decision ---> memory store
                                                         about 3 facts       ADD / UPDATE /
                                                                             DELETE / NOOP

A few new words, each in one line.

  • The session store holds the turns of the chat that is happening now. It is the short-term side, and Chapter 4 covers it.
  • The embedder turns a sentence into a list of numbers so that sentences with similar meaning can be found. Chapter 7 explains it.
  • The filter on user_id makes sure the librarian only opens Maya's box, never Tom's.
  • The queue is a waiting line for background work.
  • The update decision chooses, for each new fact, whether to add it, change an old card, remove an old card or do nothing. Chapter 10 is about that choice.

One request, millisecond by millisecond

A latency budget is the time each step is allowed to take. Here is the life of one message, with the budgets from the class notes.

StepWhat happensTime
1The message travels from the app to the orchestrator1 to 5 ms
2The orchestrator fetches this session's turnsabout 1 ms
3The question is embedded10 to 30 ms
4The memory store is searched, filtered to this user1 to 10 ms
5The prompt is assembled: system text, memories, history, messageunder 1 ms
6The assistant LLM produces its first token200 to 500 ms
7The answer streams to the user50 or more tokens per second
8After the session: the transcript is put on the queuebackground
9The extractor writes facts and decides the operationsabout 1 s per call
10The facts are written to the storeabout 10 ms per write

The system text in step 5 is the app's fixed instruction to the model, such as "You are Maya's personal assistant".

A millisecond (ms) is one thousandth of a second. Add up what memory costs the waiting user. It is steps 3 and 4.

best case     10 + 1   =  11 ms
worst case    30 + 10  =  40 ms
target                    under 50 ms

the model's first token   200 to 500 ms
memory's share, worst case against a 500 ms first token
              40 / (40 + 500)  =  about 7 percent of the wait

Time to first token is the wait before the first piece of the answer appears. Next to it, 40 milliseconds of recall cannot be felt. That is why the target is 50 milliseconds: recall must be invisible beside the model.

Step through the request below. Watch which steps the user waits for, and which begin only after the answer has left. The widget's clock counts every step before the model, steps 1 to 5, so it shows a slightly larger range than our sum of steps 3 and 4.

best case     1 + 1 + 10 + 1 + 0    =  13 ms
worst case    5 + 1 + 30 + 10 + 1   =  47 ms     still under 50 ms

Try it yourself: the read path, timed

The class experiment retrieval-at-inference runs the whole read path on a laptop and times each step. It stores 8 facts about Maya and asks one question. Here are the parts of its output that matter to us. The cosine score beside each fact says how close its meaning is to the question. Higher is closer, and Chapter 7 shows how it is computed.

    question: "Any ideas for a weekend activity with the family?"

    selected, top 3 by cosine score (filter: user_id = maya):
        0.436  Maya has two kids, aged 6 and 9.
        0.382  Maya is training for the Berlin marathon in September.
        0.373  Maya's partner Sam has a birthday on June 18.

    step                          measured
    embed the question                44.7 ms
    vector search, top 3 of 8         32.1 ms
    LLM call with 3 memories       2,266.0 ms

    Retrieval adds 44.7 + 32.1 = 76.8 ms before a 2,266 ms LLM call:
    about 3.4% of the wait, and about 95 extra prompt tokens.
recall      44.7 + 32.1              =  76.8 ms
share       76.8 / 2,266             =  about 3.4 percent

tokens      3 dated facts               71
            fixed wrapper text          24
            total        71 + 24     =  95 extra prompt tokens

The laptop run takes 76.8 ms, over the 50 ms target. On the laptop, each step is a separate web request, and the store had only just started. A production system keeps the embedder and the store running and ready, which is how the class arrives at its budget. Even at 76.8 ms, recall is about 3.4 percent of the wait.

The three facts cost 71 tokens, close to our budget of 75. The other 24 tokens are the fixed lines around them, such as "You are Maya's personal assistant". They do not grow with the box.

ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/retrieval-at-inference
pip install -r requirements.txt
python main.py

Common questions

Why does the read path not use an LLM to pick the memories?

Because an LLM call takes hundreds of milliseconds and the read path has fifty. Recall uses arithmetic on numbers, which takes a few milliseconds. The expensive thinking was already done by the secretary at write time.

Why write once per session and not after every message?

A whole session gives the secretary more context. "She moved there for the job" only makes sense next to the earlier lines. One call per session is also cheaper: 400,000 calls a day and not 2,400,000.

What if the background write fails?

The transcript waits on the queue, so the work can be tried again. The user's answer was already sent and is not affected. The only cost of a delay is that a new fact reaches the box a little later.

Is the memory store a special kind of database?

It is an ordinary database that can also search by meaning. Products such as Qdrant and pgvector do this. Our own library uses SQLite, a database that lives in a single file. Chapter 7 opens the box and counts the bytes.

Who decides what is worth a card?

The extractor LLM, guided by a written instruction called a prompt. It keeps durable facts and drops small talk, jokes and one-off requests. Chapter 6 shows the prompt and a real run.

Does recall run on every message, even "thanks"?

In the simple design, yes. It costs about 40 milliseconds and 75 tokens, so skipping it saves little. Some systems skip recall for very short messages. It is an optimisation, not a requirement.

Carry this

  • The box is the store, the secretary is the extractor, the librarian is retrieval.
  • The loop is recall, answer, remember. The model in the middle never changes.
  • Read urgent, write lazy: recall is synchronous and under 50 ms, remembering is asynchronous and may take seconds.
  • Recall costs 11 to 40 ms (embed 10 to 30, search 1 to 10) beside a first token of 200 to 500 ms.
  • The secretary keeps about 5 percent: 1,500 tokens in, about 75 out.

Check yourself

1. A teammate proposes running the extractor before the answer, so that memory is always up to date. The extractor takes about 1 second. What happens to the user's wait?

Answer

It grows by about 1 second on every message, from a first token at 200 to 500 ms to one at 1,200 to 1,500 ms. And nothing is gained, because facts from this session are already on the desk in the session's turns.

2. Embedding takes 25 ms and search takes 8 ms. The model's first token arrives after 300 ms. What share of the wait is recall?

Answer
recall    25 + 8           =  33 ms
total     33 + 300         =  333 ms
share     33 / 333         =  about 10 percent

3. Which steps of the loop touch the memory store, and which of them does the user wait for?

Answer

Recall reads the store and remember writes to it. The user waits only for recall. Remember runs in the background after the answer is sent.