Carebun Interactive/Memory Memory labCheat sheet

Part III · Journeyman

Chapter 13

The write path nobody waits for

Recall happens before the answer and must be fast. Remembering happens after the answer and may take seconds. One experiment shows 58 % of the wait disappear.

10 min read · interactive

On this page
  1. Two jobs with opposite needs
  2. The same two calls, a very different wait
  3. The queue and the workers
  4. How busy are the two paths?
  5. The price: freshness
  6. When things go wrong
  7. The write path in hanumemAI
  8. Try it yourself
  9. Common questions
  10. Carry this
  11. Check yourself

In this chapter, we will learn when an assistant should do its remembering. We will see that reading memory and writing memory have opposite needs, and measure what happens when the two are mixed up. Then we will build the queue that lets the slow work happen after the user already has the answer.

Two jobs with opposite needs

Every message touches memory twice. Before answering, the assistant recalls: it reads the cards that match the question. After answering, it remembers: it writes new cards from what was said.

A synchronous step is one the user waits for. An asynchronous step is one that runs in the background while the user carries on.

Recall must be synchronous. The answer cannot begin until the cards are on the desk. So recall has a strict budget: under 50 ms, which is invisible next to the 200 to 500 ms a model takes to produce its first word.

Remembering has no such pressure. The cat's name that Maya mentioned a moment ago is not needed for the answer she is waiting for right now. It is needed next week. So remembering is asynchronous.

Recall (read path)Remember (write path)
Whenbefore the answerafter the answer is sent
Does the user wait?yesno
Time allowedunder 50 msseconds are fine
Uses an LLM?no, only the small embedding modelyes, the extractor
Characterfast and simpleslow and careful

The same two calls, a very different wait

The class experiment async-write sends one message that holds both a fact worth keeping and a question.

Maya:  By the way, I just adopted a cat named Miso! Anyway: give me one tip
       for keeping a flat tidy with kids. Two sentences.

This turn needs two LLM calls: one to answer, one to extract the fact about Miso. The only question is the order. In run (b) the extraction runs in a background thread, which is a second line of work inside the same program.

(a) BLOCKING write: extract first, then answer

    extraction call :    1,605 ms
    answer call     :    1,746 ms
    user-perceived latency = 1,605 + 1,746 = 3,351 ms (the user waits for BOTH)

(b) ASYNC write: answer immediately, extraction in a background thread

    answer call     :    1,393 ms
    user-perceived latency = 1,393 ms (the answer call alone)

    background extraction: 1,859 ms, finished 571 ms after the answer was done
      stored memory: Maya has a cat named Miso.

Latency is the time the user waits. The difference is large.

blocking        3,351 ms
async           1,393 ms
saved           3,351 - 1,393 = 1,958 ms
share saved     1,958 / 3,351 = 58 % of the wait gone

The same two calls were made. The same fact was stored. The same answer was given. The user simply stopped waiting for work that was never for them. Exact timings change from run to run, because this is a small model on a laptop. The shape does not change.

Set how long the extraction and the answer take. Compare the two timelines, and find the moment the new fact becomes readable.

The queue and the workers

A queue is a waiting line for jobs. When the assistant has sent its reply, it drops a small note in the line: "extract memories from this session, for this user". Then it turns to the next message. It does not wait to see the job done.

A worker is a program that takes jobs from the queue and does them, one after another. Here the worker runs the extractor, checks for duplicates and writes the new cards.

READ PATH   the user waits                        about 28 per second

  message --> recall (under 50 ms) --> assistant LLM --> answer streams to the user
                                              |
                                              |  after the response is sent
                                              v
WRITE PATH  nobody waits                          about 4.6 per second

  [ queue ] --> worker --> extractor LLM --> dedup and supersede --> memory store
   the          takes       about 3 facts                           writes,
   transcript   a job       about 1 s                               about 10 ms

Once per session, not once per message

The class design enqueues one job when a session ends, holding the whole transcript. Why not one job per message?

  • It costs a sixth as much. A session has 6 user messages. One extraction for all of them is one call where there would have been six.
  • The extractor sees the whole story. "I love blue. Actually, orange." becomes one clean fact. Extracting per message would write "loves blue" and then have to correct it.

Some products do extract after every message, and show a small "memory updated" notice at once. That buys instant freshness at about six times the extractor calls. It is a product choice, not a mistake.

How busy are the two paths?

daily active users                               200,000
sessions per day            200,000 x 2       =   400,000
user messages per day       400,000 x 6       = 2,400,000
seconds in a day                                   86,400

reads  (one per message)    2,400,000 / 86,400 = about 28 per second    peak 3x = about 83
writes (one per session)      400,000 / 86,400 = about 4.6 per second   peak 3x = about 14

There are 6 reads for every write. And each read is allowed 50 ms, where each write may take a second or more. The path that is used most has the least time. That is the right way round, and it is only possible because the slow work was moved out of it.

The queue also absorbs rushes. If a thousand sessions end in the same minute, the line grows longer for a while and the workers catch up. No user feels it, because no user is waiting.

The price: freshness

Freshness is how soon a new fact can be read back from the store. With an asynchronous write there is a short gap between saying something and the store knowing it. In the experiment the Miso fact became readable about two seconds after the answer call started.

answer finished                  1,393 ms
extraction finished after that     571 ms
fact readable      1,393 + 571 = 1,964 ms     the program prints 1,963; it rounds each part

Is that gap a problem? Picture Maya asking, in the very next message, "what did I say my cat was called?" The store may not have the card yet.

It does not matter. That question is answered from the context window, the short-term memory of Chapter 4. Everything said in this session is still on the desk. The long-term store is for next week, and by next week the card is written. The class sets the target plainly: memory must be fresh by the next session, and background extraction lands in seconds.

asked in the SAME session   -->  the context window answers   (always fresh)
asked in a LATER session    -->  the memory store answers     (written long ago)

When things go wrong

Background work fails in ways the user never sees, which makes it easy to ignore. Four questions need an answer.

What if the extractor call fails?

A retry is a second attempt at a job that failed. The model provider may be slow or briefly down. The worker waits a moment and tries again, a few times, waiting longer each time. hanumemAI retries an LLM call up to 4 times by default, after waits of 1, 2, 4 and 8 seconds.

What if a job runs twice?

Retries create a new risk. Suppose the write succeeded but the worker crashed before it could say so. The job runs again. Will Maya get two Miso cards?

An operation is idempotent when doing it twice has the same result as doing it once. Pressing a lift button twice does not call two lifts. The write path gets this property from the hash of Chapter 10: a fact whose text is already stored for this user is skipped. Run the job twice and the second run adds nothing.

What if two jobs for the same user overlap?

Maya says "I live in London" in the morning and "I moved to Paris" in the evening. If the evening job were processed first, the store would end the day believing in London. So writes for one user must be done in the order they happened. Different users can be processed side by side, because their cards never touch.

What if a worker dies with a job in its hands?

A proper queue hands out a job but does not forget it until the worker reports success. If the worker goes silent, the queue gives the job to another worker after a time limit. And a job that fails again and again is moved aside to a separate line, often called a dead-letter queue, for a person to look at. It must not block the jobs behind it.

The write path in hanumemAI

The library gives the two styles side by side. add() waits for the write. add_async() queues it and returns at once.

from hmem import Memory, Config

memory = Memory(Config(db_path="app.sqlite"))

def handle_turn(user_id, session, message, today):
    # 1. recall: synchronous, no LLM
    block, rows = memory.prompt_block(message, user_id=user_id, budget_tokens=1500, ref_date=today)

    # 2. answer: your model, your prompt
    reply = llm.chat(system="Known facts about the user:\n" + block,
                     messages=session + [{"role": "user", "content": message}])

    # 3. remember: queued, returns immediately
    memory.add_async(session + [{"role": "user", "content": message},
                                {"role": "assistant", "content": reply}],
                     user_id=user_id, observed_at=today)
    return reply

memory.flush()   # before shutdown: wait for the queued writes to finish

add_async() returns a future, which is a small ticket for a result that does not exist yet. You may ignore it. Or you may call future.result() to wait for that one write, which is useful in tests.

Behind it is a single background writer for each Memory object. One writer means writes happen in the order they were queued, so the London and Paris problem cannot occur. The price is that all users share that one writer, so their writes wait in one line.

This example queues a write after every turn, which keeps the code short. Each job re-reads the session so far, and the dedup checks skip the facts already stored. To pay for one extraction per session, as the class design does, call add_async() once, when the session ends.

Try it yourself

Run the class experiment async-write and compare your two waits with ours. It needs Ollama, a program that runs small models on a laptop.

ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/async-write
pip install -r requirements.txt
python main.py

Common questions

If writing is asynchronous, can the assistant say "I'll remember that" honestly?

Yes. The fact is in the queue and will be written within seconds. What the assistant should not claim is that the store already holds it. Some products show the new memory to the user once the write has landed, which is also a useful safety habit.

How does the system know a session has ended?

Usually by silence. If no message arrives for some minutes, the session is treated as finished and its transcript is queued. A long session can also be cut into parts, so that a user who chats for three hours does not wait three hours to be remembered.

Why not make recall asynchronous too?

Because the answer depends on it. If the allergy card arrives after the model has started writing, the snack suggestion has already been made. Recall has to finish first, which is why it uses no LLM and has a 50 ms budget.

Does the background call not cost the same money?

Yes, exactly the same. Moving the write to the background saves waiting time, not money. The savings in money come from extracting once per session and not once per message.

What if the extractor is slow, say 7 seconds on a small local model?

Then it takes 7 seconds, and nobody notices. That is the freedom the write path has. It lets us choose a careful prompt and a cheap model without touching the user's wait.

Can I lose a memory this way?

Only if a job is lost before it runs. A durable queue with retries makes that rare. And hanumemAI saves the raw turns of the conversation before it calls the extractor. A session whose extraction failed can be passed to add() again later.

Carry this

  • Recall is synchronous: under 50 ms, no LLM. Remembering is asynchronous: seconds, with an LLM, after the answer is sent.
  • In the experiment the same two calls took 3,351 ms when blocking and 1,393 ms when asynchronous: 58 % of the wait gone.
  • Extract once per session: about 4.6 writes a second against about 28 reads a second.
  • The price is freshness. Same-session questions are answered from the context window, so nothing is lost.
  • Make the write path safe with retries, idempotent writes, order per user and a queue that does not forget unfinished jobs.

Check yourself

1. An answer call takes 900 ms and an extraction call takes 1,500 ms. What does the user wait with a blocking write, and with an asynchronous write? What share of the wait is saved?

Answer
blocking   1,500 + 900     = 2,400 ms
async                        900 ms
saved      2,400 - 900     = 1,500 ms
share      1,500 / 2,400   = 62.5 %

2. A service has 500,000 daily active users with 2 sessions a day and 6 messages per session. How many reads and writes per second, on average?

Answer
sessions   500,000 x 2          = 1,000,000 a day
messages   1,000,000 x 6        = 6,000,000 a day
writes     1,000,000 / 86,400   = about 11.6 per second
reads      6,000,000 / 86,400   = about 69 per second

3. A worker writes Maya's new fact, then crashes before confirming. The queue gives the job to another worker. Why does Maya not end up with two copies?

Answer

Because the write is idempotent. The second worker extracts the same fact, the hash check finds the text already stored for this user, and the candidate is skipped. If the wording differs a little, the similarity check of 0.95 catches it.