Part III · Journeyman
Chapter 13
The write path nobody waits for
Recall happens before the answer and must be fast. Remembering happens after the answer and may take seconds. One experiment shows 58 % of the wait disappear.
On this page
In this chapter, we will learn when an assistant should do its remembering. We will see that reading memory and writing memory have opposite needs, and measure what happens when the two are mixed up. Then we will build the queue that lets the slow work happen after the user already has the answer.
Two jobs with opposite needs
Every message touches memory twice. Before answering, the assistant recalls: it reads the cards that match the question. After answering, it remembers: it writes new cards from what was said.
A synchronous step is one the user waits for. An asynchronous step is one that runs in the background while the user carries on.
Recall must be synchronous. The answer cannot begin until the cards are on the desk. So recall has a strict budget: under 50 ms, which is invisible next to the 200 to 500 ms a model takes to produce its first word.
Remembering has no such pressure. The cat's name that Maya mentioned a moment ago is not needed for the answer she is waiting for right now. It is needed next week. So remembering is asynchronous.
| Recall (read path) | Remember (write path) | |
|---|---|---|
| When | before the answer | after the answer is sent |
| Does the user wait? | yes | no |
| Time allowed | under 50 ms | seconds are fine |
| Uses an LLM? | no, only the small embedding model | yes, the extractor |
| Character | fast and simple | slow and careful |
The same two calls, a very different wait
The class experiment async-write sends one message that holds both a fact worth keeping and
a question.
Maya: By the way, I just adopted a cat named Miso! Anyway: give me one tip
for keeping a flat tidy with kids. Two sentences.
This turn needs two LLM calls: one to answer, one to extract the fact about Miso. The only question is the order. In run (b) the extraction runs in a background thread, which is a second line of work inside the same program.
(a) BLOCKING write: extract first, then answer
extraction call : 1,605 ms
answer call : 1,746 ms
user-perceived latency = 1,605 + 1,746 = 3,351 ms (the user waits for BOTH)
(b) ASYNC write: answer immediately, extraction in a background thread
answer call : 1,393 ms
user-perceived latency = 1,393 ms (the answer call alone)
background extraction: 1,859 ms, finished 571 ms after the answer was done
stored memory: Maya has a cat named Miso.
Latency is the time the user waits. The difference is large.
blocking 3,351 ms async 1,393 ms saved 3,351 - 1,393 = 1,958 ms share saved 1,958 / 3,351 = 58 % of the wait gone
The same two calls were made. The same fact was stored. The same answer was given. The user simply stopped waiting for work that was never for them. Exact timings change from run to run, because this is a small model on a laptop. The shape does not change.
Set how long the extraction and the answer take. Compare the two timelines, and find the moment the new fact becomes readable.
The queue and the workers
A queue is a waiting line for jobs. When the assistant has sent its reply, it drops a small note in the line: "extract memories from this session, for this user". Then it turns to the next message. It does not wait to see the job done.
A worker is a program that takes jobs from the queue and does them, one after another. Here the worker runs the extractor, checks for duplicates and writes the new cards.
READ PATH the user waits about 28 per second
message --> recall (under 50 ms) --> assistant LLM --> answer streams to the user
|
| after the response is sent
v
WRITE PATH nobody waits about 4.6 per second
[ queue ] --> worker --> extractor LLM --> dedup and supersede --> memory store
the takes about 3 facts writes,
transcript a job about 1 s about 10 ms
Once per session, not once per message
The class design enqueues one job when a session ends, holding the whole transcript. Why not one job per message?
- It costs a sixth as much. A session has 6 user messages. One extraction for all of them is one call where there would have been six.
- The extractor sees the whole story. "I love blue. Actually, orange." becomes one clean fact. Extracting per message would write "loves blue" and then have to correct it.
Some products do extract after every message, and show a small "memory updated" notice at once. That buys instant freshness at about six times the extractor calls. It is a product choice, not a mistake.
How busy are the two paths?
daily active users 200,000 sessions per day 200,000 x 2 = 400,000 user messages per day 400,000 x 6 = 2,400,000 seconds in a day 86,400 reads (one per message) 2,400,000 / 86,400 = about 28 per second peak 3x = about 83 writes (one per session) 400,000 / 86,400 = about 4.6 per second peak 3x = about 14
There are 6 reads for every write. And each read is allowed 50 ms, where each write may take a second or more. The path that is used most has the least time. That is the right way round, and it is only possible because the slow work was moved out of it.
The queue also absorbs rushes. If a thousand sessions end in the same minute, the line grows longer for a while and the workers catch up. No user feels it, because no user is waiting.
The price: freshness
Freshness is how soon a new fact can be read back from the store. With an asynchronous write there is a short gap between saying something and the store knowing it. In the experiment the Miso fact became readable about two seconds after the answer call started.
answer finished 1,393 ms extraction finished after that 571 ms fact readable 1,393 + 571 = 1,964 ms the program prints 1,963; it rounds each part
Is that gap a problem? Picture Maya asking, in the very next message, "what did I say my cat was called?" The store may not have the card yet.
It does not matter. That question is answered from the context window, the short-term memory of Chapter 4. Everything said in this session is still on the desk. The long-term store is for next week, and by next week the card is written. The class sets the target plainly: memory must be fresh by the next session, and background extraction lands in seconds.
asked in the SAME session --> the context window answers (always fresh) asked in a LATER session --> the memory store answers (written long ago)
When things go wrong
Background work fails in ways the user never sees, which makes it easy to ignore. Four questions need an answer.
What if the extractor call fails?
A retry is a second attempt at a job that failed. The model provider may be slow or briefly down. The worker waits a moment and tries again, a few times, waiting longer each time. hanumemAI retries an LLM call up to 4 times by default, after waits of 1, 2, 4 and 8 seconds.
What if a job runs twice?
Retries create a new risk. Suppose the write succeeded but the worker crashed before it could say so. The job runs again. Will Maya get two Miso cards?
An operation is idempotent when doing it twice has the same result as doing it once. Pressing a lift button twice does not call two lifts. The write path gets this property from the hash of Chapter 10: a fact whose text is already stored for this user is skipped. Run the job twice and the second run adds nothing.
What if two jobs for the same user overlap?
Maya says "I live in London" in the morning and "I moved to Paris" in the evening. If the evening job were processed first, the store would end the day believing in London. So writes for one user must be done in the order they happened. Different users can be processed side by side, because their cards never touch.
What if a worker dies with a job in its hands?
A proper queue hands out a job but does not forget it until the worker reports success. If the worker goes silent, the queue gives the job to another worker after a time limit. And a job that fails again and again is moved aside to a separate line, often called a dead-letter queue, for a person to look at. It must not block the jobs behind it.
The write path in hanumemAI
The library gives the two styles side by side. add() waits for the write.
add_async() queues it and returns at once.
from hmem import Memory, Config
memory = Memory(Config(db_path="app.sqlite"))
def handle_turn(user_id, session, message, today):
# 1. recall: synchronous, no LLM
block, rows = memory.prompt_block(message, user_id=user_id, budget_tokens=1500, ref_date=today)
# 2. answer: your model, your prompt
reply = llm.chat(system="Known facts about the user:\n" + block,
messages=session + [{"role": "user", "content": message}])
# 3. remember: queued, returns immediately
memory.add_async(session + [{"role": "user", "content": message},
{"role": "assistant", "content": reply}],
user_id=user_id, observed_at=today)
return reply
memory.flush() # before shutdown: wait for the queued writes to finish
add_async() returns a future, which is a small ticket for a result that does not exist yet.
You may ignore it. Or you may call future.result() to wait for that one write, which is useful
in tests.
Behind it is a single background writer for each Memory object. One writer means writes
happen in the order they were queued, so the London and Paris problem cannot occur. The price is that all
users share that one writer, so their writes wait in one line.
This example queues a write after every turn, which keeps the code short. Each job re-reads the session
so far, and the dedup checks skip the facts already stored. To pay for one extraction per session, as the
class design does, call add_async() once, when the session ends.
Try it yourself
Run the class experiment async-write and compare your two waits with ours. It needs Ollama,
a program that runs small models on a laptop.
ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/async-write
pip install -r requirements.txt
python main.py
Common questions
If writing is asynchronous, can the assistant say "I'll remember that" honestly?
Yes. The fact is in the queue and will be written within seconds. What the assistant should not claim is that the store already holds it. Some products show the new memory to the user once the write has landed, which is also a useful safety habit.
How does the system know a session has ended?
Usually by silence. If no message arrives for some minutes, the session is treated as finished and its transcript is queued. A long session can also be cut into parts, so that a user who chats for three hours does not wait three hours to be remembered.
Why not make recall asynchronous too?
Because the answer depends on it. If the allergy card arrives after the model has started writing, the snack suggestion has already been made. Recall has to finish first, which is why it uses no LLM and has a 50 ms budget.
Does the background call not cost the same money?
Yes, exactly the same. Moving the write to the background saves waiting time, not money. The savings in money come from extracting once per session and not once per message.
What if the extractor is slow, say 7 seconds on a small local model?
Then it takes 7 seconds, and nobody notices. That is the freedom the write path has. It lets us choose a careful prompt and a cheap model without touching the user's wait.
Can I lose a memory this way?
Only if a job is lost before it runs. A durable queue with retries makes that rare. And hanumemAI saves
the raw turns of the conversation before it calls the extractor. A session whose extraction failed can be
passed to add() again later.
Carry this
- Recall is synchronous: under 50 ms, no LLM. Remembering is asynchronous: seconds, with an LLM, after the answer is sent.
- In the experiment the same two calls took 3,351 ms when blocking and 1,393 ms when asynchronous: 58 % of the wait gone.
- Extract once per session: about 4.6 writes a second against about 28 reads a second.
- The price is freshness. Same-session questions are answered from the context window, so nothing is lost.
- Make the write path safe with retries, idempotent writes, order per user and a queue that does not forget unfinished jobs.
Check yourself
1. An answer call takes 900 ms and an extraction call takes 1,500 ms. What does the user wait with a blocking write, and with an asynchronous write? What share of the wait is saved?
Answer
blocking 1,500 + 900 = 2,400 ms async 900 ms saved 2,400 - 900 = 1,500 ms share 1,500 / 2,400 = 62.5 %
2. A service has 500,000 daily active users with 2 sessions a day and 6 messages per session. How many reads and writes per second, on average?
Answer
sessions 500,000 x 2 = 1,000,000 a day messages 1,000,000 x 6 = 6,000,000 a day writes 1,000,000 / 86,400 = about 11.6 per second reads 6,000,000 / 86,400 = about 69 per second
3. A worker writes Maya's new fact, then crashes before confirming. The queue gives the job to another worker. Why does Maya not end up with two copies?
Answer
Because the write is idempotent. The second worker extracts the same fact, the hash check finds the text already stored for this user, and the candidate is skipped. If the wording differs a little, the similarity check of 0.95 catches it.