Part V · Practice
Chapter 22
Build an assistant that remembers
Sixty lines of Python, four sessions in Maya's life, and the real output: what worked, what the box kept, and two honest flaws.
On this page
In this chapter, we will build a working assistant with long-term memory and run it. We will use hanumemAI, the library from Chapter 20. Every line of output in this chapter is real: we ran the program and pasted what it printed, including the two places where it did not behave well.
The plan
An assistant with memory is a loop of three steps: recall, answer, remember. We met this loop in Chapter 3. Here it becomes code.
for every message:
1. recall ask the memory layer for a small, dated block of facts (fast, no LLM call)
2. answer send system + block + this session + message to the model
at the end of the session:
3. remember hand the whole session to the memory layer (slow, in the background)
We will play four sessions from Maya's life. They are weeks and months apart. Each one is a new chat with an empty context window. The only thing that travels from one session to the next is the box. In this run the assistant first meets Maya on January 14, 2026, so its dates differ from the examples of Chapter 11.
| Session | Date | What it tests |
|---|---|---|
| 1 | January 14 | Writing the first cards: diet, allergy, kids, home, a birthday |
| 2 | March 2 | A fact that changes (the move to Paris) and a standing instruction |
| 3 | March 20 | The question from Chapter 1: a snack for the school trip |
| 4 | June 10 | A question about the past, and a secret that must not be kept |
Set up
Setting up means installing the library and giving it a key to a model provider. The
library needs Python and one API key for any provider that speaks the OpenAI protocol. By default it uses
small Qwen models through OpenRouter, a service that sells access to many models with one key. The package
imports as hmem until the rename to hanumemAI ships.
The command uv sync installs what the library depends on. uv is a package
manager for Python.
cd ~/apps/hmem
uv sync
export OPENROUTER_API_KEY=sk-or-...
The code
The program has three parts: the memory layer with its test data, one turn, and the loop over sessions. First the memory layer and the four sessions.
from hmem import Config, Memory
memory = Memory(Config(db_path="family.sqlite"))
llm = memory.llm # the same small model answers; any chat model works here
USER = "maya"
SESSIONS = [
("2026-01-14", ["I'm vegetarian, and I have a serious peanut allergy. I have two kids, aged 6 and 9.",
"I live in London. My partner Sam has a birthday on June 18."]),
("2026-03-02", ["Big news: I just moved from London to Paris for a job at a payments startup.",
"Keep your answers short and direct, please."]),
("2026-03-20", ["Suggest a snack for my kids' school trip."]),
("2026-06-10", ["Where did I live before Paris, and when did I move?",
"My wifi password is sunflower42. Any gift ideas for next week?"]),
]
Now one turn. Read the numbered comments: 1 and 2 are here, and 3 is in the loop that follows. They are the whole book in a few lines.
def handle_turn(message, session, today):
# 1. recall: synchronous, no LLM call, a dated block capped in tokens
block, rows = memory.prompt_block(message, user_id=USER, budget_tokens=600, ref_date=today)
system = ("You are Maya's personal assistant. Today is " + today + ".\n"
"Known facts about the user (dated; rows marked outdated are history):\n"
+ (block or "(nothing yet)") +
"\nPersonalize the answer using these facts where relevant. "
"Do not recite them unless asked. Answer in one to three sentences.")
# 2. answer: your model, your prompt
reply = llm.complete([{"role": "system", "content": system}] + session
+ [{"role": "user", "content": message}], max_tokens=200)
return reply
And the loop over sessions. Step 3 happens once per session, after the last reply.
for today, messages in SESSIONS:
session = [] # a new chat: an empty context window
for msg in messages:
reply = handle_turn(msg, session, today)
session += [{"role": "user", "name": "Maya", "content": msg},
{"role": "assistant", "content": reply}]
# 3. remember: once per session, off the critical path
memory.add_async(session, user_id=USER, observed_at=today, session_id=f"s-{today}")
memory.flush() # the demo waits so the next session can read; a server would not
Three details deserve a second look.
user_idis on every call. It is the privacy wall from Chapter 15. There is no way to search without it.budget_tokens=600is a budget in tokens, not in rows. The rows in the block can never add up to more than it, however large the box becomes. The running profile, at most 120 words, is placed on top and is not counted in the 600.observed_attells the secretary what day it is. Without it, "last month" cannot be turned into a date.
Run it: the real output
cd ~/apps/hmem
uv run python ~/apps/interactive/books/memory/examples/family_assistant.py
Session 1: an empty box
=== session on 2026-01-14 (a new chat, an empty context window) ===
Maya: I'm vegetarian, and I have a serious peanut allergy. I have two kids, aged 6 and 9.
Assistant: Noted. I'll keep your vegetarian diet, severe peanut allergy, and the ages of your
two children in mind for any future recommendations or meal planning.
[recall: 0 rows, 0 tokens, 1 ms]
Maya: I live in London. My partner Sam has a birthday on June 18.
Assistant: Noted. I'll keep your London location and Sam's June 18 birthday in mind for any
future recommendations or planning.
[recall: 0 rows, 0 tokens, 0 ms]
Recall returns nothing, because nothing has been written yet. The assistant says "Noted", exactly as it did in Chapter 1. The difference is what happens after the session: the secretary reads it and writes cards.
The line in square brackets is printed by our program after every reply. It shows how many rows the librarian returned, how many tokens the memory block took, and how long recall took.
Session 2: life changes
=== session on 2026-03-02 (a new chat, an empty context window) ===
Maya: Big news: I just moved from London to Paris for a job at a payments startup.
Assistant: Congratulations on the new role at the payments startup! Since you are now in Paris,
I'll update your location and keep an eye out for vegetarian-friendly spots that
strictly avoid peanuts, which is crucial given your allergy. Let me know if you need
recommendations for family activities for your 6 and 9-year-olds in their new city.
[recall: 8 rows, 245 tokens, 116 ms]
Maya: Keep your answers short and direct, please.
Assistant: Congratulations on the move to Paris and your new role! I've updated your location and
will ensure future recommendations are vegetarian, peanut-free, and suitable for your
children.
[recall: 8 rows, 245 tokens, 172 ms]
Almost seven weeks have passed and this is a new chat. The assistant knows about the diet, the allergy and the children. It read them from 245 tokens of memory.
Why 8 rows, and why 245 tokens?
hanumemAI stores the cards and also the raw turns they came from. Both can be returned as rows.
cards written after session 1 4 (diet and allergy, children, London, Sam's birthday) raw turns of session 1 4 (2 messages from Maya, 2 replies) rows returned 4 + 4 = 8 session 1 as plain text 410 characters / 4 = about 102 tokens the memory block 245 tokens
So in this tiny demo the memory block is larger than the session it came from. It holds each fact twice, as a card and as a raw turn, and a short profile of Maya on top. Memory does not win on day one. It wins because the block has a ceiling, 600 tokens of rows plus the short profile, while a history that is resent grows with every session and has none.
Session 3: the question from Chapter 1
=== session on 2026-03-20 (a new chat, an empty context window) ===
Maya: Suggest a snack for my kids' school trip.
Assistant: Try cheese sticks and fruit slices. They are vegetarian, peanut-free, and easy for
kids to eat.
[recall: 14 rows, 508 tokens, 150 ms]
This is the moment the book was written for. In Chapter 1 the answer was peanut butter energy balls. Here the allergy card was on the desk, and the answer is safe. Notice two more things. The answer is short and direct, as the rule card written in session 2 asks. And Maya did not mention food, allergies or her instruction. The assistant used the memories at the right moment, without being asked, which is the second functional requirement from Chapter 5.
Session 4: the past, and a secret
=== session on 2026-06-10 (a new chat, an empty context window) ===
Maya: Where did I live before Paris, and when did I move?
Assistant: You lived in London before moving to Paris in March 2026.
[recall: 16 rows, 560 tokens, 131 ms]
Maya: My wifi password is sunflower42. Any gift ideas for next week?
Assistant: I cannot store or use your Wi-Fi password for security reasons. For Sam's birthday on
June 18, consider a vegetarian-friendly dining experience in Paris or a tech gadget
from his favorite payments startup brand.
[recall: 16 rows, 560 tokens, 125 ms]
The first answer is only possible because the London card was closed, not deleted. A design that overwrites "London" with "Paris" cannot answer this question at all. This is Chapter 11 at work.
The second answer connects "next week" to Sam's birthday on June 18. Maya did not say whose gift it was. The date in her question and the date on the card met on the desk.
What the box holds at the end
=== what the box holds now ===
[semantic imp 8] Maya is vegetarian and has a serious peanut allergy.
[semantic imp 7] Maya has two children aged 6 and 9.
[semantic imp 6] Maya's partner Sam has a birthday on June 18.
[episodic imp 9] Maya moved from London to Paris in March 2026 for a job at a payments startup.
[procedural imp 6] Maya prefers short and direct answers from the assistant.
=== closed cards (history, never deleted) ===
Maya lives in London. [2026-01-14 -> 2026-03-02]
stats: {'facts': 6, 'episodes': 14, 'vector_bytes': 215040, 'shard': 0}
after delete_user: {'facts': 0, 'episodes': 0, 'vector_bytes': 0, 'shard': 0}
Four sessions became five open cards and one closed card. "imp" is the importance that the extractor gave the card, from 1 to 10.
facts 5 open + 1 closed = 6 episodes 7 messages from Maya + 7 replies = 14 (the raw turns)
There is no card for the wifi password. Credentials are what the write policy refuses
(Chapter 15). Our program does not print the refusals,
so the output shows the result and not the reason. And one call, delete_user, emptied the box
completely, vectors included.
Reading the output like an engineer
Reading like an engineer means comparing what the system did with what it promised. Let us check the run against the earlier chapters.
| Promise | What we saw | Verdict |
|---|---|---|
| Self-contained cards with names and dates | "Maya moved from London to Paris in March 2026 ..." | Kept |
| Smalltalk and one-off requests get no card | No card for the snack or gift requests | Kept |
| A change closes the old card | London closed on 2026-03-02, still readable | Kept |
| Secrets are refused | No fact about the wifi password | Kept, with a gap (see below) |
| Read cost has a ceiling | 245, 508, 560 tokens: growing, and under the 600 budget | Kept |
| One attribute per card | "vegetarian and peanut allergy" is one card | Not by default |
| Read path under 50 ms | 116 to 172 ms | Missed |
Why is the read path over budget?
The search inside the box takes a few milliseconds. The rest is one network call to a hosted embedding model, from a laptop, for every question.
measured recall time 116 to 172 ms
search inside SQLite about 5 ms (measured in our experiments)
the embedding call 116 - 5 = 111 ms to 172 - 5 = 167 ms
(a remote endpoint; about 150 ms in our experiments)
To meet the 50 ms target of Chapter 8, embed near the application: a local model, or an embedding service in the same data centre. The design is not the problem. The distance is.
The compound card
"Maya is vegetarian and has a serious peanut allergy" is one card holding two attributes. If Maya stops being vegetarian, closing that card would close the allergy too. We found this bug with a golden set (Chapter 17). The fix is one setting:
memory = Memory(Config(db_path="family.sqlite", atomic_facts=True)) # one attribute per card
It is off by default because it wrote 37 percent more cards and cost 2 to 3 points on the benchmarks (E0004). For an assistant that handles allergies, turn it on. Correctness beats a score.
The gap in the write policy
The stats line says episodes: 14. hanumemAI stores the raw turns as well as the facts,
because the exact words help the reader. The write policy guards the facts. It does not remove the
raw turn, so the sentence with the password is still in the box as an episode until the user is deleted.
In session 4 the assistant even said "I cannot store your Wi-Fi password". That was the model talking. It
does not know what the memory layer keeps.
We say this plainly because a book about memory should not hide it. The setting
store_episodes=False is only half a fix. With it, raw turns are no longer searchable and never
reach a prompt. But the library still saves the text of every turn in a side table, which the extractor
uses to understand the next session. Until the library closes the gap, an application that may receive
secrets should remove them before the text reaches the memory layer:
clean = redact_secrets(session) # your own function: mask passwords, keys, card numbers
memory.add_async(clean, user_id=USER, observed_at=today)
From demo to service
A production service keeps the same three steps and changes what surrounds them.
| In the demo | In production | Chapter |
|---|---|---|
flush() after each session | Never wait. A queue and workers write in the background | 13 |
| One SQLite file | shards=16, one file per bucket of users | 16 |
| float32 vectors | vector_dtype="int8" or vector_backend="turbovec" | 16 |
| A remote embedding call | An embedding service next to the application | 8 |
| Nobody looks at the box | A page where the user sees, edits and deletes memories | 15 |
| We read the output by eye | A golden set replayed on every change | 17 |
| Nothing expires | memory.forget(user_id) every night | 12 |
memory = Memory(Config(db_path="data/memory", shards=16, vector_dtype="int8", atomic_facts=True))
# the user's own controls
memory.get_all("maya") # see every card
memory.update(card_id, "Maya lives in Lyon.", "maya") # edit by superseding: history survives
memory.delete(card_id, "maya") # close one card: it leaves search, it is not erased
memory.delete_user("maya") # the right to be forgotten
Common questions
Can I use a different model for the answers?
Yes. hanumemAI never answers the user. It returns a block of text, and you place it in the prompt of any model you like. In our measurements a stronger answer model was the cheapest large gain: LongMemEval rose from 81.0 to 88.6 with the same memories and the same 2,950 tokens.
How often should I call add()?
Once per session, at its end, with the whole session. The extractor works best when it sees the whole conversation. In experiment 0020, extracting in windows of ten turns gave the same score with 38 percent more cards, and it took 2.6 times as long.
What if the user asks about something they said one minute ago?
That is answered from the context window, because the turn is still on the desk. Long-term memory is for other sessions. Nothing is lost while the background write is still running.
The run cost money. How much?
We did not meter this run, so we give a ceiling and not a bill. The run made seven short answers, a few extraction and profile calls and the embedding calls, all on small models. The default model costs $0.03 per million tokens read and $0.13 per million tokens written.
even 100,000 tokens, all at the higher price 100,000 / 1,000,000 x $0.13 = $0.013 about one cent
The whole experiment series behind the library, 121 runs on three benchmarks, cost about 25 dollars.
Why did the move get the category "episodic"?
The extractor saw it as an event with a date. It is also the fact that says where Maya lives now, so "semantic" would be just as fair. The labels are the extractor's judgement, and they are not perfect. This matters because only episodic cards decay. If your application relies on decay, test the labels with a golden set.
Will my output be the same as yours?
The cards should be very close. The wording of the answers will differ. Even at temperature 0 the provider is not fully deterministic: when we re-ran an identical configuration, 63 of 385 answers changed and the score moved by one point.
Carry this
- An assistant with memory is three steps: recall (fast, every message), answer, remember (slow, once per session).
- The snack question from Chapter 1 now gets a safe answer, from 508 tokens of memory. The ceiling, 600 tokens of rows plus a profile of at most 120 words, holds however long the history grows.
- "Where did I live before?" works only because the old card was closed and not deleted.
- Memory puts true sentences on the desk. The model can still recite them or blend them badly. Those are reading errors.
- Know the limits of your library: here, compound cards by default, raw turns kept as written, and a remote embedding call that breaks the 50 ms budget. Remove secrets before the text reaches memory.
Check yourself
1. In session 3 Maya did not mention her allergy. Trace how the word "peanut-free" reached the answer.
Answer
In session 1 she stated the allergy. After the session the extractor wrote the card "Maya is vegetarian and has a serious peanut allergy". In session 3 the question about a snack was embedded, the search inside Maya's box ranked that card high, and the card was placed in the prompt block. The model read it and chose a safe snack.
2. The demo calls flush() after every session. Why must a
real server not do this inside the request?
Answer
flush() waits for the extraction to finish, which takes
seconds. The user would wait for the secretary. The write path belongs in the background: the answer is
already on the screen, and the new cards are only needed by the next session.
3. Recall took about 130 ms. You must reach 50 ms. What do you change first, and what do you leave alone?
Answer
Move the embedding step close to the application, with a local model or a service in the same data centre. It is nearly all of the time. Leave the search alone: it already takes only a few milliseconds.