Carebun Interactive/Memory Memory labCheat sheet

Part V · Practice

Chapter 22

Build an assistant that remembers

Sixty lines of Python, four sessions in Maya's life, and the real output: what worked, what the box kept, and two honest flaws.

12 min read

On this page
  1. The plan
  2. Set up
  3. The code
  4. Run it: the real output
  5. Reading the output like an engineer
  6. From demo to service
  7. Common questions
  8. Carry this
  9. Check yourself

In this chapter, we will build a working assistant with long-term memory and run it. We will use hanumemAI, the library from Chapter 20. Every line of output in this chapter is real: we ran the program and pasted what it printed, including the two places where it did not behave well.

The plan

An assistant with memory is a loop of three steps: recall, answer, remember. We met this loop in Chapter 3. Here it becomes code.

for every message:
    1. recall      ask the memory layer for a small, dated block of facts     (fast, no LLM call)
    2. answer      send  system + block + this session + message  to the model
at the end of the session:
    3. remember    hand the whole session to the memory layer                 (slow, in the background)

We will play four sessions from Maya's life. They are weeks and months apart. Each one is a new chat with an empty context window. The only thing that travels from one session to the next is the box. In this run the assistant first meets Maya on January 14, 2026, so its dates differ from the examples of Chapter 11.

SessionDateWhat it tests
1January 14Writing the first cards: diet, allergy, kids, home, a birthday
2March 2A fact that changes (the move to Paris) and a standing instruction
3March 20The question from Chapter 1: a snack for the school trip
4June 10A question about the past, and a secret that must not be kept

Set up

Setting up means installing the library and giving it a key to a model provider. The library needs Python and one API key for any provider that speaks the OpenAI protocol. By default it uses small Qwen models through OpenRouter, a service that sells access to many models with one key. The package imports as hmem until the rename to hanumemAI ships.

The command uv sync installs what the library depends on. uv is a package manager for Python.

cd ~/apps/hmem
uv sync
export OPENROUTER_API_KEY=sk-or-...

The code

The program has three parts: the memory layer with its test data, one turn, and the loop over sessions. First the memory layer and the four sessions.

from hmem import Config, Memory

memory = Memory(Config(db_path="family.sqlite"))
llm = memory.llm          # the same small model answers; any chat model works here
USER = "maya"

SESSIONS = [
    ("2026-01-14", ["I'm vegetarian, and I have a serious peanut allergy. I have two kids, aged 6 and 9.",
                    "I live in London. My partner Sam has a birthday on June 18."]),
    ("2026-03-02", ["Big news: I just moved from London to Paris for a job at a payments startup.",
                    "Keep your answers short and direct, please."]),
    ("2026-03-20", ["Suggest a snack for my kids' school trip."]),
    ("2026-06-10", ["Where did I live before Paris, and when did I move?",
                    "My wifi password is sunflower42. Any gift ideas for next week?"]),
]

Now one turn. Read the numbered comments: 1 and 2 are here, and 3 is in the loop that follows. They are the whole book in a few lines.

def handle_turn(message, session, today):
    # 1. recall: synchronous, no LLM call, a dated block capped in tokens
    block, rows = memory.prompt_block(message, user_id=USER, budget_tokens=600, ref_date=today)

    system = ("You are Maya's personal assistant. Today is " + today + ".\n"
              "Known facts about the user (dated; rows marked outdated are history):\n"
              + (block or "(nothing yet)") +
              "\nPersonalize the answer using these facts where relevant. "
              "Do not recite them unless asked. Answer in one to three sentences.")

    # 2. answer: your model, your prompt
    reply = llm.complete([{"role": "system", "content": system}] + session
                         + [{"role": "user", "content": message}], max_tokens=200)
    return reply

And the loop over sessions. Step 3 happens once per session, after the last reply.

for today, messages in SESSIONS:
    session = []                      # a new chat: an empty context window
    for msg in messages:
        reply = handle_turn(msg, session, today)
        session += [{"role": "user", "name": "Maya", "content": msg},
                    {"role": "assistant", "content": reply}]

    # 3. remember: once per session, off the critical path
    memory.add_async(session, user_id=USER, observed_at=today, session_id=f"s-{today}")
    memory.flush()                    # the demo waits so the next session can read; a server would not

Three details deserve a second look.

  • user_id is on every call. It is the privacy wall from Chapter 15. There is no way to search without it.
  • budget_tokens=600 is a budget in tokens, not in rows. The rows in the block can never add up to more than it, however large the box becomes. The running profile, at most 120 words, is placed on top and is not counted in the 600.
  • observed_at tells the secretary what day it is. Without it, "last month" cannot be turned into a date.

Run it: the real output

cd ~/apps/hmem
uv run python ~/apps/interactive/books/memory/examples/family_assistant.py

Session 1: an empty box

=== session on 2026-01-14 (a new chat, an empty context window) ===
  Maya:       I'm vegetarian, and I have a serious peanut allergy. I have two kids, aged 6 and 9.
  Assistant:  Noted. I'll keep your vegetarian diet, severe peanut allergy, and the ages of your
              two children in mind for any future recommendations or meal planning.
              [recall: 0 rows, 0 tokens, 1 ms]
  Maya:       I live in London. My partner Sam has a birthday on June 18.
  Assistant:  Noted. I'll keep your London location and Sam's June 18 birthday in mind for any
              future recommendations or planning.
              [recall: 0 rows, 0 tokens, 0 ms]

Recall returns nothing, because nothing has been written yet. The assistant says "Noted", exactly as it did in Chapter 1. The difference is what happens after the session: the secretary reads it and writes cards.

The line in square brackets is printed by our program after every reply. It shows how many rows the librarian returned, how many tokens the memory block took, and how long recall took.

Session 2: life changes

=== session on 2026-03-02 (a new chat, an empty context window) ===
  Maya:       Big news: I just moved from London to Paris for a job at a payments startup.
  Assistant:  Congratulations on the new role at the payments startup! Since you are now in Paris,
              I'll update your location and keep an eye out for vegetarian-friendly spots that
              strictly avoid peanuts, which is crucial given your allergy. Let me know if you need
              recommendations for family activities for your 6 and 9-year-olds in their new city.
              [recall: 8 rows, 245 tokens, 116 ms]
  Maya:       Keep your answers short and direct, please.
  Assistant:  Congratulations on the move to Paris and your new role! I've updated your location and
              will ensure future recommendations are vegetarian, peanut-free, and suitable for your
              children.
              [recall: 8 rows, 245 tokens, 172 ms]

Almost seven weeks have passed and this is a new chat. The assistant knows about the diet, the allergy and the children. It read them from 245 tokens of memory.

Why 8 rows, and why 245 tokens?

hanumemAI stores the cards and also the raw turns they came from. Both can be returned as rows.

cards written after session 1     4    (diet and allergy, children, London, Sam's birthday)
raw turns of session 1            4    (2 messages from Maya, 2 replies)
rows returned                     4 + 4 = 8

session 1 as plain text           410 characters / 4 = about 102 tokens
the memory block                  245 tokens

So in this tiny demo the memory block is larger than the session it came from. It holds each fact twice, as a card and as a raw turn, and a short profile of Maya on top. Memory does not win on day one. It wins because the block has a ceiling, 600 tokens of rows plus the short profile, while a history that is resent grows with every session and has none.

Session 3: the question from Chapter 1

=== session on 2026-03-20 (a new chat, an empty context window) ===
  Maya:       Suggest a snack for my kids' school trip.
  Assistant:  Try cheese sticks and fruit slices. They are vegetarian, peanut-free, and easy for
              kids to eat.
              [recall: 14 rows, 508 tokens, 150 ms]

This is the moment the book was written for. In Chapter 1 the answer was peanut butter energy balls. Here the allergy card was on the desk, and the answer is safe. Notice two more things. The answer is short and direct, as the rule card written in session 2 asks. And Maya did not mention food, allergies or her instruction. The assistant used the memories at the right moment, without being asked, which is the second functional requirement from Chapter 5.

Session 4: the past, and a secret

=== session on 2026-06-10 (a new chat, an empty context window) ===
  Maya:       Where did I live before Paris, and when did I move?
  Assistant:  You lived in London before moving to Paris in March 2026.
              [recall: 16 rows, 560 tokens, 131 ms]
  Maya:       My wifi password is sunflower42. Any gift ideas for next week?
  Assistant:  I cannot store or use your Wi-Fi password for security reasons. For Sam's birthday on
              June 18, consider a vegetarian-friendly dining experience in Paris or a tech gadget
              from his favorite payments startup brand.
              [recall: 16 rows, 560 tokens, 125 ms]

The first answer is only possible because the London card was closed, not deleted. A design that overwrites "London" with "Paris" cannot answer this question at all. This is Chapter 11 at work.

The second answer connects "next week" to Sam's birthday on June 18. Maya did not say whose gift it was. The date in her question and the date on the card met on the desk.

What the box holds at the end

=== what the box holds now ===
  [semantic   imp 8] Maya is vegetarian and has a serious peanut allergy.
  [semantic   imp 7] Maya has two children aged 6 and 9.
  [semantic   imp 6] Maya's partner Sam has a birthday on June 18.
  [episodic   imp 9] Maya moved from London to Paris in March 2026 for a job at a payments startup.
  [procedural imp 6] Maya prefers short and direct answers from the assistant.

=== closed cards (history, never deleted) ===
  Maya lives in London.   [2026-01-14 -> 2026-03-02]

stats: {'facts': 6, 'episodes': 14, 'vector_bytes': 215040, 'shard': 0}
after delete_user: {'facts': 0, 'episodes': 0, 'vector_bytes': 0, 'shard': 0}

Four sessions became five open cards and one closed card. "imp" is the importance that the extractor gave the card, from 1 to 10.

facts       5 open + 1 closed                 =  6
episodes    7 messages from Maya + 7 replies  = 14    (the raw turns)

There is no card for the wifi password. Credentials are what the write policy refuses (Chapter 15). Our program does not print the refusals, so the output shows the result and not the reason. And one call, delete_user, emptied the box completely, vectors included.

Reading the output like an engineer

Reading like an engineer means comparing what the system did with what it promised. Let us check the run against the earlier chapters.

PromiseWhat we sawVerdict
Self-contained cards with names and dates"Maya moved from London to Paris in March 2026 ..."Kept
Smalltalk and one-off requests get no cardNo card for the snack or gift requestsKept
A change closes the old cardLondon closed on 2026-03-02, still readableKept
Secrets are refusedNo fact about the wifi passwordKept, with a gap (see below)
Read cost has a ceiling245, 508, 560 tokens: growing, and under the 600 budgetKept
One attribute per card"vegetarian and peanut allergy" is one cardNot by default
Read path under 50 ms116 to 172 msMissed

Why is the read path over budget?

The search inside the box takes a few milliseconds. The rest is one network call to a hosted embedding model, from a laptop, for every question.

measured recall time            116 to 172 ms
search inside SQLite            about 5 ms          (measured in our experiments)
the embedding call              116 - 5 = 111 ms  to  172 - 5 = 167 ms
                                (a remote endpoint; about 150 ms in our experiments)

To meet the 50 ms target of Chapter 8, embed near the application: a local model, or an embedding service in the same data centre. The design is not the problem. The distance is.

The compound card

"Maya is vegetarian and has a serious peanut allergy" is one card holding two attributes. If Maya stops being vegetarian, closing that card would close the allergy too. We found this bug with a golden set (Chapter 17). The fix is one setting:

memory = Memory(Config(db_path="family.sqlite", atomic_facts=True))   # one attribute per card

It is off by default because it wrote 37 percent more cards and cost 2 to 3 points on the benchmarks (E0004). For an assistant that handles allergies, turn it on. Correctness beats a score.

The gap in the write policy

The stats line says episodes: 14. hanumemAI stores the raw turns as well as the facts, because the exact words help the reader. The write policy guards the facts. It does not remove the raw turn, so the sentence with the password is still in the box as an episode until the user is deleted. In session 4 the assistant even said "I cannot store your Wi-Fi password". That was the model talking. It does not know what the memory layer keeps.

We say this plainly because a book about memory should not hide it. The setting store_episodes=False is only half a fix. With it, raw turns are no longer searchable and never reach a prompt. But the library still saves the text of every turn in a side table, which the extractor uses to understand the next session. Until the library closes the gap, an application that may receive secrets should remove them before the text reaches the memory layer:

clean = redact_secrets(session)       # your own function: mask passwords, keys, card numbers
memory.add_async(clean, user_id=USER, observed_at=today)

From demo to service

A production service keeps the same three steps and changes what surrounds them.

In the demoIn productionChapter
flush() after each sessionNever wait. A queue and workers write in the background13
One SQLite fileshards=16, one file per bucket of users16
float32 vectorsvector_dtype="int8" or vector_backend="turbovec"16
A remote embedding callAn embedding service next to the application8
Nobody looks at the boxA page where the user sees, edits and deletes memories15
We read the output by eyeA golden set replayed on every change17
Nothing expiresmemory.forget(user_id) every night12
memory = Memory(Config(db_path="data/memory", shards=16, vector_dtype="int8", atomic_facts=True))

# the user's own controls
memory.get_all("maya")                                   # see every card
memory.update(card_id, "Maya lives in Lyon.", "maya")   # edit by superseding: history survives
memory.delete(card_id, "maya")                           # close one card: it leaves search, it is not erased
memory.delete_user("maya")                               # the right to be forgotten

Common questions

Can I use a different model for the answers?

Yes. hanumemAI never answers the user. It returns a block of text, and you place it in the prompt of any model you like. In our measurements a stronger answer model was the cheapest large gain: LongMemEval rose from 81.0 to 88.6 with the same memories and the same 2,950 tokens.

How often should I call add()?

Once per session, at its end, with the whole session. The extractor works best when it sees the whole conversation. In experiment 0020, extracting in windows of ten turns gave the same score with 38 percent more cards, and it took 2.6 times as long.

What if the user asks about something they said one minute ago?

That is answered from the context window, because the turn is still on the desk. Long-term memory is for other sessions. Nothing is lost while the background write is still running.

The run cost money. How much?

We did not meter this run, so we give a ceiling and not a bill. The run made seven short answers, a few extraction and profile calls and the embedding calls, all on small models. The default model costs $0.03 per million tokens read and $0.13 per million tokens written.

even 100,000 tokens, all at the higher price
100,000 / 1,000,000 x $0.13 = $0.013      about one cent

The whole experiment series behind the library, 121 runs on three benchmarks, cost about 25 dollars.

Why did the move get the category "episodic"?

The extractor saw it as an event with a date. It is also the fact that says where Maya lives now, so "semantic" would be just as fair. The labels are the extractor's judgement, and they are not perfect. This matters because only episodic cards decay. If your application relies on decay, test the labels with a golden set.

Will my output be the same as yours?

The cards should be very close. The wording of the answers will differ. Even at temperature 0 the provider is not fully deterministic: when we re-ran an identical configuration, 63 of 385 answers changed and the score moved by one point.

Carry this

  • An assistant with memory is three steps: recall (fast, every message), answer, remember (slow, once per session).
  • The snack question from Chapter 1 now gets a safe answer, from 508 tokens of memory. The ceiling, 600 tokens of rows plus a profile of at most 120 words, holds however long the history grows.
  • "Where did I live before?" works only because the old card was closed and not deleted.
  • Memory puts true sentences on the desk. The model can still recite them or blend them badly. Those are reading errors.
  • Know the limits of your library: here, compound cards by default, raw turns kept as written, and a remote embedding call that breaks the 50 ms budget. Remove secrets before the text reaches memory.

Check yourself

1. In session 3 Maya did not mention her allergy. Trace how the word "peanut-free" reached the answer.

Answer

In session 1 she stated the allergy. After the session the extractor wrote the card "Maya is vegetarian and has a serious peanut allergy". In session 3 the question about a snack was embedded, the search inside Maya's box ranked that card high, and the card was placed in the prompt block. The model read it and chose a safe snack.

2. The demo calls flush() after every session. Why must a real server not do this inside the request?

Answer

flush() waits for the extraction to finish, which takes seconds. The user would wait for the secretary. The write path belongs in the background: the answer is already on the screen, and the new cards are only needed by the next session.

3. Recall took about 130 ms. You must reach 50 ms. What do you change first, and what do you leave alone?

Answer

Move the embedding step close to the application, with a local model or a service in the same data centre. It is nearly all of the time. Leave the search alone: it already takes only a few milliseconds.