Part V · Practice
Chapter 24
Questions readers ask
Sixty questions that remain after the last chapter, each answered in a few plain sentences, with a number where one exists and the chapter that teaches it.
On this page
In this chapter, we will learn the short answers to sixty questions that readers still ask after the last chapter. Find your question under one of the ten headings, read the answer, and follow the link to the chapter that teaches it in full.
The answers use the same card box, the same secretary and librarian, and the same Maya as the rest of the book. Where an answer depends on a measurement, it names the measurement. Where we do not know, it says so.
The basics
What is agent memory, in one sentence?
It is a small set of facts about one user, kept outside the model, from which only the few relevant ones are placed in the prompt. The model stays frozen. The facts live in a database that we can read, correct and delete. Chapter 1 builds this sentence piece by piece.
Why does the model forget me when I open a new chat?
Because an LLM is stateless. Its knowledge is a frozen file, and a conversation is only ever placed on its desk, the context window. When the chat ends the desk is cleared. Inside one chat it seems to remember only because the app resends the whole conversation with every message. Chapter 1 shows it fail in a real experiment.
What are the secretary, the librarian and the box?
The box is the memory store. The secretary is the extractor: it reads a chat once, at the end, and writes cards. The librarian is retrieval: before each answer it finds the few cards that match the question. The loop is recall, answer, remember. Chapter 3 introduces all three.
What is the difference between short-term and long-term memory?
Short-term memory is the context window: the turns of this chat, plus a rolling summary when the chat grows long. Long-term memory is the card box, which survives between chats. A summary may lose details. The box must not. Chapter 4 draws the map.
What are semantic, episodic and procedural memories?
They are the three colours of card. A fact (semantic) such as "Maya is vegetarian" stays until something replaces it. An event (episodic) such as "dentist on Tuesday" fades about 90 days after its last use. A rule (procedural) such as "answer short and direct" stays until the user edits it. The colour decides how long a card lives. Chapter 4 has an exercise for sorting them.
How many memories does a normal user have?
The class sizes the store at about 200 memories per user. A session produces about 3 candidate facts, but most of them are duplicates or updates of something already known, so the box grows slowly. Without consolidation the same user would hold 2 x 3 x 730 = 4,380 cards after two years, 22 times more. Chapter 12 has the arithmetic.
Memory, RAG and context windows
Is a 1,000,000 or 10,000,000 token window the end of memory?
No. A bigger window moves the day the history stops fitting. It does not shrink the bill, because every message pays for everything on the desk. One year of ordinary chat is 730 sessions of 1,500 tokens, which is 1,095,000 tokens. At $3 per million input tokens a full 1,000,000 token prompt costs $3 for a single message. Memory stays at about 75 tokens. Chapter 2 tests the bigger window with numbers.
How is memory different from RAG?
RAG, short for retrieval-augmented generation, searches documents that somebody wrote, such as a handbook, and mostly only reads. Memory searches facts that the system itself wrote down about one user, and those facts change when the user's life changes. So memory must also write, update, forget and respect a privacy wall per user. The search techniques are shared. The write path is what makes memory its own subject. Chapter 1 introduces the difference.
Why not fine-tune a small model for each user?
Because a model's numbers are not a database. We cannot read one fact out of them, correct it, or delete it when the user asks. Training is also slow and costly, and a million users would mean a million models. Chapter 2 calls this the fix where everything breaks.
Does the model train on my memories?
No. Memories are text placed in the prompt at the moment of the question. Nothing is written into the model's frozen file. Whether a company later uses conversations to train a future model is a separate policy question, and it is not how memory works. Chapter 1 explains the frozen file and the desk.
If prompt caching makes repeated tokens cheap, is resending history fine?
Prompt caching is a discount that a provider gives when the start of a prompt, the prefix, is the same as in the previous call. It lowers the price of a prefix that repeats exactly. It does not make the window larger, and it does not help the model read a long history well. The history still grows every session while memory stays flat. Caching is a good discount on the system prompt. It is not a memory. Chapter 2 has the comparison.
Can I use a rolling summary as my long-term memory?
For one session, yes. Across sessions, no. A summary is lossy: each compression pass may drop an exact detail, and a summary of summaries loses more. In the class experiment a 14-message chat went from 818 prompt tokens to 125, which is 6.5 times smaller, and the budget figure survived only by luck. Durable facts belong on cards. Chapter 4 shows the experiment.
Writing memories
Why not simply store everything the user says?
Because storage is not the problem. Retrieval is. Every extra card competes for the same few places on the desk, and stale or trivial cards push the right one out. The secretary keeps about 5 percent: 1,500 tokens of chat become about 75 tokens of facts. Chapter 6 explains what deserves a card.
What if the extractor gets a fact wrong?
Then the error is invisible and lasting, because nobody reads the box and every later answer trusts it. Three defences help. Keep the source of every card, so a wrong card can be traced to its sentence. Keep the raw turns. Let the user see, edit and delete every card. Chapter 6 and Chapter 15 cover them.
How often should extraction run: every message or every session?
Once per session, after the reply is sent. That is about 4.6 writes per second in the class sizing, against about 28 reads per second. In our experiments one call over the whole session did as well as 10-turn windows, which wrote 38 percent more facts and took 2.6 times longer to ingest. Chapter 13 covers the queue.
Does a stronger extractor model give better memories?
Not in our measurements. A stronger extractor wrote fewer, more cautious facts and lost 8 points on the BEAM development set (65 to 57). Every change to the extraction prompt also lost. The extractor needs to be small, disciplined about JSON, and run at temperature 0. Chapter 14 explains the three model roles.
What does it mean that a card must stand on its own?
A card is read months later, with no conversation around it. So it must carry real names and real dates: "Maya moved from London to Paris in March 2026", not "she moved last week". A card that says "she" or "last week" is useless the day after it is written. Chapter 6 has examples.
Should one card hold one fact or several?
One attribute per card is the safe rule when updates must be exact. Our golden set found the reason: "I live in London and I'm vegetarian" became one card, and the later move to Paris closed the vegetarian part with it. Atomic facts fixed the bug but grew the store by 37 percent and cost 2 to 3 benchmark points, so in hanumemAI it is an option you switch on. Chapter 11 tells the story.
Why is the write path asynchronous?
Because nobody should wait for the secretary. In the class experiment the user waited 3,351 ms when extraction ran first and 1,393 ms when it ran in the background. So 1,958 of 3,351 ms, or 58 percent of the wait, disappeared. The price is freshness: a new fact is readable a few seconds later. Questions in the same chat are answered from the context window, so nothing is lost. Chapter 13 shows both timelines.
Does it work when the user writes in another language, say French?
We have not measured it. All three benchmarks are in English, so every score in this book is an English score. The models and the embeddings are multilingual, so extraction and search by meaning should carry over. Two parts are English only today: the parser for time phrases such as "last week", and the phrases that mean "do not remember this". Test with your own golden set before you rely on it.
How are sarcasm, hypotheticals and retractions kept out?
By the extractor's judgement, and that is a weak point. "I could murder a pizza" and "if I moved to Rome" should produce no card, and a good extractor usually writes none, but nothing guarantees it. A retraction with no replacement, such as "forget what I said about the marathon", has no card to supersede the old one. The app should offer the user a way to close a card directly. We did not measure any of these cases. Put them in your golden set (Chapter 17).
Finding memories
What is an embedding?
It is a list of numbers that places a sentence on a map of meaning. Sentences with similar meaning land close together. Cosine similarity measures the angle between two such arrows: 1 means the same direction, 0 means unrelated. The class uses 768 numbers per sentence, which is 3,072 bytes. Chapter 7 has a map you can drag.
Why only the top 3? Is that a law?
It is a dial. Three cards of about 25 tokens give the flat 75 token budget that makes the cost argument. In our measurements recall rose with k, from 83.0 percent at k = 15 to 93.3 percent at k = 100. But each extra point cost more tokens. Going from k = 30 to k = 50 lowered accuracy slightly, by 0.26 points. hanumemAI uses k = 30 inside a token budget. Chapter 8 has the curve.
Why is similarity alone not enough?
Similarity finds the topic. It cannot tell a birthday from printer ink. The class adds recency and importance, each normalised to the range 0 to 1, so that an old but important card is rescued and fresh trivia is pushed down. Chapter 9 works one row out by hand.
What are dense search, BM25 and the entity boost?
They are three ways to search the same box. Dense search matches meaning. BM25 matches exact words and gives more weight to rare ones. The entity boost matches names, so a question about Sam finds the cards that mention Sam. Our best measured mix was 0.8 dense, 0.1 BM25 and 0.1 entity, with 89.05 percent evidence recall. Chapter 9 has the sweep.
Should I use a vector store or a graph?
They answer different questions. Vectors find what is similar. A graph with typed edges follows how things relate and how they changed. Graph systems pay for that with several model calls per write. hanumemAI takes a middle road: plain rows in SQLite, with a pointer from each replaced card to its replacement, and one hop along that pointer at read time. Chapter 19 puts the designs side by side.
Should I add a reranker or rewrite the query with an LLM?
Measure before you add one. We tried a rewrite in the style called HyDE: a model first writes what the answer might look like, and the search uses that text too. It cost one extra model call and about a second per search, and gave no gain. We have not measured a reranker yet. Once evidence recall reached about 88 percent, retrieval changes stopped moving the score. The read path in hanumemAI uses no LLM at all. Chapter 21 lists what never helped.
Why cap the read in tokens and not in rows?
Because rows differ in length. On BEAM, 30 rows of long assistant messages were 12,380 tokens per question. A cap of 4,000 tokens cut that to 4,227, which is 66 percent less, and the score rose from 46 to 49. A row count hides the real cost. A token budget does not. Chapter 8 explains it.
How fast must recall be?
Under 50 ms. Embedding the question takes 10 to 30 ms and the filtered search 1 to 10 ms. The model needs 200 to 500 ms before its first word, so recall is invisible next to it. Chapter 3 walks one request millisecond by millisecond.
How do I guarantee that a safety-critical fact, such as an allergy, is always on the desk?
Search alone cannot promise it. In Chapter 8 the allergy card
misses the top 3 for a dinner question. Three things help. Use a larger k inside a token budget: hanumemAI
defaults to k = 30. Keep the running profile on, because it is placed in every prompt, and an allergy
belongs in it. And write a golden-set check that fails when the allergy is missing for a food question.
hanumemAI also has an importance weight, w_importance, but it defaults to 0 and we never
measured it.
Time, change and forgetting
What if two facts conflict?
First ask whether one replaces the other. "Lives in Paris" replaces "lives in London", so the old card is closed with a date and pointed at the new one. If both could be true, both stay. If the conflict is real and unresolved, the honest answer names both statements with their dates and asks which is correct. That answer format took BEAM's contradiction questions from 20 to 80 percent on the development set. Chapter 10 and Chapter 11 cover it.
Why not overwrite the old fact?
Because overwriting answers "where do I live?" and destroys "where did I use to live?". The benchmarks have whole categories for the second kind of question. Superseding keeps both: a new row, the old row closed, and a pointer between them. Chapter 11 calls it git for facts.
What are the two clocks?
Valid time is when a fact was true in the world. Transaction time is when the system learned it. Maya
may tell us on 2 March 2026 that she moved on 1 March. With both clocks the box can answer "what was true on that
day" and "what did the assistant believe on that day". In hanumemAI these are as_of_event and
as_of. Chapter 11 has the code.
Is forgetting not the same as losing data?
No. Forgetting is precision. With five address cards in an append-only box, two of the top three places went to stale cards, a precision at 3 of 1/3. Four tools keep the top three correct, in this order: dedup, supersede, decay, erase. Chapter 12 has the simulation.
What expires, and what never does?
Events expire about 90 days after their last use. Facts never expire, they are replaced when life changes. Rules never expire, the user edits them. In hanumemAI an expired event is closed, not erased, so it remains reachable through a point-in-time query. Chapter 12 has the policy table.
How does recency decay work?
Recency is 0.995 to the power of the hours since the card was last used. After 2,000 hours, about 83 days, the score is 0.000044. That is why recency alone would bury a partner's birthday, and why importance and relevance must be added to it. Chapter 9 has the arithmetic and Chapter 12 has the curve.
Privacy and safety
How are users kept apart?
By a user_id filter on every query. In a shared collection that filter is the whole privacy
boundary. If it is forgotten nothing crashes and no error is logged: the search simply returns another
person's life. So make it structural. hanumemAI can route each user to one of N SQLite files by a hash of the
user id. Chapter 15 shows the leak.
What must never be stored?
Financial details, credentials, and anything the user asked not to remember. The audit trail names the kind of information, never the secret itself. A policy can also refuse too much: our first patterns matched bare words such as "token" and "salary", silently refused ordinary facts, and cost 4 points on a LongMemEval sample. Chapter 15 has the details.
If I delete a chat, are its memories deleted?
Not by that act alone. The cards extracted from a chat live in the box, separate from the chat. That is why a product must let the user see every memory and delete each one. Chapter 15 lists the four user controls.
What is memory poisoning?
It is a prompt injection that gets saved. Hidden text in a web page tells the assistant to store an instruction, and from then on the instruction rides into every session. The defences are to extract only from the user's own words, keep a source on every card, show new memories to the user, and audit the writes. Chapter 15 walks through a real case.
Are the vectors personal data too?
Yes. Text can be reconstructed from an embedding, so erasure must reach the vector store, the caches and the logs. In the class experiment one filtered delete removed all of one user's rows in 4 ms and left the other user untouched. Chapter 15 covers the right to be forgotten.
What about memories of other people, such as Sam and the children?
Maya's box holds facts about people who never agreed to it. hanumemAI does nothing special here: a card
about Sam belongs to Maya's user_id and is erased only with her box or card by card. This is
general practice, not something we measured: store facts about third parties only when they serve the
user's own requests, never share them between users, and check the privacy law that applies to you. Facts
about children deserve extra care.
Who can read the box? Is it encrypted?
hanumemAI does not encrypt anything. It writes ordinary SQLite files, and anyone who can read the files can read the memories. Encryption of the disk, access rules for staff and an audit trail of who opened what are the job of the application and its hosting. Remember that vectors are personal data too (Chapter 15).
Can the user see and correct the running profile?
Not through a finished call yet. The profile is a paragraph of at most 120 words, rewritten by a model after each session from the facts that were just stored. Facts refused by the write policy never reach it. The store can read and replace it, but hanumemAI has no user-facing call for it today, and a wrong sentence in it is repeated in every prompt until the next rewrite. An application should show the profile next to the cards.
Cost and scale
What does the memory layer cost per day?
At the class baseline, about $312 a day with a small assistant model and about $825 with a frontier one. Resending ten sessions of history would cost $5,400 or $108,000. Memory is 312 / 5,400 = 5.8 percent or 825 / 108,000 = 0.76 percent of that. Chapter 16 derives every line.
Where does the money go?
With a small assistant, mostly to the write path: $126 for extraction and $127.80 for update decisions, out of $312. Embeddings are $1.56 and the store about $30. With a frontier assistant the injected tokens become the largest line, at $540. Chapter 16 has the table.
How big is the store?
One memory is about 3.3 KB, and 92 percent of it is the vector (3,072 of 3,322 bytes). Two hundred memories for each of 1,000,000 users is 200,000,000 memories, about 664 GB. Storing each number as one byte brings a full copy to 204 GB, and three replicas to 612 GB. Chapter 5 and Chapter 16 derive it.
Does traffic follow registered users or active users?
Traffic follows active users: 200,000 of them produce about 28 reads and 4.6 writes per second. Storage follows registered users, because a user who left still has a box. Chapter 5 teaches the derivation.
Can I change the embedding model later?
You can, but it costs more than it looks. Every stored vector must be computed again. In hanumemAI the extractor also sees the nearest existing memories, so a new embedder changes every extraction prompt, and the whole ingest is paid again. That is why the class calls the embedder married to the store. Chapter 14 has our measurement.
What is the daily bill at hanumemAI's real budget of about 1,470 tokens, not the class's 75?
Only the injected-tokens line changes. It grows about twenty times, and memory still costs about one tenth of resending history.
messages per day 2,400,000 injected tokens 2,400,000 x 1,470 = 3,528,000,000 a day small assistant, $0.15 per M 3,528 x $0.15 = $529.20 frontier assistant, $3 per M 3,528 x $3 = $10,584 resending 10 sessions $5,400 or $108,000 share 529.20 / 5,400 = 9.8% 10,584 / 108,000 = 9.8%
The other lines of the bill in Chapter 16 stay as they are.
How do I change the embedding model on a live store?
Every vector must be made again, because vectors from two models cannot be compared. hanumemAI has no tool for this yet. The general practice, which we did not measure, is to build the new vectors next to the old ones in the background, switch a user over when all of that user's vectors are ready, and delete the old ones afterwards. Only the text is embedded again. No extraction call is repeated.
How do I fill the box from chat history I already have?
Replay the old sessions through add() in the order they happened, each with its own
observed_at date. Order matters, because a later fact must close an earlier one. At the class
prices one session costs $0.000315 to extract, so a year of one user's history costs about 23 cents.
one year of sessions 730 extraction per session $0.000315 one user, one year 730 x $0.000315 = about $0.23
Measuring
How do I test a memory system?
In two layers. Layer 1 asks whether the memory that must be retrieved reached the prompt. It is deterministic and costs nothing to run. Layer 2 replays session transcripts and checks that the write path did the expected operations. Run both after every change to a prompt, a model or a weight. Chapter 17 shows how to build the golden sets.
What should I watch in production?
Eight dials. Four describe the box: the mix of write operations, memories per user, the injection rate and the contradiction rate. Two are clocks: extraction latency and read latency. Two are about people: how often users correct the assistant about themselves, and whether deletions happen on time. Memory fails quietly, so the dials are how you hear it. Chapter 17 explains each one.
What do LoCoMo, LongMemEval and BEAM measure?
LoCoMo asks whether memory survives a long conversation. LongMemEval asks whether memory stays correct when facts change. BEAM asks whether memory holds at extreme length, up to a million tokens and beyond. Chapter 18 describes each exam.
Can I compare two published scores?
Only when both name their answer model, their judge, their question count and their context tokens per question. The judge alone moved identical LoCoMo answers by about 5 points, from 88.5 to 94.0. On LongMemEval a 100-question sample scored 91 where the full 500 scored 81.0. A score without those details is a claim, not a result. Chapter 18 shows how to read one.
My change gained one point. Is it real?
Probably not yet. The same configuration, answered again, scored 93.51 and then 92.47, so run-to-run noise is about one point even at temperature 0. Gains under 2 points need a second set of questions that agrees. Chapter 17 and Chapter 21 explain the noise bands.
How do I measure answer quality in production, where no gold answers exist?
With signals, not scores. Count how often a user corrects the assistant about themselves, edits or deletes a memory, or gives a thumbs down to a personalised answer. Each week, read the worst cases and name the layer that failed: extraction, retrieval, consolidation or reading. Then add the case to the golden set, so that it has a right answer from now on (Chapter 17).
Choosing and building
What should I build first?
The store with its user_id filter, then the read path, then the write path, then the golden
set. Start with a handful of hand-written cards so that you can see recall work before any extractor
exists. Add consolidation and decay once the box has real traffic.
Chapter 22 builds the loop end to end.
When should I not add memory?
When every session stands alone, as in a one-off translation tool. When the facts you need already live in your own database, such as an order history, which you should query directly. When you cannot offer the user a way to see and delete what is stored. Memory adds a write path, a privacy duty and a new way to be wrong. Add it when users come back and repeat themselves. Chapter 5 helps you scope it.
Build my own, or use a library?
Build the small version once, to understand it. Then choose by what you must own. Mem0 gives two function calls and leaves contradictions for you to resolve. HydraDB resolves them at write time with a heavier engine. hanumemAI sits between them in one SQLite file. Chapter 19 compares what you get and what you inherit.
Can memory be shared between agents, or across a team?
The class places shared memory out of scope, and the book follows it. Three things change when you add
it. The privacy wall needs a second key, such as a team or a project, and rules for who may read what.
Provenance must record which agent or person wrote each card. Conflicts become common, because two writers
can disagree. In hanumemAI a composite id such as user:project gives separate boxes, which is
scoping, not sharing. Chapter 5 lists the
scope.
Do I need a profile, a collection, or both?
Both. A profile is one short document per user that is always injected. A collection is many discrete cards from which the top few are retrieved. In our experiments a running profile of about 120 words on top of the collection was worth 2.3 points, and a 200-word profile lost 2.6. Chapter 7 has the comparison.
Which model should answer the questions?
Start small and measure. When the right memories are on the desk and the answers are still wrong, a stronger reader is the next step. On LongMemEval the same memories scored 81.0 with a small reader and 88.6 with a stronger one. Chapter 14 explains when to change the reader.
hanumemAI
What is hanumemAI?
It is the memory library built beside this book: one Python package, one SQLite file, and any
OpenAI-compatible model. One extractor call per write, no model call per read, two clocks on every row. It
imports as hmem until the rename ships.
Chapter 20 follows one fact through it.
Is hanumemAI state of the art?
On one benchmark it is ahead, on one it is level, and on one it is behind. On LoCoMo it scores 94.2 with a small reader and 96.0 with a stronger one, under the Mem0 paper's judge, at about 1,470 tokens per question. The Mem0 platform reports 92.5 at about 7,000 tokens, with other models. On BEAM at one million tokens it scores 64.0 with the stronger reader, level with Mem0's self-reported 64.1, at about half the tokens (3,494 against about 6,700). On LongMemEval it is behind: 88.6 against a self-reported 94.4 for the Mem0 platform and a reported 95.6 for Agent Zero. The other systems used other models and judges, so none of this is a ranking. Chapter 18 has every number with its models.
What does hanumemAI not do yet?
Four things. The background writer is a single worker inside one process, and a worker that serves many
processes from a shared queue is not built. Raw turns are saved even when the fact policy refuses a
secret. Switching store_episodes off stops them from being searched, but their text is still
kept in a side table. So remove secrets before the text reaches the library
(Chapter 22). The largest single index we measured held
300,000 vectors, and the largest store held one million memories for 5,000 users, both on
one laptop. Larger figures are arithmetic. The recency term could not be tested offline, because every
row is created and touched at the same moment. Chapter 16 and
Chapter 15 say so where it matters.
How was hanumemAI tuned?
By a loop: one idea, one run, keep or discard, one line in the log. The log holds 121 runs: 16 kept, 29 discarded and 76 informational, for about $25 in total. The answer prompt and the running profile gave most of the gain. Every change to the extraction prompt lost. Chapter 21 tells the story, mistakes included.
Can hanumemAI run without the internet?
Yes. A local embedding model runs on the processor of the laptop, and the extractor can point at any local server that speaks the OpenAI protocol. The tests run with mocks only. Chapter 20 shows the settings.
Where do I go from here?
Build the small assistant in Chapter 22. Move every slider in Chapter 23. Then close the book and try Chapter 25 from memory.