Part III · Journeyman
Chapter 17
Monitoring and evaluation
Eight dials to watch, two layers of tests, and a routine that turns every wrong answer into the name of the part that failed.
On this page
In this chapter, we will learn how to know whether a memory system is working. We will choose the numbers to watch in production and build a small exam called a golden set. We will learn a routine that takes any wrong answer and names the part that caused it. We will also learn how much a score can move by luck, so that we do not celebrate noise.
Why memory fails quietly
When a web page breaks, we see an error. When memory breaks, we see a perfectly polite answer that happens to ignore what the user said last month. Nothing crashes. The secretary's mistakes are invisible and they are permanent, because a wrong card stays in the box until someone finds it.
So we cannot wait for errors. We have to go looking.
The dashboard: eight metrics
A metric is a number we record continuously, so that a change in it tells us something changed in the system. The class lists eight for a memory layer.
| Metric | What it counts | What a change warns about |
|---|---|---|
| Write operation mix | the share of ADD, UPDATE, DELETE and NOOP, the four operations of Chapter 10 | all ADD means nothing is being merged; all NOOP means the extractor has stopped finding anything |
| Memories per user | how the store grows for each user | steady growth past the planned 200 means duplicates or no forgetting |
| Injection rate | memories used per answer | near zero means retrieval finds nothing; always at the maximum means we are stuffing the prompt |
| Contradiction rate | pairs of open memories that disagree | the update step is failing: London and Paris are both current |
| Extraction latency | time from session end to memory stored | the queue is falling behind, so this morning's fact is missing this afternoon |
| Read-path latency | embed plus search, per message | above 50 ms the read path is over the budget set in Chapter 5 |
| User corrections | "no, I moved", edits, deleted memories | the most honest signal we have: the user is telling us a memory is wrong |
| Timely deletion | time from an erase request to erased | a broken promise, and often a broken law |
A healthy system has a shape. Here is what the write mix should roughly look like for a user we already know well.
heavy user, 2 sessions/day x 3 candidates = 6 candidates per day two years 6 x 730 = 4,380 candidates planned store 200 memories absorbed by NOOP and UPDATE 1 - 200/4,380 = about 95% of candidates
This quick estimate ignores that events expire. Chapter 12 counts the expiry and arrives at about 93 percent. Either way, more than nine in ten.
For a new user almost everything is an ADD. After a year, most candidates are things we already know. If the ADD share stays high for old users, the store is filling with duplicates and stale facts, the problem of Chapter 12.
The garage: two layers of evaluation
Evaluation is running the system on fixed inputs with known correct results. The class splits it into two layers, one for each path.
layer 1 did the right memory reach the prompt? READ PATH
golden set: (question, memory that MUST be retrieved)
no LLM needed: deterministic, fast, nearly free
layer 2 did extraction and consolidation do the right ops? WRITE PATH
golden set: (session transcript, expected ADD / UPDATE / NOOP list)
replayed on every extractor or prompt change
Layer 1: did the right card reach the desk?
We store a fixed set of memories, ask a fixed question, and check whether one particular memory is in the top results. Either it is or it is not. No model judges anything, so the test gives the same result every time and costs almost nothing.
| Question | Memory that must be retrieved |
|---|---|
| What should I cook for dinner tonight? | Maya is vegetarian. |
| Suggest a snack for the school trip. | Maya has a serious peanut allergy. |
| What should I get Sam as a birthday present? | Maya's partner Sam has a birthday on June 18. |
| Any good bakeries near me? | Maya lives in Paris. |
The widget below runs 8 such pairs over 5 questions. Here k is the number of memories placed in the prompt. With k = 3, 7 of the 8 pass: the dinner question finds the vegetarian card and misses the peanut allergy card, which is ranked fourth. Change k and the extra weight, and watch which pairs pass and what the tokens cost.
Layer 2: did the secretary do the right thing?
Here we replay a whole session through the real write path and check what happened to the store.
| Session says | Expected operation | Check |
|---|---|---|
| "I moved from London to Paris" | UPDATE, or a new row that closes London | no current fact contains "London" |
| "I'm vegetarian" (already stored) | NOOP | the number of facts did not grow |
| "I adopted a cat named Miso" | ADD | a current fact contains "Miso" |
| "Mondays should be illegal" | nothing | no fact contains "Monday" |
| "Our wifi password is sunflower42" | refused | no fact contains "sunflower42" |
This layer calls the extractor, so it costs a little. Our 23-check set takes about ten model calls and about one cent. It runs every time the extraction prompt, the extractor model or the policy changes.
How to build a golden set
- Start with 20 pairs, by hand. Use the questions your users really ask. Twenty is enough to catch most regressions. A regression is something that used to work and no longer does.
- Add every real failure. When a user corrects the assistant, turn that moment into a new pair. The set grows exactly where the system is weak.
- Cover each kind of difficulty. A changed fact, a question about the past, a date ("last March"), a rare name, a fact that needs two memories, and a question with no answer at all.
- Include things that must not happen. Smalltalk that must not be stored, a secret that must be refused, another user's fact that must not appear.
- Freeze it. Changing the exam and the student in the same week tells us nothing.
Classifying a wrong answer
Every wrong answer has an address. When an answer is wrong, we do three lookups and give it one of five labels. We used this routine after every benchmark run.
| Label | What happened | How we see it | Where to fix |
|---|---|---|---|
| NOT STORED | the fact never became a memory | search the store: it is not there | extraction |
| NOT RETRIEVED | it is stored, but was not in the top k | evidence recall is 0 for this question | retrieval: signals, weights, k |
| IGNORED | it was on the desk and the answer missed it | read the exact prompt the model saw | answer prompt, too much context, or the answer model |
| WRONG DATE | right fact, wrong time arithmetic | the answer names the wrong day or month | date handling |
| JUDGE | the answer is right and the grader marked it wrong | read the grader's reason | note it, do not chase it |
Sample twenty wrong answers, label each, and count. The biggest pile is the next piece of work. "Memory feels off" starts an argument. "Twelve of twenty are IGNORED" starts a fix.
Evidence recall: a free metric
Evidence recall is the share of the needed source messages that appear among the sources of the retrieved memories. Benchmarks mark which turns hold the answer. Every hanumemAI memory remembers the turns it came from. Comparing the two lists needs no model call, so it costs nothing.
LoCoMo, 385 questions, free sweep over k (no answers generated, $0) k = 15 evidence recall 83.0% 875 tokens per question k = 30 evidence recall 87.6% 1,489 tokens per question k = 50 evidence recall 90.5% 2,324 tokens per question k = 100 evidence recall 93.3% 4,431 tokens per question
More recall is not always a better answer. Experiment 0008 answered the questions at k = 30 and at k = 50. It ran after a change that had shortened the prompt, so its token counts are lower than in the sweep.
k = 30 k = 50 evidence recall 87.3 90.2 90.2 - 87.3 = +2.9 points tokens 1,250 2,092 2,092 / 1,250 = 1.67, so 67% more score 91.17 90.91 90.91 - 91.17 = -0.26 points
A small reader is distracted by a crowded desk. The free metric tells us what reached the desk. Only an answered exam tells us what was used.
The improvement loop
users chat | v signals: corrections, memory edits, deletions, thumbs on personalized answers | v triage: the worst memories, every week | v which layer failed? |-- bad extraction --> fix the extractor prompt |-- bad retrieval --> fix scoring or k |-- bad update --> fix consolidation | v replay the golden sets -- no regression --> ship --> users chat
The step that people skip is the replay. A fix for one failure often breaks something that used to work. The golden sets are how we find out before the users do.
Noise: how much can a score move by luck?
Noise is the amount a score changes when nothing in the system changed.
We ran the same configuration twice on 385 LoCoMo questions, with fresh model calls, at temperature 0. Temperature 0 is the setting that asks a model for its most likely answer every time.
run 1 93.51 run 2, same code 92.47 difference 93.51 - 92.47 = 1.04 points answers that changed 63 of 385 = 16%
Sixty-three answers came back different from an unchanged system. Hosted models are not perfectly repeatable. So a gain of one point is not evidence of anything. Our rules became two. Keep a change that gains 2 points or more. Test any gain between 1 and 2 on a second set of questions, a confirmation split, where it must point the same way.
Small test sets are worse. With 50 questions, one question is 2 points. A rule that looked like +5 on 100 questions lost 5 on another set (A0002). And a 100-question sample of LongMemEval scored 91 while the full 500 scored 81 with the same code, because the sample held fewer of the hardest questions.
The public benchmarks
Our own golden set tells us about our users. Public benchmarks let us compare with other systems. There are three, and Chapter 18 is about them.
| Benchmark | The question it asks |
|---|---|
| LoCoMo | Does memory survive a long conversation? |
| LongMemEval | Does memory stay correct when facts change? |
| BEAM | Does memory hold at extreme length? |
Try it yourself
Run the layer-2 golden set. It makes about ten model calls and costs about a cent.
# from the root of the hanumem research repository
export OPENROUTER_API_KEY=sk-or-...
uv run python scripts/golden.py # prints what was stored, then PASS or FAIL per check
The script turns atomic_facts on, so it should end with 23 of 23.
Then run a layer-1 check of your own. It needs a store with a few memories in it. The quickstart of
Chapter 20 creates quickstart.sqlite.
from hmem import Memory, Config
m = Memory(Config(db_path="quickstart.sqlite"))
golden = [
("Any good bakeries near me?", "Paris"),
("What should I get Sam for his birthday?", "June 18"),
]
passed = 0
for question, must_contain in golden:
rows = m.search(question, user_id="maya", k=5)
ok = any(must_contain in r.text for r in rows)
passed += ok
print("PASS" if ok else "FAIL", question)
print(f"layer 1: {passed} of {len(golden)}")
Add a pair that fails. Then change k until it passes, and check what that did to the
number of tokens.
Common questions
How big should a golden set be?
Start with 20 pairs and grow to 50 or 100. Coverage matters more than size. A set of 100 easy questions about stable facts will pass forever and teach nothing. One changed fact, one date question and one secret are worth more.
Why not judge everything with an LLM?
An LLM judge costs money, takes time and is itself noisy. The same LoCoMo answers scored 88.5 under one judge model and 94.0 under another. Layer 1 needs no judge at all, so we run it on every change and keep the judged exams for the changes that pass.
How do I measure the contradiction rate?
Sample users, take their open memories, and look for pairs about the same attribute with different values: two home cities, two employers. A simple version groups memories by entity and asks a small model whether any two disagree. The trend matters more than the exact figure.
Users rarely correct the assistant. Is that signal useful?
It is rare and precise. One "no, I moved last year" names a wrong memory, the question that exposed it and the right value. Turn each one into a golden pair. Deleted memories and edits are the same signal from the settings page.
My score went up 1.5 points. Do I keep the change?
Not yet. Run it on a second set of questions. If it gains there too, keep it. If it loses, the first result was luck. Also look at the cost: a gain that doubles the tokens per question is usually not worth keeping.
What is a good injection rate?
It depends on the product, so watch the change rather than the level. The class design injects the top 3 for every message. If the average number of memories that pass a relevance threshold drops to zero after a release, retrieval broke.
Carry this
- Memory fails quietly. Watch eight dials: write mix, memories per user, injection rate, contradictions, extraction latency, read latency, user corrections, timely deletion.
- Layer 1 asks whether the right memory reached the prompt. Layer 2 asks whether the write path did the right operations. Replay both after every change.
- Every wrong answer gets one label: NOT STORED, NOT RETRIEVED, IGNORED, WRONG DATE or JUDGE. Fix the biggest pile.
- A 23-check golden set found a real bug that three benchmarks missed.
- The same system scored 93.51 and 92.47. Gains under 2 points need a second set of questions.
Check yourself
1. The assistant recommends a steakhouse to Maya. You find "Maya is vegetarian." in the store, and it was in the prompt the model saw. Which label, and where do you look?
Answer
IGNORED. The memory was stored and retrieved, so extraction and retrieval are fine. Look at the answer prompt, at how many other memories were crowding the desk, and at the answer model.
2. Your test set has 50 questions. A change moves the score from 78 to 82. How many questions changed, and is that enough?
Answer
one question 100 / 50 = 2 points the gain (82 - 78) / 2 = 2 questions
Two questions is inside the noise. Confirm on a second sample before keeping it.
3. After a release, the share of ADD operations for users older than one year rises from 5 percent to 60 percent. What probably broke?
Answer
Deduplication or the update decision. Known facts are being written again as new rows instead of being absorbed as NOOP or UPDATE. Memories per user and the contradiction rate will rise next.