Part IV · Ninja
Chapter 18
Benchmarks are claims
Three public exams for memory, how a machine grades them, and why a score means nothing until it names its models, its judge and its token cost.
On this page
- Why we need an exam at all
- LoCoMo: does memory survive a long conversation?
- LongMemEval: does memory stay correct when facts change?
- BEAM: does memory hold at extreme length?
- How a machine grades an answer
- The five dials behind every score
- What an honest score looks like
- hanumemAI's results
- Next to other systems
- Try it yourself: read a score like an examiner
- Common questions
- Carry this
- Check yourself
In this chapter, we will learn how memory systems are measured. We will meet the three public exams and see how a second AI model grades the answers. We will learn why the same system can score 88 or 94 without changing a single line of code. Then we will read hanumemAI's own results, including the places where other systems score higher.
Why we need an exam at all
A benchmark is a fixed set of conversations and questions with known correct answers. Because the set is fixed, two systems can be compared on the same work, and one system can be compared with itself before and after a change.
Chapter 17 built our own small golden sets. Those tell us whether our system works for our users. Public benchmarks answer a different question: how does this design compare with everybody else's? Three of them matter, and each one tests a different weakness.
LoCoMo: does memory survive a long conversation?
LoCoMo is a set of ten very long conversations between two people, with about 1,540 questions that are normally scored. It comes from the paper "Evaluating Very Long-Term Conversational Memory of LLM Agents" (2024).
Each conversation has about 300 turns spread over up to 35 sessions. Every session has a date. The questions fall into five categories.
| Category | What it asks | Example shape |
|---|---|---|
| Single-hop | The answer is in one turn | "What is the name of Caroline's dog?" |
| Multi-hop | Combine two or more turns | "Which of her hobbies began after the move?" |
| Temporal | Dates and durations | "How long after X did Y happen?" |
| Open-domain | The conversation plus common sense | "Would she enjoy a camping trip?" |
| Adversarial | The question's premise is false | Most leaderboards leave this one out |
Almost everyone reports the first four categories. In the paper, humans scored about 88 and GPT-4 Turbo with the whole conversation on its desk scored about 52. Temporal questions were the biggest gap.
LongMemEval: does memory stay correct when facts change?
LongMemEval is a set of 500 questions, each hidden inside its own haystack, the paper's word for a pile of chat sessions. It comes from the paper "Benchmarking Chat Assistants on Long-Term Interactive Memory" (2024). We use the S version: about 115,000 tokens and about 40 sessions for every question.
| Question type | What it tests |
|---|---|
| Single-session, user | Recall something the user said |
| Single-session, assistant | Recall something the assistant said |
| Single-session, preference | Apply a preference the user stated |
| Multi-session | Combine facts from different sessions |
| Knowledge update | The fact changed. Use the newest version |
| Temporal reasoning | Dates and durations |
| Abstention | The answer is not in the history. Say so |
This is the exam for Chapter 11. A system that overwrites old facts, or that cannot tell which of two facts is newer, loses the knowledge-update questions.
BEAM: does memory hold at extreme length?
BEAM is a set of 100 conversations at four sizes: about 128,000 tokens, 500,000 tokens, 1 million tokens and 10 million tokens. It comes from the paper "Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs" (ICLR 2026). It has 2,000 questions over ten abilities. The dataset calls its smallest size the 100K split, so our tables say BEAM-100K.
information extraction multi-session reasoning knowledge update temporal reasoning summarization preference following abstention contradiction resolution event ordering instruction following
BEAM is graded differently. Every question has a rubric, a short list of points that a good answer must contain. The judge checks the points one at a time and gives each one 0, 0.5 or 1. In our tables, a question counts as passed when its mean rubric score is 0.5 or more, and we print the mean rubric score next to the pass rate. In the paper, models that read the whole conversation with no memory system score about 0.30 to 0.36 on a scale of 0 to 1. That is 30 to 36 on our scale of 0 to 100.
How a machine grades an answer
An LLM judge is a second model that reads the question, the correct answer and the system's answer, and decides whether they agree.
We need a judge because exact matching is too strict. The known correct answer is called the gold answer. If the gold answer is "7 May 2023" and the system says "on May 7th, 2023", a string comparison marks it wrong. A judge marks it right.
question ─┐ gold ─┼─► judge model + judge prompt ─► CORRECT or WRONG answer ─┘
But the judge is a model too, with its own habits. One judge accepts a loose paraphrase. Another wants the exact detail. This brings us to the most important lesson of the chapter.
We repeated the double grading on our final run. The widget below shows that pair.
earlier run 94.0 - 88.5 = 5.5 points final run 94.22 - 87.73 = 6.49 points open-domain 80.21 - 55.21 = 25 points, the largest gap of any category
Look at the two judges side by side, category by category. Watch which category moves most.
The five dials behind every score
A benchmark score is produced by a whole pipeline, and only one part of that pipeline is the memory system. At least five dials move the number without touching memory at all.
| Dial | What it changes | What we measured |
|---|---|---|
| The answer model | Who reads the memories and writes the answer | LongMemEval, all 500 questions: 81.0 with qwen3.7-flash, 88.6 with qwen3.8-flash, same memories |
| The judge | Who grades | LoCoMo: 88.5 and 94.0 on identical answers |
| The context budget | How much memory is placed on the desk | Ours: about 1,470 tokens per LoCoMo question. Mem0 reports a top-200 budget, the 200 best-matching memories, of about 7,000 tokens |
| The question set | All questions, or a sample | LongMemEval: 91 on a 100-question sample, 81.0 on all 500 |
| Luck | The provider is not fully repeatable | The same configuration scored 93.51 and then 92.47 |
Samples flatter
Our 100-question sample of LongMemEval scored 91. The full 500 scored 81.0 with the same code. The sample happened to contain fewer questions of the two hardest types. In the full set, multi-session has 121 questions and we scored 65 percent on them. Temporal has 127 questions and we scored 72 percent. A score must say how many questions it covers.
Noise is about one point
This is the noise test of Chapter 17. We ran one configuration twice with fresh answers, and 63 of the 385 answers came back different.
run 1 93.51 run 2, same configuration 92.47 difference 93.51 - 92.47 = 1.04 points, from nothing at all
So on that set, any difference under one point is noise. Any gain under two points needs a second, separate set of questions to confirm it.
What an honest score looks like
A score is a claim until it names its answer model, its judge, its question count and its context tokens per question. With those four, a reader can decide what the number is worth.
claim "We score 92.5 on LoCoMo."
report LoCoMo, all 1,540 questions, categories 1 to 4
memory layer qwen3.7-flash extraction, qwen3-embedding-4b
answer model qwen3.7-flash
judge gpt-4o-mini
context about 1,470 tokens per question
score 94.2
hanumemAI's results
Here is every headline run. The memory layer is the same in all of them: qwen3.7-flash writes the memories and qwen3-embedding-4b makes the vectors. "Flash" means qwen3.7-flash. The other answer model, qwen3.8-flash, is a stronger and more expensive model.
| Benchmark | Questions | Answer model | Judge | Score | Context tokens per question |
|---|---|---|---|---|---|
| LoCoMo | 1,540 | flash | gpt-4o-mini | 94.2 | ~1,470 |
| LoCoMo | 1,540 | qwen3.8-flash | gpt-4o-mini | 96.0 | ~1,470 |
| LoCoMo | 1,540 | flash | flash | 87.7 | ~1,470 |
| LoCoMo | 1,540 | qwen3.8-flash | flash | 90.8 | ~1,470 |
| LongMemEval-S | 500 | flash | flash, paper's prompts | 81.0 | ~2,950 |
| LongMemEval-S | 500 | qwen3.8-flash | flash, paper's prompts | 88.6 | ~2,950 |
| BEAM-100K | 400 | flash | flash, rubric | 65.8 | ~3,760 |
| BEAM-100K | 400 | qwen3.8-flash | flash, rubric | 72.5 | ~3,760 |
| BEAM-500K | 700 | flash | flash, rubric | 61.3 | ~3,400 |
| BEAM-1M | 200 | flash | flash, rubric | 56.5 | ~3,500 |
| BEAM-1M | 200 | qwen3.8-flash | flash, rubric | 64.0 | ~3,500 |
The BEAM rows do not cover the same share of the benchmark. BEAM-100K is 20 conversations and BEAM-500K is 35. BEAM-1M is the first 10 conversations only. Each conversation has 20 questions.
The full LoCoMo run, by category, under both judges. Evidence recall is the free retrieval metric of Chapter 17.
| Judge | Overall | Multi-hop | Temporal | Open-domain | Single-hop | Evidence recall |
|---|---|---|---|---|---|---|
| qwen3.7-flash | 87.73 | 88.30 | 86.92 | 55.21 | 91.56 | 88.75 |
| gpt-4o-mini | 94.22 | 94.33 | 92.83 | 80.21 | 96.31 | 88.68 |
And LongMemEval, all 500 questions, by type:
| Answer model | Overall | Single-session user | Multi-session | Knowledge update | Temporal |
|---|---|---|---|---|---|
| qwen3.7-flash | 81.00 | 95.31 | 65.29 | 91.67 | 72.44 |
| qwen3.8-flash | 88.60 | 96.88 | 81.82 | 91.67 | 83.46 |
Next to other systems
These numbers come from the other systems' own papers and posts. We did not run them. They use other answer models, other judges and other budgets, so this is context and not a leaderboard. "Self-reported" marks a number that a vendor published about its own product.
| Benchmark | System | Score | Stack and status |
|---|---|---|---|
| LoCoMo | Mem0 paper (2025) | 66.9 | gpt-4o-mini, about 7,000 tokens, reported |
| LoCoMo | Mem0 platform (2026) | 92.5 | gpt-5, top-200, about 7,000 tokens, self-reported |
| LoCoMo | Agent Zero (2026) | 93.6 | reported |
| LoCoMo | hanumemAI | 94.2 / 96.0 | flash / qwen3.8-flash, gpt-4o-mini judge, about 1,470 tokens, measured by us |
| LongMemEval-S | Zep | 71.2 | reported |
| LongMemEval-S | Paper's best pipeline | ~75 | gpt-4o, reported |
| LongMemEval-S | hanumemAI | 81.0 / 88.6 | flash / qwen3.8-flash, about 2,950 tokens, measured by us |
| LongMemEval-S | HydraDB | 85.8 / 90.8 | gpt-4o-mini / Gemini 3 Pro, self-reported |
| LongMemEval-S | SodaMem | 92.8 | about 18,000 tokens, reported |
| LongMemEval-S | Mem0 platform (2026) | 94.4 | gpt-5, self-reported |
| LongMemEval-S | Agent Zero | 95.6 | reported |
| BEAM-1M | Mem0 platform (2026) | 64.1 | gpt-5 stack, about 6,700 tokens, self-reported |
| BEAM-1M | hanumemAI | 56.5 / 64.0 | flash / qwen3.8-flash, about 3,500 tokens, measured by us |
Where others score higher
On LongMemEval, hanumemAI is not the top number. Agent Zero reports 95.6, the Mem0 platform reports 94.4 with gpt-5, SodaMem reports 92.8 with about 18,000 tokens of context, and HydraDB reports 90.8 with Gemini 3 Pro. Our best is 88.6. Our evidence recall on that set is 98.1 percent, so the right memories do reach the desk. The misses are reading mistakes on multi-session and temporal questions. They are made by a small answer model in a single pass: one model call per question, with no second look.
Where hanumemAI leads
On LoCoMo, under the same judge model the Mem0 paper used, hanumemAI scores 94.2 with a flash answerer and 96.0 with qwen3.8-flash. It does so with about 1,470 tokens per question against about 7,000. The answer models differ, so even this is not a like-for-like race.
context per question hanumem 1,470 Mem0 top-200 about 7,000 ratio 7,000 / 1,470 = about 4.8 times less context
On BEAM at one million tokens, hanumemAI's 64.0 is level with Mem0's self-reported 64.1, with about half the context (3,500 tokens against 6,700).
The honest framing
hanumemAI's claim is this: the best score per context token, with small and cheap models. It is not "the highest number on every exam". Every point in the chart below is a score next to its cost. Choose a benchmark, hover a point to see its model stack, and read the table under the chart. Systems that publish no token count are in the table only.
Try it yourself: read a score like an examiner
Take any memory product's website. Find its benchmark number. Then try to fill in this card.
benchmark and version .................... questions scored ....... of ....... answer model .................... judge model and prompt .................... context tokens per question .................... who ran it the vendor / a third party per-category scores shown yes / no
Every line you cannot fill is a dial that may have been turned. If you have the hanumemAI research repository, you can also reproduce a row yourself. The free mode measures retrieval only and costs nothing.
scripts/bench.sh --free # evidence recall only, no answer or judge calls, $0
scripts/bench.sh # LoCoMo, 385 questions, a few cents with Qwen flash
scripts/bench.sh --full # all 1,540 questions
Common questions
Is a higher benchmark score always a better memory system?
No. A higher score may come from a stronger answer model, a kinder judge or five times more context. Ask for the score and its token cost together. A system that needs 18,000 tokens per question costs about twelve times more on every message than one that needs 1,470, because 18,000 / 1,470 = 12.2.
Why not grade with exact matching and avoid the judge problem?
Because correct answers come in many wordings. Exact matching would mark "May 7th" wrong when the gold answer says "7 May". A judge model is the practical choice. The fix is to keep the judge fixed, name it, and never change it in the middle of a series of experiments.
What is evidence recall, and why do you call it free?
The benchmark marks which turns contain the answer. Evidence recall is the share of those turns that our search returned. It needs no answer model and no judge, so it costs nothing to compute, and it measures the memory system alone. On LoCoMo ours is about 88.7 percent. On LongMemEval it is 98.1 percent.
Why does hanumemAI publish a score of 87.7 when it could publish only 94.2?
Both are true readings of the same answers under two judges. Publishing only the kinder one would turn a dial in our favour and hide it. The 94.2 is comparable with the Mem0 paper because it uses that paper's judge model. The 87.7 shows what a stricter judge says.
Should I choose a memory system by its benchmark rank?
Use the benchmark to shortlist, then test on your own data. Build a golden set of 50 to 100 questions from your real users, as in Chapter 17. Your users' questions are the only exam that matters in the end.
How can I tell whether a vendor's number is believable?
Look at the per-category scores. A system that keeps two clocks should win on temporal questions. A system with a strong profile should win on open-domain questions. When the gains match the mechanism, the number means something. When a system wins every category equally, check how the test was run.
Carry this
- Three exams, three weaknesses: LoCoMo for long conversations, LongMemEval for facts that change, BEAM for extreme length.
- A score is a claim until it names its answer model, its judge, its question count and its context tokens per question.
- The judge alone moved identical answers from 88.5 to 94.0. A sample scored 91 where the full set scored 81.0. Run-to-run noise is about one point.
- hanumemAI leads on score per token: LoCoMo 94.2 / 96.0 at about 1,470 tokens. On LongMemEval others report higher numbers than our 88.6.
Check yourself
1. System A reports 92 on LoCoMo with 7,000 context tokens per question. System B reports 90 with 1,500. The input price is $3 per million tokens and a product sends 2,400,000 messages a day. What does each system's context cost per day?
Answer
A 2,400,000 x 7,000 = 16,800,000,000 tokens x $3 / 1,000,000 = $50,400 per day B 2,400,000 x 1,500 = 3,600,000,000 tokens x $3 / 1,000,000 = $10,800 per day
Two points of score cost $39,600 a day. And two points is close to the noise band.
2. A change raises our score from 91.2 to 91.9 on 385 questions. Should we keep it?
Answer
Not on that evidence. The gain is 0.7 points and the measured noise is about 1 point. We keep it only if it also makes the system cheaper or simpler, or if a separate set of questions shows a clear gain.
3. Why did the 100-question LongMemEval sample score 91 while the full 500 scored 81.0?
Answer
The sample contained too few questions of the two hardest types. In the full set, multi-session (121 questions, 65 percent) and temporal (127 questions, 72 percent) make up about half of all questions and pull the average down.