Carebun Interactive/Memory Memory labCheat sheet

Part IV · Ninja

Chapter 18

Benchmarks are claims

Three public exams for memory, how a machine grades them, and why a score means nothing until it names its models, its judge and its token cost.

12 min read · interactive

On this page
  1. Why we need an exam at all
  2. LoCoMo: does memory survive a long conversation?
  3. LongMemEval: does memory stay correct when facts change?
  4. BEAM: does memory hold at extreme length?
  5. How a machine grades an answer
  6. The five dials behind every score
  7. What an honest score looks like
  8. hanumemAI's results
  9. Next to other systems
  10. Try it yourself: read a score like an examiner
  11. Common questions
  12. Carry this
  13. Check yourself

In this chapter, we will learn how memory systems are measured. We will meet the three public exams and see how a second AI model grades the answers. We will learn why the same system can score 88 or 94 without changing a single line of code. Then we will read hanumemAI's own results, including the places where other systems score higher.

Why we need an exam at all

A benchmark is a fixed set of conversations and questions with known correct answers. Because the set is fixed, two systems can be compared on the same work, and one system can be compared with itself before and after a change.

Chapter 17 built our own small golden sets. Those tell us whether our system works for our users. Public benchmarks answer a different question: how does this design compare with everybody else's? Three of them matter, and each one tests a different weakness.

LoCoMo: does memory survive a long conversation?

LoCoMo is a set of ten very long conversations between two people, with about 1,540 questions that are normally scored. It comes from the paper "Evaluating Very Long-Term Conversational Memory of LLM Agents" (2024).

Each conversation has about 300 turns spread over up to 35 sessions. Every session has a date. The questions fall into five categories.

CategoryWhat it asksExample shape
Single-hopThe answer is in one turn"What is the name of Caroline's dog?"
Multi-hopCombine two or more turns"Which of her hobbies began after the move?"
TemporalDates and durations"How long after X did Y happen?"
Open-domainThe conversation plus common sense"Would she enjoy a camping trip?"
AdversarialThe question's premise is falseMost leaderboards leave this one out

Almost everyone reports the first four categories. In the paper, humans scored about 88 and GPT-4 Turbo with the whole conversation on its desk scored about 52. Temporal questions were the biggest gap.

LongMemEval: does memory stay correct when facts change?

LongMemEval is a set of 500 questions, each hidden inside its own haystack, the paper's word for a pile of chat sessions. It comes from the paper "Benchmarking Chat Assistants on Long-Term Interactive Memory" (2024). We use the S version: about 115,000 tokens and about 40 sessions for every question.

Question typeWhat it tests
Single-session, userRecall something the user said
Single-session, assistantRecall something the assistant said
Single-session, preferenceApply a preference the user stated
Multi-sessionCombine facts from different sessions
Knowledge updateThe fact changed. Use the newest version
Temporal reasoningDates and durations
AbstentionThe answer is not in the history. Say so

This is the exam for Chapter 11. A system that overwrites old facts, or that cannot tell which of two facts is newer, loses the knowledge-update questions.

BEAM: does memory hold at extreme length?

BEAM is a set of 100 conversations at four sizes: about 128,000 tokens, 500,000 tokens, 1 million tokens and 10 million tokens. It comes from the paper "Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs" (ICLR 2026). It has 2,000 questions over ten abilities. The dataset calls its smallest size the 100K split, so our tables say BEAM-100K.

information extraction     multi-session reasoning     knowledge update
temporal reasoning         summarization               preference following
abstention                 contradiction resolution    event ordering
instruction following

BEAM is graded differently. Every question has a rubric, a short list of points that a good answer must contain. The judge checks the points one at a time and gives each one 0, 0.5 or 1. In our tables, a question counts as passed when its mean rubric score is 0.5 or more, and we print the mean rubric score next to the pass rate. In the paper, models that read the whole conversation with no memory system score about 0.30 to 0.36 on a scale of 0 to 1. That is 30 to 36 on our scale of 0 to 100.

How a machine grades an answer

An LLM judge is a second model that reads the question, the correct answer and the system's answer, and decides whether they agree.

We need a judge because exact matching is too strict. The known correct answer is called the gold answer. If the gold answer is "7 May 2023" and the system says "on May 7th, 2023", a string comparison marks it wrong. A judge marks it right.

question  ─┐
gold      ─┼─►  judge model + judge prompt  ─►  CORRECT or WRONG
answer    ─┘

But the judge is a model too, with its own habits. One judge accepts a loose paraphrase. Another wants the exact detail. This brings us to the most important lesson of the chapter.

We repeated the double grading on our final run. The widget below shows that pair.

earlier run    94.0  - 88.5  = 5.5 points
final run      94.22 - 87.73 = 6.49 points
open-domain    80.21 - 55.21 = 25 points, the largest gap of any category

Look at the two judges side by side, category by category. Watch which category moves most.

The five dials behind every score

A benchmark score is produced by a whole pipeline, and only one part of that pipeline is the memory system. At least five dials move the number without touching memory at all.

DialWhat it changesWhat we measured
The answer modelWho reads the memories and writes the answerLongMemEval, all 500 questions: 81.0 with qwen3.7-flash, 88.6 with qwen3.8-flash, same memories
The judgeWho gradesLoCoMo: 88.5 and 94.0 on identical answers
The context budgetHow much memory is placed on the deskOurs: about 1,470 tokens per LoCoMo question. Mem0 reports a top-200 budget, the 200 best-matching memories, of about 7,000 tokens
The question setAll questions, or a sampleLongMemEval: 91 on a 100-question sample, 81.0 on all 500
LuckThe provider is not fully repeatableThe same configuration scored 93.51 and then 92.47

Samples flatter

Our 100-question sample of LongMemEval scored 91. The full 500 scored 81.0 with the same code. The sample happened to contain fewer questions of the two hardest types. In the full set, multi-session has 121 questions and we scored 65 percent on them. Temporal has 127 questions and we scored 72 percent. A score must say how many questions it covers.

Noise is about one point

This is the noise test of Chapter 17. We ran one configuration twice with fresh answers, and 63 of the 385 answers came back different.

run 1                       93.51
run 2, same configuration   92.47
difference                  93.51 - 92.47 = 1.04 points, from nothing at all

So on that set, any difference under one point is noise. Any gain under two points needs a second, separate set of questions to confirm it.

What an honest score looks like

A score is a claim until it names its answer model, its judge, its question count and its context tokens per question. With those four, a reader can decide what the number is worth.

claim     "We score 92.5 on LoCoMo."

report    LoCoMo, all 1,540 questions, categories 1 to 4
          memory layer   qwen3.7-flash extraction, qwen3-embedding-4b
          answer model   qwen3.7-flash
          judge          gpt-4o-mini
          context        about 1,470 tokens per question
          score          94.2

hanumemAI's results

Here is every headline run. The memory layer is the same in all of them: qwen3.7-flash writes the memories and qwen3-embedding-4b makes the vectors. "Flash" means qwen3.7-flash. The other answer model, qwen3.8-flash, is a stronger and more expensive model.

BenchmarkQuestionsAnswer modelJudgeScoreContext tokens per question
LoCoMo1,540flashgpt-4o-mini94.2~1,470
LoCoMo1,540qwen3.8-flashgpt-4o-mini96.0~1,470
LoCoMo1,540flashflash87.7~1,470
LoCoMo1,540qwen3.8-flashflash90.8~1,470
LongMemEval-S500flashflash, paper's prompts81.0~2,950
LongMemEval-S500qwen3.8-flashflash, paper's prompts88.6~2,950
BEAM-100K400flashflash, rubric65.8~3,760
BEAM-100K400qwen3.8-flashflash, rubric72.5~3,760
BEAM-500K700flashflash, rubric61.3~3,400
BEAM-1M200flashflash, rubric56.5~3,500
BEAM-1M200qwen3.8-flashflash, rubric64.0~3,500

The BEAM rows do not cover the same share of the benchmark. BEAM-100K is 20 conversations and BEAM-500K is 35. BEAM-1M is the first 10 conversations only. Each conversation has 20 questions.

The full LoCoMo run, by category, under both judges. Evidence recall is the free retrieval metric of Chapter 17.

JudgeOverallMulti-hopTemporalOpen-domainSingle-hopEvidence recall
qwen3.7-flash87.7388.3086.9255.2191.5688.75
gpt-4o-mini94.2294.3392.8380.2196.3188.68

And LongMemEval, all 500 questions, by type:

Answer modelOverallSingle-session userMulti-sessionKnowledge updateTemporal
qwen3.7-flash81.0095.3165.2991.6772.44
qwen3.8-flash88.6096.8881.8291.6783.46

Next to other systems

These numbers come from the other systems' own papers and posts. We did not run them. They use other answer models, other judges and other budgets, so this is context and not a leaderboard. "Self-reported" marks a number that a vendor published about its own product.

BenchmarkSystemScoreStack and status
LoCoMoMem0 paper (2025)66.9gpt-4o-mini, about 7,000 tokens, reported
LoCoMoMem0 platform (2026)92.5gpt-5, top-200, about 7,000 tokens, self-reported
LoCoMoAgent Zero (2026)93.6reported
LoCoMohanumemAI94.2 / 96.0flash / qwen3.8-flash, gpt-4o-mini judge, about 1,470 tokens, measured by us
LongMemEval-SZep71.2reported
LongMemEval-SPaper's best pipeline~75gpt-4o, reported
LongMemEval-ShanumemAI81.0 / 88.6flash / qwen3.8-flash, about 2,950 tokens, measured by us
LongMemEval-SHydraDB85.8 / 90.8gpt-4o-mini / Gemini 3 Pro, self-reported
LongMemEval-SSodaMem92.8about 18,000 tokens, reported
LongMemEval-SMem0 platform (2026)94.4gpt-5, self-reported
LongMemEval-SAgent Zero95.6reported
BEAM-1MMem0 platform (2026)64.1gpt-5 stack, about 6,700 tokens, self-reported
BEAM-1MhanumemAI56.5 / 64.0flash / qwen3.8-flash, about 3,500 tokens, measured by us

Where others score higher

On LongMemEval, hanumemAI is not the top number. Agent Zero reports 95.6, the Mem0 platform reports 94.4 with gpt-5, SodaMem reports 92.8 with about 18,000 tokens of context, and HydraDB reports 90.8 with Gemini 3 Pro. Our best is 88.6. Our evidence recall on that set is 98.1 percent, so the right memories do reach the desk. The misses are reading mistakes on multi-session and temporal questions. They are made by a small answer model in a single pass: one model call per question, with no second look.

Where hanumemAI leads

On LoCoMo, under the same judge model the Mem0 paper used, hanumemAI scores 94.2 with a flash answerer and 96.0 with qwen3.8-flash. It does so with about 1,470 tokens per question against about 7,000. The answer models differ, so even this is not a like-for-like race.

context per question     hanumem 1,470      Mem0 top-200  about 7,000
ratio                    7,000 / 1,470  =  about 4.8 times less context

On BEAM at one million tokens, hanumemAI's 64.0 is level with Mem0's self-reported 64.1, with about half the context (3,500 tokens against 6,700).

The honest framing

hanumemAI's claim is this: the best score per context token, with small and cheap models. It is not "the highest number on every exam". Every point in the chart below is a score next to its cost. Choose a benchmark, hover a point to see its model stack, and read the table under the chart. Systems that publish no token count are in the table only.

Try it yourself: read a score like an examiner

Take any memory product's website. Find its benchmark number. Then try to fill in this card.

benchmark and version        ....................
questions scored             ....... of .......
answer model                 ....................
judge model and prompt       ....................
context tokens per question  ....................
who ran it                   the vendor / a third party
per-category scores shown    yes / no

Every line you cannot fill is a dial that may have been turned. If you have the hanumemAI research repository, you can also reproduce a row yourself. The free mode measures retrieval only and costs nothing.

scripts/bench.sh --free      # evidence recall only, no answer or judge calls, $0
scripts/bench.sh             # LoCoMo, 385 questions, a few cents with Qwen flash
scripts/bench.sh --full      # all 1,540 questions

Common questions

Is a higher benchmark score always a better memory system?

No. A higher score may come from a stronger answer model, a kinder judge or five times more context. Ask for the score and its token cost together. A system that needs 18,000 tokens per question costs about twelve times more on every message than one that needs 1,470, because 18,000 / 1,470 = 12.2.

Why not grade with exact matching and avoid the judge problem?

Because correct answers come in many wordings. Exact matching would mark "May 7th" wrong when the gold answer says "7 May". A judge model is the practical choice. The fix is to keep the judge fixed, name it, and never change it in the middle of a series of experiments.

What is evidence recall, and why do you call it free?

The benchmark marks which turns contain the answer. Evidence recall is the share of those turns that our search returned. It needs no answer model and no judge, so it costs nothing to compute, and it measures the memory system alone. On LoCoMo ours is about 88.7 percent. On LongMemEval it is 98.1 percent.

Why does hanumemAI publish a score of 87.7 when it could publish only 94.2?

Both are true readings of the same answers under two judges. Publishing only the kinder one would turn a dial in our favour and hide it. The 94.2 is comparable with the Mem0 paper because it uses that paper's judge model. The 87.7 shows what a stricter judge says.

Should I choose a memory system by its benchmark rank?

Use the benchmark to shortlist, then test on your own data. Build a golden set of 50 to 100 questions from your real users, as in Chapter 17. Your users' questions are the only exam that matters in the end.

How can I tell whether a vendor's number is believable?

Look at the per-category scores. A system that keeps two clocks should win on temporal questions. A system with a strong profile should win on open-domain questions. When the gains match the mechanism, the number means something. When a system wins every category equally, check how the test was run.

Carry this

  • Three exams, three weaknesses: LoCoMo for long conversations, LongMemEval for facts that change, BEAM for extreme length.
  • A score is a claim until it names its answer model, its judge, its question count and its context tokens per question.
  • The judge alone moved identical answers from 88.5 to 94.0. A sample scored 91 where the full set scored 81.0. Run-to-run noise is about one point.
  • hanumemAI leads on score per token: LoCoMo 94.2 / 96.0 at about 1,470 tokens. On LongMemEval others report higher numbers than our 88.6.

Check yourself

1. System A reports 92 on LoCoMo with 7,000 context tokens per question. System B reports 90 with 1,500. The input price is $3 per million tokens and a product sends 2,400,000 messages a day. What does each system's context cost per day?

Answer
A   2,400,000 x 7,000 = 16,800,000,000 tokens   x $3 / 1,000,000 = $50,400 per day
B   2,400,000 x 1,500 =  3,600,000,000 tokens   x $3 / 1,000,000 = $10,800 per day

Two points of score cost $39,600 a day. And two points is close to the noise band.

2. A change raises our score from 91.2 to 91.9 on 385 questions. Should we keep it?

Answer

Not on that evidence. The gain is 0.7 points and the measured noise is about 1 point. We keep it only if it also makes the system cheaper or simpler, or if a separate set of questions shows a clear gain.

3. Why did the 100-question LongMemEval sample score 91 while the full 500 scored 81.0?

Answer

The sample contained too few questions of the two hardest types. In the full set, multi-session (121 questions, 65 percent) and temporal (127 questions, 72 percent) make up about half of all questions and pull the average down.