Carebun Interactive/Memory Memory labCheat sheet

Part IV · Ninja

Chapter 21

What 121 experiments taught us

One idea, one run, keep or discard: the loop that tuned hanumemAI for about 25 dollars, what moved the score, what never did, and the mistakes we made on the way.

14 min read · interactive

On this page
  1. The loop
  2. The ledger
  3. The climb on LoCoMo
  4. What moved the score
  5. What never helped
  6. The mistakes
  7. What is still open
  8. Try it yourself: run one turn of the loop
  9. Common questions
  10. Carry this
  11. Check yourself

In this chapter, we will learn how hanumemAI was tuned: by a simple loop of small experiments, each one kept or thrown away. We will follow the score as it climbed, list the ideas that worked and the many that did not, and admit the mistakes. The method is worth more than any single result, because it works for any system you will ever build.

The loop

An experiment loop is a fixed routine: pick one idea, make the change, run the exam, then keep the change or discard it, and write one line in the log. The shape comes from Andrej Karpathy's autoresearch project. We adapted it to a memory library.

Four words return on every page of this chapter. The first two come from Chapter 18. The exam is a benchmark: LoCoMo, LongMemEval or BEAM. The judge is the model that grades the answers. The reader is the answer model, the one that reads the memories and replies. Evidence recall is the share of the needed source messages that the search returned. It is counted without any model, so it is free.

   ┌─► 1. start from the last kept version (the baseline)
   │   2. pick ONE idea and edit the code
   │   3. run the tests. They must pass
   │   4. run the exam (the benchmark)
   │   5. write one line in results.tsv: score, recall, tokens, cost, what changed
   │   6. keep (the change becomes the baseline) or discard (restore the baseline)
   └── 7. go to 1

Rule 1. The exam is frozen

The questions, the judge prompt and the judge model never change inside a series of experiments. The answer prompt is part of the system, so it may change. The judge is the exam, so it may not. A student who can edit the exam always passes.

Rule 2. Cost is part of the score

Every run reports its accuracy next to its context tokens per question. A change that adds 1 point and doubles the tokens is discarded. A change that halves the tokens at an equal score is kept.

Rule 3. Know the noise

We ran one configuration twice with fresh answers. It scored 93.51 and then 92.47. So the noise on that set is about 1 point, and the keep rule follows from it.

Result on the 385-question setDecision
Score up by 2.0 or moreKeep
Score up by 1.0 to 2.0Keep only if evidence recall also rose, or if a second, separate set of questions agrees
Score within 1.0, tokens down 15 percent or more, or less codeKeep
Anything elseDiscard

Rule 4. Simpler wins

Half a point gained with 100 lines of special cases is a discard. An equal score with less code is a keep.

The ledger

The ledger is the log file, results.tsv, with one line for every run. Each line has a status: keep, discard or info. An info line is a run that tests no idea, such as a measurement of the full set, a diagnostic or a repair.

Every run has an id. A plain number such as 0014 is a run of the first series, on LoCoMo. A letter in front names a later series: B for BEAM, L for LongMemEval, A for the answer prompt, E for extraction and embedding, R for the read path, C for the citation lock and P for product features.

runs logged in results.tsv          121
   kept                              16
   discarded                         29
   info                              76      16 + 29 + 76 = 121
product changes kept                 13      (features with no score, logged as info)
spent in total                       about $25

Of the 45 runs that ended in a decision (16 + 29), 29 were discards. That is normal, and it is the point. Most failures cost a few cents, and each one closed a door for good. Browse all of them here. Filter by kept, discarded or info, and click any one to read what was tried and what happened.

The instrument adds up the cost column of the log, so with every run shown it prints $26.43. Some entries in that column are rounded estimates. The metered spend was $25.22.

The climb on LoCoMo

The climb is the list of kept experiments, in order, with the score each one reached. These are the six keeps on the LoCoMo development set. The development set is the part of the exam we allow ourselves to tune on: 385 questions, with qwen3.7-flash as answer model and judge.

ExperimentChangeScoreTokens per question
0001The first working version86.491,489
0002A new answer prompt. Retrieval unchanged90.391,489
0003Time quota: reserved places for rows on a named date90.911,489
0006Stop injecting "rule" rows into every read91.171,250
0013Fusion weights 0.8 / 0.1 / 0.191.171,218
0014A running profile of 120 words93.511,440
total gain        93.51 - 86.49 = 7.02 points
from the answer prompt alone    90.39 - 86.49 = 3.90 points
from the running profile        93.51 - 91.17 = 2.34 points
the two together                3.90 + 2.34   = 6.24 of the 7.02 points
where the other 0.78 went
time quota (0003)               90.91 - 90.39 = +0.52
repair of the judge (0004b)     91.69 - 90.91 = +0.78   same answers, graded again
no rule rows (0006)             91.17 - 91.69 = -0.52   kept for 16 percent fewer tokens
fusion weights (0013)           91.17 - 91.17 =  0.00   kept for +1.7 points of evidence recall
sum                             0.52 + 0.78 - 0.52 + 0.00 = 0.78

Two ideas gave almost nine tenths of the gain (6.24 / 7.02 = 0.89). Neither one touched the search. Two of the six keeps did not raise the score at all. They were kept because they made the same score cheaper, or made the free metric better.

What moved the score

A change moved the score when the gain was larger than the noise and held on the other benchmarks. Five changes did.

The answer prompt came first (0002)

The first prompt allowed the model to say "I don't know". On LoCoMo's four scored categories, every question has an answer, so each "I don't know" was a free miss. The new prompt said: read every memory, prefer the speaker's own words, count relative dates from the message date, and always commit to the best answer. It was worth 3.9 points before any work on retrieval.

A date in the question is a channel, not a boost (0003)

When the question names a date, we first tried adding 0.15 to the score of rows on that date. Nothing changed. Then we reserved 10 of the 30 places for the best rows on that date. That moved the score by 0.52. This is HydraDB's channel quota (Chapter 19).

The running profile (0014)

One paragraph of about 120 words per user, updated once per session and placed at the top of the block. It answers "what kind of person is she" questions that 30 separate rows never quite add up to. It cost 18 percent more tokens (1,218 to 1,440) and gave 2.34 points. When we hid the profile from the extractor and showed it only to the reader, the score fell by 2.1 (0015). The extractor writes better facts when it knows the person.

A cap in tokens, not in rows (B0002)

BEAM's assistant messages run to 500 words. Thirty such rows were 12,380 tokens per question. A cap of 4,000 tokens cut that to 4,227 and the score went up, from 46 to 49.

Answer rules that depend on the question (B0005, B0006, L0003, A0001)

On BEAM the climb was 46, 49, 56, 65, and most of it came from how the answer is shaped. A question about a contradiction wants both statements with their dates. A summary question wants a paragraph. A question about a total wants the sum computed.

The last rule is about length. We asked for "one to three sentences with every specific detail" in place of "one short sentence" (A0001). LongMemEval rose by 4 points, from 78 to 82. The BEAM rubric mean rose by 8 points, from 52.3 to 60.3. The rubric mean is the average grade over all the points a good answer must contain, so it rewards an answer that is partly right.

The measured trade between recall and tokens

We measured how much of the evidence the search returns as k, the number of rows, grows. This costs nothing, because it needs no answer model and no judge.

kEvidence recallTokens per question
1583.0 %875
3087.6 %1,489
5090.5 %2,324
10093.3 %4,431

More rows always find more evidence. But accuracy did not follow. We paid for two scored runs to see it. They were made later than the free sweep, on baselines that used fewer tokens, so their token counts differ from the table above.

k = 30 to k = 50 (0008)    recall  90.23 - 87.34 = +2.89
                           score   90.91 - 91.17 = -0.26
                           tokens  2,092 / 1,250 = 1.67, so 67 percent more
k = 30 to k = 20 (0018)    score   89.09 - 93.51 = -4.42
                           tokens  1,038 / 1,440 = 0.72, so 28 percent fewer

A small answer model reads 30 rows better than 50, and 20 rows are too few. In the instrument below, switch between the two views. Evidence recall climbs with every extra row. Answer accuracy does not.

What never helped

A negative result is an idea that was tried fairly and lost, or changed nothing. The numbers in the table are changes in score, in points, on the named set.

IdeaExperimentsResult
Rewriting the extraction prompt, the instructions of the secretary0007, 0016, E0002Lost 1.6, 3.4 and 5 (BEAM). More facts crowd the block
A stronger extractor modelE0001Fewer facts, and 8 points lost on BEAM
Favouring raw turns, or favouring facts0009, 0010-2.08 and -0.26. Equal footing is right
Ordering the block by relevance0011Identical score. Order does not matter to the reader
One date header per day0012-0.52. A date on every line is worth its tokens
Rank fusion (RRF), which mixes the three searches by rank position, in place of a weighted sum of scores00132 points of recall worse (85.31 against 87.34)
Extraction in 10-turn windows0020Same score, 38 percent more facts, 2.6 times slower
Asking the model to write a hypothetical answer before searching0022No gain, one more LLM call
"List the relevant lines, then answer" in one call0005+0.26, inside the noise, 3 times the output tokens
Dropping a fact when its source message is presentR0001-7 on LongMemEval-100, -3 on BEAM-400
Citation lock: the reader must name the lines it used, or decline to answerC0001-8 on LongMemEval-100, -3.5 on BEAM-400
Cutting long messages at 1,200 charactersB0010-7 on BEAM. The instructions live in the long turns
Recency termR0002Identical recall at every weight. It cannot be tested offline

The pattern is clear. The write side reached a local best early, and the search stopped being the limit once evidence recall passed about 88 percent. A local best is a setting where every small change makes things worse, although a very different design might still be better.

On LongMemEval the evidence recall is 98.1 percent and the score is 81.0. The memories arrive. The small model misreads them. With qwen3.8-flash as reader the same memories score 88.6.

The mistakes

A mistake of the loop is an error in the method, not a failed idea. The loop made six that mattered. Each one became a rule.

What went wrongEntryThe rule it left
A change was kept with a failing test, because a ; stood where && belongedP0002bTests gate the keep. Always chain with &&
An experiment edited a benchmark file. The discard restored the library but not that file, and the next headline run crashedB0007bThe exam is read-only. Revert it by hand if touched
The write policy matched bare words such as "token", "pin" and "salary". It refused ordinary facts and cost 4 pointsP0009bMatch the secret's value, not the word. Product changes are cross-checked too
A rule looked like +5 on 100 questions, then lost 5 on a larger setA0002Small sets lie. Use the large sets
The judge's reply was cut at 200 tokens, and 3 unfinished judgments were counted wrong0004bEvery "judge error" is a scoring bug until proven otherwise
Rules tuned on BEAM cost 8 points on LongMemEvalL0001bEvery keep is checked on all three benchmarks

Small sets lie

The A0002 story deserves a closer look. A rule for ordering questions scored +5 on a 100-question set and +2 on a second 100-question set. It was kept. On the larger sets it lost: LongMemEval-100 went from 88 to 83, and the BEAM rubric mean went from 60.9 to 59.0. It was reverted.

on 100 questions, 1 question          = 1 point
a "+5" result                         = 5 questions
answers that came back with different words when we ran the same
configuration twice (the noise test of Rule 3):
   63 of 385    63 / 385 = 0.16      = about 16 in every 100

Most of those 16 stay right or stay wrong. A few change sides. Five questions is inside what that luck can do. After that day the large sets became the everyday test. They are cheap, because ingest is cached. Ingest is the step that reads every conversation of a benchmark and writes its memories. With it cached, LoCoMo 385 questions costs about $0.03 to re-score, LongMemEval-100 about $0.008, and BEAM-400 about $0.06.

What is still open

An open problem is a weakness that we measured and did not fix. There are three.

  • Multi-session and temporal reading on LongMemEval. With the flash reader: 65 percent and 72 percent. With qwen3.8-flash: 82 percent and 83 percent. The reader is the ceiling.
  • Ordering, contradiction and summarization at one million tokens. On BEAM-1M with the flash reader: event ordering 20 percent, contradiction 20 percent, summarization 25 percent.
  • Summaries. Neither a larger budget (B0008) nor spreading the rows over all days (B0009) helped. The rubric asks for early details that the extracted facts do not carry.

The next ideas on our list are three. A cheap reader in two passes: the first call lists the useful lines and the second answers from those lines only. A local reranker, a small model that sorts the found rows once more before the block is cut. And sending only the hard questions to the stronger model.

Try it yourself: run one turn of the loop

With the research repository and an API key, one full turn of the loop looks like this. The first run is always the untouched baseline.

scripts/diff.sh                                   # must print nothing: we start at the baseline
# edit ONE knob, for example k = 30 to k = 40 in hmem/config.py
uv run ruff check hmem tests && uv run pytest -q    # the code checker and the tests must pass
scripts/bench.sh --free > run.log 2>&1            # free: evidence recall and tokens only
grep -E "^(evidence_recall|ctx_tokens_mean|RESULT):" run.log
scripts/discard.sh                                # or scripts/keep.sh "k=40" if it earned it

You do not need our repository to use the method. For any project, write down three things before you change anything: the exam, the noise, and the keep rule. Then keep a log with one line per run. A line in our log looks like this. The first four columns are the experiment, the score, the evidence recall and the tokens per question. Then come four scores by category (left out here), the cost in dollars, the decision and the description.

0018  89.09  84.83  1038  ...  0.0217  discard  k=20 with the profile: -4.42 (89.09) for
      -28% tokens (1440->1038); recall 84.8. The profile does not replace retrieved rows

Common questions

Why change only one thing at a time?

Because then the result has one explanation. Experiment B0004 changed three answer rules at once. BEAM went up 4 points and LoCoMo went down 2.6. We had to split it (B0005) to learn that two rules helped and the third, which allowed abstaining, did the damage.

Is tuning on a benchmark not a form of cheating?

It can be. Three habits keep it honest. We tune on a development part (3 of the 10 LoCoMo conversations) and never read the questions of the rest. We check every keep on all three benchmarks, so a rule that fits one exam and hurts another is caught. And we publish the full-set numbers with their cost.

How could 121 runs cost only about 25 dollars?

Two reasons. The models are small: qwen3.7-flash costs $0.03 per million input tokens. And every LLM call and every embedding is cached, so a repeated call is free. A read-side experiment re-uses all the stored memories and pays only for new answers, a few cents. A write-side experiment pays for the whole ingest again, about $1.50.

Why did a stronger extractor make things worse?

The stronger model was more careful and wrote fewer facts (355 fell to 328 on the BEAM set). The prompt had been tuned with the smaller model, which writes more freely. Fewer facts meant fewer ways to find an answer. A prompt and its model are a pair.

If the reader is the limit, why not always use the bigger model?

Cost. On the LoCoMo development set (385 questions) the stronger reader cost about 4 times more for 2 points. On LongMemEval it is worth 7.6 points, so there it may be worth paying. The choice belongs to the application, which is why hanumemAI reports both.

Can an AI agent run this loop by itself?

Yes. This one was run by an AI agent following a written program, with a budget guard that stops any run that overspends. The human set the rules and read the log in the morning. The mistakes in the table above are the agent's, and so are the keeps.

Carry this

  • One idea, one run, keep or discard, one log line. Freeze the exam. Count the cost. Know the noise.
  • LoCoMo went from 86.49 to 93.51. The answer prompt (+3.90) and the running profile (+2.34) gave 6.24 of the 7.02 points.
  • Every change to the extraction prompt lost. Search stopped being the limit at about 88 percent recall. The reader is the ceiling.
  • Small sets lie at the level of 2 to 5 points. Check every keep on all three benchmarks.
  • Failed experiments are results. Write them down.

Check yourself

1. An experiment raises the score from 93.0 to 94.2 and the tokens from 1,440 to 2,900. Keep or discard?

Answer
gain          94.2 - 93.0   = 1.2 points
token cost    2,900 / 1,440 = 2.01 times

Discard. The gain is close to the noise and the context has doubled. Cost is part of the score.

2. With k = 50 the search found more evidence than with k = 30, yet the score fell. How can that be?

Answer

Finding the evidence and reading it are two jobs. The extra 20 rows brought a little more evidence and much more text that was not needed. A small answer model reads a short block more reliably than a long one.

3. Why is a write-side experiment about 15 times more expensive than a read-side one?

Answer
read side, all three sets    0.03 + 0.008 + 0.06 = about $0.10
write side                   about $1.50
ratio                        1.50 / 0.10         = 15 times

A read-side change re-uses the stored memories and pays only for new answers and judgments. A write-side change alters what the extractor sees, so its cached calls no longer match and every conversation must be extracted again.