Part II · Apprentice
Chapter 10
Update or add: consolidation
A store that only grows stops being true. Four verbs keep it honest, one experiment shows them at work, and then the industry changes its mind.
On this page
In this chapter, we will learn how a memory store stays true while its owner's life keeps changing. We will meet the four verbs of consolidation, watch a real model use them on four new facts, and count what that decision costs. Then we will see why the newest systems, including ours, no longer let the model delete.
The problem: a store that only grows
A memory store must not just grow. It must stay true.
In Chapter 6 the secretary learned to write cards. Every session she reads the transcript and writes about three of them. So far, so good.
Now Maya moves from London to Paris. The secretary writes a new card: "Maya moved from London to Paris last month." The old card, "Maya lives in London", is still in the box. Nobody touched it.
Next week Maya asks for a good bakery nearby. The librarian pulls the cards about where she lives and finds two. One says London. One says Paris. Both look equally confident. The assistant has to guess.
The step that keeps the store true is called consolidation.
The four verbs
Consolidation is one decision, made for every new fact, with exactly four possible answers. The class calls them the four verbs.
| Verb | When | What happens to the store |
|---|---|---|
| ADD | nothing similar is stored | a new memory is written |
| NOOP | the fact is already stored | nothing; the store is unchanged |
| UPDATE | the fact replaces memory N | memory N is rewritten |
| DELETE | the fact cancels memory N | memory N is removed |
NOOP is short for "no operation". It is the quiet hero of the four. After the first few weeks, most of what a person says about themselves is something the assistant already knows. NOOP is what stops "Maya is vegetarian" from being stored again and again.
candidate fact
|
v
find the 2 most similar stored memories (same user only)
|
v
LLM decides, one call
|
+-- nothing similar ---------> ADD as a new memory
+-- already stored ----------> NOOP
+-- replaces memory N -------> UPDATE memory N
+-- cancels memory N --------> DELETE memory N
Notice the second box. Before the write path writes, it reads. It runs a small search to find the neighbours of the new fact. That search obeys the same rule as every other search in this book: it looks only inside one user's memories.
The update-decision prompt
The update-decision prompt is the instruction that turns a language model into the careful secretary. It is small, and its shape matters more than its wording.
What goes in: one candidate fact, plus the 2 most similar stored memories, each with its id and its similarity score.
What comes out: strict JSON holding exactly one operation.
{"op": "ADD" | "UPDATE" | "DELETE" | "NOOP", "target_id": ..., "memory": ...}
Why so strict? Because a program reads the answer, not a person. If the model replies with a friendly paragraph, the program cannot act on it. One object, three fields, nothing else. The model runs at temperature 0, which means it is asked to be as repeatable as it can be.
Here is a real input and output, so the shape is clear.
candidate: "Maya moved from London to Paris last month."
stored: id 2 score 0.846 Maya lives in London.
id 3 score 0.636 Maya is training for the Berlin marathon in September.
decision: {"op": "UPDATE", "target_id": 2, "memory": "Maya lives in Paris."}
Now it is your turn to be the judge. For each of the six new facts, read its two nearest stored memories and pick a verb. Then press the button and compare your choices with the model's. The first four candidates are the measured decisions of the experiment below. The last two are illustrative.
Try it yourself: four candidates, four decisions
The class experiment update-or-add starts with four memories about Maya and then feeds in
four new facts, one at a time. Here is the store at the start.
id memory 0 Maya is vegetarian. 1 Maya has a serious peanut allergy. 2 Maya lives in London. 3 Maya is training for the Berlin marathon in September.
And here is what the model decided for each candidate. The scores are cosine similarity, the closeness measure from Chapter 7, where 1.000 means identical.
candidate 1: "Maya moved from London to Paris last month."
retrieved score 0.846 id 2 Maya lives in London.
retrieved score 0.636 id 3 Maya is training for the Berlin marathon in September.
decision {"op": "UPDATE", "target_id": 2, "memory": "Maya lives in Paris."}
candidate 2: "Maya is vegetarian."
retrieved score 1.000 id 0 Maya is vegetarian.
retrieved score 0.706 id 1 Maya has a serious peanut allergy.
decision {"op": "NOOP", "target_id": null, "memory": null}
candidate 3: "Maya adopted a cat named Miso."
retrieved score 0.613 id 2 Maya lives in Paris.
retrieved score 0.612 id 0 Maya is vegetarian.
decision {"op": "ADD", "target_id": null, "memory": "Maya adopted a cat named Miso."}
candidate 4: "Maya injured her knee and is no longer training for the marathon."
retrieved score 0.761 id 3 Maya is training for the Berlin marathon in September.
retrieved score 0.618 id 1 Maya has a serious peanut allergy.
decision {"op": "UPDATE", "target_id": 3,
"memory": "Maya injured her knee and is no longer training for the marathon."}
Look at candidate 3. The nearest memory scores only 0.613, and it is about Paris, not about cats. Nothing in the box is really about a pet, so the model adds a new card. Look at candidate 2. A score of 1.000 means the same sentence is already stored, so the model does nothing.
The store at the end:
id memory 0 Maya is vegetarian. 1 Maya has a serious peanut allergy. 2 Maya lives in Paris. 3 Maya injured her knee and is no longer training for the marathon. 4 Maya adopted a cat named Miso.
Compare the two ways this could have gone.
without consolidation 4 stored + 4 new = 8 memories
includes London and Paris side by side (a contradiction)
includes "Maya is vegetarian." twice (a duplicate)
with consolidation 4 stored + 1 ADD = 5 memories
every one of them currently true
To run it, you need Ollama and a Qdrant container:
ollama pull qwen2.5:7b-instruct
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/update-or-add
pip install -r requirements.txt
python main.py
What the decision costs
Consolidation is a second LLM call, and it is paid once per candidate fact, not once per session. That detail makes it the busiest model in the whole write path.
The experiment measured the four decisions above.
4 decisions 1,486 prompt tokens + 104 completion tokens, 9.5 s of LLM time
per decision 1,486 / 4 = about 372 prompt tokens
9.5 s / 4 = about 2.4 s, a small model on a laptop
The class notes budget 550 prompt tokens for the same call, because a production rules prompt is longer than this teaching one. Now let us scale it with the class numbers. The prices are the class's prices for a small model, per million tokens (M): $0.15 for tokens sent in and $0.60 for tokens the model writes out.
sessions per day 200,000 users x 2 = 400,000
extraction calls per day one per session = 400,000
candidates per session about 3
update-decision calls per day 400,000 x 3 = 1,200,000
one decision 550 in x $0.15/M = $0.0000825
40 out x $0.60/M = $0.000024
total = $0.0001065
per day 1,200,000 x $0.0001065 = $127.80
Extraction itself is sized the same way, with 1,500 tokens in and 150 out per session.
one extraction 1,500 in x $0.15/M = $0.000225
150 out x $0.60/M = $0.00009
total = $0.000315
per day 400,000 x $0.000315 = $126.00
write path with decisions $126.00 + $127.80 = $253.80
calls with decisions 400,000 + 1,200,000 = 1,600,000 4 times 400,000
So the decision step roughly doubles the write-path bill, and it makes four times as many LLM calls. Nobody waits for these calls, because the write path runs in the background. But somebody pays for them.
The twist: newer systems stopped deleting
An ADD-only write path never updates or deletes a memory at write time. Here the story takes a turn. Mem0 is a widely used open-source memory library, and its original paper described exactly the design above: extract candidates, then judge each one with the four verbs. In its 2026 rewrite, the four verbs are gone from the write path. The new version is ADD-only, with one LLM call per write.
Why would anyone remove the part that keeps the store true? Two reasons.
- A wrong DELETE destroys real data. When a model decides "this cancels that, erase it", a bad guess is permanent. There is no undo for a row that no longer exists. The same is true of an UPDATE that overwrites: the old sentence is gone.
- Cost and delay. One call per candidate, after the extraction call, is the larger half of the write-path bill, as we just counted.
Look again at candidate 1 in the experiment. After the UPDATE, the store says "Maya lives in Paris." That is true. But ask "where did Maya live before?" and the store has no answer. London was overwritten. The update kept the store true and made it forget its own history.
What hanumemAI does instead
hanumemAI folds the decision into the extraction call and never destroys a row. It works in three parts.
- The extractor sees the neighbours. Before the one LLM call, the write path finds the 10 existing facts most similar to the new messages and shows them to the extractor, with their ids.
- The extractor may say "this replaces that". Each new fact can carry a
supersedeslist holding the ids of the older facts it replaces. The store then closes the old row with a date and a pointer to the new row. The old row is not deleted. Chapter 11 is about exactly this. - Duplicates are caught by arithmetic, not by a model. A hash is a short fingerprint computed from a text: the same text always gives the same fingerprint. It catches a sentence that is identical, character for character. A cosine check catches a sentence that is nearly identical: if the new fact scores 0.95 or more against a neighbour, it is skipped.
Here is how the four verbs map onto this design.
| Class verb | hanumemAI equivalent | Who decides |
|---|---|---|
| ADD | insert a new row | the extractor |
| NOOP | hash match, or cosine 0.95 or higher: skip | arithmetic |
| UPDATE | insert a new row, close the old one through supersedes | the extractor |
| DELETE | not available to the model; only the user or the app can remove a memory | a person |
from hmem import Memory, Config
m = Memory(Config(db_path="app.sqlite"))
m.add("I live in London.", user_id="maya", observed_at="2025-01-14")
result = m.add("Big news: I just moved from London to Paris.",
user_id="maya", observed_at="2026-03-02")
print(result["added"]) # ids of the new facts
print(result["superseded"]) # ids of the older facts that were closed, not deleted
print(result["skipped"]) # how many candidates the dedup refused
When is each design right?
Neither design is wrong. They suit different stores.
| Four verbs, second LLM call | ADD-only with supersede and dedup | |
|---|---|---|
| LLM calls per session | 1 + about 3 | 1 (hanumemAI adds 1 for its running profile) |
| Store size | smallest: every row currently true | larger: closed rows are kept |
| History ("where did I live before?") | lost on UPDATE and DELETE | kept |
| A wrong decision | permanent | repairable: the old row still exists |
| Good for | a small profile a person can read at a glance; teaching | questions about the past; audits; large scale |
If your product only ever asks "what is true now?" and you want the smallest, tidiest store, the four verbs are a fine choice. If your users will ask about the past, or if you must be able to explain and undo every change, do not let a model overwrite or delete.
Common questions
Why show the model only 2 similar memories and not the whole store?
Because the prompt would grow with the store, and so would the bill. Two neighbours keep the decision at a few hundred tokens whether Maya has 20 memories or 2,000. The price is that a related memory ranked third is never seen. hanumemAI shows 10 neighbours for that reason, and can afford to, because it makes one extraction call per session and not one call per candidate.
What if the model picks the wrong verb?
With the four verbs, a wrong ADD leaves a duplicate, which is mild. A wrong NOOP loses a new fact. A wrong UPDATE or DELETE destroys a true memory, and that cannot be undone. This uneven risk is the main reason the newer systems removed UPDATE and DELETE from the model's hands.
Is a similarity score of 0.846 high enough to mean "the same topic"?
The score alone does not decide anything. It only chooses which two memories the model gets to read. The model then reads the actual sentences. "Lives in London" and "moved to Paris" score 0.846 because both are about where Maya lives, and the model sees that one replaces the other.
Why 0.95 for the duplicate check, and not 0.80?
In the experiment, "Maya moved from London to Paris last month" scores 0.846 against "Maya lives in
London". With a threshold of 0.80 the move would be thrown away as a duplicate, because 0.846 is above
0.80. That is the opposite of what we want. The threshold must be high enough that only rewordings of the same fact cross it. It is a
knob, dedup_cosine, and setting it to 1.0 turns the check off.
If nothing is deleted, does the librarian not find the old London card?
Not by default. A closed row is left out of ordinary search. It comes back only as the history of a current fact that was found, marked as outdated, or when the app asks for the state on a past date. Chapter 11 shows how.
Does the decision call slow down the answer?
No. The whole write path runs after the reply has been sent. Chapter 13 explains how. The decision costs money, not waiting time.
Carry this
- A store must stay true, not just grow. Before the write path writes, it reads the nearest stored memories.
- Four verbs: ADD (new), NOOP (already known), UPDATE (replaces), DELETE (cancels). One strict JSON object per candidate.
- In the experiment, 4 stored plus 4 new became 5 true memories, not 8 with a contradiction and a duplicate.
- The decision is one more LLM call per candidate: 1,200,000 calls and $127.80 a day in the class sizing.
- hanumemAI makes the decision inside the extraction call: the extractor names what a new fact
supersedes, a hash and a 0.95 cosine check replace NOOP, and the model can never delete.
Check yourself
1. The store holds "Maya's partner Sam has a birthday on June 18." A new candidate arrives: "Sam's birthday is June 18." Which verb, and why?
Answer
NOOP. The fact is already stored in other words. Adding it would put two cards with the same meaning in the box, and they would take two of the three places in a top 3 search.
2. A service has 50,000 sessions a day and the extractor averages 4 candidates per session. How many update-decision calls a day does the four-verb design make, and what do they cost at $0.0001065 each?
Answer
calls per day 50,000 x 4 = 200,000 cost per day 200,000 x $0.0001065 = $21.30
3. After an UPDATE turns "lives in London" into "lives in Paris", Maya asks "where did I live before Paris?" What does the store answer, and what would fix it?
Answer
Nothing. The London sentence was overwritten. The fix is to add the Paris fact as a new row and close the London row with a date and a pointer, so that both survive. That is superseding, and it is the subject of the next chapter.