Carebun Interactive/Memory Memory labCheat sheet

Part IV · Ninja

Chapter 19

How others do it: Mem0, HydraDB, Zep

Two opposite bets on the same problem, the research systems around them, and what hanumemAI took from each.

13 min read · interactive

On this page
  1. Mem0, as the paper described it
  2. Mem0, as it ships in 2026
  3. HydraDB: memory is a database problem
  4. Zep and Graphiti: four timestamps
  5. The research ideas around them
  6. The memories inside products
  7. Side by side
  8. What hanumemAI borrowed
  9. Try it yourself: one message, three systems
  10. Common questions
  11. Carry this
  12. Check yourself

In this chapter, we will learn how the best-known memory systems work inside. We will open Mem0 twice, once as its paper described it and once as it ships today. We will see how HydraDB and Zep treat memory as a database problem. Then we will meet the research ideas around them and list what hanumemAI borrowed from each.

Mem0, as the paper described it

Mem0 is the most used open-source memory library. Its goal is to make memory a drop-in part: two function calls, add() and search().

The Mem0 paper (April 2025) described a write path in two phases. It is the design from Chapter 10.

phase 1  extraction
         running summary + last 10 messages + the new message pair
              ─► LLM ─► candidate facts

phase 2  update, once for EACH candidate fact
         the 10 most similar stored memories
              ─► LLM ─► ADD | UPDATE | DELETE | NOOP ─► vector store

That is two kinds of LLM call for every message pair: one to extract, and one per fact to decide. On LoCoMo the paper reported 66.9. Pasting the whole conversation scored 72.9. So full history was more accurate, and Mem0 still won the argument on cost:

Full historyMem0 (paper)
LoCoMo score, reported72.966.9
Tokens per questionabout 26,000about 7,000
Slowest answers (p95)about 17 sabout 1.4 s

p95 means that 95 of every 100 answers were faster than this.

Mem0, as it ships in 2026

In 2026 Mem0 rewrote its write path to be ADD-only: one LLM call per write, and nothing is updated or deleted at write time. Its own code calls this the "V3 phased batch pipeline".

0  fetch the last 10 messages of this session
1  embed the new messages, fetch the 10 most similar stored memories
2  ONE LLM call: the additive extraction prompt
3  embed every extracted text
4  hash each text, skip it if it is already stored
5  insert the new memories
6  find the entities (names, places, titles) and link them to the memories
7  save the raw messages

Why did the four verbs go? Our sources give three reasons. A model that decides "this contradicts that, erase it" destroys real data when it guesses wrong. Two LLM calls per message pair cost twice as much as one. And conflict resolution moved to Mem0's paid platform, where a background job marks old memories as superseded and merges near-duplicates.

Rich memories

The new extraction prompt is the heart of Mem0. Its rules are worth learning.

  • Each memory is 15 to 80 words, self-contained, with no pronouns.
  • Capture the change, not only the result: "switched from almond milk to oat milk after an almond sensitivity", not "prefers oat milk".
  • Turn every relative date ("last week") into a real date, counted from the day the message was said.
  • Keep every name, title and number. "Assistant manager", not "manager".
  • When in doubt, extract. A slightly repeated memory costs less than a missing one.

Search with no LLM

Mem0's read path makes no LLM call. It runs three searches and adds the scores, the same three signals we met in Chapter 9.

dense     search by meaning (cosine similarity of vectors)
bm25      search by exact words
entity    a boost for memories linked to a name in the question

combined = (cosine + sigmoid(bm25) + entity boost) / the highest possible total

A word-match score has no upper limit. The sigmoid is an S-shaped function that squeezes it into the range 0 to 1, so that it can be added to the other two. Dividing by the highest possible total keeps the final score between 0 and 1 as well.

What Mem0 is good at, and what it hands back to you

You getYou inherit
Two function calls, any model, any vector store"Lives in Lisbon" and "moved to Berlin" stay side by side in the open-source version
One LLM call per write, zero per readA consolidation job you must build yourself, or buy
Nothing on the write path can destroy dataNo time model in the open-source code
A carefully written, open extraction promptNo typed links between facts

On its managed platform Mem0 reports LoCoMo 92.5, LongMemEval 94.4 and BEAM-1M 64.1. These are self-reported, with gpt-5. The LoCoMo number uses a top-200 budget: the 200 best-matching memories, about 7,000 tokens per question.

HydraDB: memory is a database problem

HydraDB's main claim is that when an agent repeats stale information, the fault is usually in the store and not in the model. It lists four kinds of failure: the model invents something, the agent reads stale context, state is lost between runs, and two agents overwrite each other. Three of the four belong to the store.

So HydraDB asks for the guarantees that databases have given for decades.

IdeaIn plain words
Append-onlyNever change a record. Add a new one that points at the old one.
BitemporalTwo clocks on every fact: when it was true, and when we learned it.
ProvenanceEvery fact knows the sentence it came from and the model that wrote it.
Typed edgesA link that says how two things relate: caused by, prefers, lives in.
Channel quotasEach search signal gets reserved places in the final list.

Typed edges: when similar is not relevant

stored
  m1  "Customer reported Error 503 after the March deploy."
  m2  "The March deploy migrated auth to the new gateway."
  m3  "Slow dashboards are usually a caching issue."

question
  "why is the app behaving strangely for this customer?"

ranking by similarity
  m3   0.62   sounds relevant, and is a red herring
  m1   0.44
  m2   0.11   the memory we needed, ranked last

The answer needs m1 joined to m2. But m2 shares no words and no meaning with the question. A better embedding cannot fix that, because the link is a fact about the world: the error was caused by the deploy. A typed edge stores that link, so the search can walk from m1 to m2.

The price

HydraDB is closed source, so everything here is its own description. It spends several LLM calls on every write. It rewrites each fragment to stand alone, because by its own count about 40 percent of naively cut chunks cannot be found by any question. It writes down what each fragment implies and embeds that too. It extracts entities and typed edges. Writes are slow and smart so that reads can follow the structure.

The read path is long too. A small model rewrites the question in several ways. Several searches are combined. The search then walks the typed edges from the first hits, and the shortlist is reranked, which means sorted again by a more careful scorer. HydraDB claims about 200 ms for all of this, against the 50 ms budget of this book.

HydraDB reports 85.8 on LongMemEval-S with gpt-4o-mini and 90.8 with Gemini 3 Pro. Both are self-reported.

Zep and Graphiti: four timestamps

Zep's Graphiti is a graph of facts in which every link carries four timestamps. A graph stores things as points, and the relations between them as links.

t_valid     t_invalid     t_created     t_expired

Our sources name the four and do not define them one by one. The names follow the two clocks of Chapter 11: valid and invalid for the world, created and expired for the system. When a fact changes, Graphiti marks the old link as invalid. It does not delete it. Zep reported 71.2 on LongMemEval-S.

The research ideas around them

Generative Agents: three scores added together

The Generative Agents paper added two terms to similarity: recency and importance. The class notes use the same three terms, as we saw in Chapter 9. Recency is 0.995 raised to the hours since the memory was last used. Importance is a rating from 1 to 10, given when the memory is written.

LIGHT: a running profile

LIGHT is the system proposed in the BEAM paper. It keeps three memories: retrieved segments of the conversation, the recent turns, and a scratchpad of the most important facts. The paper reports gains of 3.5 to 12.7 points over its baselines. hanumemAI's running profile is that scratchpad.

SodaMem: links that say "contradicts" and "updates"

SodaMem marks the relation between facts with typed links such as contradicts and updates. It reports 92.8 on LongMemEval-S with about 18,000 tokens of context per question.

Hindsight: a summary per entity

Hindsight keeps a summary for each person or thing the conversations mention. That costs one more LLM call for each entity in each session. We did not run this experiment, for that reason.

Agent Zero: the citation lock

Agent Zero makes its reader cite the memory lines behind each answer, and abstain when it cannot cite any. It reports 93.6 on LoCoMo and 95.6 on LongMemEval-S. It reads in several passes.

The memories inside products

ChatGPT, Claude and Gemini each ship a memory. We cannot see their code. What we can see is the control they give the user, as listed in the class notes.

ControlWho offers it
See every memoryChatGPT, Claude
Edit or delete one memoryChatGPT, Claude, Gemini
Memory off for one chatChatGPT Temporary Chat, Claude incognito, Gemini Temporary
A separate memory per projectClaude

Side by side

Class designMem0 paperMem0 open source, 2026HydraDBhanumemAI
Write pathextract, then decideextract, then decideextract onlyrewrite, imply, extract edgesextract only, with supersedes
When a fact changesUPDATE or DELETE the rowUPDATE or DELETE the rowboth stay opennew edge, old edge closednew row, old row closed
Time modelcreated and last usednot described in our sourcesnonetwo clockstwo clocks
Read pathsimilarity, top 3not described in our sourcesmeaning + words + namesmany signals, graph walk, rerankmeaning + words + names, time quota, one hop
LLM calls per write1 + 1 per fact1 + 1 per fact1several1, plus 1 short profile update per session
LLM calls per read0not described in our sources01 small call, to rewrite the question0

Choose which systems to compare below. The widget adds two rows to this table: whether each design can answer "where did I live before?", and where it keeps its data.

What hanumemAI borrowed

FromWe tookWe changed
Class notesCategories, importance, 90-day decay for events, the user_id wall, the write policyNo separate update call
Mem0, 2026ADD-only extraction, rich self-contained memories, hash dedup, entity links, three-signal searchThe extractor may say "this replaces that". Entities come from the same LLM call, and word search comes from SQLite
HydraDB, ZepTwo clocks, superseded_by, source ids on every row, reserved places for dated rowsPlain SQLite rows, no graph engine
LongMemEval paperFacts as keys, raw turns as values, time-aware questions
LIGHTThe running profile120 words, updated once per session

Try it yourself: one message, three systems

Take this message and write, on paper, what each system stores. The store already holds m17, "Maya lives in London."

Maya:  Big news: I just moved from London to Paris for a job at a payments startup.
Mem0 paper      phase 1 proposes "Maya moved from London to Paris"
                phase 2 decides UPDATE m17
                store: 1 row, "Maya lives in Paris"          London is gone

Mem0 2026       ADD m88 "Maya moved from London to Paris in March 2026
                for a job at a payments startup"
                store: 2 rows, both open                      m17 still says London

hanumem         ADD the new row with supersedes: [m17]
                m17 gets valid_to = 2026-03-02 and superseded_by = the new row
                store: 2 rows, 1 open and 1 closed            history is kept

Now ask each store "where did Maya live before Paris?" Only the stores that kept London can answer. Only the store that also closed London can answer "where does she live now?" without a guess.

Common questions

Which one should I use?

Match the system to the questions your users ask. If memory means preferences and facts about one person, and you want it working soon, an ADD-only library such as Mem0 or hanumemAI fits. If your questions are about how things connect across many documents and agents, a graph system such as HydraDB fits. If your users' facts change and you must answer about the past, you need two clocks.

Why did Mem0 remove UPDATE and DELETE when the class teaches them?

The four verbs are the clearest way to learn consolidation, so the class teaches them. In production, a wrong DELETE destroys a true fact forever and a second LLM call doubles the write cost. Closing a row with a date keeps the benefit of UPDATE and makes every mistake reversible.

Is a graph database required for good memory?

Not for personal memory. hanumemAI keeps two clocks and a superseded_by pointer in ordinary SQLite rows and follows the pointer one step at read time. A graph pays off when answers need several steps across typed links, as in the Error 503 example.

Why are so many scores marked self-reported?

Because the vendor ran the test on its own system with its own choice of answer model, judge and budget. That does not make the number false. It makes it a claim, as Chapter 18 explained. The only numbers in this book that we measured are hanumemAI's.

Does Mem0 really make no LLM call when searching?

Yes, in the open-source code we read. Its search is vector math, word statistics and addition. The same is true of hanumemAI. The model that answers the user is the application's own call, and it is not part of the memory layer.

What is the difference between an entity link and a typed edge?

An entity link says "these memories mention Sam". A typed edge says how two things relate: "Error 503 was caused by the March deploy". The first helps find memories about a name. The second lets a search follow a chain of reasons.

Carry this

  • Mem0's paper used two LLM calls per write and four verbs. Shipping Mem0 is ADD-only with one call, and its search uses no LLM.
  • HydraDB and Zep treat memory as a database: append-only, two clocks, provenance, typed edges.
  • Cheap writes leave contradictions for later. Smart writes cost several LLM calls. Every system chooses a point between the two.
  • hanumemAI takes the extractor and three-signal search from Mem0, and the two clocks and the supersede pointer from HydraDB and Zep.
  • Ideas built for multi-pass readers lost points on a single-pass small reader.

Check yourself

1. A session has 6 message pairs, and each pair yields 1 candidate fact. How many LLM calls does the write path make under the Mem0 paper's design, and under ADD-only?

Answer
paper      6 extraction calls + 6 x 1 decision calls = 12 calls
ADD-only   6 extraction calls                        =  6 calls

hanumemAI extracts once per session, so it makes 1 extraction call plus 1 profile update.

2. Why can a better embedding model not find memory m2 in the Error 503 example?

Answer

Embeddings place text by meaning. The question and m2 mean different things. They are linked only by a fact about the world, the cause. A relationship needs a stored link, not a better sense of meaning.

3. In an ADD-only store with no supersede, what goes wrong when a user moves house four times?

Answer

Five address facts stay open. They are nearly identical sentences, so their scores are nearly equal, and the top results may show an old address first. This is the dilution we saw in Chapter 12.