Part IV · Ninja
Chapter 20
hanumemAI: the hybrid design
One Python package, one SQLite file, and one fact followed from the moment Maya says it to the moment the assistant reads it.
On this page
In this chapter, we will learn how hanumemAI is built. It is the memory library that this book's ideas were tested in. We will see the whole design in one diagram, follow one fact through every step of it, and read the real code an application writes to use it.
What hanumemAI is
hanumemAI is a plug-and-play memory library for AI agents: one Python package, one SQLite file underneath, and any OpenAI-compatible model on top. OpenAI-compatible means a model server that accepts requests in the format of OpenAI's API. Most providers and local servers do.
A note on the name. The library is being renamed to hanumemAI. Until the rename ships, the package imports
as hmem, so every code example in this book uses from hmem import ....
It is called a hybrid because it joins two schools from Chapter 19. From Mem0 it takes a cheap write path and a search that uses no LLM. From HydraDB and Zep it takes two clocks on every row and the rule that nothing is overwritten.
The design in one diagram
WRITE one LLM call per add(), seconds are fine READ no LLM, tens of milliseconds
messages question
─► save the raw turns (indexed too) ─► find dates and names in it
─► fetch the 10 most similar stored facts ─► search by meaning ┐
─► extractor LLM writes rich, dated facts ─► search by words ├─► add the scores
with names, importance, category, ─► search by names ┘
source turns and "supersedes" ─► reserve places for rows on a named date
─► drop duplicates (hash, then similarity) ─► hop: the replaced fact behind each hit
─► insert, and close the rows that were replaced ─► a dated block, capped in tokens
The left side is the secretary from Chapter 6. The right side is the librarian from Chapter 8. The write side may take seconds, because nobody waits for it. The read side must be quick, because the user is waiting.
One fact, from Maya's mouth to the prompt
Let us follow one fact. The store already holds a card from January 2025: "Maya lives in London and
works as a nurse at St Thomas'." Call it card a1. (This walk uses the sample data of the
library's own documentation, where Maya is a nurse. In the rest of the book she is a data engineer. Only
the job differs.) On 2 March 2026 Maya writes:
Maya: Big news: I just moved from London to Paris for a job at a payments startup.
Step through the journey here, one step per click, and watch the data change. The widget cuts the journey into eleven small steps, and its sample values are illustrative. Below, we group the same journey into eight stages.
Stage 1. Save the raw turn
The message itself is stored and indexed, with its date. Facts are good keys for finding things. The original words are what the answer model likes to read. We keep both.
Stage 2. Fetch similar facts
The new message is turned into a vector, and the 10 most similar stored facts are fetched. Card
a1 about London is among them.
Stage 3. One extractor call
The extractor sees the date, the speakers, the profile so far, the last turns, the 10 similar facts and the new message. It returns strict JSON.
{"memory": [{
"text": "Maya moved from London to Paris in March 2026 for a job at a payments startup.",
"category": "semantic",
"importance": 9,
"event_date": "2026-03",
"entities": ["Maya", "London", "Paris"],
"sources": ["t1"],
"supersedes": ["a1"]
}]}
This output is an illustration of the format. Read it field by field. The text stands alone: it names
Maya, both cities and the month. Importance is 9 because the prompt rates life events such as a move at 9
or 10. sources points at the turn it came from. And
supersedes says that card a1 is now out of date.
Stage 4. Drop duplicates
The text is hashed. If the same hash is already stored, the fact is skipped. Then its vector is compared with the stored ones, and a fact that is 0.95 similar or more to an existing one is skipped too.
Stage 5. Insert, and close the old card
a1 "Maya lives in London and works as a nurse at St Thomas'."
valid_from 2025-01-14 valid_to 2026-03-02 superseded_by b7
b7 "Maya moved from London to Paris in March 2026 for a job at a payments startup."
valid_from 2026-03-02 valid_to (open) event_date 2026-03-01 source_ids [t1]
The extractor knew only the month, so the store keeps the first day of that month as the event date.
Nothing was deleted. Card a1 was closed with a date and a pointer. Once per session the
120-word profile is updated as well.
Stage 6. A question arrives
Maya: Any good bakeries near me?
The librarian runs three searches over Maya's open rows only, and adds the scores with the weights the experiments kept.
score = 0.8 x meaning + 0.1 x words + 0.1 x names (each scaled to 0..1)
Stage 7. Hop over the chain
Card b7 is a hit. The librarian looks for cards whose superseded_by points at
b7 and brings card a1 along as history. Closed rows never compete in the search. They only travel with an open
row that replaced them.
Stage 8. The dated prompt block
Profile so far: Maya is a nurse who moved from London to Paris in March 2026 for a payments startup; vegetarian; partner Sam (birthday June 18). (January 14, 2025) Maya lives in London and works as a nurse at St Thomas'. [outdated as of March 02, 2026] (March 02, 2026) [message] Maya: Big news: I just moved from London to Paris for a job at a payments startup. (March 02, 2026) Maya moved from London to Paris in March 2026 for a job at a payments startup.
The date in round brackets is the day the memory was recorded. This block goes into the assistant's prompt. From these few lines the model can answer "where do I live", "where did I live before" and "when did I move".
What came from where
| Piece | File | Borrowed from |
|---|---|---|
| ADD-only extraction, rich self-contained memories, entity links, hash dedup | extract.py, prompts.py | Mem0 v3 |
| Supersede with timestamps, provenance on every row, references resolved at write time | store.py, extract.py | HydraDB, Zep/Graphiti |
| Facts as keys and raw turns as values, time-aware questions | retrieve.py, timeparse.py | LongMemEval paper |
| Three signals added together, a damped entity boost (a name found in hundreds of memories counts for less), a score you can explain | retrieve.py | Mem0 scoring |
| Categories, importance, decay of old events, the user_id wall | everywhere | the class notes |
| Compressed vectors behind the same index | store.py | turbovec (TurboQuant) |
The module layout
Every module is under about 300 lines, with one exception, so a person can read the whole library in an afternoon.
| Module | Lines | Job |
|---|---|---|
config.py | 138 | Every knob and its default |
llm.py | 288 | Calls any OpenAI-compatible model, with a cache and a cost meter |
embed.py | 169 | Turns text into vectors, remote or local |
store.py | 412 | The box: SQLite rows, word index, name index, vector index |
extract.py | 220 | The secretary: the write path and the write policy |
retrieve.py | 193 | The librarian: the read path |
timeparse.py | 112 | Turns "last summer" into a date range |
prompts.py | 92 | The extractor, profile and answer prompts |
memory.py | 188 | The facade: the functions an application calls |
cli.py | 99 | The same functions from a terminal |
138 + 288 + 169 + 412 + 220 + 193 + 112 + 92 + 188 + 99 + 10 (__init__) = 1,921 lines
The store is the one module over the target, at 412 lines. The library's code imports numpy, openai and sqlite3, which comes with Python. It loads turbovec only when compressed vectors are switched on. fastembed, for local embeddings, is an optional extra.
The facade: what an application writes
A facade is one small front door to a larger system. An application never touches the
extractor or the store. It calls a handful of functions on one Memory object.
from hmem import Memory, Config
m = Memory(Config(db_path="app.sqlite")) # the API key is read from the environment
# remember: one LLM call, and add_async() never blocks the answer
m.add([{"role": "user", "name": "Maya",
"content": "I just moved from London to Paris for a payments job"}],
user_id="maya", observed_at="2026-03-02")
# recall: a dated block for the prompt, capped in tokens
block, rows = m.prompt_block("any good bakeries near me?", user_id="maya", budget_tokens=500)
paris = next(r for r in rows if r.kind == "fact" and "Paris" in r.text)
m.history(paris.id) # the chain: London (2025-01 .. 2026-03) then Paris
m.search("where did Maya live", user_id="maya", as_of="2025-06-01") # time travel
m.forget("maya") # expire events unused for 90 days
m.delete_user("maya") # the privacy wall is also the delete button
| Function | What it does |
|---|---|
add, add_async, flush | Write a session. The async form queues it and returns at once |
search | The ranked rows, with as_of and as_of_event for the two clocks |
prompt_block | The same rows as dated text, cut at a token budget |
get_all, history | What the user would see: every open fact, and the chain behind one |
update, delete | Edit one memory, or take it out of search. Both work by closing a row, so history stays. Neither erases, and both refuse another user's id |
forget, delete_user | Decay of old events, and full erasure of one user |
as_of asks what the store believed on a date. as_of_event asks what was true
in the world on that date.
Every write and every search takes a user_id. There is no call that searches all users.
history and delete take a memory id, and also need the user_id when
the store is sharded.
The knobs that mattered
All knobs live in config.py. The defaults are the values the experiments kept.
| Knob | Default | What the experiments said |
|---|---|---|
llm_model | qwen/qwen3.7-flash | A stronger extractor wrote fewer facts and lost 8 points on BEAM (E0001) |
embed_model | qwen/qwen3-embedding-4b | 2,560 dimensions. Scores level with the larger model, vectors 37 percent smaller, calls about 10 times faster (E0003) |
profile_summary | True | The running profile: +2.3 on LoCoMo (0014) |
k | 30 | 20 loses 4 points. 50 loses 0.3 and costs 67 percent more tokens (0008, 0018) |
read_budget_tokens | 4000 | BEAM went from 46 to 49 while tokens fell from 12.4K to 4.2K (B0002) |
w_dense / w_bm25 / w_entity | 0.8 / 0.1 / 0.1 | +1.7 evidence recall at equal accuracy (0013) |
time_quota | 10 | Rows on a date named in the question get reserved places (0003) |
procedural_always | False | 16 percent fewer tokens at equal score (0006) |
hop_superseded | True | Closed rows arrive only through the hop. When they competed in the main search, BEAM fell from 49 to 46 (B0003) |
atomic_facts | False | One attribute per fact. Fixes a real supersede bug at a cost of 2 to 3 points |
shards | 1 | More than 1 sends each user to one of N SQLite files |
vector_backend | numpy | turbovec is 8 times smaller at 98.4 percent recall@10 |
Offline mode
hanumemAI can run with no network at all. The embeddings can be made on the laptop's processor, and the extractor can point at any local server that speaks the OpenAI format.
uv sync --extra local # installs fastembed
from hmem import Memory, Config
m = Memory(Config(
embed_model="local/BAAI/bge-small-en-v1.5", # fastembed on the CPU, 384 dimensions
base_url="http://localhost:11434/v1", # a local OpenAI-compatible server
llm_model="your-local-model", # the name your local server gives its model
))
The local embedding model takes about 1.4 ms per text on a laptop. The tests use mock models only, so they run offline and cost nothing.
Try it yourself: the quickstart
The repository has a 32-line example that adds two sessions, asks a question, travels back in time and prints the history of one fact.
curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
uv run pytest # offline tests, no key needed
export OPENROUTER_API_KEY=sk-or-...
uv run python examples/quickstart.py
The part that shows both clocks is this:
for r in m.search("where does Maya live", user_id="maya", k=3, as_of="2025-06-01"):
print(f"{r.text} [valid {r.valid_from[:10]} .. {r.valid_to or 'now'}]")
Asked as of June 2025, the store answers London, because in June 2025 that was what it believed. The same library has a command line:
uv run hmem --db quickstart.sqlite facts maya
uv run hmem --db quickstart.sqlite search maya "where does Maya live" -k 5
uv run hmem --db quickstart.sqlite block maya "any bakeries near me?" --budget 500
Chapter 22 builds a full assistant on these calls.
Common questions
Why SQLite and not a vector database?
One user has a few hundred memories. Searching a few hundred vectors exactly takes well under a millisecond, so no special index is needed. SQLite also gives word search (FTS5), ordinary filters and one file that is easy to copy, back up and delete. At large scale hanumemAI splits users over many SQLite files and compresses the vectors, as Chapter 16 showed.
Why store the raw messages when we already extract facts?
Because they do different jobs. The fact is the resolved, dated, stand-alone version. The message is the exact words. When we removed facts whose source message was already in the block, we lost 7 points on a LongMemEval sample and 3 on BEAM (experiment R0001).
Does hanumemAI need Qwen?
No. Qwen through OpenRouter is the default because it is cheap and the experiments ran on it. Any model
behind an OpenAI-compatible endpoint works. Set HMEM_LLM, HMEM_EMBED,
HMEM_BASE_URL and HMEM_API_KEY.
What does one write cost?
One extraction call and one short profile update per session, plus the embedding calls. The integration guide in the repository puts this at well under one cent per session with flash-class models.
What if the extractor guesses "supersedes" wrongly?
The old row is closed, not destroyed. It still exists with its text, its dates and its sources. It can be
found with history() and with as_of searches. Nothing is lost, so a wrong close
can be repaired by writing the fact again. This is the reason hanumemAI never deletes on the write path.
Is hanumemAI an agent?
No. It does not answer the user and it does not plan. It hands your application a block of dated text. Your own model and your own prompt write the answer.
Carry this
- hanumemAI is one Python package over SQLite. It imports as
hmemuntil the rename ships. - Write: one extractor call, ADD-only, and a changed fact closes the old row with
valid_toandsuperseded_by. - Read: no LLM. Meaning, words and names are added with weights 0.8, 0.1 and 0.1, and the replaced fact travels with the hit.
- The application calls a facade:
add,search,prompt_block,history,forget,delete_user. Every call takes a user_id.
Check yourself
1. After Maya's move, how many rows about where she lives are in the store, and how many of them can win a normal search?
Answer
Two rows are stored. One is open (Paris) and one is closed (London).
Only the open row competes in a normal search. The closed row is brought along as history by the hop, or
found with an as_of search.
2. A row scores 0.70 on meaning, 0.20 on words and 1.00 on names, all scaled to 0..1. What is its fused score?
Answer
0.8 x 0.70 + 0.1 x 0.20 + 0.1 x 1.00 = 0.56 + 0.02 + 0.10 = 0.68
3. Why does the write path fetch the 10 most similar stored facts before calling the extractor?
Answer
For two reasons. The extractor can skip a fact that is already stored
with the same details. And it can name the stored facts that the new fact replaces, which is how
supersedes gets filled.