Carebun Interactive/Memory Memory labCheat sheet

Part IV · Ninja

Chapter 20

hanumemAI: the hybrid design

One Python package, one SQLite file, and one fact followed from the moment Maya says it to the moment the assistant reads it.

13 min read · interactive

On this page
  1. What hanumemAI is
  2. The design in one diagram
  3. One fact, from Maya's mouth to the prompt
  4. What came from where
  5. The module layout
  6. The facade: what an application writes
  7. The knobs that mattered
  8. Offline mode
  9. Try it yourself: the quickstart
  10. Common questions
  11. Carry this
  12. Check yourself

In this chapter, we will learn how hanumemAI is built. It is the memory library that this book's ideas were tested in. We will see the whole design in one diagram, follow one fact through every step of it, and read the real code an application writes to use it.

What hanumemAI is

hanumemAI is a plug-and-play memory library for AI agents: one Python package, one SQLite file underneath, and any OpenAI-compatible model on top. OpenAI-compatible means a model server that accepts requests in the format of OpenAI's API. Most providers and local servers do.

A note on the name. The library is being renamed to hanumemAI. Until the rename ships, the package imports as hmem, so every code example in this book uses from hmem import ....

It is called a hybrid because it joins two schools from Chapter 19. From Mem0 it takes a cheap write path and a search that uses no LLM. From HydraDB and Zep it takes two clocks on every row and the rule that nothing is overwritten.

The design in one diagram

WRITE  one LLM call per add(), seconds are fine     READ  no LLM, tens of milliseconds

messages                                             question
  ─► save the raw turns (indexed too)                  ─► find dates and names in it
  ─► fetch the 10 most similar stored facts            ─► search by meaning  ┐
  ─► extractor LLM writes rich, dated facts            ─► search by words    ├─► add the scores
     with names, importance, category,                 ─► search by names    ┘
     source turns and "supersedes"                     ─► reserve places for rows on a named date
  ─► drop duplicates (hash, then similarity)           ─► hop: the replaced fact behind each hit
  ─► insert, and close the rows that were replaced     ─► a dated block, capped in tokens

The left side is the secretary from Chapter 6. The right side is the librarian from Chapter 8. The write side may take seconds, because nobody waits for it. The read side must be quick, because the user is waiting.

One fact, from Maya's mouth to the prompt

Let us follow one fact. The store already holds a card from January 2025: "Maya lives in London and works as a nurse at St Thomas'." Call it card a1. (This walk uses the sample data of the library's own documentation, where Maya is a nurse. In the rest of the book she is a data engineer. Only the job differs.) On 2 March 2026 Maya writes:

Maya:  Big news: I just moved from London to Paris for a job at a payments startup.

Step through the journey here, one step per click, and watch the data change. The widget cuts the journey into eleven small steps, and its sample values are illustrative. Below, we group the same journey into eight stages.

Stage 1. Save the raw turn

The message itself is stored and indexed, with its date. Facts are good keys for finding things. The original words are what the answer model likes to read. We keep both.

Stage 2. Fetch similar facts

The new message is turned into a vector, and the 10 most similar stored facts are fetched. Card a1 about London is among them.

Stage 3. One extractor call

The extractor sees the date, the speakers, the profile so far, the last turns, the 10 similar facts and the new message. It returns strict JSON.

{"memory": [{
  "text": "Maya moved from London to Paris in March 2026 for a job at a payments startup.",
  "category": "semantic",
  "importance": 9,
  "event_date": "2026-03",
  "entities": ["Maya", "London", "Paris"],
  "sources": ["t1"],
  "supersedes": ["a1"]
}]}

This output is an illustration of the format. Read it field by field. The text stands alone: it names Maya, both cities and the month. Importance is 9 because the prompt rates life events such as a move at 9 or 10. sources points at the turn it came from. And supersedes says that card a1 is now out of date.

Stage 4. Drop duplicates

The text is hashed. If the same hash is already stored, the fact is skipped. Then its vector is compared with the stored ones, and a fact that is 0.95 similar or more to an existing one is skipped too.

Stage 5. Insert, and close the old card

a1  "Maya lives in London and works as a nurse at St Thomas'."
    valid_from 2025-01-14    valid_to 2026-03-02    superseded_by b7

b7  "Maya moved from London to Paris in March 2026 for a job at a payments startup."
    valid_from 2026-03-02    valid_to (open)        event_date 2026-03-01 source_ids [t1]

The extractor knew only the month, so the store keeps the first day of that month as the event date. Nothing was deleted. Card a1 was closed with a date and a pointer. Once per session the 120-word profile is updated as well.

Stage 6. A question arrives

Maya:  Any good bakeries near me?

The librarian runs three searches over Maya's open rows only, and adds the scores with the weights the experiments kept.

score = 0.8 x meaning  +  0.1 x words  +  0.1 x names       (each scaled to 0..1)

Stage 7. Hop over the chain

Card b7 is a hit. The librarian looks for cards whose superseded_by points at b7 and brings card a1 along as history. Closed rows never compete in the search. They only travel with an open row that replaced them.

Stage 8. The dated prompt block

Profile so far: Maya is a nurse who moved from London to Paris in March 2026 for a payments
startup; vegetarian; partner Sam (birthday June 18).

(January 14, 2025) Maya lives in London and works as a nurse at St Thomas'. [outdated as of March 02, 2026]
(March 02, 2026) [message] Maya: Big news: I just moved from London to Paris for a job at a payments startup.
(March 02, 2026) Maya moved from London to Paris in March 2026 for a job at a payments startup.

The date in round brackets is the day the memory was recorded. This block goes into the assistant's prompt. From these few lines the model can answer "where do I live", "where did I live before" and "when did I move".

What came from where

PieceFileBorrowed from
ADD-only extraction, rich self-contained memories, entity links, hash dedupextract.py, prompts.pyMem0 v3
Supersede with timestamps, provenance on every row, references resolved at write timestore.py, extract.pyHydraDB, Zep/Graphiti
Facts as keys and raw turns as values, time-aware questionsretrieve.py, timeparse.pyLongMemEval paper
Three signals added together, a damped entity boost (a name found in hundreds of memories counts for less), a score you can explainretrieve.pyMem0 scoring
Categories, importance, decay of old events, the user_id walleverywherethe class notes
Compressed vectors behind the same indexstore.pyturbovec (TurboQuant)

The module layout

Every module is under about 300 lines, with one exception, so a person can read the whole library in an afternoon.

ModuleLinesJob
config.py138Every knob and its default
llm.py288Calls any OpenAI-compatible model, with a cache and a cost meter
embed.py169Turns text into vectors, remote or local
store.py412The box: SQLite rows, word index, name index, vector index
extract.py220The secretary: the write path and the write policy
retrieve.py193The librarian: the read path
timeparse.py112Turns "last summer" into a date range
prompts.py92The extractor, profile and answer prompts
memory.py188The facade: the functions an application calls
cli.py99The same functions from a terminal
138 + 288 + 169 + 412 + 220 + 193 + 112 + 92 + 188 + 99 + 10 (__init__) = 1,921 lines

The store is the one module over the target, at 412 lines. The library's code imports numpy, openai and sqlite3, which comes with Python. It loads turbovec only when compressed vectors are switched on. fastembed, for local embeddings, is an optional extra.

The facade: what an application writes

A facade is one small front door to a larger system. An application never touches the extractor or the store. It calls a handful of functions on one Memory object.

from hmem import Memory, Config

m = Memory(Config(db_path="app.sqlite"))       # the API key is read from the environment

# remember: one LLM call, and add_async() never blocks the answer
m.add([{"role": "user", "name": "Maya",
        "content": "I just moved from London to Paris for a payments job"}],
      user_id="maya", observed_at="2026-03-02")

# recall: a dated block for the prompt, capped in tokens
block, rows = m.prompt_block("any good bakeries near me?", user_id="maya", budget_tokens=500)

paris = next(r for r in rows if r.kind == "fact" and "Paris" in r.text)
m.history(paris.id)                   # the chain: London (2025-01 .. 2026-03) then Paris
m.search("where did Maya live", user_id="maya", as_of="2025-06-01")    # time travel
m.forget("maya")                      # expire events unused for 90 days
m.delete_user("maya")                 # the privacy wall is also the delete button
FunctionWhat it does
add, add_async, flushWrite a session. The async form queues it and returns at once
searchThe ranked rows, with as_of and as_of_event for the two clocks
prompt_blockThe same rows as dated text, cut at a token budget
get_all, historyWhat the user would see: every open fact, and the chain behind one
update, deleteEdit one memory, or take it out of search. Both work by closing a row, so history stays. Neither erases, and both refuse another user's id
forget, delete_userDecay of old events, and full erasure of one user

as_of asks what the store believed on a date. as_of_event asks what was true in the world on that date.

Every write and every search takes a user_id. There is no call that searches all users. history and delete take a memory id, and also need the user_id when the store is sharded.

The knobs that mattered

All knobs live in config.py. The defaults are the values the experiments kept.

KnobDefaultWhat the experiments said
llm_modelqwen/qwen3.7-flashA stronger extractor wrote fewer facts and lost 8 points on BEAM (E0001)
embed_modelqwen/qwen3-embedding-4b2,560 dimensions. Scores level with the larger model, vectors 37 percent smaller, calls about 10 times faster (E0003)
profile_summaryTrueThe running profile: +2.3 on LoCoMo (0014)
k3020 loses 4 points. 50 loses 0.3 and costs 67 percent more tokens (0008, 0018)
read_budget_tokens4000BEAM went from 46 to 49 while tokens fell from 12.4K to 4.2K (B0002)
w_dense / w_bm25 / w_entity0.8 / 0.1 / 0.1+1.7 evidence recall at equal accuracy (0013)
time_quota10Rows on a date named in the question get reserved places (0003)
procedural_alwaysFalse16 percent fewer tokens at equal score (0006)
hop_supersededTrueClosed rows arrive only through the hop. When they competed in the main search, BEAM fell from 49 to 46 (B0003)
atomic_factsFalseOne attribute per fact. Fixes a real supersede bug at a cost of 2 to 3 points
shards1More than 1 sends each user to one of N SQLite files
vector_backendnumpyturbovec is 8 times smaller at 98.4 percent recall@10

Offline mode

hanumemAI can run with no network at all. The embeddings can be made on the laptop's processor, and the extractor can point at any local server that speaks the OpenAI format.

uv sync --extra local        # installs fastembed
from hmem import Memory, Config

m = Memory(Config(
    embed_model="local/BAAI/bge-small-en-v1.5",   # fastembed on the CPU, 384 dimensions
    base_url="http://localhost:11434/v1",         # a local OpenAI-compatible server
    llm_model="your-local-model",                 # the name your local server gives its model
))

The local embedding model takes about 1.4 ms per text on a laptop. The tests use mock models only, so they run offline and cost nothing.

Try it yourself: the quickstart

The repository has a 32-line example that adds two sessions, asks a question, travels back in time and prints the history of one fact.

curl -LsSf https://astral.sh/uv/install.sh | sh
uv sync
uv run pytest                                  # offline tests, no key needed
export OPENROUTER_API_KEY=sk-or-...
uv run python examples/quickstart.py

The part that shows both clocks is this:

for r in m.search("where does Maya live", user_id="maya", k=3, as_of="2025-06-01"):
    print(f"{r.text}  [valid {r.valid_from[:10]} .. {r.valid_to or 'now'}]")

Asked as of June 2025, the store answers London, because in June 2025 that was what it believed. The same library has a command line:

uv run hmem --db quickstart.sqlite facts maya
uv run hmem --db quickstart.sqlite search maya "where does Maya live" -k 5
uv run hmem --db quickstart.sqlite block maya "any bakeries near me?" --budget 500

Chapter 22 builds a full assistant on these calls.

Common questions

Why SQLite and not a vector database?

One user has a few hundred memories. Searching a few hundred vectors exactly takes well under a millisecond, so no special index is needed. SQLite also gives word search (FTS5), ordinary filters and one file that is easy to copy, back up and delete. At large scale hanumemAI splits users over many SQLite files and compresses the vectors, as Chapter 16 showed.

Why store the raw messages when we already extract facts?

Because they do different jobs. The fact is the resolved, dated, stand-alone version. The message is the exact words. When we removed facts whose source message was already in the block, we lost 7 points on a LongMemEval sample and 3 on BEAM (experiment R0001).

Does hanumemAI need Qwen?

No. Qwen through OpenRouter is the default because it is cheap and the experiments ran on it. Any model behind an OpenAI-compatible endpoint works. Set HMEM_LLM, HMEM_EMBED, HMEM_BASE_URL and HMEM_API_KEY.

What does one write cost?

One extraction call and one short profile update per session, plus the embedding calls. The integration guide in the repository puts this at well under one cent per session with flash-class models.

What if the extractor guesses "supersedes" wrongly?

The old row is closed, not destroyed. It still exists with its text, its dates and its sources. It can be found with history() and with as_of searches. Nothing is lost, so a wrong close can be repaired by writing the fact again. This is the reason hanumemAI never deletes on the write path.

Is hanumemAI an agent?

No. It does not answer the user and it does not plan. It hands your application a block of dated text. Your own model and your own prompt write the answer.

Carry this

  • hanumemAI is one Python package over SQLite. It imports as hmem until the rename ships.
  • Write: one extractor call, ADD-only, and a changed fact closes the old row with valid_to and superseded_by.
  • Read: no LLM. Meaning, words and names are added with weights 0.8, 0.1 and 0.1, and the replaced fact travels with the hit.
  • The application calls a facade: add, search, prompt_block, history, forget, delete_user. Every call takes a user_id.

Check yourself

1. After Maya's move, how many rows about where she lives are in the store, and how many of them can win a normal search?

Answer

Two rows are stored. One is open (Paris) and one is closed (London). Only the open row competes in a normal search. The closed row is brought along as history by the hop, or found with an as_of search.

2. A row scores 0.70 on meaning, 0.20 on words and 1.00 on names, all scaled to 0..1. What is its fused score?

Answer
0.8 x 0.70 + 0.1 x 0.20 + 0.1 x 1.00 = 0.56 + 0.02 + 0.10 = 0.68

3. Why does the write path fetch the 10 most similar stored facts before calling the extractor?

Answer

For two reasons. The extractor can skip a fact that is already stored with the same details. And it can name the stored facts that the new fact replaces, which is how supersedes gets filled.