Part II · Apprentice
Chapter 7
Where memories live
Open the box and look at one card, field by field. Learn what an embedding is with a map and two arrows, and count the bytes all the way to 664 GB.
On this page
In this chapter, we will learn where a memory is kept and what it looks like when it is stored. We will read one row field by field, understand embeddings without any mathematics beyond multiplication, and count the bytes of a store that serves a million users.
One row in the memory store
A memory is one row in a database: a short sentence, plus the fields that say whose it is, what kind it is, where it came from and when.
id 7f3a... user_id maya text "Maya is vegetarian." this one ~4 tok, ~25 tok on average category semantic semantic | episodic | procedural embedding [0.031, -0.118, ...] 768 floats = 3,072 B created_at 2026-03-02T09:14:00Z updated_at 2026-03-02T09:14:00Z last_accessed 2026-06-09T18:02:00Z source session_18241 total ~3.3 KB (sized at 25 tok, ~100 B)
Let us go through the fields one at a time. Each exists because something breaks without it.
| Field | What it holds | What breaks without it |
|---|---|---|
id | a unique name for the row | we cannot update or delete one memory |
user_id | whose memory this is | Tom sees Maya's life |
text | the fact, as a sentence | nothing to show the model |
category | semantic, episodic or procedural | we cannot let events fade and keep facts |
embedding | the sentence's place on the map | we cannot search by meaning |
created_at | when the row was written; never changes | no history, no audit |
updated_at | when the text last changed | we cannot tell a fresh fact from an old one |
last_accessed | when the row was last retrieved | no recency score, no decay |
source | the session it came from | a wrong card cannot be traced |
Two of the dates are easy to confuse. created_at is stamped once and never moves.
last_accessed is restamped every time the row is retrieved. For a fact that is written once
and read often, the two drift months apart. That gap is what the recency score in
Chapter 9 is built on.
What is an embedding?
An embedding is a list of numbers that gives a sentence its place on the map of meaning.
A small model, called an embedding model, reads a sentence and returns the list. The class uses a model that returns 768 numbers. Another word for such a list is a vector.
"Maya is vegetarian." ---> [ 0.031, -0.118, 0.224, ... 768 numbers ] "Maya does not eat meat." ---> [ 0.029, -0.109, 0.231, ... ] nearly the same place "Tom lives in Austin." ---> [-0.204, 0.077, -0.015, ... ] somewhere else
A place on a paper map needs 2 numbers, one for east and west and one for north and south. A place in a room needs 3. The map of meaning needs 768, because meaning has many more directions than a room: food or not food, past or future, person or place, and hundreds more that have no names.
Nobody can picture 768 directions. We do not need to. The rule that matters works the same in 2 directions as in 768: sentences with similar meaning get nearby places.
Cosine similarity: the angle between two arrows
Cosine similarity is a score for how closely two vectors point in the same direction.
Draw each sentence as an arrow from the centre of the map to its place. Now look at the angle between two arrows.
same direction angle 0 degrees cosine 1.0 same meaning close angle 45 degrees cosine 0.71 related at right angles angle 90 degrees cosine 0.0 unrelated opposite angle 180 degrees cosine -1.0 opposite
Only the direction counts, not the length of the arrow. A long sentence and a short sentence about the same thing point the same way.
Drag the question arrow around the map below. Watch the scores change and see which three of Maya's memories win.
The vector database
A vector database is a database that can answer one special question quickly: which stored vectors are nearest to this one?
An ordinary database finds rows where a field equals something. A vector database finds rows whose embedding is closest to the question's embedding. It also stores the other fields of the row, which it calls the payload.
query vector ---> filtered search ---> one collection, ---> maya's facts only
filter: user_id = maya ALL users' memories, top 3, 1 to 10 ms
payload-partitioned
One collection, not one per user
All users share one collection, and every search carries a filter on
user_id. This pattern is called multitenancy.
The other choice, one collection per user, sounds tidier. With a million users it means a million
collections to create, watch and back up. So the class design keeps one collection and puts an index on
the user_id field. An index is a lookup table: it finds all rows of one user without reading
the others. The filter removes other users' rows before any scoring happens. That is what
"payload-partitioned" means in the diagram: one collection, divided by a payload field.
Which database?
The right database depends on how many rows one search must cover. Three options are common.
| Option | What it is | When it fits |
|---|---|---|
| Qdrant | a dedicated vector database | a large shared store; the class experiments use it |
| pgvector | vector search inside PostgreSQL | you already run Postgres and want one system |
| SQLite + numpy | a plain file, exact search in memory | what hanumemAI uses; one user has only hundreds of rows |
The last row surprises people. A memory store is not one huge search. It is a very small search, done a million times. One user has about 200 memories. Exact search means the question is compared with every one of them, with no shortcut. In hanumemAI's scale test on a laptop, exact search over 10,000 vectors of 4,096 numbers took 0.17 ms. A search over 200 is 50 times smaller.
Counting the bytes
One memory costs about 3.3 KB, and almost all of it is the vector.
A float32 is the usual way a computer stores a number with decimals. It takes 4 bytes. We write B for bytes.
one number in the vector (a float32) 4 bytes vector 768 numbers x 4 B = 3,072 B text about 25 tokens = ~100 B payload ids, dates, category = ~150 B one memory 3,072 + 100 + 150 = 3,322 B = about 3.3 KB vector share 3,072 / 3,322 = 92 percent vector / text 3,072 / 100 = about 30 times
The sentence we care about is 100 bytes. Its address on the map is 30 times larger. Now multiply.
average user 200 memories one user 200 x 3,322 B = 664,400 B = about 0.66 MB all users 1,000,000 x 200 = 200,000,000 memories total 200,000,000 x 3,322 B = 664,400,000,000 B = about 664 GB
One user is trivial. A million users is a real database. Change the numbers below to size your own store. Try switching the vectors from float32 to int8, a whole number stored in 1 byte, and watch the total fall.
How that shrinking works, and what it costs in accuracy, is the subject of Chapter 16.
Profile or collection?
A profile is one document per user that is rewritten as we learn more. A collection is many separate facts, searched on demand. They are the two basic shapes a memory can take.
PROFILE: one mutable document per user COLLECTION: many discrete facts +----------------------------+ +------------------------------+ | Maya: vegetarian, peanut | | "Maya is vegetarian." | | allergy, 2 kids (6, 9), | | "Maya has a peanut allergy." | | lives in Paris, partner | | "Maya lives in Paris." | | Sam (b-day Jun 18), ... | | "...197 more rows..." | +----------------------------+ +------------------------------+ inject the WHOLE thing every turn retrieve top 3 per question ~300 to 1,500 tokens, always ~75 tokens, flat forever trivial to view and edit needs retrieval + scoring rewrites lose detail over time grows without bound
The profile is a page of notes about a person. It is simple, and it is always on the desk. But each rewrite squeezes it, and details fall out. The collection keeps every detail, but something must find the right cards.
Try it yourself: one collection, two users
The class experiment memory-store puts 12 memories for two users into one Qdrant
collection.
ollama pull nomic-embed-text
docker run -p 6333:6333 qdrant/qdrant
cd memory_classnotes/experiments/memory-store
pip install -r requirements.txt
python main.py
Maya and Tom ask the same question. Each search carries its own filter.
(b) Maya asks: "what should I cook for dinner tonight" (filter user_id=maya, top 3)
score 0.434 [semantic] Maya is vegetarian.
score 0.386 [semantic] Maya is training for the Berlin marathon in September.
score 0.357 [semantic] Maya's partner Sam has a birthday on June 18.
her diet memory ranks first: the assistant can plan dinner around it.
(c) Tom asks the SAME question (filter user_id=tom, top 3)
score 0.394 [semantic] Tom has a golden retriever named Biscuit.
score 0.365 [semantic] Tom is allergic to shellfish.
score 0.324 [semantic] Tom lives in Austin.
his shellfish allergy surfaces, and none of maya's 9 facts can appear:
the filter excludes them before scoring.
The word "vegetarian" does not appear in Maya's question. The map of meaning put "cook for dinner" near "vegetarian" anyway. That is what embeddings buy us.
Now the same search with the filter forgotten.
(d) The bug: same question, but the code forgot the filter (top 3, all users)
score 0.434 user=maya Maya is vegetarian.
score 0.394 user=tom Tom has a golden retriever named Biscuit.
score 0.386 user=maya Maya is training for the Berlin marathon in September.
cross-user leak: the filter IS the privacy boundary.
Whoever asked, they now see another user's life in their results.
Nothing crashed. No warning appeared. The search simply returned the nearest cards from the whole box. Finally, the timings.
(e) Timing and scale
embed the query: 49.1 ms filtered search over 12 memories: 16.1 ms
These are laptop numbers, with one network round trip per step. The class budgets 10 to 30 ms to embed and 1 to 10 ms to search in production, where the services sit side by side and the index is warm.
Common questions
Do I need to understand the 768 numbers?
No. Nobody can read them, and no single number has a meaning on its own. Treat the vector as an address. All we ever do with it is compare it with other addresses.
Can I change the embedding model later?
Yes, but every stored vector must be computed again. Two models draw two different maps, and an address from one map means nothing on the other. The question and the memories must always be embedded by the same model. In hanumemAI the change cost more than new vectors. The extractor is shown the stored facts nearest to each new session, and a new map picks different neighbours. So every extraction call had to be paid again too (experiment E0003).
Are more dimensions better?
More dimensions can hold finer differences, but they cost more bytes and the gain flattens. hanumemAI moved from a 4,096 number model to a 2,560 number one (E0003). The scores stayed inside the noise, and the embedding calls became about 10 times faster.
smaller by (4,096 - 2,560) / 4,096 = about 37 percent
Why not search the text with keywords and skip vectors?
Keywords miss meaning. "What should I cook tonight" shares no word with "Maya is vegetarian". But keywords are better at exact names and rare words, so good systems use both. Chapter 9 shows how they are combined.
Is the vector personal data?
Yes. Text can be partly reconstructed from its embedding by what is called an inversion attack. When a user asks to be deleted, the vectors must go too, along with caches and logs.
Is a vector database required at all?
Not at small scale. With 200 memories per user, exact search over a plain file is fast and simple, which is why hanumemAI starts with SQLite. A dedicated vector database earns its place when one search must cover very many rows, or when many machines share one store.
Carry this
- A memory is a row: text, user_id, category, embedding, three dates and a source.
- An embedding is an address on a map of meaning. Cosine similarity is the angle between two arrows: 1 is the same direction, 0 is unrelated.
- One collection for all users, with a
user_idfilter on every search. The filter is the privacy wall. - 768 x 4 = 3,072 bytes of vector plus about 250 bytes of text and payload is about 3.3 KB. Times 200 memories, times a million users: 664 GB.
- Keep the collection and add a short profile of about 120 words. It was worth 2.3 points.
Check yourself
1. An embedding model returns 1,024 numbers, each stored as a float32. Text and payload take 250 bytes. How large is one memory, and how large is the store for 500,000 users with 100 memories each?
Answer
vector 1,024 x 4 B = 4,096 B one memory 4,096 + 250 = 4,346 B memories 500,000 x 100 = 50,000,000 total 50,000,000 x 4,346 B = 217,300,000,000 B = about 217 GB
2. Two sentences have a cosine similarity of 0. What does that say about them?
Answer
Their arrows are at right angles, so the sentences are unrelated in meaning. It does not mean they are opposites. Opposite arrows would score close to -1.
3. In the experiment with the missing filter, why did no error appear?
Answer
Because a search without a filter is a perfectly valid search. The database returned the nearest vectors in the whole collection, as asked. Isolation is not something the database enforces by itself. Our code must send the filter every time, which is why later chapters make it structural.