Part II · Apprentice
Chapter 9
Scoring beyond similarity
Similarity finds the topic. It cannot tell a birthday from printer ink. We add recency and importance, then search by meaning, words and names at once.
On this page
In this chapter, we will learn why the most similar memory is not always the right one. We will build a score from three parts, work one row out by hand, and then give the librarian three ways to search instead of one. At the end we will see which of these ideas survived measurement.
Why similarity alone is not enough
Similarity measures whether a memory is about the same topic as the question. It does not measure whether the memory is current, or whether it matters.
In Chapter 8 the librarian ranked cards by cosine similarity only. That fails in two ways.
- It cannot see time. "Maya lives in London" and "Maya lives in Paris" are almost the same sentence. The class simulation in Chapter 12 puts such sentences within 0.01 of each other. Which one ranks first is luck.
- It cannot see weight. "Maya has a serious peanut allergy" and "Maya asked for a pasta recipe yesterday" are both about food. Only one of them can send a child to hospital.
Three parts of one score
The retrieval score adds three components: relevance, recency and importance. This recipe comes from a research project called Generative Agents, in which simulated people had to remember their own lives.
score = relevance + recency + importance (each min-max normalized to [0,1]) relevance cosine(query, memory) recency 0.995 ^ hours_since_last_access exponential decay per hour importance LLM-rated 1 to 10 at write time
Relevance
Relevance is the cosine similarity between the question and the memory. It is the score we already know. It keeps the answer on topic.
Recency
Recency is a score that starts at 1 and shrinks a little with every hour since the memory was last used.
Each hour the score is multiplied by 0.995, so it loses half a percent. That sounds like nothing. It compounds.
after 1 hour 0.995 ^ 1 = 0.995 after 26 hours 0.995 ^ 26 = 0.878 after 120 hours 0.995 ^ 120 = 0.548 5 days: about half after 300 hours 0.995 ^ 300 = 0.222 after 1,000 hours 0.995 ^ 1000 = 0.0067 6 weeks: almost nothing after 2,000 hours 0.995 ^ 2000 = 0.000044 83 days: gone
Note the word accessed. The clock restarts each time a memory is retrieved. A fact that keeps being useful stays fresh. A fact nobody asks for fades.
Importance
Importance is a number from 1 to 10 that says how much a fact matters, given once, when the
card is written. The extractor rates it in the same call that writes the fact. Printer ink gets a
1. A peanut allergy should get a 9 or 10. The small class prompt in
Chapter 6 asks only for a fact and a category. A
production prompt, such as hanumemAI's, adds an importance field to the same JSON.
It is decided at write time because that is when a model is reading the conversation anyway. At read time there is no model call to spare.
Min-max normalisation
Min-max normalisation rescales a set of numbers so that the smallest becomes 0 and the largest becomes 1.
We need it because the three parts live on different scales. Cosine scores sit between about 0.3 and 0.7. Importance runs from 1 to 10. Added raw, importance would drown the others.
normalized = (value - smallest) / (largest - smallest) cosine scores 0.35, 0.40, 0.69 smallest 0.35, largest 0.69 0.35 -> (0.35 - 0.35) / 0.34 = 0.00 0.40 -> (0.40 - 0.35) / 0.34 = 0.15 0.69 -> (0.69 - 0.35) / 0.34 = 1.00
Try it yourself: a gift for Sam
The class experiment memory-scoring ranks eight of Maya's memories three ways. It needs
only the embedding model, because the importance numbers are fixed by hand in this experiment.
ollama pull nomic-embed-text
cd memory_classnotes/experiments/memory-scoring
pip install -r requirements.txt
python main.py
First, relevance alone.
(a) Ranking by RELEVANCE alone: cosine(query vector, memory vector)
query: "What should I get Sam as a birthday present?"
rank rel fact
1 0.6942 Maya's partner Sam has a birthday on June 18.
2 0.6627 Maya said Sam loves vinyl records and old jazz.
3 0.4227 Maya asked for a pasta recipe yesterday.
4 0.4058 Maya has two kids, aged 6 and 9.
5 0.3880 Maya mentioned the printer is out of ink.
6 0.3744 Maya likes hiking on weekends.
7 0.3721 Maya is vegetarian.
8 0.3475 Maya works as a data engineer at a fintech company.
The first two are right. The third is a pasta recipe. Relevance found the topic and then ran out of ideas. Now recency alone.
(b) Ranking by RECENCY alone: recency = 0.995 ^ hours_since_last_access
rank recency hours fact
1 0.995000 1 Maya mentioned the printer is out of ink.
2 0.877809 26 Maya asked for a pasta recipe yesterday.
3 0.547986 120 Maya is vegetarian.
4 0.222292 300 Maya has two kids, aged 6 and 9.
5 0.029933 700 Maya likes hiking on weekends.
6 0.000543 1500 Maya works as a data engineer at a fintech company.
7 0.000046 1992 Maya said Sam loves vinyl records and old jazz.
8 0.000044 2000 Maya's partner Sam has a birthday on June 18.
Printer ink wins. The birthday is dead last. Recency alone retrieves trivia. Now all three together.
(c) Combined score: min-max normalize each component to [0, 1] across the
8 memories, then total = relevance + recency + importance.
fact rel rec imp total
Maya's partner Sam has a birthday on June 18. 1.00 0.00 1.00 2.00 <- inject
Maya said Sam loves vinyl records and old jazz. 0.91 0.00 0.71 1.62 <- inject
Maya is vegetarian. 0.07 0.55 0.71 1.33 <- inject
Maya has two kids, aged 6 and 9. 0.17 0.22 0.86 1.25
Maya asked for a pasta recipe yesterday. 0.22 0.88 0.14 1.24
Maya mentioned the printer is out of ink. 0.12 1.00 0.00 1.12
Maya works as a data engineer at a fintech company. 0.00 0.00 0.57 0.57
Maya likes hiking on weekends. 0.08 0.03 0.43 0.54
The birthday, the vinyl records and the vegetarian diet. Those are the three facts a gift question needs, and the last one matters if the gift is a dinner. The printer note has the best recency of all and still falls to rank 6.
One row, worked out by hand
To check a score, normalise each part and add the three results. Let us do it for the top row: "Maya's partner Sam has a birthday on June 18."
relevance cosine = 0.6942 smallest 0.3475, largest 0.6942 among the 8
(0.6942 - 0.3475) / (0.6942 - 0.3475) = 1.00
recency 0.995 ^ 2000 = 0.000044 (2,000 hours is about 83 days)
smallest 0.000044, largest 0.995000
(0.000044 - 0.000044) / (0.995000 - 0.000044) = 0.00
importance 8 on the scale of 1 to 10 smallest 1, largest 8 among the 8
(8 - 1) / (8 - 1) = 1.00
total 1.00 + 0.00 + 1.00 = 2.00
Recency gave this memory nothing at all. Relevance and importance rescued it. That is why one component is never enough.
Now take the controls. Give each of the three parts a weight, and watch the eight memories reorder. Set recency high and see the printer climb. Then change the decay rate and see how quickly old memories vanish.
Three ways to search
A memory can be found by its meaning, by its exact words or by the names in it. Relevance so far has meant one thing: cosine similarity between vectors. That is search by meaning. It has blind spots, and two older kinds of search cover them.
By meaning: dense vectors
Dense search compares embeddings, so it finds memories that mean the same thing even when the words differ. "What should I cook tonight" finds "Maya is vegetarian". The vectors are called dense because every one of their numbers is filled in.
Its weakness is exactness. To an embedding, "Poppy" and "Daisy" are both just pet names and land in almost the same place. So do "March 14" and "March 15".
By exact words: BM25
BM25 is a keyword search that scores a memory by the words it shares with the question, and counts rare words more than common ones.
The idea is simple. If the question contains "the", a match tells us nothing, because every card contains "the". If it contains "Miso", a match tells us almost everything, because one card in two hundred has that word. BM25 also stops a long card from winning just because it has more words.
question: "what is the name of the cat?" word cards that contain it what a match is worth the nearly all almost nothing name a few a little cat one or two a lot
By names: the entity boost
An entity is a named thing: a person, a place, a company, a pet. At write time the extractor lists the entities on each card, in one more field of its JSON. At read time, names in the question are looked up in that list, and every card that mentions them gets a boost.
There is a catch. If Sam appears on 60 cards, a match on "Sam" says very little. So the boost is damped: the more cards share a name, the less each one gains.
| Signal | Good at | Blind to |
|---|---|---|
| Dense (meaning) | paraphrase, topic, "dinner" finds "vegetarian" | exact names, numbers, dates |
| BM25 (words) | rare words, codes, exact phrases | any rewording |
| Entity (names) | everything about one person or place | questions with no name in them |
Using several signals together is called hybrid search.
Fusing the signals
Fusion is the step that turns several rankings into one. There are two common ways.
A weighted sum
Normalise each signal to the range 0 to 1, multiply each by a weight, and add.
score = 0.8 x dense + 0.1 x bm25 + 0.1 x entity a card with dense 0.70, bm25 1.00, entity 1.00 score 0.8 x 0.70 + 0.1 x 1.00 + 0.1 x 1.00 = 0.56 + 0.10 + 0.10 = 0.76 a card with dense 0.80, bm25 0.00, entity 0.00 score 0.8 x 0.80 = 0.64
The first card is less similar in meaning, but it holds the exact word and the exact name. It wins.
Reciprocal rank fusion
Reciprocal rank fusion (RRF) ignores the scores and uses only the position of each card in each list. A card earns 1 / (60 + rank) from every list it appears in, and the points are added. The 60 is the customary constant. It keeps a first place in one list from outweighing everything else.
a card ranked 1st by dense and 3rd by BM25 1 / (60 + 1) + 1 / (60 + 3) = 0.0164 + 0.0159 = 0.0323 a card ranked 2nd by dense only 1 / (60 + 2) = 0.0161
RRF needs no normalising and no weights, which is why it is popular. But it throws information away. A card that wins by a mile and a card that wins by a hair both earn the same points for first place.
What the lab found
Every idea in this chapter was measured, and several did not survive. We measured fusion on the LoCoMo development set, the 385 test questions introduced in Chapter 6. The measure is evidence recall: the share of the turns that hold the answer which were actually retrieved. It needs no model call, so every run was free.
| Dense | BM25 | Entity | Evidence recall | Note |
|---|---|---|---|---|
| 0.45 | 0.45 | 0.1 | 85.64 % | |
| 0.5 | 0.4 | 0.1 | 86.21 % | |
| 0.6 | 0.4 | 0.0 | 86.47 % | no entity boost |
| 0.6 | 0.3 | 0.1 | 87.34 % | where we started |
| 0.6 | 0.3 | 0.2 | 87.86 % | |
| 0.8 | 0.2 | 0.0 | 88.40 % | no entity boost |
| 0.7 | 0.2 | 0.1 | 88.52 % | |
| 0.7 | 0.2 | 0.2 | 88.52 % | |
| 0.9 | 0.1 | 0.0 | 88.71 % | no entity boost |
| 0.75 | 0.15 | 0.1 | 88.96 % | |
| 0.8 | 0.1 | 0.1 | 89.05 % | best; kept |
| reciprocal rank fusion | 85.31 % | worst | ||
best weights against the start 89.05 - 87.34 = +1.71 points RRF against the start 85.31 - 87.34 = -2.03 points RRF against the best 85.31 - 89.05 = -3.74 points
Three lessons sit in this table. Meaning should carry most of the weight. More weight on keywords hurt. And a small entity boost still helped: at dense 0.8, moving 0.1 of weight from keywords to names raised recall from 88.40 to 89.05. Click a row of measured weights below and read its measured recall.
Common questions
Who decides a fact's importance, and can it be wrong?
The extractor decides, at write time, and yes, it can be wrong. It is a judgement by a small model. That is why importance is one part of the score and never the only one. In hanumemAI its weight is 0 by default, and it is yours to turn up.
Why multiply by 0.995 every hour and not some other number?
It is the value used in the Generative Agents work, and it is a starting point, not a law. It halves a score in under 6 days.
0.995 ^ 138 = 0.50 138 / 24 = 5.75 days
A system where users return weekly might want a slower decay. hanumemAI measures the decay in days, and a score halves after 90 of them.
Does recency mean that old facts are forgotten?
No. Recency only lowers a score. The memory is still stored, and relevance and importance can lift it back, as they did for the birthday. Removing memories is a separate decision, covered in Chapter 12.
Should the three parts have equal weight?
The class formula adds them equally, and the experiment shows that this already works. Production systems tune the weights on their own data. There is no universal answer, only a measured one.
If RRF lost here, why is it so widely used?
Because it needs no tuning and no normalising, so it is a safe default when nothing has been measured. On our data a tuned weighted sum beat it by almost 4 points of recall. If you have test questions, measure both.
Is a reranking model worth adding on top?
A reranker is a second model that re-reads the top candidates and reorders them. It can help, and it costs time on a path with 50 ms to spend. In our experiments the remaining errors were mostly reading errors by the answering model, not ranking errors, so we did not add one.
Carry this
- Score = relevance + recency + importance, each normalised to the range 0 to 1.
- Recency is 0.995 to the power of hours since last access. After 2,000 hours it is 0.000044.
- Relevance alone ranks a pasta recipe third. Recency alone ranks printer ink first. Together they pick the birthday, the vinyl records and the diet.
- Search three ways: meaning, exact words and names. Then fuse.
- Measured best weights: dense 0.8, BM25 0.1, entity 0.1, recall 89.05 percent. RRF scored 85.31.
Check yourself
1. A memory was last accessed 300 hours ago. What is its raw recency score, and roughly how many days is that?
Answer
recency 0.995 ^ 300 = 0.222 days 300 / 24 = 12.5 days
After less than two weeks a memory has lost more than three quarters of its recency.
2. Three memories have importance 2, 5 and 8. What are their normalised importance scores?
Answer
smallest 2, largest 8, range 6 2 -> (2 - 2) / 6 = 0.00 5 -> (5 - 2) / 6 = 0.50 8 -> (8 - 2) / 6 = 1.00
3. A user asks "what did the vet say about Miso?" Which of the three search signals is most likely to find the card, and why?
Answer
BM25 and the entity boost. "Miso" is a rare word and a name, so a match is strong evidence. Dense search may place the question near any card about pets or health. This is the case hybrid search exists for.