Carebun Interactive/Memory Memory labCheat sheet

Part II · Apprentice

Chapter 9

Scoring beyond similarity

Similarity finds the topic. It cannot tell a birthday from printer ink. We add recency and importance, then search by meaning, words and names at once.

12 min read · interactive

On this page
  1. Why similarity alone is not enough
  2. Three parts of one score
  3. Try it yourself: a gift for Sam
  4. One row, worked out by hand
  5. Three ways to search
  6. Fusing the signals
  7. What the lab found
  8. Common questions
  9. Carry this
  10. Check yourself

In this chapter, we will learn why the most similar memory is not always the right one. We will build a score from three parts, work one row out by hand, and then give the librarian three ways to search instead of one. At the end we will see which of these ideas survived measurement.

Why similarity alone is not enough

Similarity measures whether a memory is about the same topic as the question. It does not measure whether the memory is current, or whether it matters.

In Chapter 8 the librarian ranked cards by cosine similarity only. That fails in two ways.

  • It cannot see time. "Maya lives in London" and "Maya lives in Paris" are almost the same sentence. The class simulation in Chapter 12 puts such sentences within 0.01 of each other. Which one ranks first is luck.
  • It cannot see weight. "Maya has a serious peanut allergy" and "Maya asked for a pasta recipe yesterday" are both about food. Only one of them can send a child to hospital.

Three parts of one score

The retrieval score adds three components: relevance, recency and importance. This recipe comes from a research project called Generative Agents, in which simulated people had to remember their own lives.

score = relevance + recency + importance     (each min-max normalized to [0,1])

relevance   cosine(query, memory)
recency     0.995 ^ hours_since_last_access      exponential decay per hour
importance  LLM-rated 1 to 10 at write time

Relevance

Relevance is the cosine similarity between the question and the memory. It is the score we already know. It keeps the answer on topic.

Recency

Recency is a score that starts at 1 and shrinks a little with every hour since the memory was last used.

Each hour the score is multiplied by 0.995, so it loses half a percent. That sounds like nothing. It compounds.

after     1 hour       0.995 ^ 1      =  0.995
after    26 hours      0.995 ^ 26     =  0.878
after   120 hours      0.995 ^ 120    =  0.548      5 days: about half
after   300 hours      0.995 ^ 300    =  0.222
after 1,000 hours      0.995 ^ 1000   =  0.0067     6 weeks: almost nothing
after 2,000 hours      0.995 ^ 2000   =  0.000044   83 days: gone

Note the word accessed. The clock restarts each time a memory is retrieved. A fact that keeps being useful stays fresh. A fact nobody asks for fades.

Importance

Importance is a number from 1 to 10 that says how much a fact matters, given once, when the card is written. The extractor rates it in the same call that writes the fact. Printer ink gets a 1. A peanut allergy should get a 9 or 10. The small class prompt in Chapter 6 asks only for a fact and a category. A production prompt, such as hanumemAI's, adds an importance field to the same JSON.

It is decided at write time because that is when a model is reading the conversation anyway. At read time there is no model call to spare.

Min-max normalisation

Min-max normalisation rescales a set of numbers so that the smallest becomes 0 and the largest becomes 1.

We need it because the three parts live on different scales. Cosine scores sit between about 0.3 and 0.7. Importance runs from 1 to 10. Added raw, importance would drown the others.

normalized = (value - smallest) / (largest - smallest)

cosine scores  0.35, 0.40, 0.69      smallest 0.35, largest 0.69
  0.35   ->   (0.35 - 0.35) / 0.34   =  0.00
  0.40   ->   (0.40 - 0.35) / 0.34   =  0.15
  0.69   ->   (0.69 - 0.35) / 0.34   =  1.00

Try it yourself: a gift for Sam

The class experiment memory-scoring ranks eight of Maya's memories three ways. It needs only the embedding model, because the importance numbers are fixed by hand in this experiment.

ollama pull nomic-embed-text
cd memory_classnotes/experiments/memory-scoring
pip install -r requirements.txt
python main.py

First, relevance alone.

(a) Ranking by RELEVANCE alone: cosine(query vector, memory vector)
    query: "What should I get Sam as a birthday present?"

    rank     rel  fact
       1  0.6942  Maya's partner Sam has a birthday on June 18.
       2  0.6627  Maya said Sam loves vinyl records and old jazz.
       3  0.4227  Maya asked for a pasta recipe yesterday.
       4  0.4058  Maya has two kids, aged 6 and 9.
       5  0.3880  Maya mentioned the printer is out of ink.
       6  0.3744  Maya likes hiking on weekends.
       7  0.3721  Maya is vegetarian.
       8  0.3475  Maya works as a data engineer at a fintech company.

The first two are right. The third is a pasta recipe. Relevance found the topic and then ran out of ideas. Now recency alone.

(b) Ranking by RECENCY alone: recency = 0.995 ^ hours_since_last_access

    rank    recency   hours  fact
       1   0.995000       1  Maya mentioned the printer is out of ink.
       2   0.877809      26  Maya asked for a pasta recipe yesterday.
       3   0.547986     120  Maya is vegetarian.
       4   0.222292     300  Maya has two kids, aged 6 and 9.
       5   0.029933     700  Maya likes hiking on weekends.
       6   0.000543    1500  Maya works as a data engineer at a fintech company.
       7   0.000046    1992  Maya said Sam loves vinyl records and old jazz.
       8   0.000044    2000  Maya's partner Sam has a birthday on June 18.

Printer ink wins. The birthday is dead last. Recency alone retrieves trivia. Now all three together.

(c) Combined score: min-max normalize each component to [0, 1] across the
    8 memories, then total = relevance + recency + importance.

    fact                                                   rel   rec   imp  total
    Maya's partner Sam has a birthday on June 18.         1.00  0.00  1.00   2.00  <- inject
    Maya said Sam loves vinyl records and old jazz.       0.91  0.00  0.71   1.62  <- inject
    Maya is vegetarian.                                   0.07  0.55  0.71   1.33  <- inject
    Maya has two kids, aged 6 and 9.                      0.17  0.22  0.86   1.25
    Maya asked for a pasta recipe yesterday.              0.22  0.88  0.14   1.24
    Maya mentioned the printer is out of ink.             0.12  1.00  0.00   1.12
    Maya works as a data engineer at a fintech company.   0.00  0.00  0.57   0.57
    Maya likes hiking on weekends.                        0.08  0.03  0.43   0.54

The birthday, the vinyl records and the vegetarian diet. Those are the three facts a gift question needs, and the last one matters if the gift is a dinner. The printer note has the best recency of all and still falls to rank 6.

One row, worked out by hand

To check a score, normalise each part and add the three results. Let us do it for the top row: "Maya's partner Sam has a birthday on June 18."

relevance    cosine = 0.6942       smallest 0.3475, largest 0.6942 among the 8
             (0.6942 - 0.3475) / (0.6942 - 0.3475)          =  1.00

recency      0.995 ^ 2000 = 0.000044     (2,000 hours is about 83 days)
             smallest 0.000044, largest 0.995000
             (0.000044 - 0.000044) / (0.995000 - 0.000044)  =  0.00

importance   8 on the scale of 1 to 10   smallest 1, largest 8 among the 8
             (8 - 1) / (8 - 1)                              =  1.00

total        1.00 + 0.00 + 1.00                             =  2.00

Recency gave this memory nothing at all. Relevance and importance rescued it. That is why one component is never enough.

Now take the controls. Give each of the three parts a weight, and watch the eight memories reorder. Set recency high and see the printer climb. Then change the decay rate and see how quickly old memories vanish.

A memory can be found by its meaning, by its exact words or by the names in it. Relevance so far has meant one thing: cosine similarity between vectors. That is search by meaning. It has blind spots, and two older kinds of search cover them.

By meaning: dense vectors

Dense search compares embeddings, so it finds memories that mean the same thing even when the words differ. "What should I cook tonight" finds "Maya is vegetarian". The vectors are called dense because every one of their numbers is filled in.

Its weakness is exactness. To an embedding, "Poppy" and "Daisy" are both just pet names and land in almost the same place. So do "March 14" and "March 15".

By exact words: BM25

BM25 is a keyword search that scores a memory by the words it shares with the question, and counts rare words more than common ones.

The idea is simple. If the question contains "the", a match tells us nothing, because every card contains "the". If it contains "Miso", a match tells us almost everything, because one card in two hundred has that word. BM25 also stops a long card from winning just because it has more words.

question: "what is the name of the cat?"

word       cards that contain it      what a match is worth
the        nearly all                 almost nothing
name       a few                      a little
cat        one or two                 a lot

By names: the entity boost

An entity is a named thing: a person, a place, a company, a pet. At write time the extractor lists the entities on each card, in one more field of its JSON. At read time, names in the question are looked up in that list, and every card that mentions them gets a boost.

There is a catch. If Sam appears on 60 cards, a match on "Sam" says very little. So the boost is damped: the more cards share a name, the less each one gains.

SignalGood atBlind to
Dense (meaning)paraphrase, topic, "dinner" finds "vegetarian"exact names, numbers, dates
BM25 (words)rare words, codes, exact phrasesany rewording
Entity (names)everything about one person or placequestions with no name in them

Using several signals together is called hybrid search.

Fusing the signals

Fusion is the step that turns several rankings into one. There are two common ways.

A weighted sum

Normalise each signal to the range 0 to 1, multiply each by a weight, and add.

score = 0.8 x dense + 0.1 x bm25 + 0.1 x entity

a card with   dense 0.70,  bm25 1.00,  entity 1.00
score         0.8 x 0.70 + 0.1 x 1.00 + 0.1 x 1.00   =  0.56 + 0.10 + 0.10  =  0.76

a card with   dense 0.80,  bm25 0.00,  entity 0.00
score         0.8 x 0.80                              =  0.64

The first card is less similar in meaning, but it holds the exact word and the exact name. It wins.

Reciprocal rank fusion

Reciprocal rank fusion (RRF) ignores the scores and uses only the position of each card in each list. A card earns 1 / (60 + rank) from every list it appears in, and the points are added. The 60 is the customary constant. It keeps a first place in one list from outweighing everything else.

a card ranked 1st by dense and 3rd by BM25
  1 / (60 + 1)  +  1 / (60 + 3)   =  0.0164 + 0.0159   =  0.0323

a card ranked 2nd by dense only
  1 / (60 + 2)                     =  0.0161

RRF needs no normalising and no weights, which is why it is popular. But it throws information away. A card that wins by a mile and a card that wins by a hair both earn the same points for first place.

What the lab found

Every idea in this chapter was measured, and several did not survive. We measured fusion on the LoCoMo development set, the 385 test questions introduced in Chapter 6. The measure is evidence recall: the share of the turns that hold the answer which were actually retrieved. It needs no model call, so every run was free.

DenseBM25EntityEvidence recallNote
0.450.450.185.64 %
0.50.40.186.21 %
0.60.40.086.47 %no entity boost
0.60.30.187.34 %where we started
0.60.30.287.86 %
0.80.20.088.40 %no entity boost
0.70.20.188.52 %
0.70.20.288.52 %
0.90.10.088.71 %no entity boost
0.750.150.188.96 %
0.80.10.189.05 %best; kept
reciprocal rank fusion85.31 %worst
best weights against the start      89.05 - 87.34   =  +1.71 points
RRF against the start               85.31 - 87.34   =  -2.03 points
RRF against the best                85.31 - 89.05   =  -3.74 points

Three lessons sit in this table. Meaning should carry most of the weight. More weight on keywords hurt. And a small entity boost still helped: at dense 0.8, moving 0.1 of weight from keywords to names raised recall from 88.40 to 89.05. Click a row of measured weights below and read its measured recall.

Common questions

Who decides a fact's importance, and can it be wrong?

The extractor decides, at write time, and yes, it can be wrong. It is a judgement by a small model. That is why importance is one part of the score and never the only one. In hanumemAI its weight is 0 by default, and it is yours to turn up.

Why multiply by 0.995 every hour and not some other number?

It is the value used in the Generative Agents work, and it is a starting point, not a law. It halves a score in under 6 days.

0.995 ^ 138    =  0.50
138 / 24       =  5.75 days

A system where users return weekly might want a slower decay. hanumemAI measures the decay in days, and a score halves after 90 of them.

Does recency mean that old facts are forgotten?

No. Recency only lowers a score. The memory is still stored, and relevance and importance can lift it back, as they did for the birthday. Removing memories is a separate decision, covered in Chapter 12.

Should the three parts have equal weight?

The class formula adds them equally, and the experiment shows that this already works. Production systems tune the weights on their own data. There is no universal answer, only a measured one.

If RRF lost here, why is it so widely used?

Because it needs no tuning and no normalising, so it is a safe default when nothing has been measured. On our data a tuned weighted sum beat it by almost 4 points of recall. If you have test questions, measure both.

Is a reranking model worth adding on top?

A reranker is a second model that re-reads the top candidates and reorders them. It can help, and it costs time on a path with 50 ms to spend. In our experiments the remaining errors were mostly reading errors by the answering model, not ranking errors, so we did not add one.

Carry this

  • Score = relevance + recency + importance, each normalised to the range 0 to 1.
  • Recency is 0.995 to the power of hours since last access. After 2,000 hours it is 0.000044.
  • Relevance alone ranks a pasta recipe third. Recency alone ranks printer ink first. Together they pick the birthday, the vinyl records and the diet.
  • Search three ways: meaning, exact words and names. Then fuse.
  • Measured best weights: dense 0.8, BM25 0.1, entity 0.1, recall 89.05 percent. RRF scored 85.31.

Check yourself

1. A memory was last accessed 300 hours ago. What is its raw recency score, and roughly how many days is that?

Answer
recency     0.995 ^ 300   =  0.222
days        300 / 24      =  12.5 days

After less than two weeks a memory has lost more than three quarters of its recency.

2. Three memories have importance 2, 5 and 8. What are their normalised importance scores?

Answer
smallest 2, largest 8, range 6
2   ->   (2 - 2) / 6   =  0.00
5   ->   (5 - 2) / 6   =  0.50
8   ->   (8 - 2) / 6   =  1.00

3. A user asks "what did the vet say about Miso?" Which of the three search signals is most likely to find the card, and why?

Answer

BM25 and the entity boost. "Miso" is a rare word and a name, so a match is strong evidence. Dense search may place the question near any card about pets or health. This is the case hybrid search exists for.