Part II · Apprentice
Chapter 12
Forgetting and decay
A memory that never forgets gets worse, not better. Four tools keep the store small and the top three cards correct: dedup, supersede, decay and erase.
On this page
In this chapter, we will learn why an assistant that remembers everything becomes a worse assistant. We will count how fast a store grows when nothing is removed, and watch stale cards push the right answer out of the top three. Then we will learn the four tools that keep a memory small, true and legal.
Why forget at all?
Forgetting is the planned hiding or removal of memories that no longer help. It sounds backwards. We have spent eleven chapters teaching the assistant to remember. Now we will teach it to forget.
Start with arithmetic. Take one active user and a store that only ever adds.
sessions per day 2 candidate facts per session about 3 candidates per day 2 x 3 = 6 candidates per year 6 x 365 = 2,190 candidates in two years 2,190 x 2 = 4,380 planned store size about 200 memories per user append-only against plan 4,380 / 200 = about 22x bigger
Twenty-two times bigger than planned, for one user, in two years. And among those 4,380 cards, five say where Maya lives. Four of the five are stale.
The loud reason: storage
A store that only adds grows without limit. The class has a simulation called
forgetting. It needs no model. It follows one active user for 24 months under two
policies.
- Policy A, append-only. Every candidate becomes a new memory.
- Policy B, consolidate and decay. 45 % of candidates merge into a memory that already exists. Of the ones that are added, 40 % are events, and an event expires 90 days after it was last used. The simulation never reads a card back, so here that means 90 days after it was written.
month append-only consolidate+decay
3 540 297
6 1,080 475
9 1,620 653
12 2,160 832
15 2,700 1,010
18 3,240 1,188
21 3,780 1,366
24 4,320 1,544
The simulation uses 30-day months, so its two-year total is 4,320 where our calendar arithmetic above gave 4,380. Here is where each column comes from.
policy A 2 x 3 x 30 = 180 new memories a month, for ever
policy B 2 x 3 x 0.55 x 30 = 99 added a month
after day 90: 99 x 0.40 = about 40 events expire each month
net growth 99 - 40 = about 59 a month
Now the storage. One memory is about 3,322 bytes, as we sized it in Chapter 7.
policy A 4,320 x 3,322 B = 14,351,040 B = about 14.4 MB per user
x 1,000,000 users = about 14.4 TB
policy B 1,544 x 3,322 B = 5,129,168 B = about 5.1 MB per user
x 1,000,000 users = about 5.1 TB
saved 14.4 - 5.1 = 9.3 TB policy B is 64 % smaller
These totals assume that every one of a million users chats twice a day for two years. That is an upper bound, not a forecast. It shows the direction, and the direction is clear.
Move the merge rate, the share of events and the expiry time. Watch the two lines pull apart over 24 months.
How a heavy user still lands at 200 memories
Even policy B ends at 1,544 memories, far above the 200 we planned for. Is the plan wrong? No. The simulation holds the merge rate at 45 % for two years, and that is pessimistic. Let us solve for the merge rate m that would land this user at 200.
facts and rules kept 24 months x 180 x (1 - m) x 0.60 = 2,592 x (1 - m)
events kept 3 months x 180 x (1 - m) x 0.40 = 216 x (1 - m)
total = 2,808 x (1 - m)
set the total to 200 1 - m = 200 / 2,808 = 0.0712
m = about 93 %
A 93 % merge rate sounds extreme. It is what a mature profile looks like. After a year, almost everything Maya mentions about herself is something the assistant already knows, or a change to it. New facts become rare. Duplicates and updates become the rule.
The quiet reason: retrieval dilution
Storage is cheap, and it gets cheaper every year. So here is the reason that matters more.
Retrieval dilution is what happens when stale or repeated memories take the places of useful ones in the top results. The librarian hands over three cards. If two of them are wrong, the model reads wrong things first.
Maya has moved four times. An append-only store kept all five address cards. The scores and the years below are illustrative. They come from the class simulation, not from a real search and not from the timeline of Chapter 11.
query: "where does Maya live?" top 3 from the append-only store rank score memory status 1 0.89 Maya lives in London. stale (written 2024) 2 0.88 Maya lives in Paris. current 3 0.88 Maya lives in Bristol. stale (written 2021)
These sentences are nearly the same, so their scores sit within 0.01 of each other. Which one comes first is luck.
precision at 3 = correct cards in the top 3 / 3 = 1 / 3
Now the same question against a store where each move replaced the last.
rank score memory status 1 0.89 Maya lives in Paris. current
One address card is visible, and it is the right one. Forgetting is not data loss. It is precision.
The policy, by colour of card
Cards come in three colours, and the colour decides how long a card lives.
| Category | Policy | Why |
|---|---|---|
| semantic (facts) | never expire; replace on change | "vegetarian" does not age out |
| episodic (events) | expire about 90 days after last access | "dentist on Tuesday" decays to noise |
| procedural (rules) | never expire; the user can edit them | standing orders stay in force |
| everything | remove duplicates at write time | duplicates dilute the top results |
Note the words "last access". The clock does not start when the card was written. It restarts each time the card is used. An event that Maya keeps asking about keeps earning its place. An event that nobody has touched for three months fades.
Four tools, in order
Forgetting is not one thing. It is four tools, and we reach for them in this order, from the gentlest to the strongest.
1. Dedup: never store the duplicate
Dedup, short for deduplication, means refusing to store the same fact twice. It is the cheapest forgetting there is, because the card is never written. hanumemAI makes two checks, and neither needs a model.
- The text has the same hash as a stored memory, which means it is identical: skip.
- The text has a cosine similarity of 0.95 or more to a stored memory, which means it is a rewording: skip.
2. Supersede: close the card that changed
This is Chapter 11. The London card is closed and points at the Paris card. Ordinary search reads only open cards, so the librarian sees one address. From her side the old card is forgotten. From the auditor's side it is still there. Both are satisfied.
3. Decay: let unused events fade
Decay is the slow loss of a memory's strength when nobody uses it. The class gives it a formula, the same one that scores recency in Chapter 9.
recency = 0.995 ^ hours_since_last_access after 1 hour 0.995 ^ 1 = 0.995 after 1 day 0.995 ^ 24 = 0.887 after 1 week 0.995 ^ 168 = 0.431 after 1 month 0.995 ^ 720 = 0.027 after 2,000 hours 0.995 ^ 2,000 = 0.000044 (about 83 days)
Losing half a percent each hour sounds gentle. Repeated two thousand times, it leaves almost nothing. Change the base and the hours, and see how quickly the curve reaches the floor.
The curve is for ranking. Expiry is simpler. A nightly job closes events whose last access is older than a fixed limit, 90 days by default. It is a cut-off, not a score.
4. Erase: the delete button
Erasure is the removal of a memory because a person asked for it. The first three tools
serve the quality of answers. This one serves the user's rights. Every row carries a
user_id, so one call removes everything about one person. In the class experiment that call
took Maya from 3 memories to 0 in about 4 ms, while Tom kept his 2.
A short word on human memory
People forget too, and mostly it helps. You do not remember what you ate for lunch on a Tuesday three years ago, and you are no worse for it. You do remember your best friend's name, because you use it every week. Memories that are used stay strong. Memories that are never used fade.
The 90-day rule on events copies that pattern in the plainest way possible. We should not stretch the comparison further. The assistant's forgetting is a policy that we chose, that we can read, and that we can change.
Try it yourself
The simulation uses only the Python standard library. It needs no model and no database.
cd memory_classnotes/experiments/forgetting
pip install -r requirements.txt # installs nothing
python main.py
In hanumemAI, the last two tools are one line each.
from hmem import Memory, Config
m = Memory(Config(db_path="app.sqlite", episodic_ttl_days=90))
# run nightly: closes episodic rows not used for 90 days, returns how many
expired = m.forget("maya")
# hide one memory, at the user's request: the row is closed, not erased
m.delete(mem_id, user_id="maya")
# the right to be forgotten: every row, vector, raw turn and profile of one user
m.delete_user("maya")
forget() only ever touches rows in the episodic category. Two similar words need care here.
Episodic is the category of event cards, which decay. Episodes are the raw
turns that store_episodes keeps, and hanumemAI files them in the episodic category too. Facts and rules are never expired by time. It closes rows, in the same way superseding
does. So does delete(). delete_user() is different: it removes the rows
themselves.
Common questions
Is forgetting not dangerous? What if the expired card was important?
That is why only events expire by time, and only when nobody has used them for about three months. The peanut allergy is a fact, so it never expires. If an event matters enough to be asked about, the asking itself restarts its clock.
Why 90 days?
It is a sensible starting point from the class, not a law. It is a knob,
episodic_ttl_days. A travel assistant may want a year, because people ask about last summer's
trip. A support assistant may want 30 days. Choose it by watching what your users ask about.
Why measure from last access and not from when the card was written?
Because age alone says little. A two-year-old card that was used yesterday is clearly still useful. A two-week-old card that has never been read may be noise. Use is the better signal.
If storage is cheap, why not keep everything and let search sort it out?
Because search hands over a fixed, small number of cards. Every stale or repeated card that ranks well pushes a useful one out. The address example shows it: five cards stored, two of the top three wrong. The cost of keeping everything is paid in wrong answers, not in disk space.
Does an expired event disappear for ever?
In hanumemAI, forget() closes the row and does not destroy it. It leaves ordinary search and
can still be reached by a question about a past date. Only delete_user() removes rows from the
disk.
What is the difference between decay and the recency score?
They read the same field, the time of last access, for two jobs. The recency score is a curve, used while reading, to rank cards. Expiry is a cut-off, used at night, to close event cards. The class curve loses half a percent per hour. hanumemAI's own recency term halves every 90 days (Chapter 9).
Carry this
- Append-only means 6 candidates a day, 2,190 a year, 4,380 in two years: 22 times the planned 200.
- Over 24 months the simulation holds 1,544 memories with consolidation and decay against 4,320 without: 64 % smaller.
- The quiet cost is retrieval dilution: five address cards gave a precision at 3 of 1/3.
- Four tools in order: dedup, supersede, decay, erase. Facts and rules never expire. Events fade about 90 days after last use.
- Forgetting is not data loss. It is precision.
Check yourself
1. A user has 3 sessions a day and the extractor finds 2 candidates per session. With a merge rate of 50 %, how many memories are added in a 30-day month?
Answer
candidates per day 3 x 2 = 6 added per day 6 x (1 - 0.5) = 3 added per month 3 x 30 = 90
2. A card was last used 48 hours ago. What is its recency score with a base of 0.995?
Answer
0.995 ^ 48 = about 0.786
Two days without use costs about a fifth of the score.
3. Maya deletes last week's conversation about a surprise party. Is the party forgotten?
Answer
Not necessarily. Any memory extracted from that conversation is a separate row and is still in the store. The product must remove the extracted memories as well, along with their vectors and any cached copies.