Part I · Novice
Chapter 4
Four kinds of memory
The sticky note and the card box: short-term memory, and the three colours of long-term memory, with a real experiment on rolling summaries.
On this page
In this chapter, we will learn the four kinds of memory an assistant needs. One is short-term and lives for a single chat. Three are long-term and live in the card box. We will see how a long chat is kept small with a rolling summary, and we will practise sorting sentences into the right kind.
The map
An assistant's memory has four kinds: one short-term and three long-term.
the memory of an agent
+------------------------------------------------+
SHORT-TERM | working memory = the context window itself |
(this | the live turns of this conversation |
session) | + a rolling summary when it overflows |
+------------------------------------------------+
LONG-TERM | semantic facts: "Maya is vegetarian" |
(across | episodic experiences: "last week we |
sessions) | planned the Lisbon trip" |
| procedural instructions: "answer short |
| and direct" |
+------------------------------------------------+
The names come from the study of human memory. We do not need the psychology. We need the four boxes on this map, because each one is stored, searched and forgotten in its own way.
Where the kinds sit in an agent
An agent uses short-term memory to think and long-term memory to remember. The message enters working memory, the relevant cards are retrieved, the model answers, and afterwards new cards are extracted and consolidated. The class keeps tools and planning out of scope, and so do we.
user message --> working memory --> retrieve from --> LLM --> answer
(context window) long-term |
^ v
| extract + consolidate
| |
long-term memory <---+
Short-term memory
Short-term memory, also called working memory, is the context window itself: the turns of the conversation that is happening now.
We met it in Chapter 1. The app resends the conversation so far with every message, so the model can see what was said a minute ago. Nothing needs to be extracted or searched. It is simply on the desk.
For a normal session this is cheap. For a long one it is not.
tokens per turn 1,500 / 6 = 250 turn 1 1 x 250 = 250 tokens on the desk turn 6 6 x 250 = 1,500 tokens turn 40 40 x 250 = 10,000 tokens
At turn 40, every new message pays for 10,000 tokens of old turns. This is the resend problem of Chapter 2, inside one chat.
The rolling summary
A rolling summary replaces the older turns of a long chat with a short summary, and keeps the most recent turns word for word.
turn 40 [ system | memory | t1 t2 t3 ............ 40 x 250 ............ t40 ] about 10,000 tokens
|
v summarise turns 1 to 35, keep the last 5
turn 40' [ system | memory | summary (about 150 tokens) | t36 t37 t38 t39 t40 ] about 1,400 tokens
summary = 150 tokens last 5 turns 5 x 250 = 1,250 tokens total 150 + 1,250 = 1,400 tokens before 10,000, after 1,400 10,000 / 1,400 = about 7 times smaller
The rule is simple. A threshold is a limit that we choose. When the history crosses it, say at 10,000 tokens, summarise the older turns. The recent turns stay exact, because the next question is most likely about them.
The summary is written by one extra LLM call. After that, each new turn pushes the oldest exact turn into the summary, so the desk stays near 1,400 tokens.
Move the sliders. Watch the line without a summary climb, and the line with a summary drop when the threshold is crossed and then stay flat.
Long-term memory has three colours
Long-term memory is what survives the end of the session. It comes in three kinds: semantic, episodic and procedural. In our card box, each kind has its own colour.
Semantic: facts
A semantic memory is a fact that is true until something replaces it. "Maya is vegetarian." "Maya has a serious peanut allergy." "Maya lives in Paris."
Facts do not fade with age. The peanut allergy is as true after two years as on the day Maya said it. A fact leaves the box only when a newer fact takes its place, as when Paris replaced London.
Episodic: events
An episodic memory is something that happened at a point in time. "Last week Maya planned a trip to Lisbon." "Maya has a dentist appointment on Tuesday."
An event always has a date. Most events matter for a while and then turn into noise. The dentist appointment is useful this week and useless next year. So event cards are allowed to fade when nobody asks about them for about 90 days.
Procedural: rules
A procedural memory is a standing instruction about how the assistant should behave. "Answer short and direct." "Use metric units."
Rules do not describe Maya's life. They describe how to talk to her. They never fade, and the user can edit them at any time.
| Colour | Kind | Example | How long it lives |
|---|---|---|---|
| Facts | semantic | Maya is vegetarian. | until replaced |
| Events | episodic | Maya planned a Lisbon trip last week. | fades about 90 days after it was last used |
| Rules | procedural | Answer short and direct. | until the user changes it |
Why the colour matters
The kind of a memory decides how long it lives. That is its main job in the system.
Here is the practical surprise. The three kinds do not need three databases. They live in one box, and the
colour is one small field on each card, called category. The librarian searches all cards the
same way. The colour is read when it is time to forget, which is the subject of
Chapter 12.
Two cases that fool people
A birthday is semantic, not episodic. "Sam has a birthday on June 18" contains a date, so it looks like an event. But it comes back every year. It is a standing fact about Sam. If we let it fade after 90 days, the assistant would forget the birthday before it arrives.
A training plan is semantic too. "Maya is training for the Berlin marathon" describes her life over months. It is not one thing that happened on one day.
The test is this: did it happen once, at one moment? Then it is an event. Is it true over a stretch of time, or does it repeat? Then it is a fact.
And a fourth pile: nothing
Most of what is said deserves no card at all. "Mondays should be illegal." "What time is it in Tokyo?" "Pouring rain today." These are small talk and one-off requests. They belong on the sticky note and go in the bin with it.
Secrets get no card either, for a different reason. A password is durable, but it must never be stored. Chapter 15 explains the write policy that refuses it.
Sort the ten sentences below into the three colours or the bin. Then check your answers and read the reason for each.
Try it yourself: a chat compressed to a summary
The class experiment short-term-memory takes a 14-message chat about a family trip to Lisbon.
It asks a follow-up question twice: once with the full history, once with only a summary.
(a) Follow-up asked with the FULL 14-message history in the prompt
Book in Belem for its flat terrain, making it easier for the kids, and proximity to
attractions like the Jerónimos Monastery and Belém Tower.
prompt tokens: 818 (14 messages of trip planning plus the question)
(b) One LLM call compresses the chat into a short-term summary
- Budget: 120 euros per night
- Neighborhoods: Alfama or Belem (Belem is easier for kids)
- Food: Time Out Market, Ao 26 Vegan Food Project, Jardim das Cerejas
- Day trip: Sintra (40 min train from Rossio)
- Rain gear: Light jackets for all
- No car needed; public transport sufficient
summary length: 84 tokens (resp.usage.completion_tokens)
(c) SAME follow-up, prompt is ONLY a system message with the summary
We should book in Belem because it's more kid-friendly and offers easy access to the
recommended attractions. Plus, its proximity to the Time Out Market makes for
convenient dining options.
prompt tokens: 125. Picked: Belem, full-history run picked: Belem.
full history 818 tokens summary only 125 tokens ratio 818 / 125 = 6.5 times smaller
Both runs choose Belem. The summary did the job at less than one sixth of the tokens. The summary itself is 84 tokens. The other 41 of the 125 are the instruction around it and the question (125 - 84 = 41).
The experiment then tests the weak point. It asks for an exact detail, using the summary alone.
(e) Lossiness check: a detail question against the summary ONLY
Q: What exact nightly budget did I mention? One sentence.
The nightly budget mentioned is 120 euros per night.
the exact 120 euro figure survived this summary. Summaries are lossy by
design: any detail can vanish in a compression pass, which is why
long-term memory stores discrete facts instead.
The figure survived this time. Nothing guarantees it will survive the next summary, or the one after that.
ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/short-term-memory
pip install -r requirements.txt
python main.py
Common questions
If the summary works so well, why not use summaries for long-term memory too?
Two reasons. A summary is rewritten again and again, and each rewrite can drop a detail. And a summary is injected whole, even when the question needs one line of it. Separate facts keep their details and can be picked one at a time.
Where is short-term memory stored?
In a session store, a small fast database that keeps the turns of each open chat. Redis is a common choice. The turns are deleted some hours after the chat ends. They do not go into the card box.
What happens if the secretary picks the wrong colour?
The card is still found by search, because search does not look at the colour. The harm comes later: a fact marked as an event may fade when it should not. In our own experiments a small model often labelled ordinary advice as a rule. So we stopped placing rule cards in every prompt (experiment 0006). The prompt shrank by 16 percent, from 1,489 to 1,250 tokens per question, and the test score did not change beyond the noise of a repeated run. (Chapter 6 introduces the test, and Chapter 17 the noise.) Labels from an LLM are useful and imperfect.
Should rules be placed in every prompt?
In principle, yes. "Answer short and direct" applies to every answer, and a rule is only a few tokens. In practice it depends on how well the rules are labelled. If only real standing instructions carry the label, always inject them. If the label is noisy, let rules compete in search like other cards.
Do all memory systems use these three kinds?
Most use some form of them, with different names. Some add more, such as a profile of the user or a summary of each session. The three kinds are the shared vocabulary, and an interviewer will expect them.
Can one sentence produce two cards?
Yes. "I moved from London to Paris last month for a new job" holds a fact (lives in Paris), an event (the move in a given month) and a second fact (the new job). How finely to split is a real design choice, and we return to it in Chapter 6.
Carry this
- Short-term memory is the context window: this chat's turns. Long-term memory is the card box.
- A rolling summary keeps a long chat small: 10,000 tokens become about 1,400 (a 150 token summary plus the last 5 turns).
- Summaries are lossy. That is acceptable for a session and not for the box.
- Three colours: facts (semantic) live until replaced, events (episodic) fade after about 90 days unused, rules (procedural) stay until the user edits them.
- One store, with the kind as a field on each card. The kind decides lifetime, not search.
Check yourself
1. A chat has reached 60 turns of 250 tokens. We summarise all but the last 5 turns into 150 tokens. How large is the history before and after?
Answer
before 60 x 250 = 15,000 tokens after 150 + (5 x 250) = 1,400 tokens
The result after the summary is the same 1,400 tokens as at turn 40. That is the point: it stays flat.
2. Which kind is each of these? (a) "Maya's partner Sam has a birthday on June 18." (b) "Maya visited the dentist on Tuesday." (c) "Always use metric units."
Answer
(a) Semantic: it repeats every year, so it is a standing fact. (b) Episodic: it happened once, at one moment. (c) Procedural: it tells the assistant how to behave.
3. Why does long-term memory store separate facts and not one summary per user?
Answer
A summary loses details each time it is rewritten, and it must be injected whole. Separate facts keep exact details and let the librarian hand over only the few that match the question.