Part II · Apprentice
Chapter 6
The secretary: fact extraction
One model reads the chat once and decides what is worth a card. We watch it turn 221 tokens into 39, and learn why its mistakes are the ones nobody sees.
On this page
In this chapter, we will learn how a memory system decides what to remember. We will read the prompt that does the deciding and watch it work on a real conversation. Then we will see why this one step sets the quality of everything that comes after it.
What is fact extraction?
Fact extraction is one model call that reads a finished conversation and writes down the few facts worth keeping for future conversations.
It is the first step of the write path, the road a memory travels on its way into the box. The model that does it is called the extractor. It is a different job from answering the user, and it usually gets a smaller, cheaper model.
session transcript extractor LLM candidate facts
about 1,500 tokens ---> temperature 0 ---> 0 to N of them, about 3 on average
strict JSON about 25 tokens each
|
v
dropped: smalltalk, jokes,
one-off requests, the weather
Look at the two arrows leaving the extractor. One goes to the box. The other goes to the bin. Deciding which is which is the whole job.
What deserves a card?
A fact deserves a card when it will still be useful in a future conversation. Everything else is ephemera, which means things that matter for a moment and then do not.
| Maya says | Card? | Why |
|---|---|---|
| "I'm vegetarian, for any future food suggestions." | Yes | True next month, and it changes answers |
| "My partner Sam has a birthday on June 18." | Yes | Comes back every year |
| "Keep your answers short and direct." | Yes | A standing instruction shapes every answer |
| "What time is it in Tokyo right now?" | No | A one-off request, useless tomorrow |
| "Mondays should be illegal." | No | A joke |
| "Pouring rain today, so gloomy." | No | Weather chat |
The test is simple to say and hard to apply: would a friend who knows Maya well remember this in a month? A friend remembers the allergy. A friend does not remember the rain.
The extraction prompt
The extractor is steered by a prompt, a short set of written instructions that goes in front of the transcript. Here is the one from the class notes.
SYSTEM = (
"You extract durable facts about the user from a chat transcript: "
"only facts worth remembering in future sessions. Return a STRICT "
'JSON array of {"fact": "...", "category": "..."}, category one of: '
"semantic (facts), episodic (events), procedural. One short sentence each. "
"Extract as many as the transcript truly contains, possibly none: "
"return an empty array if nothing is worth remembering. "
"Standing instructions about how to answer ARE durable. Do NOT "
"store smalltalk, jokes, or one-off requests: those are ephemeral."
)
resp = client.chat.completions.create(
model=MODEL, temperature=0,
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": transcript}])
Four choices in this small prompt are worth slowing down for.
Strict JSON
JSON is a plain text format that programs can read without guessing. The extractor's reader is not a person. It is our code, which must put each fact into a database row. If the model replies "Sure! Here are some facts I found", the code breaks. So the prompt demands a JSON array and nothing else.
Temperature 0
Temperature is a dial for how adventurous the model's word choices are. At a high setting the model is creative. At 0 it takes the most likely word every time. A secretary should not be creative. We want the same transcript to produce the same cards.
"Possibly none"
The prompt allows an empty answer, so the extractor is never forced to invent. A model that is asked for facts will try hard to find some. If a session is pure smalltalk and the prompt demands three facts, we get three invented or worthless cards. The words "possibly none" and "return an empty array" give the model permission to write nothing. The average is about 3 facts per session, but the count follows the content. It is not a quota.
Categories
The category is the colour of the card. Every card gets one: semantic for facts, episodic for events, procedural for standing instructions. We met them in Chapter 4. The colour decides later how long the card lives, so it is set here, at the moment of writing.
Play the secretary
Before we see what the model did, try the job yourself. Here is one real session of 14 messages. Mark each line keep or drop, then reveal what the extractor kept, and compare the token counts.
Try it yourself: one session, four cards
The class experiment fact-extraction runs this step. It needs a local model and nothing
else. Its prompt is a longer version of the one above. It adds an example for each category, and it asks
that every fact starts with the user's name, Maya.
ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/fact-extraction
pip install -r requirements.txt
python main.py
The transcript is 14 messages, durable facts mixed with ephemera.
1 user I just moved from London to Paris for a new role at a payments startup.
2 assistant Congratulations on the move and the new role! How can I help today?
3 user Quick one: what time is it in Tokyo right now?
4 assistant Tokyo is 7 or 8 hours ahead of Paris depending on the season: ...
5 user Ha. Honestly, Mondays should be illegal. Worst invention ever.
6 assistant A classic Monday complaint. Hang in there, the week gets better.
7 user Rephrase this email politely: 'The dashboard is broken again, fix it.'
8 assistant 'Hi team, the dashboard seems down. Could someone take a look? Thanks!'
9 user Thanks. By the way, I'm vegetarian, for any future food suggestions.
10 assistant Noted: vegetarian suggestions only from here on.
11 user My partner Sam has a birthday on June 18, I need to plan something.
12 assistant June 18 for Sam, noted. Happy to brainstorm gift ideas anytime.
13 user Pouring rain today, so gloomy. Also: keep your answers short and direct.
14 assistant Understood: short and direct answers from now on.
messages: 14
size: 166 words / 0.75 = about 221 tokens
One call later, this is what the extractor wrote.
category fact tokens
------------ --------------------------------------------------------------------- ------
semantic Maya moved from London to Paris for a new role at a payments startup. ~19
semantic Sam has a birthday on June 18. ~9
semantic Maya is vegetarian. ~4
procedural Maya prefers short, direct answers. ~7
4 facts stored, about 39 tokens total
ephemera that must produce NO memory (substring scan over the facts):
Tokyo time question 'tokyo' appears in 0 facts OK, dropped
Monday joke 'monday' appears in 0 facts OK, dropped
email rephrase request 'email' appears in 0 facts OK, dropped
rain smalltalk 'rain' appears in 0 facts OK, dropped
Four cards kept. Four pieces of ephemera dropped. Notice message 13: one sentence held both rain chat and a standing instruction, and the extractor split them correctly.
Notice also that the birthday is semantic, not episodic. A birthday comes back every year, so it is a stable fact about Sam, not one event that happened once.
Write once, read forever
The transcript is read once, by the extractor. The facts are read in every future session. This asymmetry is why extraction is worth paying for.
transcript about 221 tokens memory store 19 + 9 + 4 + 7 = 39 tokens compression 221 / 39 = about 5.7x smaller over 100 future sessions replay the transcript 100 x 221 = 22,100 prompt tokens inject the facts 100 x 39 = 3,900 prompt tokens
The extraction call cost 527 prompt tokens, one time. The saving repeats in every session after it. And real sessions are far less dense than this demo. The class baseline is 1,500 tokens of transcript for about 3 facts of about 25 tokens each:
baseline session 1,500 tokens baseline facts 3 x 25 = 75 tokens compression 1,500 / 75 = 20x smaller
The call took 7.2 seconds on a laptop with a small local model. The class budgets about 1 second on a hosted model. Either way nobody waits, because extraction runs after the answer has been sent. Chapter 13 is about that.
A card must stand on its own
A self-contained fact is a sentence that makes full sense with nothing around it.
A card will be read months later, alone, next to cards from other conversations. The words "she", "there" and "last week" point at things that are no longer on the desk.
| Weak card | Strong card | What was fixed |
|---|---|---|
| She moved. | Maya moved from London to Paris in March 2026. | name, places, date |
| His birthday is next month. | Maya's partner Sam has a birthday on June 18. | who, exact date |
| Went there last week. | Maya visited Lisbon in the week of June 15, 2026. | "there" and "last week" resolved |
The extractor is the only one who can fix these. It is the only one who sees the whole conversation and the date on which it happened. So the rule is: resolve every name and every date at write time. "Last month", written in April 2026, becomes "March 2026" on the card.
The rough edge in the experiment
Look again at the second card: "Sam has a birthday on June 18." The experiment's prompt asked for facts that start with the user's name, and even gave "Maya's partner Sam has a birthday on June 18" as an example. The model dropped "Maya's partner" anyway. That drift is normal for a small model.
The fix that worked was not more prompt text. Extra instructions only made this model drop facts. The fix
is structural: give the JSON its own subject field, so the model fills in a slot instead of
being asked to shape a sentence.
Why the secretary's mistakes are the dangerous ones
An extraction error is invisible and permanent.
When the assistant gives a bad answer, the user sees it and complains. When the extractor misses a fact, nothing happens. No error appears. No log turns red. The fact is simply not in the box, and the transcript it came from will not be read again. Weeks later the assistant suggests peanut butter, and by then nobody can tell which session went wrong.
kind of mistake who notices when bad answer the user at once bad retrieval the user, maybe when the answer is off missed fact nobody never, or weeks later wrong fact on a card nobody it is trusted and injected for months
Two defences exist, and hanumemAI uses both. (hanumemAI is the memory library built alongside this book. Chapter 20 introduces it.) First, every card keeps a pointer to the turns it came from, so a wrong card can be traced to its sentence. Second, the raw turns are stored and searchable too, so what the secretary missed can still be found. We come back to both in Chapter 20.
What the lab found
A lab result is a score on a benchmark, a public test made of long conversations and questions about them. We use three: LoCoMo, LongMemEval and BEAM. Chapter 18 describes them. The score is the percentage of questions answered correctly.
Unless we say otherwise, a score in this book is on the LoCoMo development set of 385 questions, with qwen3.7-flash as extractor, answer model and judge. The judge is the model that marks each answer right or wrong. The same setup run twice can differ by about 1 point (experiment 0019), so we call a smaller difference noise.
We ran many extraction experiments while building hanumemAI. Most of them lost. The losses teach more than the wins.
Two prompt rewrites that made the extractor write more facts also lost. The first raised the facts from 484 to 662 and the score fell from 91.17 to 89.61 (0007). The second raised the facts on BEAM from 355 to 402 and the BEAM score fell from 65 to 60 (E0002). The pattern is steady: more cards is not better. Each extra card competes for the few places on the desk.
Common questions
Why not store the whole conversation and skip extraction?
Storing it is fine, and hanumemAI does keep the raw turns. Reading it is the problem. A transcript is 20 times larger than its facts, full of "she" and "last week", and most of it is greetings. Facts are the index. The best results we measured come from keeping both: facts to find things, raw turns for the exact words.
Does extraction run after every message?
No. In the class design it runs once per session, after the session ends. One call reads the whole transcript. That is cheaper, and the extractor decides better when it can see the full conversation.
What if the extractor invents a fact?
It can happen. The defence is provenance, which means a record of where something came from: each card records which turns it came from. A card with no supporting sentence can be found and removed, and the user can always view and delete cards.
Should I use my best model as the extractor?
Not by default. Extraction is a narrow, repetitive task that runs on every session, so cost matters, and our experiment E0001 shows a stronger model can even score lower. Pick a small model that follows JSON reliably, then measure before changing it.
Does the assistant's side of the chat produce facts too?
It can. "The assistant recommended the Belem neighbourhood" is useful later. But text that came from a web page or a tool must never become a card, because that is how attackers plant false memories. Chapter 15 covers this.
How many facts should a session produce?
As many as it truly contains. The average is about 3. A session of smalltalk gives 0, and the fact-dense demo above gives 4. If your average climbs far above that, the extractor is probably writing ephemera.
Carry this
- The extractor is one model call per session: temperature 0, strict JSON, permission to return nothing.
- 221 tokens of transcript became 4 facts and 39 tokens, 5.7 times smaller. Over 100 sessions that is 3,900 tokens instead of 22,100.
- A card must stand alone: real names, real dates, no "she" and no "last week".
- Extraction errors are invisible and permanent, so keep the source of every card and keep the raw turns.
- More cards is not better. One call, one job, whole session, one attribute per fact when updates must be exact.
Check yourself
1. Maya writes: "Ugh, traffic was awful. Anyway, my daughter started swimming lessons on Saturdays." What should the extractor store?
Answer
One card, something like "Maya's daughter has swimming lessons on Saturdays." The traffic complaint is ephemera. The card names Maya, because "my daughter" would mean nothing when read alone.
2. A session is 1,800 tokens and yields 4 facts of 25 tokens each. How much smaller is the memory than the transcript, and how many tokens are saved over 50 future sessions if all four facts are injected each time?
Answer
facts 4 x 25 = 100 tokens compression 1,800 / 100 = 18x smaller replay 50 x 1,800 = 90,000 tokens inject 50 x 100 = 5,000 tokens saved 90,000 - 5,000 = 85,000 tokens
3. Why did asking the extractor for a session summary in the same call lower the score?
Answer
The call had two jobs, and the first one suffered. The regular facts dropped from 403 to 332 and the score fell by about 3.4 points (experiment 0016). Keep the extraction call single-purpose. If a summary is needed, use a separate call.