Carebun Interactive/Memory Memory labCheat sheet

Part II · Apprentice

Chapter 6

The secretary: fact extraction

One model reads the chat once and decides what is worth a card. We watch it turn 221 tokens into 39, and learn why its mistakes are the ones nobody sees.

10 min read · interactive

On this page
  1. What is fact extraction?
  2. What deserves a card?
  3. The extraction prompt
  4. Play the secretary
  5. Try it yourself: one session, four cards
  6. Write once, read forever
  7. A card must stand on its own
  8. Why the secretary's mistakes are the dangerous ones
  9. What the lab found
  10. Common questions
  11. Carry this
  12. Check yourself

In this chapter, we will learn how a memory system decides what to remember. We will read the prompt that does the deciding and watch it work on a real conversation. Then we will see why this one step sets the quality of everything that comes after it.

What is fact extraction?

Fact extraction is one model call that reads a finished conversation and writes down the few facts worth keeping for future conversations.

It is the first step of the write path, the road a memory travels on its way into the box. The model that does it is called the extractor. It is a different job from answering the user, and it usually gets a smaller, cheaper model.

session transcript          extractor LLM             candidate facts
about 1,500 tokens   --->   temperature 0     --->    0 to N of them, about 3 on average
                            strict JSON               about 25 tokens each
                                 |
                                 v
                            dropped: smalltalk, jokes,
                            one-off requests, the weather

Look at the two arrows leaving the extractor. One goes to the box. The other goes to the bin. Deciding which is which is the whole job.

What deserves a card?

A fact deserves a card when it will still be useful in a future conversation. Everything else is ephemera, which means things that matter for a moment and then do not.

Maya saysCard?Why
"I'm vegetarian, for any future food suggestions."YesTrue next month, and it changes answers
"My partner Sam has a birthday on June 18."YesComes back every year
"Keep your answers short and direct."YesA standing instruction shapes every answer
"What time is it in Tokyo right now?"NoA one-off request, useless tomorrow
"Mondays should be illegal."NoA joke
"Pouring rain today, so gloomy."NoWeather chat

The test is simple to say and hard to apply: would a friend who knows Maya well remember this in a month? A friend remembers the allergy. A friend does not remember the rain.

The extraction prompt

The extractor is steered by a prompt, a short set of written instructions that goes in front of the transcript. Here is the one from the class notes.

SYSTEM = (
    "You extract durable facts about the user from a chat transcript: "
    "only facts worth remembering in future sessions. Return a STRICT "
    'JSON array of {"fact": "...", "category": "..."}, category one of: '
    "semantic (facts), episodic (events), procedural. One short sentence each. "
    "Extract as many as the transcript truly contains, possibly none: "
    "return an empty array if nothing is worth remembering. "
    "Standing instructions about how to answer ARE durable. Do NOT "
    "store smalltalk, jokes, or one-off requests: those are ephemeral."
)
resp = client.chat.completions.create(
    model=MODEL, temperature=0,
    messages=[{"role": "system", "content": SYSTEM},
              {"role": "user", "content": transcript}])

Four choices in this small prompt are worth slowing down for.

Strict JSON

JSON is a plain text format that programs can read without guessing. The extractor's reader is not a person. It is our code, which must put each fact into a database row. If the model replies "Sure! Here are some facts I found", the code breaks. So the prompt demands a JSON array and nothing else.

Temperature 0

Temperature is a dial for how adventurous the model's word choices are. At a high setting the model is creative. At 0 it takes the most likely word every time. A secretary should not be creative. We want the same transcript to produce the same cards.

"Possibly none"

The prompt allows an empty answer, so the extractor is never forced to invent. A model that is asked for facts will try hard to find some. If a session is pure smalltalk and the prompt demands three facts, we get three invented or worthless cards. The words "possibly none" and "return an empty array" give the model permission to write nothing. The average is about 3 facts per session, but the count follows the content. It is not a quota.

Categories

The category is the colour of the card. Every card gets one: semantic for facts, episodic for events, procedural for standing instructions. We met them in Chapter 4. The colour decides later how long the card lives, so it is set here, at the moment of writing.

Play the secretary

Before we see what the model did, try the job yourself. Here is one real session of 14 messages. Mark each line keep or drop, then reveal what the extractor kept, and compare the token counts.

Try it yourself: one session, four cards

The class experiment fact-extraction runs this step. It needs a local model and nothing else. Its prompt is a longer version of the one above. It adds an example for each category, and it asks that every fact starts with the user's name, Maya.

ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/fact-extraction
pip install -r requirements.txt
python main.py

The transcript is 14 messages, durable facts mixed with ephemera.

     1 user      I just moved from London to Paris for a new role at a payments startup.
     2 assistant Congratulations on the move and the new role! How can I help today?
     3 user      Quick one: what time is it in Tokyo right now?
     4 assistant Tokyo is 7 or 8 hours ahead of Paris depending on the season: ...
     5 user      Ha. Honestly, Mondays should be illegal. Worst invention ever.
     6 assistant A classic Monday complaint. Hang in there, the week gets better.
     7 user      Rephrase this email politely: 'The dashboard is broken again, fix it.'
     8 assistant 'Hi team, the dashboard seems down. Could someone take a look? Thanks!'
     9 user      Thanks. By the way, I'm vegetarian, for any future food suggestions.
    10 assistant Noted: vegetarian suggestions only from here on.
    11 user      My partner Sam has a birthday on June 18, I need to plan something.
    12 assistant June 18 for Sam, noted. Happy to brainstorm gift ideas anytime.
    13 user      Pouring rain today, so gloomy. Also: keep your answers short and direct.
    14 assistant Understood: short and direct answers from now on.

    messages: 14
    size: 166 words / 0.75 = about 221 tokens

One call later, this is what the extractor wrote.

    category     fact                                                                  tokens
    ------------ --------------------------------------------------------------------- ------
    semantic     Maya moved from London to Paris for a new role at a payments startup. ~19
    semantic     Sam has a birthday on June 18.                                        ~9
    semantic     Maya is vegetarian.                                                   ~4
    procedural   Maya prefers short, direct answers.                                   ~7

    4 facts stored, about 39 tokens total

    ephemera that must produce NO memory (substring scan over the facts):
      Tokyo time question      'tokyo'  appears in 0 facts   OK, dropped
      Monday joke              'monday' appears in 0 facts   OK, dropped
      email rephrase request   'email'  appears in 0 facts   OK, dropped
      rain smalltalk           'rain'   appears in 0 facts   OK, dropped

Four cards kept. Four pieces of ephemera dropped. Notice message 13: one sentence held both rain chat and a standing instruction, and the extractor split them correctly.

Notice also that the birthday is semantic, not episodic. A birthday comes back every year, so it is a stable fact about Sam, not one event that happened once.

Write once, read forever

The transcript is read once, by the extractor. The facts are read in every future session. This asymmetry is why extraction is worth paying for.

transcript                           about 221 tokens
memory store   19 + 9 + 4 + 7     =        39 tokens
compression    221 / 39           =  about 5.7x smaller

over 100 future sessions
  replay the transcript   100 x 221   =  22,100 prompt tokens
  inject the facts        100 x  39   =   3,900 prompt tokens

The extraction call cost 527 prompt tokens, one time. The saving repeats in every session after it. And real sessions are far less dense than this demo. The class baseline is 1,500 tokens of transcript for about 3 facts of about 25 tokens each:

baseline session     1,500 tokens
baseline facts       3 x 25       =  75 tokens
compression          1,500 / 75   =  20x smaller

The call took 7.2 seconds on a laptop with a small local model. The class budgets about 1 second on a hosted model. Either way nobody waits, because extraction runs after the answer has been sent. Chapter 13 is about that.

A card must stand on its own

A self-contained fact is a sentence that makes full sense with nothing around it.

A card will be read months later, alone, next to cards from other conversations. The words "she", "there" and "last week" point at things that are no longer on the desk.

Weak cardStrong cardWhat was fixed
She moved.Maya moved from London to Paris in March 2026.name, places, date
His birthday is next month.Maya's partner Sam has a birthday on June 18.who, exact date
Went there last week.Maya visited Lisbon in the week of June 15, 2026."there" and "last week" resolved

The extractor is the only one who can fix these. It is the only one who sees the whole conversation and the date on which it happened. So the rule is: resolve every name and every date at write time. "Last month", written in April 2026, becomes "March 2026" on the card.

The rough edge in the experiment

Look again at the second card: "Sam has a birthday on June 18." The experiment's prompt asked for facts that start with the user's name, and even gave "Maya's partner Sam has a birthday on June 18" as an example. The model dropped "Maya's partner" anyway. That drift is normal for a small model.

The fix that worked was not more prompt text. Extra instructions only made this model drop facts. The fix is structural: give the JSON its own subject field, so the model fills in a slot instead of being asked to shape a sentence.

Why the secretary's mistakes are the dangerous ones

An extraction error is invisible and permanent.

When the assistant gives a bad answer, the user sees it and complains. When the extractor misses a fact, nothing happens. No error appears. No log turns red. The fact is simply not in the box, and the transcript it came from will not be read again. Weeks later the assistant suggests peanut butter, and by then nobody can tell which session went wrong.

kind of mistake          who notices          when
bad answer               the user             at once
bad retrieval            the user, maybe      when the answer is off
missed fact              nobody               never, or weeks later
wrong fact on a card     nobody               it is trusted and injected for months

Two defences exist, and hanumemAI uses both. (hanumemAI is the memory library built alongside this book. Chapter 20 introduces it.) First, every card keeps a pointer to the turns it came from, so a wrong card can be traced to its sentence. Second, the raw turns are stored and searchable too, so what the secretary missed can still be found. We come back to both in Chapter 20.

What the lab found

A lab result is a score on a benchmark, a public test made of long conversations and questions about them. We use three: LoCoMo, LongMemEval and BEAM. Chapter 18 describes them. The score is the percentage of questions answered correctly.

Unless we say otherwise, a score in this book is on the LoCoMo development set of 385 questions, with qwen3.7-flash as extractor, answer model and judge. The judge is the model that marks each answer right or wrong. The same setup run twice can differ by about 1 point (experiment 0019), so we call a smaller difference noise.

We ran many extraction experiments while building hanumemAI. Most of them lost. The losses teach more than the wins.

Two prompt rewrites that made the extractor write more facts also lost. The first raised the facts from 484 to 662 and the score fell from 91.17 to 89.61 (0007). The second raised the facts on BEAM from 355 to 402 and the BEAM score fell from 65 to 60 (E0002). The pattern is steady: more cards is not better. Each extra card competes for the few places on the desk.

Common questions

Why not store the whole conversation and skip extraction?

Storing it is fine, and hanumemAI does keep the raw turns. Reading it is the problem. A transcript is 20 times larger than its facts, full of "she" and "last week", and most of it is greetings. Facts are the index. The best results we measured come from keeping both: facts to find things, raw turns for the exact words.

Does extraction run after every message?

No. In the class design it runs once per session, after the session ends. One call reads the whole transcript. That is cheaper, and the extractor decides better when it can see the full conversation.

What if the extractor invents a fact?

It can happen. The defence is provenance, which means a record of where something came from: each card records which turns it came from. A card with no supporting sentence can be found and removed, and the user can always view and delete cards.

Should I use my best model as the extractor?

Not by default. Extraction is a narrow, repetitive task that runs on every session, so cost matters, and our experiment E0001 shows a stronger model can even score lower. Pick a small model that follows JSON reliably, then measure before changing it.

Does the assistant's side of the chat produce facts too?

It can. "The assistant recommended the Belem neighbourhood" is useful later. But text that came from a web page or a tool must never become a card, because that is how attackers plant false memories. Chapter 15 covers this.

How many facts should a session produce?

As many as it truly contains. The average is about 3. A session of smalltalk gives 0, and the fact-dense demo above gives 4. If your average climbs far above that, the extractor is probably writing ephemera.

Carry this

  • The extractor is one model call per session: temperature 0, strict JSON, permission to return nothing.
  • 221 tokens of transcript became 4 facts and 39 tokens, 5.7 times smaller. Over 100 sessions that is 3,900 tokens instead of 22,100.
  • A card must stand alone: real names, real dates, no "she" and no "last week".
  • Extraction errors are invisible and permanent, so keep the source of every card and keep the raw turns.
  • More cards is not better. One call, one job, whole session, one attribute per fact when updates must be exact.

Check yourself

1. Maya writes: "Ugh, traffic was awful. Anyway, my daughter started swimming lessons on Saturdays." What should the extractor store?

Answer

One card, something like "Maya's daughter has swimming lessons on Saturdays." The traffic complaint is ephemera. The card names Maya, because "my daughter" would mean nothing when read alone.

2. A session is 1,800 tokens and yields 4 facts of 25 tokens each. How much smaller is the memory than the transcript, and how many tokens are saved over 50 future sessions if all four facts are injected each time?

Answer
facts           4 x 25           =    100 tokens
compression     1,800 / 100      =    18x smaller
replay          50 x 1,800       = 90,000 tokens
inject          50 x 100         =  5,000 tokens
saved           90,000 - 5,000   = 85,000 tokens

3. Why did asking the extractor for a session summary in the same call lower the score?

Answer

The call had two jobs, and the first one suffered. The regular facts dropped from 403 to 332 and the score fell by about 3.4 points (experiment 0016). Keep the extraction call single-purpose. If a summary is needed, use a separate call.