Carebun Interactive/Memory Memory labCheat sheet

Part III · Journeyman

Chapter 14

Models and prompts

Three model roles, four prompts, and what our experiments say about which words and which models actually move the score.

11 min read · interactive

On this page
  1. The three model roles
  2. Choosing the extractor
  3. Prompt one: extraction
  4. Prompt two: injection
  5. Prompt three: the update decision
  6. Prompt four: the answer prompt is part of the system
  7. When the prompt is full, change the reader
  8. Changing the embedder costs more than it looks
  9. Try it yourself
  10. Common questions
  11. Carry this
  12. Check yourself

In this chapter, we will learn which models a memory layer needs, what job each one does, and what we write in the prompts that drive them. We will count the calls each model receives per day, read the prompts line by line, and look at what the hanumemAI experiments found when we changed them.

The three model roles

A memory layer uses three kinds of model. An assistant model answers. A small model writes memories. An embedding model turns text into numbers for search.

Let us count how often each one is called, using the class baseline from Chapter 5.

daily active users                           200,000
sessions per day         200,000 x 2      =   400,000
user messages per day    400,000 x 6      = 2,400,000

assistant LLM      one call per message      2,400,000 calls/day
extractor LLM      one call per session        400,000 calls/day
update decision    one call per candidate
                   400,000 x 3 candidates  = 1,200,000 calls/day
write path total   400,000 + 1,200,000     = 1,600,000 calls/day
HOT   read path, 2.4M calls/day      the user is waiting
      [ assistant LLM ]              the best model we can afford

WARM  write path, 1.6M calls/day     nobody is waiting
      [ extractor LLM ]              small, tidy, 400k calls/day
      [ update-decision LLM ]        the same small model, 1.2M calls/day

BOTH  every write and every read
      [ embedding model ]            768 numbers per text, never swapped lightly

The assistant model: hot

The assistant model is on the hot path, which means a person is watching the cursor blink while it runs. It is usually the largest and most expensive model in the system. The memory layer does not choose it. The memory layer only hands it about 75 tokens of facts and gets out of the way.

The extractor and the update decision: warm

These two jobs run in the background, after the reply has been sent (Chapter 13). Nobody waits, so a second or two is fine. Because they run 1.6 million times a day, they must be cheap. One small model does both jobs.

The embedder: married to the store

An embedding model turns a sentence into a list of numbers, called a vector, so that sentences with similar meaning get similar numbers. The class uses 768 numbers per sentence.

Every memory in the store was filed using one embedding model. A question must be turned into numbers by the same model, or the numbers do not line up. So the embedder is married to the store. Changing it means re-embedding every memory we have, all 200 million of them. We choose it carefully, once.

Choosing the extractor

A good extractor is a small model that follows a JSON format without drifting, run at temperature 0.

Temperature is a dial for randomness. At 0 the model picks the most likely next token every time, so the same transcript gives close to the same facts. A secretary should be boring.

JSON-disciplined means the model returns exactly the structure we asked for, with no chatty sentence before or after it. A program reads the output, not a person. One stray word breaks the parse and the session's memories are lost.

Where it runsToolGood for
On your laptopOllamalearning, privacy, zero cost per call; the class uses qwen2.5:7b-instruct
HostedOpenRoutermany small models behind one API; about 1 second per call

LoCoMo, LongMemEval and BEAM are the three public exams for memory systems. Chapter 18 explains them. A score is the percentage of questions answered correctly. The judge is the model that marks each answer right or wrong. Unless we say otherwise, the answer model and the judge in this chapter are both qwen3.7-flash.

Prompt one: extraction

The extraction prompt tells the small model what deserves a card and what does not. This is the class version, complete.

SYSTEM = (
    "You extract durable facts about the user from a chat transcript: "
    "only facts worth remembering in future sessions. Return a STRICT "
    'JSON array of {"fact": "...", "category": "..."}, category one of: '
    "semantic (facts), episodic (events), procedural. One short sentence each. "
    "Extract as many as the transcript truly contains, possibly none: "
    "return an empty array if nothing is worth remembering. "
    "Standing instructions about how to answer ARE durable. Do NOT "
    "store smalltalk, jokes, or one-off requests: those are ephemeral."
)
resp = client.chat.completions.create(
    model=MODEL, temperature=0,
    messages=[{"role": "system", "content": SYSTEM},
              {"role": "user", "content": transcript}])

Read it as four instructions. What to keep: durable facts. The shape of the answer: strict JSON. The permission to return nothing: an empty array. What to drop: smalltalk, jokes, one-off requests. The permission to return nothing matters most. Without it, a model asked for facts will invent some.

Prompt two: injection

The injection prompt is how retrieved memories are placed on the assistant's desk.

SYSTEM:
You are Maya's personal assistant.
Known facts about the user (most relevant first, with dates):
- Maya has two kids, aged 6 and 9.            (since 2026-01)
- Maya is training for the Berlin marathon.    (since 2026-03)
- Maya's partner Sam has a birthday on June 18. (since 2026-02)
Personalize the answer using these facts where relevant.
Do not recite these facts back unless asked.

USER:
Any ideas for a weekend activity with the family?

Three details are doing work here. The facts carry dates, so the model can tell old from new. The line "where relevant" lets the model ignore a fact that does not fit. The last line stops the assistant from sounding like a stalker who lists everything it knows.

Build a prompt below. Change the number of memories and the number of history turns, and watch how small the memory share stays while history takes over the bar.

Prompt three: the update decision

The update-decision prompt shows the model one new fact next to the most similar stored facts, and asks for exactly one operation.

INPUT    one candidate fact
         + the 2 most similar stored memories, with ids and similarity scores

OUTPUT   strict JSON, one operation
         {"op": "ADD" | "UPDATE" | "DELETE" | "NOOP", "target_id": ..., "memory": ...}

We met the four operations in Chapter 10. The prompt is short, but it is called three times per session on average, so its length is a real cost. The class prices it at 550 input tokens per call.

Prompt four: the answer prompt is part of the system

The answer prompt is the set of instructions that tell the assistant how to read the memories it was handed. It is easy to treat retrieval as the whole memory system and the answer prompt as somebody else's problem. Our experiments say the opposite.

Here is what changing the answer prompt taught us, one experiment at a time.

ChangeResultLesson
"One short sentence" became "one to three sentences with every specific detail" (A0001) LongMemEval sample 78 to 82, LoCoMo 92.21 to 92.99, BEAM rubric mean (its partial-credit score) 52.3 to 60.3 Names, numbers and dates are what judges reward. They cost nothing in context tokens.
Chain-of-note: list the relevant lines, then answer (0005) +0.26, inside the noise, with 3 times the output tokens A reasoning scaffold does not turn a small model into a bigger one.
A mandatory contradiction check on yes/no questions (B0006, then L0001b) BEAM 56 to 65, but LongMemEval 78 to 70 It treated a first quote and a final price as a contradiction. A rule must also say what it does not cover.
One more rule, for indirect updates (A0003) BEAM 65.25 to 61, LongMemEval 88 to 86 The prompt saturated at about eight rules. Each new rule now costs another question type.

The third row deserves a second look. A rule that helps one benchmark can hurt another. After that result, every change we keep is checked on all three benchmarks, and we prefer rules that begin with "when the question asks for...". A conditional rule fires only where it belongs.

When the prompt is full, change the reader

Once the prompt stopped improving, we tried a stronger answer model with the same memories on the desk.

LongMemEval-S, all 500 questions, about 2,950 context tokens per question
(memory layer unchanged: qwen3.7-flash extraction, qwen3-embedding-4b vectors)

answer model qwen3.7-flash     81.0
answer model qwen3.8-flash     88.6
gain                           88.6 - 81.0 = 7.6 points

The judge on both runs was qwen3.7-flash, using the LongMemEval paper's judge prompts. Evidence recall on that run is 98.1 percent. Evidence recall is the share of the needed source messages that reached the prompt. So the right memories were already on the desk. The smaller model could not always do the date arithmetic or add up numbers from different sessions.

The gain depends on the exam.

answer model                     flash     3.8-flash   gain    context tokens
LongMemEval-S, 500 questions     81.0      88.6        7.6     about 2,950
BEAM-100K, 400 questions         65.75     72.5        6.75    about 3,760
BEAM-1M, 200 questions           56.5      64.0        7.5     about 3,490
LoCoMo, 1,540 questions          87.73     90.78       3.05    about 1,470
(full sets, judge qwen3.7-flash on every row; tokens per question, from the 3.8-flash runs)

On the two harder exams the stronger reader was worth about 7 points. That is more than any prompt rule after the first few. On LoCoMo, where the small reader already does well, it was worth 3 on the full 1,540 questions and 2 on the 385 questions of the development set (experiment 0004).

Changing the embedder costs more than it looks

We said the embedder is married to the store. We tested the divorce once (experiment E0003), moving from an 8B embedding model with 4,096 numbers per vector to a 4B model with 2,560. The "B" counts the billions of parameters in a model, a measure of its size.

vector size     4,096 -> 2,560 numbers     1 - 2,560/4,096 = 37.5% smaller
call latency    0.2 to 10 s -> about 150 ms, roughly 10x faster
LoCoMo full     88.51 -> 87.79             inside the noise
LongMemEval-100 88    -> 89                inside the noise
BEAM-400        65.25 -> 65.25             flat
re-ingest cost  $1.80 across the three sets

The scores held, so we kept the smaller model. The surprise was the bill. The extractor is shown the ten most similar existing memories. A different embedder picks different neighbours, so every extraction prompt changes, and every extraction call must be paid again. Changing the embedder re-pays the whole write path.

Try it yourself

The class experiment fact-extraction runs a longer version of the extraction prompt on a 14-message session. It keeps four facts and drops all four pieces of smalltalk.

ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/fact-extraction
pip install -r requirements.txt
python main.py
    category     fact                                                                  tokens
    ------------ --------------------------------------------------------------------- ------
    semantic     Maya moved from London to Paris for a new role at a payments startup. ~19
    semantic     Sam has a birthday on June 18.                                        ~9
    semantic     Maya is vegetarian.                                                   ~4
    procedural   Maya prefers short, direct answers.                                   ~7

Now experiment. Delete the sentence "return an empty array if nothing is worth remembering" and feed it a session of pure smalltalk. Count the facts it invents. Then raise the temperature to 0.8 and run it three times. Count how many different answers you get.

Common questions

Can one model do all three jobs?

The assistant and the extractor can be the same model, and in a small project they often are. It is simply wasteful at scale, because the write path makes 1.6 million calls a day and does not need a frontier model. The embedder is always a separate, much smaller model built for that one task.

Why is the second fact "Sam has a birthday" and not "Maya's partner Sam"?

The experiment's prompt is a longer version of the one above. It asks for sentences that start with the user's name, and the small model drifted. More prompt text made this model drop facts. The structural fix is better: give the JSON a separate subject field instead of asking for a sentence in a particular shape.

Should the extraction prompt be long and detailed?

Longer is not safer. We revised hanumemAI's extraction prompt three times, and all three lost. Two revisions (0007, E0002) produced more facts and a lower score, because extra facts pushed the original messages out of the retrieved set. The third (0016) asked the same call for a second output, a session summary. The ordinary facts got worse and LoCoMo lost 3.4 points.

How do I know whether my answer prompt or my retrieval is the problem?

Look at a wrong answer and check whether the needed fact was on the desk. If it was, the reader ignored it, and the prompt or the answer model is at fault. If it was not, retrieval is at fault. Chapter 17 turns this into a routine.

Is "never say I don't know" safe in a real product?

No. That rule was right for a benchmark where every question has an answer. When we allowed abstention, LoCoMo lost 2.6 points to eleven wrong refusals (B0004). A real assistant should be able to say it does not know. The lesson is that the prompt must match the questions it will actually face.

What does temperature 0 not guarantee?

Identical output. Hosted providers are not perfectly deterministic. We re-ran an unchanged configuration and 63 of 385 answers changed, moving the score by about one point. Treat any gain smaller than that as noise.

Carry this

  • Three roles. The assistant model is hot: 2.4M calls a day. The extractor and update decision are warm: 1.6M calls a day. The embedder serves both and is married to the store.
  • The extractor is small, returns strict JSON, runs at temperature 0, and is allowed to return nothing.
  • Injected facts carry dates, and the assistant is told not to recite them.
  • The answer prompt is part of the memory system. Ours was worth about 4 points before any retrieval work, and it saturated at about eight rules.
  • When the right memories are on the desk and answers are still wrong, a stronger reader is the next step. It took LongMemEval from 81.0 to 88.6, with the same judge and the same 2,950 context tokens.

Check yourself

1. A product has 50,000 daily active users, 3 sessions a day each, and the extractor finds 4 candidate facts per session. How many write-path LLM calls per day?

Answer
sessions per day    50,000 x 3       = 150,000     extraction calls
update decisions    150,000 x 4      = 600,000
write path total    150,000 + 600,000 = 750,000 calls/day

2. Your team wants to switch to a newer embedding model next week. Name two costs beyond the price of the new model.

Answer

Every stored memory must be re-embedded, because old and new vectors do not line up. And if the extractor is shown similar existing memories, its prompts change too, so extraction is paid again. In our case that was $1.80 for three benchmark sets.

3. A new answer rule raises your score on one test set by 9 points. What do you do before keeping it?

Answer

Run it on the other test sets. Our contradiction rule gained 9 on BEAM and lost 8 on LongMemEval. Then narrow the rule so it fires only on the question shape it was written for.