Part III · Journeyman
Chapter 14
Models and prompts
Three model roles, four prompts, and what our experiments say about which words and which models actually move the score.
On this page
- The three model roles
- Choosing the extractor
- Prompt one: extraction
- Prompt two: injection
- Prompt three: the update decision
- Prompt four: the answer prompt is part of the system
- When the prompt is full, change the reader
- Changing the embedder costs more than it looks
- Try it yourself
- Common questions
- Carry this
- Check yourself
In this chapter, we will learn which models a memory layer needs, what job each one does, and what we write in the prompts that drive them. We will count the calls each model receives per day, read the prompts line by line, and look at what the hanumemAI experiments found when we changed them.
The three model roles
A memory layer uses three kinds of model. An assistant model answers. A small model writes memories. An embedding model turns text into numbers for search.
Let us count how often each one is called, using the class baseline from Chapter 5.
daily active users 200,000
sessions per day 200,000 x 2 = 400,000
user messages per day 400,000 x 6 = 2,400,000
assistant LLM one call per message 2,400,000 calls/day
extractor LLM one call per session 400,000 calls/day
update decision one call per candidate
400,000 x 3 candidates = 1,200,000 calls/day
write path total 400,000 + 1,200,000 = 1,600,000 calls/day
HOT read path, 2.4M calls/day the user is waiting
[ assistant LLM ] the best model we can afford
WARM write path, 1.6M calls/day nobody is waiting
[ extractor LLM ] small, tidy, 400k calls/day
[ update-decision LLM ] the same small model, 1.2M calls/day
BOTH every write and every read
[ embedding model ] 768 numbers per text, never swapped lightly
The assistant model: hot
The assistant model is on the hot path, which means a person is watching the cursor blink while it runs. It is usually the largest and most expensive model in the system. The memory layer does not choose it. The memory layer only hands it about 75 tokens of facts and gets out of the way.
The extractor and the update decision: warm
These two jobs run in the background, after the reply has been sent (Chapter 13). Nobody waits, so a second or two is fine. Because they run 1.6 million times a day, they must be cheap. One small model does both jobs.
The embedder: married to the store
An embedding model turns a sentence into a list of numbers, called a vector, so that sentences with similar meaning get similar numbers. The class uses 768 numbers per sentence.
Every memory in the store was filed using one embedding model. A question must be turned into numbers by the same model, or the numbers do not line up. So the embedder is married to the store. Changing it means re-embedding every memory we have, all 200 million of them. We choose it carefully, once.
Choosing the extractor
A good extractor is a small model that follows a JSON format without drifting, run at temperature 0.
Temperature is a dial for randomness. At 0 the model picks the most likely next token every time, so the same transcript gives close to the same facts. A secretary should be boring.
JSON-disciplined means the model returns exactly the structure we asked for, with no chatty sentence before or after it. A program reads the output, not a person. One stray word breaks the parse and the session's memories are lost.
| Where it runs | Tool | Good for |
|---|---|---|
| On your laptop | Ollama | learning, privacy, zero cost per call; the class uses
qwen2.5:7b-instruct |
| Hosted | OpenRouter | many small models behind one API; about 1 second per call |
LoCoMo, LongMemEval and BEAM are the three public exams for memory systems. Chapter 18 explains them. A score is the percentage of questions answered correctly. The judge is the model that marks each answer right or wrong. Unless we say otherwise, the answer model and the judge in this chapter are both qwen3.7-flash.
Prompt one: extraction
The extraction prompt tells the small model what deserves a card and what does not. This is the class version, complete.
SYSTEM = (
"You extract durable facts about the user from a chat transcript: "
"only facts worth remembering in future sessions. Return a STRICT "
'JSON array of {"fact": "...", "category": "..."}, category one of: '
"semantic (facts), episodic (events), procedural. One short sentence each. "
"Extract as many as the transcript truly contains, possibly none: "
"return an empty array if nothing is worth remembering. "
"Standing instructions about how to answer ARE durable. Do NOT "
"store smalltalk, jokes, or one-off requests: those are ephemeral."
)
resp = client.chat.completions.create(
model=MODEL, temperature=0,
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": transcript}])
Read it as four instructions. What to keep: durable facts. The shape of the answer: strict JSON. The permission to return nothing: an empty array. What to drop: smalltalk, jokes, one-off requests. The permission to return nothing matters most. Without it, a model asked for facts will invent some.
Prompt two: injection
The injection prompt is how retrieved memories are placed on the assistant's desk.
SYSTEM: You are Maya's personal assistant. Known facts about the user (most relevant first, with dates): - Maya has two kids, aged 6 and 9. (since 2026-01) - Maya is training for the Berlin marathon. (since 2026-03) - Maya's partner Sam has a birthday on June 18. (since 2026-02) Personalize the answer using these facts where relevant. Do not recite these facts back unless asked. USER: Any ideas for a weekend activity with the family?
Three details are doing work here. The facts carry dates, so the model can tell old from new. The line "where relevant" lets the model ignore a fact that does not fit. The last line stops the assistant from sounding like a stalker who lists everything it knows.
Build a prompt below. Change the number of memories and the number of history turns, and watch how small the memory share stays while history takes over the bar.
Prompt three: the update decision
The update-decision prompt shows the model one new fact next to the most similar stored facts, and asks for exactly one operation.
INPUT one candidate fact
+ the 2 most similar stored memories, with ids and similarity scores
OUTPUT strict JSON, one operation
{"op": "ADD" | "UPDATE" | "DELETE" | "NOOP", "target_id": ..., "memory": ...}
We met the four operations in Chapter 10. The prompt is short, but it is called three times per session on average, so its length is a real cost. The class prices it at 550 input tokens per call.
Prompt four: the answer prompt is part of the system
The answer prompt is the set of instructions that tell the assistant how to read the memories it was handed. It is easy to treat retrieval as the whole memory system and the answer prompt as somebody else's problem. Our experiments say the opposite.
Here is what changing the answer prompt taught us, one experiment at a time.
| Change | Result | Lesson |
|---|---|---|
| "One short sentence" became "one to three sentences with every specific detail" (A0001) | LongMemEval sample 78 to 82, LoCoMo 92.21 to 92.99, BEAM rubric mean (its partial-credit score) 52.3 to 60.3 | Names, numbers and dates are what judges reward. They cost nothing in context tokens. |
| Chain-of-note: list the relevant lines, then answer (0005) | +0.26, inside the noise, with 3 times the output tokens | A reasoning scaffold does not turn a small model into a bigger one. |
| A mandatory contradiction check on yes/no questions (B0006, then L0001b) | BEAM 56 to 65, but LongMemEval 78 to 70 | It treated a first quote and a final price as a contradiction. A rule must also say what it does not cover. |
| One more rule, for indirect updates (A0003) | BEAM 65.25 to 61, LongMemEval 88 to 86 | The prompt saturated at about eight rules. Each new rule now costs another question type. |
The third row deserves a second look. A rule that helps one benchmark can hurt another. After that result, every change we keep is checked on all three benchmarks, and we prefer rules that begin with "when the question asks for...". A conditional rule fires only where it belongs.
When the prompt is full, change the reader
Once the prompt stopped improving, we tried a stronger answer model with the same memories on the desk.
LongMemEval-S, all 500 questions, about 2,950 context tokens per question (memory layer unchanged: qwen3.7-flash extraction, qwen3-embedding-4b vectors) answer model qwen3.7-flash 81.0 answer model qwen3.8-flash 88.6 gain 88.6 - 81.0 = 7.6 points
The judge on both runs was qwen3.7-flash, using the LongMemEval paper's judge prompts. Evidence recall on that run is 98.1 percent. Evidence recall is the share of the needed source messages that reached the prompt. So the right memories were already on the desk. The smaller model could not always do the date arithmetic or add up numbers from different sessions.
The gain depends on the exam.
answer model flash 3.8-flash gain context tokens LongMemEval-S, 500 questions 81.0 88.6 7.6 about 2,950 BEAM-100K, 400 questions 65.75 72.5 6.75 about 3,760 BEAM-1M, 200 questions 56.5 64.0 7.5 about 3,490 LoCoMo, 1,540 questions 87.73 90.78 3.05 about 1,470 (full sets, judge qwen3.7-flash on every row; tokens per question, from the 3.8-flash runs)
On the two harder exams the stronger reader was worth about 7 points. That is more than any prompt rule after the first few. On LoCoMo, where the small reader already does well, it was worth 3 on the full 1,540 questions and 2 on the 385 questions of the development set (experiment 0004).
Changing the embedder costs more than it looks
We said the embedder is married to the store. We tested the divorce once (experiment E0003), moving from an 8B embedding model with 4,096 numbers per vector to a 4B model with 2,560. The "B" counts the billions of parameters in a model, a measure of its size.
vector size 4,096 -> 2,560 numbers 1 - 2,560/4,096 = 37.5% smaller call latency 0.2 to 10 s -> about 150 ms, roughly 10x faster LoCoMo full 88.51 -> 87.79 inside the noise LongMemEval-100 88 -> 89 inside the noise BEAM-400 65.25 -> 65.25 flat re-ingest cost $1.80 across the three sets
The scores held, so we kept the smaller model. The surprise was the bill. The extractor is shown the ten most similar existing memories. A different embedder picks different neighbours, so every extraction prompt changes, and every extraction call must be paid again. Changing the embedder re-pays the whole write path.
Try it yourself
The class experiment fact-extraction runs a longer version of the extraction prompt on a
14-message session. It
keeps four facts and drops all four pieces of smalltalk.
ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/fact-extraction
pip install -r requirements.txt
python main.py
category fact tokens
------------ --------------------------------------------------------------------- ------
semantic Maya moved from London to Paris for a new role at a payments startup. ~19
semantic Sam has a birthday on June 18. ~9
semantic Maya is vegetarian. ~4
procedural Maya prefers short, direct answers. ~7
Now experiment. Delete the sentence "return an empty array if nothing is worth remembering" and feed it a session of pure smalltalk. Count the facts it invents. Then raise the temperature to 0.8 and run it three times. Count how many different answers you get.
Common questions
Can one model do all three jobs?
The assistant and the extractor can be the same model, and in a small project they often are. It is simply wasteful at scale, because the write path makes 1.6 million calls a day and does not need a frontier model. The embedder is always a separate, much smaller model built for that one task.
Why is the second fact "Sam has a birthday" and not "Maya's partner Sam"?
The experiment's prompt is a longer version of the one above. It asks for sentences that start with the
user's name, and the small model drifted. More prompt text made this model drop facts. The structural fix is better: give the JSON a separate
subject field instead of asking for a sentence in a particular shape.
Should the extraction prompt be long and detailed?
Longer is not safer. We revised hanumemAI's extraction prompt three times, and all three lost. Two revisions (0007, E0002) produced more facts and a lower score, because extra facts pushed the original messages out of the retrieved set. The third (0016) asked the same call for a second output, a session summary. The ordinary facts got worse and LoCoMo lost 3.4 points.
How do I know whether my answer prompt or my retrieval is the problem?
Look at a wrong answer and check whether the needed fact was on the desk. If it was, the reader ignored it, and the prompt or the answer model is at fault. If it was not, retrieval is at fault. Chapter 17 turns this into a routine.
Is "never say I don't know" safe in a real product?
No. That rule was right for a benchmark where every question has an answer. When we allowed abstention, LoCoMo lost 2.6 points to eleven wrong refusals (B0004). A real assistant should be able to say it does not know. The lesson is that the prompt must match the questions it will actually face.
What does temperature 0 not guarantee?
Identical output. Hosted providers are not perfectly deterministic. We re-ran an unchanged configuration and 63 of 385 answers changed, moving the score by about one point. Treat any gain smaller than that as noise.
Carry this
- Three roles. The assistant model is hot: 2.4M calls a day. The extractor and update decision are warm: 1.6M calls a day. The embedder serves both and is married to the store.
- The extractor is small, returns strict JSON, runs at temperature 0, and is allowed to return nothing.
- Injected facts carry dates, and the assistant is told not to recite them.
- The answer prompt is part of the memory system. Ours was worth about 4 points before any retrieval work, and it saturated at about eight rules.
- When the right memories are on the desk and answers are still wrong, a stronger reader is the next step. It took LongMemEval from 81.0 to 88.6, with the same judge and the same 2,950 context tokens.
Check yourself
1. A product has 50,000 daily active users, 3 sessions a day each, and the extractor finds 4 candidate facts per session. How many write-path LLM calls per day?
Answer
sessions per day 50,000 x 3 = 150,000 extraction calls update decisions 150,000 x 4 = 600,000 write path total 150,000 + 600,000 = 750,000 calls/day
2. Your team wants to switch to a newer embedding model next week. Name two costs beyond the price of the new model.
Answer
Every stored memory must be re-embedded, because old and new vectors do not line up. And if the extractor is shown similar existing memories, its prompts change too, so extraction is paid again. In our case that was $1.80 for three benchmark sets.
3. A new answer rule raises your score on one test set by 9 points. What do you do before keeping it?
Answer
Run it on the other test sets. Our contradiction rule gained 9 on BEAM and lost 8 on LongMemEval. Then narrow the rule so it fires only on the question shape it was written for.