Carebun Interactive/Memory Memory labCheat sheet

Part I · Novice

Chapter 2

Three fixes that fail

Resend the history, buy a bigger window, or train a model per user. We test each with arithmetic, and each failure hands us a rule for the design.

9 min read · interactive

On this page
  1. The workload we test against
  2. Fix 1: resend all the history
  3. Fix 2: buy a bigger window
  4. Fix 3: train a model for each user
  5. What the three failures leave behind
  6. Try it yourself: the bill for one user
  7. Common questions
  8. Carry this
  9. Check yourself

In this chapter, we will learn why the three most obvious fixes for a forgetful assistant do not survive contact with numbers. We will test each one on paper, and we will keep what each failure teaches. By the end we will hold three rules, and those three rules are the design of every memory system in this book.

The workload we test against

A workload is a description of how much a system is used. We cannot judge a fix without one. The class gives us a workload for a personal assistant, and we use it for the whole book.

registered users                 1,000,000
daily active users                 200,000
sessions per user per day                2
turns per session                        6   (6 user messages, 6 replies)
tokens per session                   1,500   (both sides of the chat)
tokens per turn        1,500 / 6  =    250

A session is one sitting, from opening the chat to closing it. A turn is one user message with its reply. Keep 1,500 tokens per session in your pocket. Almost every calculation in this book starts from it.

Fix 1: resend all the history

Resending history means pasting every past conversation into every new message. It is the first idea everyone has, and Chapter 1 showed that it works for one session. Let us see how it ages.

sessions per year        2 x 365      =       730
transcript per year      730 x 1,500  = 1,095,000 tokens
largest window today                    1,000,000 tokens

1,095,000 > 1,000,000    one year of chat does not fit

So the fix has an expiry date. But it fails long before that date, because of cost. A model charges for every token it reads, and with this fix it reads the whole past on every message.

A frontier model is one of the most capable models on sale. The class prices it at $3 per million input tokens, the tokens the model reads.

price: frontier model, $3 per million input tokens

at session 10, resend the last 10 sessions
  history   10 x 1,500   =  15,000 tokens    15,000 x $3 / 1,000,000  = $0.045
  memory    top 3 facts  =      75 tokens        75 x $3 / 1,000,000  = $0.000225

  ratio     15,000 / 75  =  200        memory is 200 times cheaper

at session 100, resend everything
  history   100 x 1,500  = 150,000 tokens                             = $0.45
  memory    top 3 facts  =      75 tokens                             = $0.000225

  ratio     150,000 / 75 = 2,000       memory is 2,000 times cheaper

Where does 75 come from? A remembered fact is about 25 tokens, and we hand over the best three. 3 x 25 = 75. That number does not change when the user has chatted for ten sessions or for ten thousand.

Memory's read cost is flat. History's read cost grows every session. A flat line and a rising line always cross, and after they cross the gap only widens.

Move the sliders. Watch what happens to the ratio when you add sessions, and notice that the memory column does not move at all.

Fix 2: buy a bigger window

A bigger context window is a bigger desk: the model can see more text at once. Windows have grown from a few thousand tokens to a million. Will the next size solve the problem?

It moves the expiry date. It does not change the bill. Every message still pays for everything placed in the window. Take the window we already have, filled to the top with Maya's history.

a full 1,000,000 token window, $3 per million input tokens

cost of one message      1,000,000 x $3 / 1,000,000  =  $3.00
Maya's messages per day  2 sessions x 6 turns        =  12
cost per day for Maya    12 x $3.00                  =  $36.00

a window ten times larger, also full
cost of one message      10,000,000 x $3 / 1,000,000 =  $30.00
years of chat it holds   10,000,000 / 1,095,000      =  about 9

Thirty-six dollars a day for one user, to answer questions like "suggest a snack". A window ten times larger would hold about nine years of chat, and each message would cost thirty dollars. A bigger window makes the fix possible. It does not make it affordable.

There are two more costs, and they are not about money.

  • Time. A model needs longer to read a long prompt. The user waits for every token on the desk, including the ones that do not matter.
  • Attention. Models find a fact more easily at the start or the end of a long prompt than in the middle. One allergy sentence buried inside a million tokens of small talk is easy to miss, and it is the one sentence we cannot afford to miss.

Fix 3: train a model for each user

Fine-tuning means training a model further on new text, so that the new knowledge is stored in the model's numbers. If the model forgets Maya because its file is frozen, why not unfreeze it and teach it about Maya?

The class notes give this fix one word for what breaks: everything. Let us list it.

What we needWhat fine-tuning per user gives
One model serving everyone1,000,000 users means 1,000,000 models to train, store and keep loaded
Memory that is fresh by the next sessiona training run after every chat, 400,000 runs a day
"I moved to Paris" replaces "lives in London"no way to find and change one fact inside billions of numbers
The user can see what is storednothing to show; the knowledge is not written anywhere as text
The user can delete one memoryno reliable way to remove one fact without retraining
training runs per day  =  sessions per day
                       =  200,000 users x 2 sessions  =  400,000

Cost alone ends the idea. But the deeper lesson is the last three rows. A model's numbers are not a database. We cannot open them and read "Maya is vegetarian". We cannot correct it. We cannot prove to Maya that we deleted it.

What the three failures leave behind

FixWhat breaksThe rule it leaves
Resend all historycost and the windowstore distilled facts, not transcripts
Bigger context windowcost still grows per messageinject a small, fixed amount
Fine-tune per usereverythingkeep memory outside the model

Put the three rules in one picture and we have the system.

session transcript          extractor LLM            memory store          next session's prompt
  about 1,500 tokens  ---->   keeps 0 to N facts  ---->  outside the   ---->   top 3 facts
                              (about 3, about           model                 about 75 tokens
                              25 tokens each)

      Rule 1: distilled            Rule 3: outside the model            Rule 2: small and fixed

The extractor is a small LLM that reads a transcript and writes down the facts worth keeping. Chapter 3 introduces it.

We did not invent this design. We were pushed into it. Every other door was closed by arithmetic.

Try it yourself: the bill for one user

The class experiment stateless-llm prints the growth of the naive fix. It needs no special hardware, and this part of it is plain arithmetic.

(e) The naive fix does not survive contact with time

    assume one chat session adds ~1,500 tokens of transcript
    resend-everything prompt size:
      after   1 sessions:   1,500 tokens   ($0.0045 per message at $3/M input)
      after  10 sessions:  15,000 tokens   ($0.0450 per message at $3/M input)
      after 100 sessions: 150,000 tokens   ($0.4500 per message at $3/M input)
      after 100 sessions the prompt is 150,000 tokens: it still fits
      a 1,000,000 token frontier window, but every message pays for
      all of it, and one year of chat (730 sessions x 1,500 =
      1,095,000 tokens) overflows even that window.
    memory fix: 3 facts, ~25 extra prompt tokens in this run, flat forever
      (the class notes budget about 75 tokens for a top 3 of
      average-length facts; these three demo facts are shorter).

In the output, $3/M means $3 per million tokens. The run measured 25 tokens for its three facts, because they are short sentences. We keep the class budget of 75 tokens in our sums, which is the cautious choice.

Now do one on paper. Take a user in their second month, at session 100, who sends 12 messages a day.

history   12 messages x $0.45      =  $5.40   per day
memory    12 messages x $0.000225  =  $0.0027 per day

Five dollars and forty cents against a quarter of a cent. Same user, same questions, same model.

Common questions

Prices keep falling. Will resending history become cheap enough?

Cheaper prices shrink both columns by the same factor, so the ratio stays. At session 100 memory is 2,000 times cheaper at any price. And the transcript still overflows the window after a year, which no price can fix.

What about prompt caching, where repeated text costs less?

Caching gives a discount on text the model has read recently. It helps inside one session. It helps much less across days, because caches expire, and the cached text still fills the window and still dilutes attention. It is a useful discount on a design that does not scale, not a replacement for memory.

Could we resend only the last few sessions?

That caps the cost, and it is a fair short-term trick. But Maya told us about her peanut allergy once, months ago. A window of recent sessions forgets exactly the old, important facts. Age is a poor measure of importance.

Could we summarise the history instead of resending it?

A summary is a distilled form, so this is a step towards Rule 1. It works well inside a session, and Chapter 4 uses it for short-term memory. For long-term memory it has a flaw: each time a summary is rewritten, details can fall out, and one big summary is injected whole even when only one line of it matters. Separate facts can be picked one at a time.

Is fine-tuning useless, then?

No. Fine-tuning is good at teaching a model a skill or a style that every user shares, such as the tone of a company's support team. It is the wrong tool for personal facts that change weekly and must be deletable.

Is 75 tokens a law?

It is a budget, not a law. It comes from three facts of about 25 tokens. Some systems inject more. Our own library injects about 1,470 tokens per question on the LoCoMo benchmark (Chapter 18), which is still far below a transcript. What matters is that the amount is chosen by us and does not grow with the user's history.

Carry this

  • One user writes 1,095,000 tokens a year. The largest window holds 1,000,000.
  • At session 10, memory is 200 times cheaper than history per message. At session 100, 2,000 times.
  • A bigger window moves the expiry date. It does not shrink the bill.
  • A model's numbers are not a database: we cannot read, edit or delete one fact in them.
  • Three rules remain: distilled facts, a small fixed injection, a store outside the model.

Check yourself

1. A user is at session 50. We resend all history at $3 per million input tokens. What does one message cost, and how many times cheaper is a 75 token memory injection?

Answer
history    50 x 1,500               =  75,000 tokens
cost       75,000 x $3 / 1,000,000  =  $0.225 per message
ratio      75,000 / 75              =  1,000 times cheaper

2. The price of input tokens drops to one tenth. What happens to the ratio between history and memory?

Answer

Nothing. Both costs fall to one tenth, so the ratio is unchanged. The ratio depends only on token counts.

3. Fine-tuning per user fails for many reasons. Which requirement does it leave behind?

Answer

Memory must live outside the model, in a store we can read, edit and delete. Cost kills the idea first, but the lasting lesson is that knowledge inside a model's numbers cannot be shown to the user or removed on request.