Carebun Interactive/Memory Memory labCheat sheet

Part I · Novice

Chapter 1

The genius with amnesia

Why the smartest AI model cannot remember your name tomorrow, shown with one real experiment and one long division.

8 min read · interactive

On this page
  1. A story to start
  2. What is an LLM?
  3. Why does the model forget?
  4. Try it yourself: watch the model forget
  5. Why not resend everything, forever?
  6. The idea that fixes it
  7. Common questions
  8. Carry this
  9. Check yourself

In this chapter, we will learn why an AI model forgets everything the moment a conversation ends. We will watch it fail in a real experiment, understand the reason in plain words, and meet the idea that fixes it. No background is needed. If you have chatted with an AI assistant once, you are ready.

A story to start

Meet Maya. On Monday she tells her AI assistant something important.

Monday, session 1
  Maya:       I'm vegetarian, and I have a serious peanut allergy.
  Assistant:  Noted!

Thursday, session 2 (a new chat)
  Maya:       Suggest a snack for my kids' school trip.
  Assistant:  How about peanut butter energy balls?

The assistant said "Noted!" and it was not lying. During Monday's chat it really did know. On Thursday the knowledge was gone. For Maya's family this is not a small bug. It is a dangerous one.

The assistant is not careless, and it is not stupid. It can write poetry, fix code and explain tax law. It simply has no memory. It is a genius with amnesia.

What is an LLM?

A Large Language Model (LLM) is a program that reads text and predicts the text that should come next. ChatGPT, Claude, Gemini and Qwen are all built on LLMs.

Two words appear on almost every page of this book, so let us learn them now.

A token is a small piece of text, roughly three quarters of a word. Models do not count words. They count tokens. "Maya is vegetarian." is about 4 tokens: three words and a full stop. We pay for an LLM by the token, both for what we send in and for what it writes back.

The context window is everything the model can see at one time. It is the model's desk. Whatever is on the desk, the model can use. Whatever is not on the desk does not exist for the model.

Why does the model forget?

An LLM is stateless, which means it keeps nothing between one call and the next.

This surprises most people, so let us go slowly.

When a model is trained, its knowledge is frozen into a very large file of numbers. After that, the file does not change while we chat. Talking to the model writes nothing into it. Our conversation is only ever placed on the desk, the context window, for the model to read.

what the model knows         =  the frozen file (training)  +  the desk (context window)
what changes when we chat    =  only the desk
what happens to the desk     =  it is cleared when the chat ends

So how does the assistant "remember" what we said five minutes ago in the same chat? It does not. The app quietly sends the whole conversation so far back to the model with every new message. The model reads it all again, from the first line, each time. It feels like memory. It is really re-reading.

When we open a new chat, the app starts with an empty desk. Monday's words are not on it. That is the entire reason Maya got peanut butter on Thursday.

Try it yourself: watch the model forget

The class that this book grew from has a small experiment called stateless-llm. It runs on a laptop with a free local model. Session 1 has already happened: Maya shared that she is vegetarian, has a peanut allergy and has two kids. Here is what session 2 prints.

(b) Session 2, NEW session, no memory: same user asks what was shared

    I'm sorry, but you haven't provided any information about your food
    allergies or the ages of your children. Could you please give me more details?

    prompt tokens: 49. The model is stateless: nothing of session 1 exists here.

The model is honest. From where it sits, Maya never said anything.

Now the experiment tries two fixes. First, paste the whole Monday conversation into the new chat. Second, paste only three short remembered facts.

(c) Session 2 with the FULL session-1 transcript resent (the naive fix)

    You have a serious peanut allergy, and you have two kids who are 6 and 9 years old.

    prompt tokens: 149

(d) Session 2 with 3 extracted memories injected instead

    You have a serious peanut allergy, and you have two kids who are 6 and 9 years old.

    prompt tokens: 74

Read those two results again. The answers are identical. The cost is not. Three small facts did the same job as the whole transcript, with half the tokens. And this was only one short session.

Prompt tokens are the tokens we send to the model. The question alone was 49 of them, so we can subtract it and see what each fix added.

full transcript added    149 - 49  =  100 tokens
three facts added         74 - 49  =   25 tokens

whole prompt              74 / 149 =  about half
ollama pull qwen2.5:7b-instruct
cd memory_classnotes/experiments/stateless-llm
pip install -r requirements.txt
python main.py

Why not resend everything, forever?

It worked for one session. Let us see what happens over a year. We will use the numbers the class uses for a typical user, and we will use them for the whole book. A session is one sitting, from opening the chat to closing it.

one chat session                 about 1,500 tokens (both sides of the chat)
sessions per day                 2
tokens per day                   2 x 1,500       =     3,000
sessions per year                2 x 365         =       730
transcript per year              730 x 1,500     = 1,095,000 tokens

largest context window today                       1,000,000 tokens
days until it is full            1,000,000 / 3,000 = about 333 days

One year of ordinary chatting is already bigger than the biggest desk that exists. The transcript stops fitting on day 334, about a month before the year ends.

Move the sliders to match your own habits, and watch the day the desk overflows.

Long before the transcript stops fitting, it becomes expensive, because every single message pays for the whole transcript again. Input tokens are the tokens the model reads. The class uses a price of $3 per million of them, which is what a top model charges.

resend everything, price $3 per million input tokens
cost of one message  =  tokens sent x $3 / 1,000,000

after   1 session      1,500 tokens    $0.0045 per message
after  10 sessions    15,000 tokens    $0.045  per message
after 100 sessions   150,000 tokens    $0.45   per message

Forty-five cents for one message, and Maya sends twelve a day: 2 sessions of 6 messages each. Chapter 2 finishes this argument and tests the other tempting fixes. For now, carry one picture: history grows every day, and so does its bill.

The idea that fixes it

Memory is a small set of facts, kept outside the model, from which only the few relevant ones are placed on the desk.

That sentence has three parts, and each one matters.

  1. Facts, not transcripts. "Maya has a serious peanut allergy" is 6 words, so about 8 tokens (6 / 0.75 = 8). A whole session is about 1,500.
  2. Outside the model. The model stays frozen. The facts live in an ordinary database, a program that stores records and finds them again. We control it, and we can inspect, correct and delete what is in it.
  3. Only the relevant few. For a snack question we hand over the allergy fact. For a gift question we hand over the birthday fact. Three facts today. Three facts in ten years.
HistoryMemory
What is sentevery old conversationthe top 3 facts
Sizegrows with every sessionabout 75 tokens (3 facts of about 25), flat
After one yeardoes not fit in the windowstill about 75 tokens

This is the doctor's patient file. It is also the card box that we open in Chapter 3, with a secretary who writes the cards and a librarian who finds them.

Common questions

The assistant I use already seems to remember me. How?

Because its makers built a memory layer around the model, exactly the kind this book teaches. The model inside is still stateless. ChatGPT, Claude and Gemini all add memory as a separate system, and all of them let you delete what is stored. Chapter 15 explains why.

Is the model not learning from my chats?

Not while you chat. A company may use conversations later to train a future model, but that is a slow, separate process, and it does not give the model a memory of you in particular.

Why can we not just build a bigger context window?

Windows are growing, and it helps. But a bigger desk does not make reading free. Every message still pays for everything on the desk. Chapter 2 shows the numbers.

Is memory the same as RAG?

They are cousins. RAG (Retrieval-Augmented Generation) searches documents that somebody wrote, such as a company handbook. Memory searches facts that the system itself wrote down about one user, and those facts change as the user's life changes. Memory has to write, update and forget. RAG mostly only reads.

Does a fact in memory always reach the model?

No, and that is on purpose. Only the facts that match the question are placed on the desk. Choosing them well is the hard part, and it takes up Chapters 8 and 9.

Carry this

  • An LLM is a frozen file plus a desk. Chatting changes only the desk, and the desk is cleared when the chat ends.
  • "Remembering" inside one chat is really re-reading the whole conversation with every message.
  • One year of ordinary chat is 1,095,000 tokens. It overflows a 1,000,000 token window.
  • Three extracted facts gave the same answer as the full transcript at half the tokens (74 against 149).
  • Memory is a layer around the model, not inside it: a few facts, kept outside, with only the relevant ones handed over.

Check yourself

1. Maya chats for 40 minutes in one session and the assistant still remembers the first thing she said. Is the model storing it?

Answer

No. The app resends the whole conversation with every message, so the first line is still on the desk. The model re-reads it each time. Nothing is stored inside the model.

2. A user has 3 sessions a day of 2,000 tokens each. After how many days is the history larger than a 1,000,000 token window?

Answer
tokens per day     3 x 2,000          =   6,000
days to overflow   1,000,000 / 6,000  = about 167 days

Less than six months.

3. In the experiment, why is the injected-memory prompt (74 tokens) smaller than the full-transcript prompt (149 tokens), although both give the same answer?

Answer

The transcript carries everything that was said, including greetings and the assistant's replies. The memories carry only the three facts that matter. The answer needs the facts, not the conversation around them.