Carebun Interactive/Memory Memory labCheat sheet

Part III · Journeyman

Chapter 16

Scaling and cost

From one laptop to a million users: where the 664 GB comes from, how rounding shrinks it, and the daily bill derived line by line.

12 min read · interactive

On this page
  1. The deployment
  2. The traffic
  3. Sizing the store
  4. Quantization: rounding the numbers
  5. One million memories on a laptop
  6. Cost per day
  7. Try it yourself
  8. Common questions
  9. Carry this
  10. Check yourself

In this chapter, we will learn what it takes to run memory for a million users. We will draw the deployment, size the store byte by byte, shrink it with a trick called quantization, and add up the daily bill one line at a time. Then we will compare that bill with the cost of having no memory at all.

The deployment

A deployment is the set of machines and services that run the system, and the way requests move between them.

                     [ load balancer ]
                             |
                  [ orchestrator pods ]          stateless: any pod can serve any user
                   |       |       |      \
                   |       |       |       \  at session end
   [ session store ]  [ embedding ]  [ memory store ]   [ extraction queue ]
     Redis              service        sharded by          |
     TTL: hours                        user hash           v
                                       3 replicas   [ extraction workers ]
                   |                        ^              |
           [ assistant LLM ]                +--------------+
             hosted or own GPUs                 writes land here
PieceJobWhy it is built this way
Load balancerspreads incoming messages over the podsno single pod is overloaded, and a dead pod is skipped
Orchestrator podsrun recall, answer, remember for each messagea pod is one running copy of the program. Pods are stateless: they keep nothing between requests, so we add pods when traffic grows and lose nothing when one dies
Session storeholds the live turns of this chatRedis, a database that keeps its data in RAM, with a TTL (time to live) of hours: short-term memory cleans itself up
Embedding serviceturns text into vectorsshared and batched, 10 to 30 ms
Memory storeholds every memorysharded by a hash of the user id, with 3 replicas
Extraction queue and workersthe write pathoff the critical path; a slow worker delays a memory, never an answer

A shard is one slice of the data, kept on its own machine. A hash is a function that turns any text into a number, and always the same number for the same text. We pick the slice by hashing the user id, so all of Maya's memories sit together and one search touches one shard. A replica is a full copy, kept so that a broken machine loses nothing and reads can be shared.

The traffic

daily active users                             200,000
sessions per day         200,000 x 2       =   400,000
user messages per day    400,000 x 6       = 2,400,000

read path    2,400,000 / 86,400 seconds    = about 28 per second
             peak 3x                       = about 83 per second
write path     400,000 / 86,400 seconds    = about 4.6 per second
             peak 3x                       = about 14 per second

Traffic is not even over the day, so the class sizes for a peak of 3 times the average. Even the peaks are small numbers for a modern database. Traffic is not the hard part of memory. Storage and model calls are.

Sizing the store

The size of the store is the number of memories multiplied by the bytes in one memory.

one memory
  text       about 25 tokens                     about   100 B
  vector     768 numbers x 4 bytes            =        3,072 B
  payload    ids, dates, category, source        about   150 B
  total                                       =        3,322 B   (about 3.3 KB)

all memories   1,000,000 users x 200          = 200,000,000
raw size       200,000,000 x 3,322 B          = about 664 GB
vectors alone  200,000,000 x 3,072 B          = about 614 GB
share          614 / 664                      = 92%

Read the last line twice. Ninety-two percent of the store is vectors. The sentences Maya actually said are a rounding error. If we want a smaller store, the vectors are the only place worth looking.

Quantization: rounding the numbers

Quantization means storing each number of a vector with fewer bits, by rounding it.

A vector holds 768 numbers such as 0.031 and -0.118. By default each one is a float32, which takes 4 bytes and keeps about seven digits. Search does not need seven digits. It only needs to know which memories are closer than others.

Think of heights. "Maya is 167.3482 cm" and "Maya is 167 cm" let us sort a classroom by height equally well. INT8 quantization does that rounding: each number becomes one of 256 levels and takes 1 byte instead of 4.

float32    768 x 4 bytes = 3,072 B per vector
INT8       768 x 1 byte  =   768 B per vector         4x smaller, recall loss about nil (class notes)

vectors    200,000,000 x 768 B     = about 154 GB     (was 614 GB)
text + payload   200,000,000 x 250 B =        50 GB   (cannot be quantized)
one full copy    154 + 50          =         204 GB
3 replicas       204 x 3           =         612 GB on disk

Three complete copies of the quantized store take less disk than one copy of the raw store.

We can round harder. Four bits give 16 levels per number, two bits give 4. Each step halves the size and costs some recall, the share of the true nearest memories that the search still finds. Choose a format below and see both numbers move.

RowsIndexBytes per vectorQuery p50 (ms)Recall@10
10,000numpy exact float3216,3840.171.000
10,000turbovec 4-bit2,0480.610.984
10,000turbovec 2-bit1,0240.280.940
100,000numpy exact float3216,3841.541.000
100,000turbovec 4-bit2,0481.440.984
100,000turbovec 2-bit1,0240.730.940
300,000numpy exact float3216,3844.931.000
300,000turbovec 4-bit2,0484.150.984
300,000turbovec 2-bit1,0242.050.941

p50 is the median: half the queries were faster than this. Three things stand out. The 4-bit index is 8 times smaller (16,384 / 2,048) and loses 1.6 points of recall. It needs no training, so new memories are indexed as they arrive. And from 100,000 rows up it is also faster than exact search.

The three scale knobs in hanumemAI

KnobWhat it does
vector_dtype="int8"stores each vector on disk as int8 codes plus a 4-byte scale, 4 times smaller; the index in RAM stays float32, so search results do not change
shards=None SQLite file per bucket of users, chosen by a hash of user_id; a query opens one file
index_cache_users=2000keeps the vector indexes of the 2,000 most recently active users in RAM and reloads the others from disk when they return (LRU: least recently used goes first)

Note what the first knob does not do. hanumemAI's int8 shrinks the rows on disk only, and search still runs on float32 numbers in RAM. The INT8 line in the sizing above is the class's figure. We did not measure a search over INT8 vectors.

The cache works because of how people behave. Only 200,000 of a million users are active on a given day, and at any one moment far fewer. We never search 200 million vectors. We search one user's 200, a million times.

One million memories on a laptop

We filled a real store with one million synthetic memories and measured the read path for one user. The shape follows the class sizing: 5,000 users with 200 memories each, 768-dimension vectors, int8 on disk, 16 shards. The text comes from templates and the vectors are random points around 40 topic centres, so no model was called. This measures the store, which is the part that must scale.

store: 5,000 users x 200 memories = 1,000,000 rows, 768-d, int8 on disk, 16 shards

rows written            1,000,000 in 96 s  (10,432 rows/s, one commit per row)
on disk                 1.38 GB  (1,376 bytes per memory, all indexes included)
vector bytes per row    772
dense search, cold      p50 0.47 ms   p95 0.55 ms   (loads the user's 200 vectors from SQLite)
dense search, warm      p50 0.015 ms   p95 0.017 ms
BM25 search             p50 5.69 ms   p95 6.17 ms
isolation check         30 of 30 returned rows belong to the asking user
delete_user             7.9 ms, rows left for that user: 0

Four things can be read from this run.

  1. Search by meaning does not slow down as the store grows. It only ever looks at one user's 200 vectors. With the index in RAM it takes 0.015 ms. For a user who has been away, loading the index from disk first takes 0.47 ms.
  2. The bytes match the arithmetic. A vector takes 772 bytes: 768 codes and a 4-byte scale. The whole row, with text, word index and name index, takes 1,376 bytes.
  3. Search by exact words is the slow part. We ran the same script again with 50 users, which is 10,000 rows, and this search took about 0.1 ms. With one million rows it takes 5.69 ms. The word index is shared by every user in a shard, so it grows with the shard. It is still inside the 50 ms budget, and more shards bring it down, but this is the first thing that would need work at a larger scale.
  4. Erasing one user took 7.9 ms, and a count afterwards returned zero rows.
measured         1,000,000 memories x 1,376 B  =  1.38 GB
class baseline   200,000,000 memories: 200 times more
scaled up        1.38 GB x 200                 =  276 GB for one copy      (arithmetic, not measured)
beyond the vector    1,376 B - 772 B           =  604 B for each memory

The class sized one INT8 copy at 204 GB with 250 bytes of payload for each memory. Our rows carry about 600 bytes beyond the vector, because the word index and the name index live in the same file. The two estimates agree on the order of magnitude, which is what a sizing exercise is for.

cd ~/apps/hmem
uv run python ~/apps/interactive/books/memory/examples/scale_million.py

Cost per day

The daily cost of memory is the sum of five lines: extraction, update decisions, embeddings, injected tokens and the store. These are the class's prices. The write path uses a small model at $0.15 per million input tokens and $0.60 per million output tokens. Embeddings cost $0.02 per million tokens. The assistant is either the same small model or a frontier model, one of the largest and most expensive, at $3 per million input tokens.

Two inputs are new. An update call reads about 550 tokens: the instructions, the candidate fact and its 2 nearest stored memories. And every user message of about 20 tokens is embedded before the search.

1. extraction       per session   1,500 tokens in  x $0.15/M = $0.000225
                                    150 tokens out x $0.60/M = $0.00009
                                                       sum   = $0.000315
                    400,000 sessions x $0.000315             = $126.00

2. update calls     per candidate   550 tokens in  x $0.15/M = $0.0000825
                                     40 tokens out x $0.60/M = $0.000024
                                                       sum   = $0.0001065
                    400,000 x 3 candidates x $0.0001065      = $127.80

3. embeddings       write  400,000 x 3 facts x 25 tokens     = 30M tokens
                    read   2,400,000 messages x 20 tokens    = 48M tokens
                    78M tokens x $0.02/M                     = $1.56

4. injected tokens  2,400,000 messages x 75 tokens           = 180M tokens
                    small assistant    180M x $0.15/M        = $27.00
                    frontier assistant 180M x $3/M           = $540.00

5. store            204 GB x 3 replicas, 3 nodes             = about $30.00

total, small assistant      126 + 127.80 + 1.56 +  27 + 30   = $312.36
total, frontier assistant   126 + 127.80 + 1.56 + 540 + 30   = $825.36

Look at where the money goes. With a small assistant, the two write-path lines are 253.80 of 312.36 dollars, which is 81 percent. The store, the part everyone worries about, is about 10 percent. Memory is a model bill, not a disk bill.

Compared with what?

A cost means nothing alone. The honest comparison is the same assistant model without memory, resending the last 10 sessions of history with every message.

history per message    10 sessions x 1,500 tokens   = 15,000 tokens
tokens per day         2,400,000 x 15,000           = 36,000,000,000  (36B)

small model     36B x $0.15/M  =   $5,400 per day     memory $312   312 / 5,400   = 5.8%
frontier model  36B x $3/M     = $108,000 per day     memory $825   825 / 108,000 = 0.76%

With a frontier model, memory costs less than one percent of resending history, and it keeps working after session 10, where the history approach has already started to forget. Change any input below and watch both bills.

Four ways to make it cheaper

  • Merge the write calls. hanumemAI makes the supersede decision inside the extraction call, which removes line 2. It adds one short profile update per session instead.
  • Extract once per session, not once per message. The bill above already does this. Extracting after each of the 6 messages would mean 6 extraction calls where we now make 1.
  • Use a small model for writing. The secretary does not need the expert's salary (Chapter 14).
  • Cap the injection in tokens, not rows. Our BEAM runs fell from 12,380 to 4,227 tokens per question with a 4,000 token cap, and the score rose from 46 to 49 (experiment B0002).

Try it yourself

Run the scale measurement on your own machine, from the root of the hanumemAI research repository. It needs the stores of a full LoCoMo run, an API key to embed the 200 questions, and up to about 6 GB of RAM.

uv sync
uv run python scripts/scale.py        # prints the table above, 2 to 5 minutes

Then turn the knobs on in an app.

from hmem import Memory, Config

memory = Memory(Config(
    db_path="data/memory",      # a directory, because shards > 1
    shards=16,                  # one SQLite file per user bucket
    vector_dtype="int8",        # 4x smaller rows on disk
    index_cache_users=2000,     # bound the RAM
))
print(memory.stats("maya"))

Common questions

Do I need any of this for my first thousand users?

No. A thousand users with 200 memories each is 200,000 memories, about 664 MB raw. One SQLite file on one machine is enough. Build the simple version, measure, and add a shard when a number tells you to.

Why does 4 times smaller vectors cost almost no accuracy?

Search compares vectors with each other. Rounding moves every vector a tiny amount, but rarely enough to change which one is nearest. And the read path fetches more candidates than it needs, 4 times k in hanumemAI, so a near miss in the first pass is usually still in the pool.

Why shard by user and not by date or by topic?

Because every query is about exactly one user. Sharding by user means one query touches one shard. Any other key would spread Maya's memories across machines and force every search to ask all of them.

What happens to memory when the extraction queue falls behind?

Answers stay fast, because the read path does not wait for the queue. New memories arrive late. A fact from this morning might be missing this afternoon. That is why extraction latency and queue depth are monitored (Chapter 17).

Our vectors have 4,096 numbers, not 768. What changes?

Only the multiplication. A float32 vector is then 16,384 bytes, and 200 million of them are about 3.3 TB. With the 4-bit index they are 2,048 bytes each, about 410 GB. Larger vectors make quantization more important, not less.

Is the $30 a day for the store realistic?

It is the class estimate for 612 GB on three nodes, and prices vary by provider. Even if it were three times higher, the conclusion would hold: the model calls are the bill.

Carry this

  • 28 reads and 4.6 writes per second, with peaks of 83 and 14. Traffic is not the hard part.
  • 664 GB raw, and 92 percent of it is vectors. INT8 brings one copy to 204 GB and three replicas to 612 GB.
  • Measured: a 4-bit index is 8 times smaller at 98.4 percent recall@10, up to 300,000 vectors on a laptop. Beyond that we calculate, we do not claim.
  • Measured: one million memories for 5,000 users take 1.38 GB. One user's search by meaning takes 0.47 ms cold and 0.015 ms warm. Search by exact words is the part that grows with the store.
  • About $312 a day with a small assistant, $825 with a frontier one. Four fifths of the small bill is the write path. Resending history costs $5,400 or $108,000 a day, so memory is 5.8 percent or 0.76 percent of that.

Check yourself

1. A smaller product has 100,000 users with 150 memories each and 768-dimension vectors. What is the raw size, and the size of one copy after INT8?

Answer
memories     100,000 x 150          = 15,000,000
raw          15,000,000 x 3,322 B   = about 49.8 GB
INT8 vectors 15,000,000 x 768 B     = about 11.5 GB
payload      15,000,000 x 250 B     = about  3.8 GB
one copy     11.5 + 3.8             = about 15.3 GB

2. Your finance team asks where to cut the memory bill. Which line do you look at first, and why not the store?

Answer

The write path: extraction and update calls are $253.80 of $312.36. Merging the update decision into the extraction call removes about $127.80. The store is about $30, so even halving it saves little.

3. The assistant model price falls from $3 to $1 per million input tokens. What is the new daily total for the frontier case?

Answer
injected tokens   180M x $1/M                    = $180
total             126 + 127.80 + 1.56 + 180 + 30 = $465.36

Resending history would cost 36B x $1/M = $36,000, so memory is about 1.3 percent of it.