Part III · Journeyman
Chapter 16
Scaling and cost
From one laptop to a million users: where the 664 GB comes from, how rounding shrinks it, and the daily bill derived line by line.
On this page
In this chapter, we will learn what it takes to run memory for a million users. We will draw the deployment, size the store byte by byte, shrink it with a trick called quantization, and add up the daily bill one line at a time. Then we will compare that bill with the cost of having no memory at all.
The deployment
A deployment is the set of machines and services that run the system, and the way requests move between them.
[ load balancer ]
|
[ orchestrator pods ] stateless: any pod can serve any user
| | | \
| | | \ at session end
[ session store ] [ embedding ] [ memory store ] [ extraction queue ]
Redis service sharded by |
TTL: hours user hash v
3 replicas [ extraction workers ]
| ^ |
[ assistant LLM ] +--------------+
hosted or own GPUs writes land here
| Piece | Job | Why it is built this way |
|---|---|---|
| Load balancer | spreads incoming messages over the pods | no single pod is overloaded, and a dead pod is skipped |
| Orchestrator pods | run recall, answer, remember for each message | a pod is one running copy of the program. Pods are stateless: they keep nothing between requests, so we add pods when traffic grows and lose nothing when one dies |
| Session store | holds the live turns of this chat | Redis, a database that keeps its data in RAM, with a TTL (time to live) of hours: short-term memory cleans itself up |
| Embedding service | turns text into vectors | shared and batched, 10 to 30 ms |
| Memory store | holds every memory | sharded by a hash of the user id, with 3 replicas |
| Extraction queue and workers | the write path | off the critical path; a slow worker delays a memory, never an answer |
A shard is one slice of the data, kept on its own machine. A hash is a function that turns any text into a number, and always the same number for the same text. We pick the slice by hashing the user id, so all of Maya's memories sit together and one search touches one shard. A replica is a full copy, kept so that a broken machine loses nothing and reads can be shared.
The traffic
daily active users 200,000
sessions per day 200,000 x 2 = 400,000
user messages per day 400,000 x 6 = 2,400,000
read path 2,400,000 / 86,400 seconds = about 28 per second
peak 3x = about 83 per second
write path 400,000 / 86,400 seconds = about 4.6 per second
peak 3x = about 14 per second
Traffic is not even over the day, so the class sizes for a peak of 3 times the average. Even the peaks are small numbers for a modern database. Traffic is not the hard part of memory. Storage and model calls are.
Sizing the store
The size of the store is the number of memories multiplied by the bytes in one memory.
one memory text about 25 tokens about 100 B vector 768 numbers x 4 bytes = 3,072 B payload ids, dates, category, source about 150 B total = 3,322 B (about 3.3 KB) all memories 1,000,000 users x 200 = 200,000,000 raw size 200,000,000 x 3,322 B = about 664 GB vectors alone 200,000,000 x 3,072 B = about 614 GB share 614 / 664 = 92%
Read the last line twice. Ninety-two percent of the store is vectors. The sentences Maya actually said are a rounding error. If we want a smaller store, the vectors are the only place worth looking.
Quantization: rounding the numbers
Quantization means storing each number of a vector with fewer bits, by rounding it.
A vector holds 768 numbers such as 0.031 and -0.118. By default each one is a float32, which takes 4 bytes and keeps about seven digits. Search does not need seven digits. It only needs to know which memories are closer than others.
Think of heights. "Maya is 167.3482 cm" and "Maya is 167 cm" let us sort a classroom by height equally well. INT8 quantization does that rounding: each number becomes one of 256 levels and takes 1 byte instead of 4.
float32 768 x 4 bytes = 3,072 B per vector INT8 768 x 1 byte = 768 B per vector 4x smaller, recall loss about nil (class notes) vectors 200,000,000 x 768 B = about 154 GB (was 614 GB) text + payload 200,000,000 x 250 B = 50 GB (cannot be quantized) one full copy 154 + 50 = 204 GB 3 replicas 204 x 3 = 612 GB on disk
Three complete copies of the quantized store take less disk than one copy of the raw store.
We can round harder. Four bits give 16 levels per number, two bits give 4. Each step halves the size and costs some recall, the share of the true nearest memories that the search still finds. Choose a format below and see both numbers move.
| Rows | Index | Bytes per vector | Query p50 (ms) | Recall@10 |
|---|---|---|---|---|
| 10,000 | numpy exact float32 | 16,384 | 0.17 | 1.000 |
| 10,000 | turbovec 4-bit | 2,048 | 0.61 | 0.984 |
| 10,000 | turbovec 2-bit | 1,024 | 0.28 | 0.940 |
| 100,000 | numpy exact float32 | 16,384 | 1.54 | 1.000 |
| 100,000 | turbovec 4-bit | 2,048 | 1.44 | 0.984 |
| 100,000 | turbovec 2-bit | 1,024 | 0.73 | 0.940 |
| 300,000 | numpy exact float32 | 16,384 | 4.93 | 1.000 |
| 300,000 | turbovec 4-bit | 2,048 | 4.15 | 0.984 |
| 300,000 | turbovec 2-bit | 1,024 | 2.05 | 0.941 |
p50 is the median: half the queries were faster than this. Three things stand out. The 4-bit index is 8 times smaller (16,384 / 2,048) and loses 1.6 points of recall. It needs no training, so new memories are indexed as they arrive. And from 100,000 rows up it is also faster than exact search.
The three scale knobs in hanumemAI
| Knob | What it does |
|---|---|
vector_dtype="int8" | stores each vector on disk as int8 codes plus a 4-byte scale, 4 times smaller; the index in RAM stays float32, so search results do not change |
shards=N | one SQLite file per bucket of users, chosen by a hash of
user_id; a query opens one file |
index_cache_users=2000 | keeps the vector indexes of the 2,000 most recently active users in RAM and reloads the others from disk when they return (LRU: least recently used goes first) |
Note what the first knob does not do. hanumemAI's int8 shrinks the rows on disk only, and
search still runs on float32 numbers in RAM. The INT8 line in the sizing above is the class's figure. We
did not measure a search over INT8 vectors.
The cache works because of how people behave. Only 200,000 of a million users are active on a given day, and at any one moment far fewer. We never search 200 million vectors. We search one user's 200, a million times.
One million memories on a laptop
We filled a real store with one million synthetic memories and measured the read path for one user. The shape follows the class sizing: 5,000 users with 200 memories each, 768-dimension vectors, int8 on disk, 16 shards. The text comes from templates and the vectors are random points around 40 topic centres, so no model was called. This measures the store, which is the part that must scale.
store: 5,000 users x 200 memories = 1,000,000 rows, 768-d, int8 on disk, 16 shards rows written 1,000,000 in 96 s (10,432 rows/s, one commit per row) on disk 1.38 GB (1,376 bytes per memory, all indexes included) vector bytes per row 772 dense search, cold p50 0.47 ms p95 0.55 ms (loads the user's 200 vectors from SQLite) dense search, warm p50 0.015 ms p95 0.017 ms BM25 search p50 5.69 ms p95 6.17 ms isolation check 30 of 30 returned rows belong to the asking user delete_user 7.9 ms, rows left for that user: 0
Four things can be read from this run.
- Search by meaning does not slow down as the store grows. It only ever looks at one user's 200 vectors. With the index in RAM it takes 0.015 ms. For a user who has been away, loading the index from disk first takes 0.47 ms.
- The bytes match the arithmetic. A vector takes 772 bytes: 768 codes and a 4-byte scale. The whole row, with text, word index and name index, takes 1,376 bytes.
- Search by exact words is the slow part. We ran the same script again with 50 users, which is 10,000 rows, and this search took about 0.1 ms. With one million rows it takes 5.69 ms. The word index is shared by every user in a shard, so it grows with the shard. It is still inside the 50 ms budget, and more shards bring it down, but this is the first thing that would need work at a larger scale.
- Erasing one user took 7.9 ms, and a count afterwards returned zero rows.
measured 1,000,000 memories x 1,376 B = 1.38 GB class baseline 200,000,000 memories: 200 times more scaled up 1.38 GB x 200 = 276 GB for one copy (arithmetic, not measured) beyond the vector 1,376 B - 772 B = 604 B for each memory
The class sized one INT8 copy at 204 GB with 250 bytes of payload for each memory. Our rows carry about 600 bytes beyond the vector, because the word index and the name index live in the same file. The two estimates agree on the order of magnitude, which is what a sizing exercise is for.
cd ~/apps/hmem
uv run python ~/apps/interactive/books/memory/examples/scale_million.py
Cost per day
The daily cost of memory is the sum of five lines: extraction, update decisions, embeddings, injected tokens and the store. These are the class's prices. The write path uses a small model at $0.15 per million input tokens and $0.60 per million output tokens. Embeddings cost $0.02 per million tokens. The assistant is either the same small model or a frontier model, one of the largest and most expensive, at $3 per million input tokens.
Two inputs are new. An update call reads about 550 tokens: the instructions, the candidate fact and its 2 nearest stored memories. And every user message of about 20 tokens is embedded before the search.
1. extraction per session 1,500 tokens in x $0.15/M = $0.000225
150 tokens out x $0.60/M = $0.00009
sum = $0.000315
400,000 sessions x $0.000315 = $126.00
2. update calls per candidate 550 tokens in x $0.15/M = $0.0000825
40 tokens out x $0.60/M = $0.000024
sum = $0.0001065
400,000 x 3 candidates x $0.0001065 = $127.80
3. embeddings write 400,000 x 3 facts x 25 tokens = 30M tokens
read 2,400,000 messages x 20 tokens = 48M tokens
78M tokens x $0.02/M = $1.56
4. injected tokens 2,400,000 messages x 75 tokens = 180M tokens
small assistant 180M x $0.15/M = $27.00
frontier assistant 180M x $3/M = $540.00
5. store 204 GB x 3 replicas, 3 nodes = about $30.00
total, small assistant 126 + 127.80 + 1.56 + 27 + 30 = $312.36
total, frontier assistant 126 + 127.80 + 1.56 + 540 + 30 = $825.36
Look at where the money goes. With a small assistant, the two write-path lines are 253.80 of 312.36 dollars, which is 81 percent. The store, the part everyone worries about, is about 10 percent. Memory is a model bill, not a disk bill.
Compared with what?
A cost means nothing alone. The honest comparison is the same assistant model without memory, resending the last 10 sessions of history with every message.
history per message 10 sessions x 1,500 tokens = 15,000 tokens tokens per day 2,400,000 x 15,000 = 36,000,000,000 (36B) small model 36B x $0.15/M = $5,400 per day memory $312 312 / 5,400 = 5.8% frontier model 36B x $3/M = $108,000 per day memory $825 825 / 108,000 = 0.76%
With a frontier model, memory costs less than one percent of resending history, and it keeps working after session 10, where the history approach has already started to forget. Change any input below and watch both bills.
Four ways to make it cheaper
- Merge the write calls. hanumemAI makes the supersede decision inside the extraction call, which removes line 2. It adds one short profile update per session instead.
- Extract once per session, not once per message. The bill above already does this. Extracting after each of the 6 messages would mean 6 extraction calls where we now make 1.
- Use a small model for writing. The secretary does not need the expert's salary (Chapter 14).
- Cap the injection in tokens, not rows. Our BEAM runs fell from 12,380 to 4,227 tokens per question with a 4,000 token cap, and the score rose from 46 to 49 (experiment B0002).
Try it yourself
Run the scale measurement on your own machine, from the root of the hanumemAI research repository. It needs the stores of a full LoCoMo run, an API key to embed the 200 questions, and up to about 6 GB of RAM.
uv sync
uv run python scripts/scale.py # prints the table above, 2 to 5 minutes
Then turn the knobs on in an app.
from hmem import Memory, Config
memory = Memory(Config(
db_path="data/memory", # a directory, because shards > 1
shards=16, # one SQLite file per user bucket
vector_dtype="int8", # 4x smaller rows on disk
index_cache_users=2000, # bound the RAM
))
print(memory.stats("maya"))
Common questions
Do I need any of this for my first thousand users?
No. A thousand users with 200 memories each is 200,000 memories, about 664 MB raw. One SQLite file on one machine is enough. Build the simple version, measure, and add a shard when a number tells you to.
Why does 4 times smaller vectors cost almost no accuracy?
Search compares vectors with each other. Rounding moves every vector a tiny amount, but rarely enough to change which one is nearest. And the read path fetches more candidates than it needs, 4 times k in hanumemAI, so a near miss in the first pass is usually still in the pool.
Why shard by user and not by date or by topic?
Because every query is about exactly one user. Sharding by user means one query touches one shard. Any other key would spread Maya's memories across machines and force every search to ask all of them.
What happens to memory when the extraction queue falls behind?
Answers stay fast, because the read path does not wait for the queue. New memories arrive late. A fact from this morning might be missing this afternoon. That is why extraction latency and queue depth are monitored (Chapter 17).
Our vectors have 4,096 numbers, not 768. What changes?
Only the multiplication. A float32 vector is then 16,384 bytes, and 200 million of them are about 3.3 TB. With the 4-bit index they are 2,048 bytes each, about 410 GB. Larger vectors make quantization more important, not less.
Is the $30 a day for the store realistic?
It is the class estimate for 612 GB on three nodes, and prices vary by provider. Even if it were three times higher, the conclusion would hold: the model calls are the bill.
Carry this
- 28 reads and 4.6 writes per second, with peaks of 83 and 14. Traffic is not the hard part.
- 664 GB raw, and 92 percent of it is vectors. INT8 brings one copy to 204 GB and three replicas to 612 GB.
- Measured: a 4-bit index is 8 times smaller at 98.4 percent recall@10, up to 300,000 vectors on a laptop. Beyond that we calculate, we do not claim.
- Measured: one million memories for 5,000 users take 1.38 GB. One user's search by meaning takes 0.47 ms cold and 0.015 ms warm. Search by exact words is the part that grows with the store.
- About $312 a day with a small assistant, $825 with a frontier one. Four fifths of the small bill is the write path. Resending history costs $5,400 or $108,000 a day, so memory is 5.8 percent or 0.76 percent of that.
Check yourself
1. A smaller product has 100,000 users with 150 memories each and 768-dimension vectors. What is the raw size, and the size of one copy after INT8?
Answer
memories 100,000 x 150 = 15,000,000 raw 15,000,000 x 3,322 B = about 49.8 GB INT8 vectors 15,000,000 x 768 B = about 11.5 GB payload 15,000,000 x 250 B = about 3.8 GB one copy 11.5 + 3.8 = about 15.3 GB
2. Your finance team asks where to cut the memory bill. Which line do you look at first, and why not the store?
Answer
The write path: extraction and update calls are $253.80 of $312.36. Merging the update decision into the extraction call removes about $127.80. The store is about $30, so even halving it saves little.
3. The assistant model price falls from $3 to $1 per million input tokens. What is the new daily total for the frontier case?
Answer
injected tokens 180M x $1/M = $180 total 126 + 127.80 + 1.56 + 180 + 30 = $465.36
Resending history would cost 36B x $1/M = $36,000, so memory is about 1.3 percent of it.