24 September 2026 · how it works
Why AI Gets Worse the Longer You Talk to It: The KV Cache Explained
The answer got worse at message forty. Nobody told the developer why.
Here's what actually happened. Every large language model runs on a 2017 paper1 by eight engineers at Google. The paper is called "Attention Is All You Need." The design has one rule baked in: to write the next word, the model checks it against every word that came before. Every token asks every other token: are you relevant to me? That check runs across the whole conversation, every single step.
This is attention. It scales as the square of the length2. Double the chat, and the work quadruples. At 1,000 tokens, the model runs one million checks. At 100,000 tokens, ten billion. That's 10,000 times more work per word written. WE's sum: 100,000² ÷ 1,000² = 10,000. No source prints that number.
Engineers built a fix: the KV cache3. Each token's data gets stored once and looked up rather than recomputed. The cost drops from square to straight-line growth3. Think of it as a notepad: instead of re-reading the whole novel to write each new sentence, the model checks its notes4. The notes grow. The novel stays closed.
But the notepad sits on a chip. And chips run out.
A 70-billion-parameter model serving 32 conversations at 8,000 tokens each needs roughly 83 GB just for the cache3, often more than the model itself weighs. A server with 80 GB of GPU memory might hold the model at 14 GB, then find it can't run several long conversations at once because the caches eat the rest4. Each new token costs a small slice of that chip. None of it returns until the session ends.
So the model at token 500 and the model at token 50,000 are running under very different conditions. The cache is the main obstacle to longer conversations or more users at once5.
Now the part that doesn't make it into the product announcement.
Even when the cache is full and nothing has been dropped, the model stops using it equally. Chroma ran tests on 18 frontier models and found every single one gets worse as the conversation grows longer6. All of them. The model has a fixed budget of attention. Spread it across 100,000 tokens and something attended closely at 1,000 tokens gets passed over at 100,0007.
There's a second problem, found by a Stanford team in 2023. Their paper, "Lost in the Middle," showed that accuracy drops more than 30% when the key fact sits in the middle of a long context8, compared to the same fact placed at the start or end. The model reads the first few messages closely. It reads the last few closely. The forty messages in between? That's where your customer's original complaint went quiet.
This sits inside a bigger bill. Running models now takes about two-thirds of all AI compute spend, up from one-third in 20239. Training was the cost in 2021. Serving is the cost now. Between 55 and 80 per cent of enterprise AI GPU spend goes on running models, not building them10. At the frontier, prices doubled between January and July 202611 as newer models replaced older ones. Longer conversations make every item on that bill worse.
By the end of 2028, enterprise AI contracts will carry a clause about context length. Not capability, not uptime. How long the conversation is allowed to run. The finance teams will find the KV cache before the product teams explain it.
And yet the announcements keep coming. Million-token windows. Two million. Each one sold as more memory. A million-token window holds more, but the model uses what it holds less reliably12. The gap between the advertised window and what actually works can reach 99% on complex tasks7. The number on the box is not the number that matters.
The smart fix isn't a bigger window. It's sending less: a retrieval system that pulls 2,000 to 5,000 precise tokens rather than flooding the model with 500,000 mixed ones13. Less context, well chosen, beats a full window, badly loaded.
The developer whose agent broke at message forty didn't get a warning. Her contract said nothing about it. The architecture solved the original problem. The notepad never forgets. It just stops reading the early pages.