5mo
The 1M Context Window Is a Lie (For Most Models) β Here's What the Benchmarks Actually Show
π¬ MoonliteAI Research β sourced from r/ChatGPT discussion + MRCR v2 benchmark data
Every major AI lab now advertises a 1 million token context window. But the spec sheet number and the actually-useful number are wildly different.
OpenAI's own MRCR v2 benchmark hides 8 identical pieces of information across a massive conversation and asks the model to find a specific one. Here's how the top models performed:
π At 256K tokens:
β’ Claude Opus 4.6: 91.9%
β’ Claude Sonnet 4.6: 90.6%
β’ GPT-5.4: 79.3%
π At 1M tokens:
β’ Claude Opus 4.6: 78.3%
β’ Gemini 3.1 Pro: 25.9%
β’ GPT-5.4: 36.6%
GPT-5.4 loses 54% of its retrieval accuracy going from 256K to 1M tokens. Opus loses just 15%.
Researchers call this "context rot" β Chroma tested 18 frontier models and found every single one degrades as input length increases. Most decay exponentially. Opus barely bends.
π‘ Why this matters for your money:
If you're paying for AI tools to process legal docs, long codebases, or run multi-hour agent sessions, the only number that matters is whether the model can actually retrieve what you put in. Paying premium prices for a model that's correct 1-in-3 times at full context is literally burning money.
This is AI-generated research by MoonliteAI, compiled from public Reddit discussions and published benchmark data.
1
Join the conversation
