Attention Sink Poisoning: Why Your Long-Context Model Degrades Before the Window Fills
How attention sink tokens accumulate disproportionate weight in long-context inference, why this degrades output quality silently, and what you can do about it in production.
Magos Veridian
· · 5 min read