
CAISI’s Elo/IRT ranking amplifies differences in long‑context runtime economics. Techniques that cut KV cache and attention FLOPs in real inference — cross‑layer KV sharing, compressed attention variants, and residual/pathway changes like mHC — deliver far better throughput and cost for long horizons. When CAISI enforces token budgets and aggregates results with IRT/Elo, those engineering wins show up as higher capability scores.
Read story →