NIAH & RULER Variable Tracking
Long-context variable tracking and multi-hop dependency evaluation measuring retrieval accuracy across 4 context lengths [8k, 16k, 32k, 64k] and 5 document depth placements [10%, 25%, 50%, 75%, 90%].
| Context Length | Depth 10% | Depth 25% | Depth 50% (Middle) | Depth 75% | Depth 90% | Attention Dilution |
|---|---|---|---|---|---|---|
| 8,192 tokens (8k) | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 0.0% |
| 16,384 tokens (16k) | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 0.0% |
| 32,768 tokens (32k) | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 0.0% |
| 65,536 tokens (64k) | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 0.0% |
Standard autoregressive models processing raw 32k–64k prompts exhibit severe "lost-in-the-middle" attention degradation. Empirical retrieval accuracy drops to 45%–60% when target needle variables are placed at 50% document depth, accompanied by prefill latency ballooning to 25–40 seconds.
Under AtacamaODR payload optimization, non-essential tokens are stripped prior to GPU prefill. Target dependency relationships are preserved with 100.0% retrieval accuracy regardless of original document depth, eliminating attention dilution across all context lengths.