Field Watch — Benchmark
RULER
01
Why it matters
A synthetic, configurable suite measuring the real usable context size of long-context models, not just their advertised context window.
02
Primary sources
03
Status
MatureLong-context benchmark
- Date observed
- Last verified
How entries are chosen, categorized, and re-checked: How the map works.