The Continuity Layer
Reference

AI Memory & Continuity Benchmarks

Public benchmarks that measure memory and long-term context — reference points, not a ranking and not endorsements. Most measure recall or long-context handling; very few attempt to measure continuity properties directly, and that gap is itself informative.

The Evidence Library shows the receipts, references, and source material. These benchmarks show how continuity claims can be evaluated against public reference points — they help discipline the work, not prove it. The Continuity Layer has not passed these benchmarks, and published figures below remain as-reported unless independently reproduced.

02

What a continuity benchmark would add

The benchmarks above are strong on recall and long-context handling. The properties still under-measured are the ones this project reads the field through: whether a changed fact is superseded with a reason, whether a system can say what was true at a given time, whether it disambiguates similar facts, and whether the same state works across models. Those are set out in the Primer’s Continuity Properties.

For the fuller sourced set — including RULER, MemoryBench, and vendor-published evaluations, each with its caveats — see the Evidence Library. Vendor-published continuity frameworks are tracked there and in Signals from the Field rather than listed here, so this page stays limited to independent, reproducible benchmarks.