AI Memory & Continuity Benchmarks
Public benchmarks that measure memory and long-term context — reference points, not a ranking and not endorsements. Most measure recall or long-context handling; very few attempt to measure continuity properties directly, and that gap is itself informative.
The Evidence Library shows the receipts, references, and source material. These benchmarks show how continuity claims can be evaluated against public reference points — they help discipline the work, not prove it. The Continuity Layer has not passed these benchmarks, and published figures below remain as-reported unless independently reproduced.
The benchmarks
Each entry: what it measures, its citation, and what it doesn’t cover. Reported figures — question counts, token lengths — are as published by each benchmark’s authors; this project has not independently reproduced them.
LOCOMO
Very long-term conversational memory across up to 35 sessions — single-hop, multi-hop, temporal, open-domain, and adversarial question answering.
Measures conversational recall, not the fuller continuity property set (update-handling, disambiguation, reconstruction, model independence).
LongMemEval
500 questions across information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention — the closest public benchmark to several continuity properties.
Results are per-model and shift as models update; it tests chat-assistant memory, not a full governed-state system.
BEAM
Automatically generated conversations up to 10M tokens — 100 conversations, 2,000 validated questions, ten memory abilities — built to force genuine retrieval-based memory rather than reward long-context stuffing.
Synthetic conversations; real-world conversational structure may differ.
What a continuity benchmark would add
The benchmarks above are strong on recall and long-context handling. The properties still under-measured are the ones this project reads the field through: whether a changed fact is superseded with a reason, whether a system can say what was true at a given time, whether it disambiguates similar facts, and whether the same state works across models. Those are set out in the Primer’s Continuity Properties.
For the fuller sourced set — including RULER, MemoryBench, and vendor-published evaluations, each with its caveats — see the Evidence Library. Vendor-published continuity frameworks are tracked there and in Signals from the Field rather than listed here, so this page stays limited to independent, reproducible benchmarks.
Continue the proof trail
Benchmarks are the yardstick. From here, follow the roles: the declared proof, the receipts, a hands-on build, and the model behind it.