LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
A 500-question benchmark testing information extraction, reasoning across sessions, time, knowledge updates and abstention.
What it contributes
Organises memory evaluation around distinct abilities and connects performance to indexing, retrieval and reading choices. It supplies a practical vocabulary for diagnosing memory failures.
Read with care
Performance on questions grounded in conversation histories does not establish relationship quality, privacy protection or safe real-world behaviour. Model comparisons reflect the versions and conditions tested.
Why it belongs here
For engineers choosing memory architectures and evaluators looking beyond simple recall.
ICLR 2025 conference paper; arXiv v2 dated 4 March 2025.
Source reading depthAbstract, version history and ICLR publication status checked. This is a source note, not a full critical review.
Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.