Research

Evaluation report · LongMemEval-S

supermemory is state of the art in agent memory.

A memory architecture for reliable recall, temporal reasoning, and knowledge updates across long-running conversations.

97%Overall Recall@20
100%Knowledge update
500Questions evaluated

01

Introduction

Large language models fundamentally suffer from forgetting. They treat every interaction as a discrete event and lack the persistent continuity required for personalized user experiences. Larger context windows help, but models still lose information in the middle of long contexts and incur high latency.

This report introduces supermemory, a memory engine designed for long-term coherence. On LongMemEval-S, it handles temporal reasoning and knowledge conflicts in conversation histories exceeding 115,000 tokens.

02

Why LongMemEval?

Many memory benchmarks do not capture the disorder of production environments. LongMemEval tests human-assistant interactions with updates, contradictions, noise, and facts distributed across sessions.

LongMemEval-S contains 500 questions across six categories and evaluates five core capabilities.

  • Information extraction. It tests literal user and assistant recall, plus implicit preferences within a session.
  • Multi-session reasoning. It requires synthesizing information spread across separate conversations.
  • Knowledge updates. It checks whether newer information supersedes obsolete facts.
  • Temporal reasoning. It tests event order, intervals, and relative timestamps.
  • Abstention. It requires recognizing when the available history cannot answer a question.

03

Architecture

supermemory reduces semantic ambiguity by coupling atomic memories with temporal metadata, relations, and their raw source chunks.

3.1Chunk-based ingestion and contextual memories

Large sessions are decomposed into semantic blocks. During indexing, the system generates atomic memories that resolve ambiguous references within each chunk. The original chunk remains attached as evidence.

3.2Relational versioning and knowledge chains

New memories are related to existing ones so facts can evolve without erasing their history.

updates
Records a contradiction or correction as a state change.
extends
Adds new detail to an existing fact without contradiction.
derives
Captures an inference formed from multiple memories.

3.3Temporal grounding

Every memory can carry both a documentDate, when the source was authored, and an eventDate, when the described event occurred. This distinction supports updates, temporal reasoning, and multi-session retrieval.

3.4Hybrid search

Semantic search identifies high-signal memories. Once a memory matches, its original source chunk is injected into the result so the answering model gets both a precise index and the underlying detail.

3.5Session-based ingestion

The dataset is ingested session by session rather than message pair by message pair, preserving the structure of each conversation while keeping sessions independently retrievable.

04

Results

supermemory reaches 97% overall Recall@20 with aggregation. It leads the strongest baseline in every category, including 100% on knowledge updates and single-session assistant recall.

LongMemEval-S Recall@20 with aggregation
CategorysupermemoryZepFull context
SSUSingle-session — User97%92.9%81.4%
SSASingle-session — Assistant100%80.4%94.6%
SSPSingle-session — Preference95%56.7%20%
KUKnowledge Update100%83.3%78.2%
TRTemporal Reasoning95%62.4%45.1%
MSMulti-session96%57.9%44.3%
ALLOverall97%71.2%60.2%

LLM-as-judge evaluation

Full results by answer model and category
SystemSSUSSASSPKUTRMSOverall
Full-contextgpt-4o81.4%94.6%20%78.2%45.1%44.3%60.2%
Zepgpt-4o92.9%80.4%56.7%83.3%62.4%57.9%71.2%
supermemorygpt-4o97%100%95%100%95%96%97%
supermemorygpt-597.14%100%76.67%87.18%81.2%75.19%84.6%
supermemorygemini-3-pro98.57%98.21%70%89.74%81.95%76.69%85.2%

SSUSingle-session user SSASingle-session assistant SSPSingle-session preference KUKnowledge update TRTemporal reasoning MSMulti-session

05

Reproducing the results

The evaluation uses the LongMemEval-S dataset and its question-specific judge prompts. The retrieval configuration is Recall@20 with aggregation, and gpt-4o judges the answers.

The ingestion pipeline, search implementation, and evaluation work are available through the supermemory GitHub organization.

Answering prompt
You are a question-answering system. Based on the retrieved context below, answer the question.

Question: ${question}
Question Date: ${questionDate}

Retrieved Context:
${retrievedContext}

Understanding the Context:
The context contains search results from a memory system. Each result has a memory, its source chunks, temporal context, and profile data when available.

How to Answer:
Start by scanning memory titles to find relevant results. Read the chunks carefully for details and evidence. Use temporal context to understand when things happened and profile data for background about the user. Synthesize information from multiple results if needed.

If the context contains enough information, provide a clear, concise answer. If it does not, respond with “I don't know” or explain what information is missing. Base the answer only on the provided context.

06

Conclusion

Reliable recall, temporal ordering, and knowledge updates are prerequisites for agentic systems. By combining atomic memories, relational versioning, temporal metadata, and source chunks, supermemory turns a stateless model into an assistant that can preserve a coherent user narrative over time.

07

Citations

  1. Liu, N. F. et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, 12, 157–173.
  2. Wu, D. et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.
  3. Maharana, A. et al. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents.
  4. Rasmussen, P. et al. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory.
  5. Keluskar, A., Bhattacharjee, A., and Liu, H. (2024). Do LLMs Understand Ambiguity in Text? IEEE International Conference on Big Data.
  6. Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 33.
  7. Barnett, S. et al. (2024). Seven Failure Points When Engineering a Retrieval Augmented Generation System. IEEE/ACM International Conference on AI Engineering.
  8. Anthropic (2024). Introducing Contextual Retrieval.