[ Research](https://supermemory.ai/research/) 

Evaluation report · LongMemEval-S

# supermemory is state of the art in agent memory.

A memory architecture for reliable recall, temporal reasoning, and knowledge updates across long-running conversations.

**97%**Overall Recall@20

**100%**Knowledge update

**500**Questions evaluated

01

## Introduction

Large language models fundamentally suffer from forgetting. They treat every interaction as a discrete event and lack the persistent continuity required for personalized user experiences. Larger context windows help, but models still lose information in the middle of long contexts and incur high latency.

This report introduces supermemory, a memory engine designed for long-term coherence. On LongMemEval-S, it handles temporal reasoning and knowledge conflicts in conversation histories exceeding 115,000 tokens.

[**Soham Daga** AI Researcher, supermemory ](https://www.linkedin.com/in/soham-daga/) [**Sreeram Sreedhar** AI Researcher, supermemory ](https://www.linkedin.com/in/sreeram-sreedhar/) [**Dhravya Shah** CEO, supermemory ](https://www.linkedin.com/in/dhravyashah/)

02

## Why LongMemEval?

Many memory benchmarks do not capture the disorder of production environments. LongMemEval tests human-assistant interactions with updates, contradictions, noise, and facts distributed across sessions.

LongMemEval-S contains 500 questions across six categories and evaluates five core capabilities.

* **Information extraction.** It tests literal user and assistant recall, plus implicit preferences within a session.
* **Multi-session reasoning.** It requires synthesizing information spread across separate conversations.
* **Knowledge updates.** It checks whether newer information supersedes obsolete facts.
* **Temporal reasoning.** It tests event order, intervals, and relative timestamps.
* **Abstention.** It requires recognizing when the available history cannot answer a question.

03

## Architecture

supermemory reduces semantic ambiguity by coupling atomic memories with temporal metadata, relations, and their raw source chunks.

### 3.1Chunk-based ingestion and contextual memories

Large sessions are decomposed into semantic blocks. During indexing, the system generates atomic memories that resolve ambiguous references within each chunk. The original chunk remains attached as evidence.

### 3.2Relational versioning and knowledge chains

New memories are related to existing ones so facts can evolve without erasing their history.

updates

Records a contradiction or correction as a state change.

extends

Adds new detail to an existing fact without contradiction.

derives

Captures an inference formed from multiple memories.

### 3.3Temporal grounding

Every memory can carry both a `documentDate`, when the source was authored, and an `eventDate`, when the described event occurred. This distinction supports updates, temporal reasoning, and multi-session retrieval.

### 3.4Hybrid search

Semantic search identifies high-signal memories. Once a memory matches, its original source chunk is injected into the result so the answering model gets both a precise index and the underlying detail.

### 3.5Session-based ingestion

The dataset is ingested session by session rather than message pair by message pair, preserving the structure of each conversation while keeping sessions independently retrievable.

04

## Results

supermemory reaches 97% overall Recall@20 with aggregation. It leads the strongest baseline in every category, including 100% on knowledge updates and single-session assistant recall.

__LongMemEval-S Recall@20 with aggregation__
| Category                       | supermemory | Zep   | Full context |
| ------------------------------ | ----------- | ----- | ------------ |
| SSUSingle-session — User       | 97%         | 92.9% | 81.4%        |
| SSASingle-session — Assistant  | 100%        | 80.4% | 94.6%        |
| SSPSingle-session — Preference | 95%         | 56.7% | 20%          |
| KUKnowledge Update             | 100%        | 83.3% | 78.2%        |
| TRTemporal Reasoning           | 95%         | 62.4% | 45.1%        |
| MSMulti-session                | 96%         | 57.9% | 44.3%        |
| ALLOverall                     | 97%         | 71.2% | 60.2%        |

### LLM-as-judge evaluation

__Full results by answer model and category__
| System                      | SSU    | SSA    | SSP    | KU     | TR     | MS     | Overall |
| --------------------------- | ------ | ------ | ------ | ------ | ------ | ------ | ------- |
| **Full-context**gpt-4o      | 81.4%  | 94.6%  | 20%    | 78.2%  | 45.1%  | 44.3%  | 60.2%   |
| **Zep**gpt-4o               | 92.9%  | 80.4%  | 56.7%  | 83.3%  | 62.4%  | 57.9%  | 71.2%   |
| **supermemory**gpt-4o       | 97%    | 100%   | 95%    | 100%   | 95%    | 96%    | 97%     |
| **supermemory**gpt-5        | 97.14% | 100%   | 76.67% | 87.18% | 81.2%  | 75.19% | 84.6%   |
| **supermemory**gemini-3-pro | 98.57% | 98.21% | 70%    | 89.74% | 81.95% | 76.69% | 85.2%   |

SSUSingle-session user SSASingle-session assistant SSPSingle-session preference KUKnowledge update TRTemporal reasoning MSMulti-session

05

## Reproducing the results

The evaluation uses the LongMemEval-S dataset and its question-specific judge prompts. The retrieval configuration is Recall@20 with aggregation, and gpt-4o judges the answers.

The ingestion pipeline, search implementation, and evaluation work are available through the [supermemory GitHub organization](https://github.com/supermemoryai).

Answering prompt

```
You are a question-answering system. Based on the retrieved context below, answer the question.

Question: ${question}
Question Date: ${questionDate}

Retrieved Context:
${retrievedContext}

Understanding the Context:
The context contains search results from a memory system. Each result has a memory, its source chunks, temporal context, and profile data when available.

How to Answer:
Start by scanning memory titles to find relevant results. Read the chunks carefully for details and evidence. Use temporal context to understand when things happened and profile data for background about the user. Synthesize information from multiple results if needed.

If the context contains enough information, provide a clear, concise answer. If it does not, respond with “I don't know” or explain what information is missing. Base the answer only on the provided context.
```

06

## Conclusion

Reliable recall, temporal ordering, and knowledge updates are prerequisites for agentic systems. By combining atomic memories, relational versioning, temporal metadata, and source chunks, supermemory turns a stateless model into an assistant that can preserve a coherent user narrative over time.

07

## Citations

1. Liu, N. F. et al. (2024). _Lost in the Middle: How Language Models Use Long Contexts._ Transactions of the Association for Computational Linguistics, 12, 157–173.
2. Wu, D. et al. (2024). [_LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory._ ](https://arxiv.org/abs/2410.10813)
3. Maharana, A. et al. (2024). [_Evaluating Very Long-Term Conversational Memory of LLM Agents._ ](https://arxiv.org/abs/2402.17753)
4. Rasmussen, P. et al. (2025). [_Zep: A Temporal Knowledge Graph Architecture for Agent Memory._ ](https://arxiv.org/abs/2501.13956)
5. Keluskar, A., Bhattacharjee, A., and Liu, H. (2024). _Do LLMs Understand Ambiguity in Text?_ IEEE International Conference on Big Data.
6. Lewis, P. et al. (2020). _Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks._ NeurIPS 33.
7. Barnett, S. et al. (2024). _Seven Failure Points When Engineering a Retrieval Augmented Generation System._ IEEE/ACM International Conference on AI Engineering.
8. Anthropic (2024). [_Introducing Contextual Retrieval._ ](https://www.anthropic.com/engineering/contextual-retrieval)
