Tasks resolved
Related bugs fixed outright by one coding agent.
SWE-ContextBench, T4Dark line: the agent without memorysupermemory
30.30%
Mem0
24.24%
Mem0 compared itself to us with its own numbers. Here are someone else’s: an independent study that ran both under the same agent, and what teams found when they switched.
Benchmark rows from SWE-ContextBench (arXiv 2602.08316), run by its authors. LongMemEval-S figures are each vendor’s own; here is how to read them.
In an independent coding study, supermemory resolved more tasks, fixed more failing tests, retrieved more past fixes, and cost less per task than Mem0. The LongMemEval scores come from each company’s own run.
Related bugs fixed outright by one coding agent.
SWE-ContextBench, T4Dark line: the agent without memorysupermemory
30.30%
Mem0
24.24%
Tests that failed before the patch and pass after it.
SWE-ContextBench, T4Dark line: the agent without memorysupermemory
55.95%
Mem0
20.24%
How often memory brought back the fix that solved a related bug.
SWE-ContextBench, T5supermemory
59.60%
Mem0
39.39%
What one task cost to run, memory included. Lower is better.
SWE-ContextBench, T4Dark line: the agent without memorysupermemory
$0.58
Mem0
$0.62
Recall, updates and time across 500 long chat histories.
Vendor-reportedsupermemoryGPT-4o
97
Mem0Self-reported
94.4
On 99 related tasks, the agent resolved 30 with supermemory, compared with 24 with Mem0 and 26 with no memory. These results describe the specific setups tested in the study.
30.30%
resolved with supermemory 4 more tasks than no memory8 of 10measures go to supermemory
supermemory ahead 8Mem0 ahead 2
Accuracy
Bugs fully fixedResolved
30.30%24.24%
+6.1 pts
Failing tests now passingFail-to-pass, by test
55.95%20.24%
2.8× more
Passing tests left intactPass-to-pass, by test
90.23%82.44%
+7.8 pts
Patches that appliedPatch applies cleanly
89.90%85.86%
+4.0 pts
Tasks with every failing test fixedFail-to-pass, by task
40.40%39.39%
+1.0 pt
Tasks with nothing brokenPass-to-pass, by task
88.89%86.87%
+2.0 pts
Retrieval
Right past fix retrievedMatched anywhere in the results
59.60%39.39%
+20.2 pts
Right fix ranked firstMatched at the top result
23.23%25.25%
Mem0 +2.0
Cost and time, lower is better
Cost per taskNo memory: $0.79
$0.58$0.62
6% less
Time per taskNo memory: 6.4 min
5.0 min4.7 min
Mem0, 19 s
Our report scores 97% with GPT-4o, 85.2% with Gemini 3 Pro, and 84.6% with GPT-5. Mem0’s page compares its 94.4% run with our Gemini 3 Pro result. Pick a model to see the category results.
97%
overall, 500 questionsGPT-4o answers in this supermemory run.
Benchmarks are an afternoon. These are the people who run memory every day, and what it does there.
Moved from Mem0“Mem0 was not great. Glad to have found Supermemory.”
Scira moved its research memory off Mem0. Indexing became reliable, recall got faster, and usage grew about 32% after launch.
Zaid MukaddamFounder, SciraRead the story“We just ditched RAG completely and went memory only through supermemory.”
Armin DaryabegiFounder, Chatarmin“Tried almost everything. The only thing that works reliably is supermemory.”
Harshil MathurFounder, Razorpay| Operation | Server | End to end |
|---|---|---|
| Search | 187ms | 356ms |
| Profile | 166ms | 248ms |
If any of these sound familiar, test both on your workload. Each answer links to the study or SDK example behind it.
Memory isn’t helping your agent finish more tasks.
30 of 99 tasks resolved, against 24 with Mem0 and 26 with no memory.The study
Your agent fixes bugs it has fixed before.
It brings back the right past fix 59.6% of the time. Mem0 does 39.4%.The study
Your context spans users, projects and documents.
Use one namespace per user or project for profiles, memories and documents.The SDK
Memory adds to what every task costs.
$0.58 a task with supermemory, $0.62 with Mem0, in the same study.Benchmarks
Use each user ID as a namespace to scope their documents, memories and profile. Both SDKs cover add, search and profile; switching means adapting payloads and responses, then backfilling existing data.
TypeScript SDKs · supermemory 5.0.1 · mem0ai 3.3.1
import { Supermemory } from "supermemory";const client = new Supermemory(); await client.add(userId, { content });const { results } = await client.search(userId, { query: q });const { profile } = await client.profile(userId);Set SUPERMEMORY_API_KEY for new Supermemory(); pass your Mem0 key as apiKey. Ingestion is asynchronous: these calls show the interfaces, not immediate read-after-write. Mem0 profiles require status: "succeeded"; Supermemory's default dynamic memory formation can take minutes.
You don’t have to. The main study is someone else’s, linked above, and every Mem0 number here was published by Mem0 or by that study. Where Mem0 comes out ahead, we show it.
They are different runs. Mem0’s figure comes from its own evaluation; ours from our report, with GPT-4o answering. Our Gemini 3 Pro and GPT-5 runs score 85.2 and 84.6. Neither has been reproduced under a matched setup, which is why the independent coding study leads this page.
Not in a form we can link here yet, so those rows are left out rather than estimated. When we publish them, they will be on this page with their method.
Yes. It runs in our cloud, in your VPC, in your own data center, and air-gapped. Talk to us about your deployment requirements.
Yes. Use each user ID as a supermemory namespace, the boundary that scopes that user’s documents, memories and profile. Both hosted SDKs support user profiles, but their payloads and response shapes differ. Update your adapter and backfill existing data; the snippets show the calls, not a drop-in migration.
Sources: SWE-ContextBench, Zhu et al., arXiv 2602.08316 v3, Tables 4 and 5, run by its authors. LongMemEval-S from supermemory.ai/research; Mem0’s figure from mem0.ai/compare/mem0-vs-supermemory (July 8, 2026). Scira’s story from our case study. Production figures from our own telemetry. Mem0 is a trademark of its owner; supermemory is not affiliated with it.