learner-1 is out.Read the note

supermemory is building the default engine for memory and continual learning for agents.

Available through our API, plugins, and MCP.

The path to useful AI is to scale in-context continuous learning that works with any model, any harness, and for all the varieties of use cases. As open-source model adoption grows and the world becomes even more multi-model, learning and memory should be separated from the model providers and be highly interoperable (hence in-context!).

Intelligence is no longer a constraint.

Pre-training and post-training, scaling on data and compute, will continue to push models’ intelligence, but will not enable them to learn and improve in real time.

We build all the hard infrastructure parts of a scalable memory system, with dense, interconnected, growing learnings, and an understanding of time.

Our model learner-1 helps any model continuously improve by extracting and dreaming on the context of any user, task, or tenant, and putting it on our vector-graph database. This is then brought to the model by injecting tokens into its context in real time.

We’re a small team based in San Francisco.

Today, supermemory already serves the memory and context for 100k+ organizations, with tens of millions of end users, processing hundreds of billions of tokens every day. It is accessible via an API, with integrations for multiple harnesses available for use.

What we do

Memory that keeps learning.
learner-1 extracts and dreams on the context of every user, task, and tenant, stores it in a vector-graph database, and brings it back to the model by injecting tokens into its context in real time.
Any model, any harness.
Learning and memory live outside the model providers, so they carry across models and harnesses and stay interoperable, in context, and yours.
The hard infrastructure, done.
Dense, interconnected, growing learnings with an understanding of time, already serving 100k+ organizations and hundreds of billions of tokens a day.

Use supermemory for yourself with our plugins, or build your agents on top of it.

In production

Recall latencyper query, at the API<300ms
Tokens processedper day100B+
Organizationstens of millions of end users100k+
BenchmarksState of the art on LongMemEval, LoCoMo and ConvoMem#1See the research

“Reduced avg response time from 40s → 12s. Using about 40–50% fewer tokens.”

Founder, Chatarmin · on X

Research

  1. 01SMFS: making agentic retrieval 55% cheaper AND more accurate
  2. 02Introducing Dynamic Dreaming: supermemory now connects the dots, for you.
  3. 03Dear reader, we just made supermemory insanely cheap... the Context Cloud
All research

Careers

We’re a small team based in San Francisco, building the memory layer for every model and every harness. If that is the problem you want to spend the next few years on, we would like to hear from you.

See open roles