[Blog](https://supermemory.ai/blog) · [Engineering](https://supermemory.ai/blog/tag/engineering)

# How AI memory actually works. A beginner's guide

How AI apps remember things between chats, why a bigger context window doesn't solve it, and the parts every memory system shares.

By Muskan Jain October 9, 2026 · 9 min read 

![A glowing card cabinet with one drawer open, index cards linked by golden threads floating out and one faded card drifting away, on a grainy blue background beside the headline How AI memory actually works.](https://supermemory.ai/_astro/cover.DdgJCty7_Zj689a.webp)

In one test, GPT-4o got 87% of questions right when it only saw the parts of a chat history that had the answer. When it saw the whole history, about 115k tokens, the same model got 60.6%.

That's from [LongMemEval](https://arxiv.org/abs/2410.10813), a memory benchmark. The answer was in the chat both times. All the extra text just made it harder to find.

![A bar chart of GPT-4o on the same 500 LongMemEval questions. It scores 87% when given only the sessions that contain the answer, and 60.6% when given the full chat history of about 115k tokens.](https://supermemory.ai/_astro/more-context.D18uw6ih_1qGC0x.webp)

## Models don't remember anything

[Anthropic's API docs](https://platform.claude.com/docs/en/build-with-claude/working-with-messages) say the API doesn't store your conversation, so your app sends the whole thing again on every turn. If you tell a chatbot your name and then ask for it, it only knows because your app sent the first message again.

![Three turns of one chat. On each turn the app sends every earlier message to the model again, and under each model it says "Keeps: nothing".](https://supermemory.ai/_astro/stateless.DtKi4Wsi_Z1HqibH.webp)

A model does know things from training, but that knowledge stops at a cutoff date.

Everything else has to fit in the context window. That's the text the model can see in one call, and it's measured in tokens (a token is usually a word or part of a word). Claude's newest models can read about 1M tokens at once.

So any app that seems to remember you is keeping notes somewhere else and pasting the right ones back in.

## Why not use a bigger window?

You'd think a 1M-token window would fix this. The test at the top already shows that fitting more into the window can make answers worse. Models also miss information buried in the middle of a long prompt ([a 2023 study](https://arxiv.org/abs/2307.03172)).

It's expensive too. A 100-turn chat is only about 50k tokens long, but every turn resends everything before it, so by the end your app has sent about 2.5M tokens.

To be fair, for short histories, putting everything in the prompt works well. Salesforce's [ConvoMem benchmark](https://arxiv.org/abs/2511.10523) found that sending the full history works best for the first 30 conversations and is still workable up to about 150\. Past that, it gets slow and expensive, and the paper recommends search-based or hybrid setups instead.

## Two ways to use raw data

Every memory system starts with raw data, like chats, documents and files. There are two basic ways to use it.

The first is to keep the raw data as it is and let the agent search it. The search can be as simple as grep, which looks for exact words. It can also be semantic search, which turns text into embeddings (lists of numbers based on meaning) and finds the closest matches. That's how "my flight got cancelled" can match "the airline scrapped my trip", even though they share almost no words. RAG (retrieval-augmented generation) is the classic version of this, where documents are split into chunks and the closest chunks get pasted into the prompt.

But raw text has no idea which fact is current, and the agent has to dig through all of it to find what it needs.

Say you tell an assistant on day 1 that you only run in adidas. On day 40 you say you switched to Puma and you're never going back. On day 60 you ask for running shoes. Search might still pull up the day 1 message, since it's a close match for a question about running shoes.

![A timeline. On day 1 the user says "I only run in adidas", on day 40 "switched to Puma, never going back", and on day 60 "recommend me running shoes". Search recommends adidas and memory recommends Puma.](https://supermemory.ai/_astro/search-vs-memory.CXTvYksl_2jJnF3.webp)

The second way is to put a model in front. The raw data goes through a model first, and the model rearranges it into something tidier, like notes, a wiki or a list of facts. That model can notice when a fact changes. The tradeoff is that everything now depends on what the model decides to keep, merge or drop.

## What every memory system has

Our founder, [Dhravya Shah](https://x.com/DhravyaShah), [wrote about](https://x.com/DhravyaShah/status/2103314339239428201) what we found after studying the memory in ChatGPT, Claude, Instinct, OpenClaw, Hermes, Muse and others. Whether a system stores markdown files, a graph or a list of facts, they all share four parts.

1. A search step. This can be RAG, grep or semantic search, and something always has to find the relevant pieces.
2. Chunking. The data gets split into pieces, because it can't all fit into one model call. That goes for search and for learning.
3. Learning outside the main conversation. The agent you're talking to usually isn't the one deciding what to remember. Learning runs on a schedule or after a trigger.
4. Code in the app that puts what was found back into the model's context.

Underneath, two things are always there. One is a model that arranges the data, in the background or in real time. The other is a store that holds it.

![Dhravya's map, redrawn. Raw data is chunked, a background job learns from it, and the result is stored. A profile and a search step then feed what gets injected into the agent's context, while the live chat goes straight in.](https://supermemory.ai/_astro/memory-pipeline.BSkW5YAD_1o7pth.webp)

## Memory is a lifecycle

First, a memory system decides what to save. The test is whether a fact would change a future answer, so "I'm vegetarian" is worth keeping and "It's raining here" isn't.

Then it keeps facts up to date. If you mention a new job, the old one should stop showing up in answers. Some facts expire, too. "I have an exam tomorrow" is wrong three weeks later.

It can also work out things you never said. If you mention you just moved to Tokyo for work, it can guess your timezone changed too.

All of this can happen in the background, before anyone asks anything, so nobody waits on it. Pulling facts back out is the only part that happens live. When a message comes in, it finds the few facts that matter for it.

![The five jobs of a memory system: write, update, forget, infer and recall. Write, update, forget and infer can run in the background, and recall happens live.](https://supermemory.ai/_astro/memory-lifecycle.BsTvKpBL_20iOpo.webp)

## Time and connections

The store has two things to keep track of. One is time, and the other is connections.

Time is about what happened when. Git works this way. Every change is saved in order, so you can see what was true before and what's true now. Some memory systems use Git directly. [Instinct](https://x.com/DhravyaShah/status/2101745550752428340), for example, keeps its memory in markdown files tracked by Git and searches them by keyword. Time tells you the Puma message came after the adidas one.

Connections are about which facts point to each other. [Obsidian](https://obsidian.md/help/links) works this way. Notes link to other notes, and you can follow the links. If a memory knows that "Alex works at Stripe" and "Alex leads a team of 5" are about the same person, it can answer questions that neither fact answers alone. A link can also say that one fact replaced another, which time alone can't tell you.

![A memory store keeps track of two things. On the left, time: a timeline of day 1 "I only run in adidas", day 40 "Switched to Puma" and day 60 "Recommend me running shoes". On the right, connections: "works at Stripe", "leads a team of 5" and "talks about payment APIs", all linked to Alex.](https://supermemory.ai/_astro/time-and-connections.Dqu3qE6p_VCwIm.webp)

## Where memory plugs in

You can connect memory to an agent in two ways. One is to give the model a memory search tool and let it decide when to use it. That only works if the model notices it's missing something. The other is hooks, where the app fetches and saves memories at set moments without waiting for the model to ask.

## How we do it at supermemory

We do the arranging with learner-1, a model we built for the job. It reads what you send, pulls out facts and links them to the ones it already has. We call this [dreaming](https://supermemory.ai/blog/introducing-dynamic-dreaming-supermemory-now-connects-the-dots-for-you/), and [OpenAI uses the same word](https://openai.com/index/chatgpt-memory-dreaming/) for the background process behind ChatGPT's memory.

We never retrain your model. The memories go into its prompt, so this works with any model.

When you send us something, like a chat or a PDF, you can read it back three ways. Documents are chunks of the original, for when you want to know what the source said. Memories are the facts learner-1 pulled out, for what's true about someone right now. The profile is a short summary your agent should always have in front of it.

![One input goes into supermemory, which chunks and indexes it first, then extracts facts and links them to old ones in the background. Three outputs come back: documents, memories and a profile.](https://supermemory.ai/_astro/one-write-three-reads.Bmf0nDnX_1jt1Be.webp)

Our store keeps track of both time and connections. Each memory is one small fact instead of a whole file, and each fact points to the facts it relates to. Here's what that looks like for a user called Alex.

* **Updates:** "Alex just started at Stripe as a PM" replaces "Alex is a software engineer at Google". Search returns the new fact, and the old one is kept as history.
* **Extends:** "Alex focuses on payments, leads a team of 5" adds detail. Both facts stay true.
* **Derives:** from the Stripe fact and "Alex often talks about payment APIs and fraud", the system guesses that Alex works on Stripe's core payments product. It trusts that guess less until it's confirmed.

The updates link is what fixes the running shoes problem from earlier.

![Five fact cards about Alex joined by updates, extends and derives arrows. The old Google fact is kept as history, and the guessed fact is marked as weighted lower until confirmed.](https://supermemory.ai/_astro/fact-graph.fofsVZSk_ZXMDfL.webp)

The profile holds long-term facts about you and what you've been up to lately. Search only has your message to go on, and "hey, what should I cook tonight?" says nothing about diet. The profile gets added whatever you ask, so the model still knows you're vegetarian.

Our Claude Code plugin uses hooks. It loads your profile when a session starts, searches memory with your prompt before Claude replies, and saves the new messages after each reply. Whenever it pulls memories into your prompt, you see a line like `◪ supermemory · recalled 5 memories (242 tok)`.

Our coding plugins also share one memory per repo. If you switch from Claude Code to Cursor halfway through a project, Cursor already knows the decisions you made.

## How do you know it works?

LongMemEval-S, the version of LongMemEval we used, gives each of its 500 questions its own chat history of about 115k tokens to search. Our [LongMemEval-S report](https://supermemory.ai/research/longmembench/) by Soham Daga, Sreeram Sreedhar and Dhravya Shah measures how often our top 20 search results contain what's needed to answer the question. With aggregation on, a setting that can merge several related memories into one result, that happens for 97% of the questions, and for 100% of the questions where a fact changed partway through the chat.

That 97% measures search. The 60.6% at the top counts how many answers GPT-4o got right after reading the whole history, so the two numbers can't be compared directly. To get a right answer, a model still has to read the search results and write the reply, so the final score depends on that model too.

We open-sourced [MemoryBench](https://github.com/supermemoryai/memorybench) so anyone can run LongMemEval on supermemory, or on their own setup. Its default run fetches 10 results with aggregation off, so it's a different test from the one in our report.

## Try it

Here's the whole thing in code, using our supermemory npm package. Use one namespace per user, project or repo, and memories stay inside it.

```
import { Supermemory } from "supermemory" // npm i supermemory

const client = new Supermemory({ apiKey: process.env.SUPERMEMORY_API_KEY })

// 1. Write: send a conversation. Facts get extracted and
//    linked in the background, which can take a few minutes.
await client.add("user_123", {
  content: "user: I just moved from Google to Stripe as a PM",
})

// 2. Recall: the profile, plus memories relevant to this question
const { profile } = await client.profile("user_123")
const { results } = await client.search("user_123", {
  query: "Suggest a talk I should give",
})

// 3. Inject: put it in the system prompt of any model
const context = [
  ...profile.static.map((m) => m.memory),
  ...profile.dynamic.map((m) => m.memory),
  ...results.map((r) => r.memory ?? r.chunk),
].filter(Boolean).join("\n")
```

If you're building an agent, [get an API key](https://console.supermemory.ai/) or try our [coding plugins](https://supermemory.ai/docs/integrations/claude-code).

## Other posts.

1. [We're open sourcing the company brain. Here's how we designed the multiplayer harness Company Brain is now open source. A walkthrough of the multiplayer harness behind its Slack experience, from proactivity and memory boundaries to approvals and recovery. NewsSep 25, 2026 ](https://supermemory.ai/blog/open-sourcing-company-brain)
2. [Jev changes a lot in memory & context engineering. Here's exactly how. We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short. EngineeringSep 24, 2026 ](https://supermemory.ai/blog/jev-memory-context-engineering)
3. [I reverse-engineered Instinct's memory. Here's exactly how it works Instinct keeps its memory as git-tracked markdown files, found with grep rather than vectors. Here is the whole system as far as black-box probing can reconstruct it, and how to rebuild it on supermemory in about 60 lines. EngineeringSep 20, 2026 ](https://supermemory.ai/blog/reverse-engineering-instinct-memory)
4. [An update to supermemory We've discontinued the supermemory company brain and Nova. Everyone who was charged has been refunded, our MCP and plugins continue to run, and we're going all in on the memory engine. NewsSep 10, 2026 ](https://supermemory.ai/blog/an-update-to-supermemory)
5. [Scaling Conversations: How Adapta Grew Usage Without Losing Context Adapta added Supermemory as a persistent memory layer so every conversation keeps its context — letting the team scale usage without losing the thread. Case StudyJun 12, 2026 ](https://supermemory.ai/blog/adapta-scaling-conversations)
6. [How Chatarmin Ditched RAG and Went Memory-Only with Supermemory Chatarmin replaced a heavy RAG pipeline with Supermemory's memory layer — cutting average AI response time from 40s to 12s and token usage by 40–50%. Case StudyJun 10, 2026 ](https://supermemory.ai/blog/chatarmin-memory-only)
7. [SMFS: making agentic retrieval 55% cheaper AND more accurate We launched SMFS.ai (Supermemory Filesystem) a few weeks ago, with a simple bet: We can redesign the filesystem specifically for agents, with special files, structures, and commands that it can use for it's tasks. Today, SMFS is used by hundreds of companies to power their agents. EngineeringMay 28, 2026 ](https://supermemory.ai/blog/smfs-making-agentic-retrieval-55-cheaper-and-more-accurate)
8. [Introducing Dynamic Dreaming: supermemory now connects the dots, for you. Dreaming is magical. TLDR: We're launching Dynamic Dreaming in supermemory today, which automatically works if you're using supermemory in any way - API, OpenClaw, Hermes agent, etc. EngineeringMay 25, 2026 ](https://supermemory.ai/blog/introducing-dynamic-dreaming-supermemory-now-connects-the-dots-for-you)
9. [Dear reader, we just made supermemory insanely cheap... the Context Cloud When I first started building supermemory, I had one goal: To build the best memory system for AI. I would talk to customers, and find out that memory was not the only thing they needed - They were all setting up 7-8 different vendors at the same time. EngineeringMay 18, 2026 ](https://supermemory.ai/blog/dear-reader-we-just-made-supermemory-insanely-cheap-the-context-cloud)
10. [Introducing @supermemory/tools v2.0.0 Today we're releasing v2.0.0\. This release unifies the API across all agents sdk integrations from AI SDK to Mastra, makes conversation identity a first-class concept, and ships with memory saving on by default. EngineeringApr 27, 2026 ](https://supermemory.ai/blog/introducing-supermemory-tools-v2-0-0)
11. [Solving the Precision-Recall Tradeoff: Search Result Aggregation When you're building memory for AI, search is your foundational layer. The way search generally works is straightforward: the user defines a query, and then sets a limit (top-K) on how many search results they want returned. Usually, this is set to 10 or 20. EngineeringApr 5, 2026 ](https://supermemory.ai/blog/solving-the-precision-recall-tradeoff-search-result-aggregation)
12. [OpenClaw Memory Problems: Why It Forgets and How to Fix It (2026) TLDR: Today, we are releasing a new version of our openclaw plugin - https://github.com/supermemoryai/openclaw-supermemory. This post is going to be a bit technical, so bear with me (or bookmark for later!) In this post, I will talk about what we do about OpenClaw memory, and how we fix it. EngineeringFeb 19, 2026 ](https://supermemory.ai/blog/why-everyone-is-complaining-about-openclaws-memory-it-sucks-and-why-supermemory-fixes-it)
13. [Stateful Coding Agents with Memory: Build Long-Running Agents (2026) We built a plugin for Claude Code and OpenCode that gives your coding agent persistent memory. It remembers your preferences, learns your codebase, and never loses context mid-conversation. The result is an agent you can run for months without starting over. EngineeringFeb 18, 2026 ](https://supermemory.ai/blog/infinitely-running-stateful-coding-agents)
14. [Clawd / Molt bot's memory SUCKS. We gave it supermemory. I'm the founder of supermemory. Clawd/Molt bot is blowing up right now, with many, many use cases. I set it up, too, and have been using it through telegram. TLDR: just go to https://supermemory.ai/docs/integrations/clawdbot to set up supermemory for your clawd bot. EngineeringJan 28, 2026 ](https://supermemory.ai/blog/clawd-molt-bots-memory-sucks-we-gave-it-supermemory)
15. [Catch up with our UNFORGETTABLE Launch Week Over the last year, one belief has guided almost everything we’ve built at Supermemory AI becomes meaningfully useful only when it remembers. Memory shouldn’t be something developers rebuild from scratch. It shouldn’t be fragile, expensive, or trapped inside a single tool. NewsJan 4, 2026 ](https://supermemory.ai/blog/catch-up-with-our-unforgettable-launch-week)
16. [Empowering the Next Generation of Founders: Supermemory Startup Program If there’s one thing we’ve learned while building Supermemory, it’s that most startups don’t fail because they didn't build features; they fail when infrastructure slows them down, or they built too slow. NewsDec 31, 2025 ](https://supermemory.ai/blog/empowering-the-next-generation-of-founders-supermemory-startup-program)
17. [Building code-chunk: AST Aware Code Chunking At Supermemory, we're building context engineering infrastructure for AI. A huge part of that is dealing with code: ingesting repos, understanding structure, and making it searchable. The problem is that most code chunking solutions are terrible. We built code-chunk to fix this. EngineeringDec 29, 2025 ](https://supermemory.ai/blog/building-code-chunk-ast-aware-code-chunking)
18. [Supermemory raises $3 million with the best memory engine for LLMs Today, I am excited to announce our first funding round to accelerate our mission of building an interoperable, scalable and reliable memory for LLMs and agents. Memory is one of the hardest challenges in AI right now. NewsOct 6, 2025 ](https://supermemory.ai/blog/supermemory-raises-3-million-and-building-the-best-memory-engine-for-llms)
19. [Mem0 vs Supermemory: Why Scira Switched Scira AI moved its production memory layer from Mem0 to Supermemory. This is what failed, what improved, and how the team evaluated the two systems. Case StudyOct 2, 2025 ](https://supermemory.ai/blog/why-scira-ai-switched)
20. [Never Record Again: How Montra Uses Supermemory to Rethink Video Creation Campbell Baron, the founder of Montra, has been making videos since he was twelve. By thirteen, he was already doing brand work. Today, he’s betting on a very different future for creators: a world where recording is the exception, and most videos are generated from scratch. Case StudyAug 21, 2025 ](https://supermemory.ai/blog/never-record-again-how-montra-uses-supermemory-to-rethink-video-creation)
21. [Unified Memory That Works Where You Work: Your Second Brain With Supermemory Hi everyone, I’m Dhravya, the founder of Supermemory. I want to start with a little story behind why this product means so much to me. You can also skip straight to what it is and how it works below. EngineeringJul 25, 2025 ](https://supermemory.ai/blog/unified-memory-that-works-where-you-work-your-second-brain-with-supermemory)
22. [Supermemory just got faster on PlanetScale What is Supermemory? Supermemory completes the missing part of the LLM puzzle: memory. Just as memory is crucial for human intelligence, it's essential for truly intelligent AI systems. EngineeringJul 18, 2025 ](https://supermemory.ai/blog/supermemory-just-got-faster-on-planetscale)
23. [Faster, smarter, reliable infinite chat: Supermemory IS context engineering. People are obsessed with prompts and prompt engineering. Sure, what you say is important, but what the model knows when you say it is the difference between a stateless text generator and an intelligent AI system. In short, context is the most crucial component. NewsJul 9, 2025 ](https://supermemory.ai/blog/faster-smarter-reliable-infinite-chat-supermemory-is-context-engineering)
24. [We solved AI API interoperability One API to rule them all, One spec to find them, One library to bring them all and in the TypeScript, bind them. When we were building the Infinite Chat API, initially, we only supported the OpenAI format. This was fine, until a lot of our customers started asking for more. EngineeringJul 7, 2025 ](https://supermemory.ai/blog/we-solved-ai-api-interoperability)
25. [The Wow Factor of Memory - How Flow Used Supermemory To Build Smarter, Stickier Products Overview: Flow is a note-taking app built around a bold vision: to create a more personal, context-aware writing experience powered by AI. At the heart of this mission is memory. Case StudyJun 14, 2025 ](https://supermemory.ai/blog/the-wow-factor-of-memory-how-flow-used-supermemory-to-build-smarter-stickier-products)
26. [The UX and technicalities of awesome MCPs Last month, we launched the Supermemory MCP, mostly to test our own infrastructure and get some initial traction. It blew up. To my absolute surprise, the initial launch itself got half a million impressions (!!!). Then, we launched and got #2 on ProductHunt too. EngineeringJun 8, 2025 ](https://supermemory.ai/blog/the-ux-and-technicalities-of-awesome-mcps)
27. [Architecting a memory engine inspired by the human brain Language is at the heart of intelligence, but what truly powers meaningful interaction is memory — the ability to accumulate, recall, and contextualize information over time. Large Language Models (LLMs) have mastered language, but memory remains their Achilles’ heel. EngineeringJun 5, 2025 ](https://supermemory.ai/blog/memory-engine)

## Start building with supermemory.

Memory and continual learning for any model, any harness. Available through our API, plugins, and MCP.

[Build with supermemory ](https://console.supermemory.ai/)
