[Blog](https://supermemory.ai/blog) · [Learning](https://supermemory.ai/blog/tag/learning)

# AI's next big thing: personalization and (super)memory.

A perspective on useful AI personalization, with the boundaries between retrieval, changing facts, profiles and durable memory made explicit.

By Dhravya Shah January 24, 2026 · 6 min read 

![Personalization with AI memory](https://supermemory.ai/_astro/cover.DZHp4uWk_ZURcBp.webp)

You are probably thinking of AI memory in the wrong way.  

  
Over the last few years, we've all seen a _lot_ of absolutely world-changing trends in AI. Things that totally changed the way we interact with computers today. The first one was **data (models start getting smarter)**, then it was **inference (everyone is able to run them)**, and vector databases/RAG. Now there are **agents everywhere (with claude code)**. So, what's next?

I believe that the next big inflection point is **memory**. And by memory, I mean truly _magical_ personalization. A feeling that the user gets, when their assistant will proactively bring up things that they would never have expected. When context lengths are no longer a concern, and people can just have long running, thoughtful conversations with their assistant.

I believe that this true personalization can be brought in some beautifully simple ways. But first, let's talk about what the industry is getting wrong.

## Retrieval alone does not define a memory lifecycle

RAG supplies retrieved information to a model before it answers. A basic vector-search example can retrieve isolated passages without resolving which fact is current. That is a limitation of that implementation, not of all RAG systems. RAG can use changing sources, timestamps, version rules and conversation history.

A memory system adds a deliberate policy for what to retain, how it changes and what should inform the next interaction.

**Memory evolves, updates, and derives/learns new information**.  

  
Let's take a scenario.

```
Day 1: "I love Adidas sneakers"
Day 30: "My Adidas broke after a month, terrible quality"
Day 31: "I'm switching to Puma"
Day 45: "What sneakers should I buy?"
```

For illustration, a similarity-only search might return the old preference even after the user changes their mind.

```
# A naive retrieval baseline sees isolated records
query = "What sneakers should I buy?"

# Semantic search finds closest match
result = vector_search(query)
# Possible stale result: "I love Adidas sneakers"

# Agent recommends Adidas 🤦
```

The assistant should use the updated preference for current questions and keep the earlier preference available when the user asks about the past.

```
# Desired behavior for a memory-aware implementation
query = "What sneakers should I buy?"

# Memory retrieval considers:
# 1. Temporal validity (Adidas preference is outdated)
# 2. Causal relationships (broke → disappointment → switch)
# 3. Current state (now prefers Puma)

# Agent correctly recommends Puma ✅
```

The same principle applies to relevance over time. An old exam worry need not dominate today's answer. That does not require pretending the event never occurred or deleting every historical record. Correction, retrieval eligibility and permanent deletion are separate operations. Supermemory's [graph-memory documentation](https://supermemory.ai/docs/concepts/graph-memory) describes updates, extensions and inferences.

## Agentic discovery has a workload-dependent cost

An agent can inspect files and search sources on demand. That can be useful when a question needs detailed evidence. It can also add model calls and latency compared with a direct lookup of a known preference.

The cost depends on how much the agent has to inspect. Measure model calls, search tools and retrieved context together. A direct preference lookup and an investigation across a large repository need different latency budgets.

Loading unnecessary history can increase input cost and distract the model, but it does not make hallucination inevitable. Likewise, compaction can retain useful preferences when its summary preserves them. Test what survives repeated compaction and use persistent records when the workflow needs independently inspectable context.

## How we approach memory at Supermemory

Memory needs to be useful without slowing down the conversation or loading every past message. How do we build that? At supermemory, we have been thinking a lot about the best way to approach this memory problem, and build it like the human brain. [supermemory.ai/blog/memory-engine](https://supermemory.ai/blog/memory-engine/)

1. **A vector-graph architecture to track knowledge change.**

![Knowledge graph of clustered nodes and connections glowing in blue, green, and purple on a dark canvas](https://supermemory.ai/_astro/image-1.BsLpsC7Z_3Dye0.webp)

Supermemory's graph links memories through update, extension and inference relationships. A container supplies scope for related context; it is not restricted to one person or proof that every extracted statement is a fact. See the documented [graph model](https://supermemory.ai/docs/concepts/graph-memory) and [container scopes](https://supermemory.ai/docs/concepts/container-tags).

**Updates: Information Changes**

```
Memory 1: "Alex works at Google as a software engineer"
Memory 2: "Alex just started at Stripe as a PM"
         ↓
Memory 2 UPDATES Memory 1
```

**Derives: Sleep-time compute for an agent**

```
Memory 1: "Alex is a PM at Stripe"
Memory 2: "Alex frequently discusses payment APIs and fraud detection"
         ↓
Derived: "Alex likely works on Stripe's core payments product"
```

The derived statement is an inference, not a confirmed job assignment. It should remain distinguishable from the source facts. Updates can make an earlier value no longer current; that is different from permanent deletion. Explicit forgetting also needs its documented retrieval and retention behavior checked.

The important lifecycle question is whether a record remains eligible for the query being asked. Do not interpret a changed preference or an expired retrieval record as proof that every copy has been erased.

**2\. Memory is not only retrieval!!! The magic of User Profiles**

Query-based retrieval can miss useful background when a message is generic. A greeting or an indirect question may not match the wording of an earlier conversation. Maintained profile context is one way to give the agent relevant background alongside retrieval; it does not guarantee that every inference is correct.

![Side-by-side chat: a generic AI returns nothing while Supermemory uses the user's CEO context for a tailored answer](https://supermemory.ai/_astro/image-3.yOXUPF9a_Z1k7toS.webp)

In this illustration, role context gives the assistant a starting point for a better follow-up. It should still ask about the user's budget and preferences. Remembering someone should make a conversation more relevant, not replace listening to what they say now.

**We built something called** [user profiles](https://supermemory.ai/docs/concepts/user-profiles), which provides maintained user context that an application can request and supply to its agent. A profile is not automatically included in every API response, and the application should select appropriate context for the task.

Think of it as a RAM layer, with both "static" (things the agent should know by DEFAULT) - like the name of the user, their age, etc, and "dynamic" - the episodic / currently-ongoing-endeavours of the user.

```
ILLUSTRATIVE STATIC CONTEXT (relatively stable; still needs correction)
- Name: Dhravya
- Location: San Francisco
- Role: Founder & CEO of Supermemory (AI memory + context platform)
- Age range: 18–25
- Interests: AI infrastructure, developer tools, vector search, PostgreSQL

ILLUSTRATIVE DYNAMIC CONTEXT (recent context, not a live personal record)
- Currently working on: a "Customer Context Graph" for consolidating customer data.
- Actively optimizing: inference cost for Claude and other LLMs.
- Exploring: migrating speech-to-speech from OpenAI to Gemini Live.
- Recent preference change: Adidas sneakers broke → switched preference to Puma.
- Mood: occasionally stressed about infra costs and scale.
```

This brings out a new way of personalization. Pair this with retrieval and evaluate whether it makes the assistant more useful. Personalization should not force unrelated or sensitive details into every answer.

**3\. Getting the best of both worlds: Hybrid retrieval**

Extracted memories are useful for concise facts, but some questions need the surrounding passage. We also return document chunks so the model can use source detail that was not captured in an extracted memory.

In Supermemory, hybrid search can return both memories and document chunks. Processing is asynchronous, so memories are not guaranteed to reflect a source change immediately. Inspect freshness and source provenance instead of assuming every extracted memory is current. The [search documentation](https://supermemory.ai/docs/recall/search) explains the supported modes and options.

Paired with selected profile context, this gives the application several sources of evidence. The application still needs to test whether the selected context is sufficient, current and authorized for the question.

X content is paused until you allow embedded content.

[Open on X ↗](https://twitter.com/DhravyaShah/status/2004702876074213425?ref%5Fsrc=twsrc%5Etfw)

> <https://twitter.com/DhravyaShah/status/2004702876074213425>

## Bringing memory into the conversation

So, that's it. That's how I think LLMs can get the perfect memory that they deserve.

A way to create beautiful user experiences, and make your agents _feel_ unforgettable.

We are building this memory engine at [supermemory.ai](https://supermemory.ai/) \- and would love for you to try it out, and add it to your own agent!

It's not just about retrieval. It's about true personalization. And it's coming to every single AI agent out there.

It's the next inflection point of AI.

## Frequently asked questions

### Can RAG support persistent memory?

Yes. RAG can retrieve changing documents and conversation records. A memory application additionally needs policies for identity, retained facts, corrections and deletion.

### Are memories always current?

No. Processing delays, stale sources and application caches can affect freshness. Verify the source revision and readiness needed for the question.

### Does a changed fact mean its history was deleted?

No. A new fact can supersede an old value while history remains available. Explicit forgetting and permanent deletion have separate behavior.

## Other posts.

1. [We're open sourcing the company brain. Here's how we designed the multiplayer harness Company Brain is now open source. A walkthrough of the multiplayer harness behind its Slack experience, from proactivity and memory boundaries to approvals and recovery. NewsSep 25, 2026 ](https://supermemory.ai/blog/open-sourcing-company-brain)
2. [Jev changes a lot in memory & context engineering. Here's exactly how. We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short. EngineeringSep 24, 2026 ](https://supermemory.ai/blog/jev-memory-context-engineering)
3. [I reverse-engineered Instinct's memory. Here's exactly how it works Instinct keeps its memory as git-tracked markdown files, found with grep rather than vectors. Here is the whole system as far as black-box probing can reconstruct it, and how to rebuild it on supermemory in about 60 lines. EngineeringSep 20, 2026 ](https://supermemory.ai/blog/reverse-engineering-instinct-memory)
4. [An update to supermemory We've discontinued the supermemory company brain and Nova. Everyone who was charged has been refunded, our MCP and plugins continue to run, and we're going all in on the memory engine. NewsSep 10, 2026 ](https://supermemory.ai/blog/an-update-to-supermemory)
5. [Scaling Conversations: How Adapta Grew Usage Without Losing Context Adapta added Supermemory as a persistent memory layer so every conversation keeps its context — letting the team scale usage without losing the thread. Case StudyJun 12, 2026 ](https://supermemory.ai/blog/adapta-scaling-conversations)
6. [How Chatarmin Ditched RAG and Went Memory-Only with Supermemory Chatarmin replaced a heavy RAG pipeline with Supermemory's memory layer — cutting average AI response time from 40s to 12s and token usage by 40–50%. Case StudyJun 10, 2026 ](https://supermemory.ai/blog/chatarmin-memory-only)
7. [SMFS: making agentic retrieval 55% cheaper AND more accurate We launched SMFS.ai (Supermemory Filesystem) a few weeks ago, with a simple bet: We can redesign the filesystem specifically for agents, with special files, structures, and commands that it can use for it's tasks. Today, SMFS is used by hundreds of companies to power their agents. EngineeringMay 28, 2026 ](https://supermemory.ai/blog/smfs-making-agentic-retrieval-55-cheaper-and-more-accurate)
8. [Introducing Dynamic Dreaming: supermemory now connects the dots, for you. Dreaming is magical. TLDR: We're launching Dynamic Dreaming in supermemory today, which automatically works if you're using supermemory in any way - API, OpenClaw, Hermes agent, etc. EngineeringMay 25, 2026 ](https://supermemory.ai/blog/introducing-dynamic-dreaming-supermemory-now-connects-the-dots-for-you)
9. [Dear reader, we just made supermemory insanely cheap... the Context Cloud When I first started building supermemory, I had one goal: To build the best memory system for AI. I would talk to customers, and find out that memory was not the only thing they needed - They were all setting up 7-8 different vendors at the same time. EngineeringMay 18, 2026 ](https://supermemory.ai/blog/dear-reader-we-just-made-supermemory-insanely-cheap-the-context-cloud)
10. [Introducing @supermemory/tools v2.0.0 Today we're releasing v2.0.0\. This release unifies the API across all agents sdk integrations from AI SDK to Mastra, makes conversation identity a first-class concept, and ships with memory saving on by default. EngineeringApr 27, 2026 ](https://supermemory.ai/blog/introducing-supermemory-tools-v2-0-0)
11. [Solving the Precision-Recall Tradeoff: Search Result Aggregation When you're building memory for AI, search is your foundational layer. The way search generally works is straightforward: the user defines a query, and then sets a limit (top-K) on how many search results they want returned. Usually, this is set to 10 or 20. EngineeringApr 5, 2026 ](https://supermemory.ai/blog/solving-the-precision-recall-tradeoff-search-result-aggregation)
12. [OpenClaw Memory Problems: Why It Forgets and How to Fix It (2026) TLDR: Today, we are releasing a new version of our openclaw plugin - https://github.com/supermemoryai/openclaw-supermemory. This post is going to be a bit technical, so bear with me (or bookmark for later!) In this post, I will talk about what we do about OpenClaw memory, and how we fix it. EngineeringFeb 19, 2026 ](https://supermemory.ai/blog/why-everyone-is-complaining-about-openclaws-memory-it-sucks-and-why-supermemory-fixes-it)
13. [Stateful Coding Agents with Memory: Build Long-Running Agents (2026) We built a plugin for Claude Code and OpenCode that gives your coding agent persistent memory. It remembers your preferences, learns your codebase, and never loses context mid-conversation. The result is an agent you can run for months without starting over. EngineeringFeb 18, 2026 ](https://supermemory.ai/blog/infinitely-running-stateful-coding-agents)
14. [Clawd / Molt bot's memory SUCKS. We gave it supermemory. I'm the founder of supermemory. Clawd/Molt bot is blowing up right now, with many, many use cases. I set it up, too, and have been using it through telegram. TLDR: just go to https://supermemory.ai/docs/integrations/clawdbot to set up supermemory for your clawd bot. EngineeringJan 28, 2026 ](https://supermemory.ai/blog/clawd-molt-bots-memory-sucks-we-gave-it-supermemory)
15. [Catch up with our UNFORGETTABLE Launch Week Over the last year, one belief has guided almost everything we’ve built at Supermemory AI becomes meaningfully useful only when it remembers. Memory shouldn’t be something developers rebuild from scratch. It shouldn’t be fragile, expensive, or trapped inside a single tool. NewsJan 4, 2026 ](https://supermemory.ai/blog/catch-up-with-our-unforgettable-launch-week)
16. [Empowering the Next Generation of Founders: Supermemory Startup Program If there’s one thing we’ve learned while building Supermemory, it’s that most startups don’t fail because they didn't build features; they fail when infrastructure slows them down, or they built too slow. NewsDec 31, 2025 ](https://supermemory.ai/blog/empowering-the-next-generation-of-founders-supermemory-startup-program)
17. [Building code-chunk: AST Aware Code Chunking At Supermemory, we're building context engineering infrastructure for AI. A huge part of that is dealing with code: ingesting repos, understanding structure, and making it searchable. The problem is that most code chunking solutions are terrible. We built code-chunk to fix this. EngineeringDec 29, 2025 ](https://supermemory.ai/blog/building-code-chunk-ast-aware-code-chunking)
18. [Supermemory raises $3 million with the best memory engine for LLMs Today, I am excited to announce our first funding round to accelerate our mission of building an interoperable, scalable and reliable memory for LLMs and agents. Memory is one of the hardest challenges in AI right now. NewsOct 6, 2025 ](https://supermemory.ai/blog/supermemory-raises-3-million-and-building-the-best-memory-engine-for-llms)
19. [Mem0 vs Supermemory: Why Scira Switched Scira AI moved its production memory layer from Mem0 to Supermemory. This is what failed, what improved, and how the team evaluated the two systems. Case StudyOct 2, 2025 ](https://supermemory.ai/blog/why-scira-ai-switched)
20. [Never Record Again: How Montra Uses Supermemory to Rethink Video Creation Campbell Baron, the founder of Montra, has been making videos since he was twelve. By thirteen, he was already doing brand work. Today, he’s betting on a very different future for creators: a world where recording is the exception, and most videos are generated from scratch. Case StudyAug 21, 2025 ](https://supermemory.ai/blog/never-record-again-how-montra-uses-supermemory-to-rethink-video-creation)
21. [Unified Memory That Works Where You Work: Your Second Brain With Supermemory Hi everyone, I’m Dhravya, the founder of Supermemory. I want to start with a little story behind why this product means so much to me. You can also skip straight to what it is and how it works below. EngineeringJul 25, 2025 ](https://supermemory.ai/blog/unified-memory-that-works-where-you-work-your-second-brain-with-supermemory)
22. [Supermemory just got faster on PlanetScale What is Supermemory? Supermemory completes the missing part of the LLM puzzle: memory. Just as memory is crucial for human intelligence, it's essential for truly intelligent AI systems. EngineeringJul 18, 2025 ](https://supermemory.ai/blog/supermemory-just-got-faster-on-planetscale)
23. [Faster, smarter, reliable infinite chat: Supermemory IS context engineering. People are obsessed with prompts and prompt engineering. Sure, what you say is important, but what the model knows when you say it is the difference between a stateless text generator and an intelligent AI system. In short, context is the most crucial component. NewsJul 9, 2025 ](https://supermemory.ai/blog/faster-smarter-reliable-infinite-chat-supermemory-is-context-engineering)
24. [We solved AI API interoperability One API to rule them all, One spec to find them, One library to bring them all and in the TypeScript, bind them. When we were building the Infinite Chat API, initially, we only supported the OpenAI format. This was fine, until a lot of our customers started asking for more. EngineeringJul 7, 2025 ](https://supermemory.ai/blog/we-solved-ai-api-interoperability)
25. [The Wow Factor of Memory - How Flow Used Supermemory To Build Smarter, Stickier Products Overview: Flow is a note-taking app built around a bold vision: to create a more personal, context-aware writing experience powered by AI. At the heart of this mission is memory. Case StudyJun 14, 2025 ](https://supermemory.ai/blog/the-wow-factor-of-memory-how-flow-used-supermemory-to-build-smarter-stickier-products)
26. [The UX and technicalities of awesome MCPs Last month, we launched the Supermemory MCP, mostly to test our own infrastructure and get some initial traction. It blew up. To my absolute surprise, the initial launch itself got half a million impressions (!!!). Then, we launched and got #2 on ProductHunt too. EngineeringJun 8, 2025 ](https://supermemory.ai/blog/the-ux-and-technicalities-of-awesome-mcps)
27. [Architecting a memory engine inspired by the human brain Language is at the heart of intelligence, but what truly powers meaningful interaction is memory — the ability to accumulate, recall, and contextualize information over time. Large Language Models (LLMs) have mastered language, but memory remains their Achilles’ heel. EngineeringJun 5, 2025 ](https://supermemory.ai/blog/memory-engine)

## Start building with supermemory.

Memory and continual learning for any model, any harness. Available through our API, plugins, and MCP.

[Build with supermemory ](https://console.supermemory.ai/)
