AI's next big thing: personalization and (super)memory.
A perspective on useful AI personalization, with the boundaries between retrieval, changing facts, profiles and durable memory made explicit.

You are probably thinking of AI memory in the wrong way.
Over the last few years, we've all seen a lot of absolutely world-changing trends in AI. Things that totally changed the way we interact with computers today. The first one was data (models start getting smarter), then it was inference (everyone is able to run them), and vector databases/RAG. Now there are agents everywhere (with claude code). So, what's next?
I believe that the next big inflection point is memory. And by memory, I mean truly magical personalization. A feeling that the user gets, when their assistant will proactively bring up things that they would never have expected. When context lengths are no longer a concern, and people can just have long running, thoughtful conversations with their assistant.
I believe that this true personalization can be brought in some beautifully simple ways. But first, let's talk about what the industry is getting wrong.
Retrieval alone does not define a memory lifecycle
RAG supplies retrieved information to a model before it answers. A basic vector-search example can retrieve isolated passages without resolving which fact is current. That is a limitation of that implementation, not of all RAG systems. RAG can use changing sources, timestamps, version rules and conversation history.
A memory system adds a deliberate policy for what to retain, how it changes and what should inform the next interaction.
Memory evolves, updates, and derives/learns new information.
Let's take a scenario.
Day 1: "I love Adidas sneakers"
Day 30: "My Adidas broke after a month, terrible quality"
Day 31: "I'm switching to Puma"
Day 45: "What sneakers should I buy?"
For illustration, a similarity-only search might return the old preference even after the user changes their mind.
# A naive retrieval baseline sees isolated records
query = "What sneakers should I buy?"
# Semantic search finds closest match
result = vector_search(query)
# Possible stale result: "I love Adidas sneakers"
# Agent recommends Adidas 🤦
The assistant should use the updated preference for current questions and keep the earlier preference available when the user asks about the past.
# Desired behavior for a memory-aware implementation
query = "What sneakers should I buy?"
# Memory retrieval considers:
# 1. Temporal validity (Adidas preference is outdated)
# 2. Causal relationships (broke → disappointment → switch)
# 3. Current state (now prefers Puma)
# Agent correctly recommends Puma ✅
The same principle applies to relevance over time. An old exam worry need not dominate today's answer. That does not require pretending the event never occurred or deleting every historical record. Correction, retrieval eligibility and permanent deletion are separate operations. Supermemory's graph-memory documentation describes updates, extensions and inferences.
Agentic discovery has a workload-dependent cost
An agent can inspect files and search sources on demand. That can be useful when a question needs detailed evidence. It can also add model calls and latency compared with a direct lookup of a known preference.
The cost depends on how much the agent has to inspect. Measure model calls, search tools and retrieved context together. A direct preference lookup and an investigation across a large repository need different latency budgets.
Loading unnecessary history can increase input cost and distract the model, but it does not make hallucination inevitable. Likewise, compaction can retain useful preferences when its summary preserves them. Test what survives repeated compaction and use persistent records when the workflow needs independently inspectable context.
How we approach memory at Supermemory
Memory needs to be useful without slowing down the conversation or loading every past message. How do we build that? At supermemory, we have been thinking a lot about the best way to approach this memory problem, and build it like the human brain. supermemory.ai/blog/memory-engine
- A vector-graph architecture to track knowledge change.

Supermemory's graph links memories through update, extension and inference relationships. A container supplies scope for related context; it is not restricted to one person or proof that every extracted statement is a fact. See the documented graph model and container scopes.
Updates: Information Changes
Memory 1: "Alex works at Google as a software engineer"
Memory 2: "Alex just started at Stripe as a PM"
↓
Memory 2 UPDATES Memory 1
Derives: Sleep-time compute for an agent
Memory 1: "Alex is a PM at Stripe"
Memory 2: "Alex frequently discusses payment APIs and fraud detection"
↓
Derived: "Alex likely works on Stripe's core payments product"
The derived statement is an inference, not a confirmed job assignment. It should remain distinguishable from the source facts. Updates can make an earlier value no longer current; that is different from permanent deletion. Explicit forgetting also needs its documented retrieval and retention behavior checked.
The important lifecycle question is whether a record remains eligible for the query being asked. Do not interpret a changed preference or an expired retrieval record as proof that every copy has been erased.
2. Memory is not only retrieval!!! The magic of User Profiles
Query-based retrieval can miss useful background when a message is generic. A greeting or an indirect question may not match the wording of an earlier conversation. Maintained profile context is one way to give the agent relevant background alongside retrieval; it does not guarantee that every inference is correct.

In this illustration, role context gives the assistant a starting point for a better follow-up. It should still ask about the user's budget and preferences. Remembering someone should make a conversation more relevant, not replace listening to what they say now.
We built something called user profiles, which provides maintained user context that an application can request and supply to its agent. A profile is not automatically included in every API response, and the application should select appropriate context for the task.
Think of it as a RAM layer, with both "static" (things the agent should know by DEFAULT) - like the name of the user, their age, etc, and "dynamic" - the episodic / currently-ongoing-endeavours of the user.
ILLUSTRATIVE STATIC CONTEXT (relatively stable; still needs correction)
- Name: Dhravya
- Location: San Francisco
- Role: Founder & CEO of Supermemory (AI memory + context platform)
- Age range: 18–25
- Interests: AI infrastructure, developer tools, vector search, PostgreSQL
ILLUSTRATIVE DYNAMIC CONTEXT (recent context, not a live personal record)
- Currently working on: a "Customer Context Graph" for consolidating customer data.
- Actively optimizing: inference cost for Claude and other LLMs.
- Exploring: migrating speech-to-speech from OpenAI to Gemini Live.
- Recent preference change: Adidas sneakers broke → switched preference to Puma.
- Mood: occasionally stressed about infra costs and scale.
This brings out a new way of personalization. Pair this with retrieval and evaluate whether it makes the assistant more useful. Personalization should not force unrelated or sensitive details into every answer.
3. Getting the best of both worlds: Hybrid retrieval
Extracted memories are useful for concise facts, but some questions need the surrounding passage. We also return document chunks so the model can use source detail that was not captured in an extracted memory.
In Supermemory, hybrid search can return both memories and document chunks. Processing is asynchronous, so memories are not guaranteed to reflect a source change immediately. Inspect freshness and source provenance instead of assuming every extracted memory is current. The search documentation explains the supported modes and options.
Paired with selected profile context, this gives the application several sources of evidence. The application still needs to test whether the selected context is sufficient, current and authorized for the question.
Bringing memory into the conversation
So, that's it. That's how I think LLMs can get the perfect memory that they deserve.
A way to create beautiful user experiences, and make your agents feel unforgettable.
We are building this memory engine at supermemory.ai - and would love for you to try it out, and add it to your own agent!
It's not just about retrieval. It's about true personalization. And it's coming to every single AI agent out there.
It's the next inflection point of AI.
Frequently asked questions
Can RAG support persistent memory?
Yes. RAG can retrieve changing documents and conversation records. A memory application additionally needs policies for identity, retained facts, corrections and deletion.
Are memories always current?
No. Processing delays, stale sources and application caches can affect freshness. Verify the source revision and readiness needed for the question.
Does a changed fact mean its history was deleted?
No. A new fact can supersede an old value while history remains available. Explicit forgetting and permanent deletion have separate behavior.