Blog·Engineering

Jev changes a lot in memory & context engineering. Here's exactly how.

We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short.

By Dhravya Shah·9 min read

Glowing golden tracks branch through a blue landscape beside the headline Jev changes memory and context engineering.

Everyone is talking about how Jev (a new system 1 model by TypeSafe AI that's really good at fast decisions) is changing computer use, browser use, labelling and many other things. It's exciting times, and it feels like a completely different paradigm of models.

I've been working on agent memory for the last 3 years, and I'm the founder of supermemory. Things like this completely change the dynamics of what we do at supermemory and how memory works in general (there's no right answer yet :)). So, we decided to explore how Jev can change, optimize, or improve our systems, or help us rethink how we do memory, and it's actually a lot of things throughout the pipeline!! Gonna be a fun and a long read — so buckle up!

How people do memory right now

Before we get into Jev-specific info, it helps to understand the general pipeline / methods people do memory, and then work backwards.

We looked at all popular memory systems right now (ChatGPT, Claude, Instinct, Openclaw, Hermes, Muse, etc.) and have been deconstructing / learning about them for a long time now. If you haven't yet, you should read my blog about Instinct.

All 'memory' pipelines that we've observed at scale, and even our customers, have a few things in common. It doesn't really matter if it's markdown-based, graph-based, fact-based etc., you can deduce it down to a few common points.

Raw data (messages, files, tool output) ────────────→ Current context
         │                                                    │
         └→ Chunking / batching [fit into one model call]      │
                       │                                      │
           Observation / learning [off-loop]                  │
           [background job, schedule, or trigger]             │
                       │                                      │
           Stored context                                     │
           [markdown files, vector DB, graph, KV]             │
                  │                    │                      │
        Summaries / profiles     Search / reads ←── query ─────┤
        [one-pager, profile]     [grep, cosine, on demand]      │
                  │                    │                      │
                  └────────────────────┴──────────────────────┤
 Harness injection [hooks, tools, system prompt, recaps] ───────┤
                                                              ↓
                                                        Agent's answer
  • Finding relevant info / retrieval. It's not always retrieval as in RAG, but there almost always is a search step of sorts, which could be grep or cosine similarity.
  • Chunking the data in order to do the observation (you can't fit the whole thing in another model, so you would batch / chunk it). This also applies to retrieval.
  • Observation / learning happens outside of the 'main loop' (the main agent is not always deciding to learn things, it's at some schedule or trigger).
  • Harness-specific logic to bring this context back into the model.

When I do these explorations, I try to forget about supermemory's specific architecture and think from first principles. And well, we found that —

Jev actually applies to all of the above

Thinking about it, we figured that Jev really does help in most of these — even if it's not practical in production in all of them yet. It decreases cost, increases accuracy, or makes things faster throughout the pipeline. It may or may not be the absolute best at everything, but this is more of an exploratory post, so we'll talk about them anyways. And I'll especially include the things that Jev is not good at.

The way to think about Jev is that it's really good at taking decisions really fast, and can output structured choices and probabilities. Not text generation.

Now, let's walk through them :)

Jev can make reranking better

We compared Jev as a reranker against other rerankers on public BEIR sets. We ran it through SciFact, NFCorpus, TREC-COVID, FiQA, SCIDOCS and some more benchmarks. Here's what we found.

  1. Jev always helped the first-stage list. On SciFact, NFCorpus, TREC-COVID, FiQA, SCIDOCS it beat base BM25 search every time — which was expected (+0.05 to +0.17 nDCG@10).

    We also tried different ways of using Jev, and found that Noul-as-a-delete-gate failed (it kept nothing, for some reason). Noul-as-a-sort and Score-10 both worked (almost the same); Score-10 was best among Jev-only methods (SciFact 0.751).

  2. It is not the quality winner #1, and it is not cheap vs BGE. But the quality is much better than average. Versus bge-reranker-base, for example: Jev wins quality on the 3 overlapping sets (mean 0.612 vs 0.564) and costs ~$0.00037/q vs ~$0.00002/q (~20×). Versus jina-reranker-turbo: Jev loses the mean (0.612 vs 0.633). Versus monoT5 / RankGPT-4 on SciFact: 0.751 vs 0.766 / 0.756.

    And it's certainly better in quality than using any standard model for reranking.

  3. Cost sits next to Voyage, not next to BGE. Jev $0.042/MTok ≈ Voyage rerank-3 $0.05/MTok. Cohere is $0.002/q (5× Jev). RankGPT-4 is ~$0.04/q. The turbopuffer blog (Jev vs Voyage/Luna) is a different bakeoff; it never included BGE.

We always maintain the SOTA cost-quality ratio reranking at supermemory, so if you're a customer you don't really have to care about this, and we will keep running evals and remain frontier!

Turbopuffer also did some work benchmarking these, but they used GPT-5.6 sol Golden set to compare against, and we used public benchmarks.

SciFact reranker quality versus price: Jev Score-10 improves on BM25 and BGE base, while monoT5 and RankGPT-4 score higher; Jev costs more than the BGE variants.

My TL;DR here is that reranker models are pretty good! Jev does score really well here, but it's a bit more expensive and I'd stick to a reranker for now. But in the decision models world, I can see it become SOTA.

Umm wait, one more thought I'll put out there: could reranking models be used as we use Jev, if you reframe the question as a ranking? They are pretty good and pretty cheap! IDK, that's a discussion for another day.

Perfectly accurate chunking, even across languages and messy data

Right now, the popular ways of chunking are still a bit iffy. It's either too expensive (embedding-based), or too deterministic (markdown heading-based), or too vibes-based (sliding window chunking, fixed-length chunking). In all of these cases what we were seeing is that the chunks always had some issues.

Eight chunking methods: fixed windows, overlap, recursive separators, sentence packing, Markdown headings, semantic embeddings, and Jev continuation and boundary decisions.

With Jev, we can ask whether each sentence continues the previous thought or starts a new one, then use those answers to choose chunk boundaries near a target size. We built an internal benchmark with messy data, markdown data, and also multilingual data.

Top-chunk answer hit rate across 72 questions: both Jev methods reach 93% with 320-character chunks, compared with 89% for recursive separators and 81% for embedding semantic chunking.

Jev is pretty clearly the SOTA at chunking.

Multilingual chunking is especially hard because the semantic boundaries are not very well defined. In many languages we put English words in the middle, or the other way around. Jev was able to perform pretty good even in these multilingual use cases :)

Answer retrieval within 640 characters across six languages: Jev boundary scores 100% for Spanish, Arabic, Hindi, Japanese, and Chinese, and 67% for English.

Shoutout to Chroma's chunking research for this way of thinking / benchmarking chunks — we took this as a reference to figure out "What is good chunking".

The caveat, again, is cost. Jev is much more expensive than other chunking methods (mainly because they are... essentially free). But Jev is the best chunker.

cost per 1k docs
 rule-based          $0.008  ██
 embedding semantic  $0.011  ██▌
 jev continuation    $0.080  ██████████████████▍
 jev boundary        $0.087  ████████████████████

Cleaning up context before observation

Right now, for pretty much all memory cases, memory generation can get a bit expensive because of two things:

  • A model has to look through pretty much all context, so inference is run twice.
  • In order to learn / register something, you need to know what's already learnt. You can't just remember a name every time you see it, you need to know if you already know the name, a different name, etc.

Most of the conversation or document is not even important for memory. It is headings, audio checks, "no action items," and talk that can stay searchable as source text.

The idea was to put Jev in front of that call. Split the document into sentences, ask a yes/no (Noul) on each: should this go to the extractor at all? Remove everything else. Basically a cheap filter before things go into the observer model.

On our internal benchmark, Jev was able to save us 58% of content tokens!! This means that we could literally spend half as much for theoretically just benefits!

The catch: however, obviously, this cut is not always free. Jev was good at removing a lot of this context, but removing some sentences independently seems to be the wrong direction, because of contextuality. One example of this is assistant turns — just trimming parts of it can damage the actual semantic meaning, which could be bad.

Evaluation of Jev's context filtering, illustrating token savings and the risk of losing important context.

At the start of a legal document, or a conversation, a certain line/topic/word may not be useful, but it could be referenced sometime later on. We could not reliably figure out how to give Jev this full context because it's classifying sentences.

Even doing it in different ways had various problems of the same nature, so, as of right now, Jev is not a very good compactor for memory!

Because we have a small specialized model doing the learnings (learner-1), observation is insanely cheap anyways, so all roads lead to using supermemory for memory :) If you're building an agent you should use it.

Having decisions in the harness made by Jev

One of the product lines we work on at supermemory is plugins for all popular agents. The most popular and my favorite one is Claude Code — check it out. It automatically injects tokens to Claude Code to drive better outputs and improve personalization without damaging the context (average tokens injected is 250 tokens, and you'll always know when it does).

Supermemory's Claude Code integration showing memory context injected into an agent session.

We tried using Jev to take decisions at a Claude Code UserPromptSubmit hook. Maybe it can figure out whether memory is needed or not for this query, but it is still a hook deciding the injection.

I have spoken a lot about my opinion on hooks vs tools for giving the model memory and context. But the TL;DR of it is that I believe models are very bad at deciding when some memory should be helpful, but are getting better at longer context.

We try to provide a magical experience with supermemory — especially in plugins where we can build them specific to the harness. You should try our Claude Code plugin!

So, we don't want the model to decide when to bring in memories. Having Jev decide feels really good! Some people mention that they sometimes don't want the model to use memories for certain questions. Because Jev is just taking the decision to search or not, the user can literally just say "without using memory, tell me..." And Jev would decide to not use the memory (without a tool call, in the UserPromptSubmit hook).

So, Jev is the perfect solution for making ad-hoc decisions on the harness, improving user love & satisfaction.

Closing it out

Jev is an incredible new model. There's a lot of hype and most of it is justified! It can completely change a lot of the parts of how people do memory. At the same time, it's not perfect for everything yet!

We have already started replacing a lot of our pipeline with decision models, and will continue doing work and research on them and writing about them! We also found ways to make supermemory much cheaper through Jev... more about that soon.

I'm biased, but if you're doing any context work (memory, retrieval, filesystems, even markdown), you should probably use supermemory!

Originally published by Dhravya Shah on X.

  1. We're open sourcing the company brain. Here's how we designed the multiplayer harnessCompany Brain is now open source. A walkthrough of the multiplayer harness behind its Slack experience, from proactivity and memory boundaries to approvals and recovery.
  2. I reverse-engineered Instinct's memory. Here's exactly how it worksInstinct keeps its memory as git-tracked markdown files, found with grep rather than vectors. Here is the whole system as far as black-box probing can reconstruct it, and how to rebuild it on supermemory in about 60 lines.
  3. An update to supermemoryWe've discontinued the supermemory company brain and Nova. Everyone who was charged has been refunded, our MCP and plugins continue to run, and we're going all in on the memory engine.
  4. Scaling Conversations: How Adapta Grew Usage Without Losing ContextAdapta added Supermemory as a persistent memory layer so every conversation keeps its context — letting the team scale usage without losing the thread.
  5. How Chatarmin Ditched RAG and Went Memory-Only with SupermemoryChatarmin replaced a heavy RAG pipeline with Supermemory's memory layer — cutting average AI response time from 40s to 12s and token usage by 40–50%.
  6. SMFS: making agentic retrieval 55% cheaper AND more accurateWe launched SMFS.ai (Supermemory Filesystem) a few weeks ago, with a simple bet: We can redesign the filesystem specifically for agents, with special files, structures, and commands that it can use for it's tasks. Today, SMFS is used by hundreds of companies to power their agents.
  7. Introducing Dynamic Dreaming: supermemory now connects the dots, for you.Dreaming is magical. TLDR: We're launching Dynamic Dreaming in supermemory today, which automatically works if you're using supermemory in any way - API, OpenClaw, Hermes agent, etc.
  8. Dear reader, we just made supermemory insanely cheap... the Context CloudWhen I first started building supermemory, I had one goal: To build the best memory system for AI. I would talk to customers, and find out that memory was not the only thing they needed - They were all setting up 7-8 different vendors at the same time.
  9. Introducing @supermemory/tools v2.0.0Today we're releasing v2.0.0. This release unifies the API across all agents sdk integrations from AI SDK to Mastra, makes conversation identity a first-class concept, and ships with memory saving on by default.
  10. Solving the Precision-Recall Tradeoff: Search Result AggregationWhen you're building memory for AI, search is your foundational layer. The way search generally works is straightforward: the user defines a query, and then sets a limit (top-K) on how many search results they want returned. Usually, this is set to 10 or 20.
  11. OpenClaw Memory Problems: Why It Forgets and How to Fix It (2026)TLDR: Today, we are releasing a new version of our openclaw plugin - https://github.com/supermemoryai/openclaw-supermemory. This post is going to be a bit technical, so bear with me (or bookmark for later!) In this post, I will talk about what we do about OpenClaw memory, and how we fix it.
  12. Stateful Coding Agents with Memory: Build Long-Running Agents (2026)We built a plugin for Claude Code and OpenCode that gives your coding agent persistent memory. It remembers your preferences, learns your codebase, and never loses context mid-conversation. The result is an agent you can run for months without starting over.
  13. Clawd / Molt bot's memory SUCKS. We gave it supermemory.I'm the founder of supermemory. Clawd/Molt bot is blowing up right now, with many, many use cases. I set it up, too, and have been using it through telegram. TLDR: just go to https://supermemory.ai/docs/integrations/clawdbot to set up supermemory for your clawd bot.
  14. Catch up with our UNFORGETTABLE Launch WeekOver the last year, one belief has guided almost everything we’ve built at Supermemory AI becomes meaningfully useful only when it remembers. Memory shouldn’t be something developers rebuild from scratch. It shouldn’t be fragile, expensive, or trapped inside a single tool.
  15. Empowering the Next Generation of Founders: Supermemory Startup ProgramIf there’s one thing we’ve learned while building Supermemory, it’s that most startups don’t fail because they didn't build features; they fail when infrastructure slows them down, or they built too slow.
  16. Building code-chunk: AST Aware Code ChunkingAt Supermemory, we're building context engineering infrastructure for AI. A huge part of that is dealing with code: ingesting repos, understanding structure, and making it searchable. The problem is that most code chunking solutions are terrible. We built code-chunk to fix this.
  17. Supermemory raises $3 million with the best memory engine for LLMsToday, I am excited to announce our first funding round to accelerate our mission of building an interoperable, scalable and reliable memory for LLMs and agents. Memory is one of the hardest challenges in AI right now.
  18. Mem0 vs Supermemory: Why Scira SwitchedScira AI moved its production memory layer from Mem0 to Supermemory. This is what failed, what improved, and how the team evaluated the two systems.
  19. Never Record Again: How Montra Uses Supermemory to Rethink Video CreationCampbell Baron, the founder of Montra, has been making videos since he was twelve. By thirteen, he was already doing brand work. Today, he’s betting on a very different future for creators: a world where recording is the exception, and most videos are generated from scratch.
  20. Unified Memory That Works Where You Work: Your Second Brain With SupermemoryHi everyone, I’m Dhravya, the founder of Supermemory. I want to start with a little story behind why this product means so much to me. You can also skip straight to what it is and how it works below.
  21. Supermemory just got faster on PlanetScaleWhat is Supermemory? Supermemory completes the missing part of the LLM puzzle: memory. Just as memory is crucial for human intelligence, it's essential for truly intelligent AI systems.
  22. Faster, smarter, reliable infinite chat: Supermemory IS context engineering.People are obsessed with prompts and prompt engineering. Sure, what you say is important, but what the model knows when you say it is the difference between a stateless text generator and an intelligent AI system. In short, context is the most crucial component.
  23. We solved AI API interoperabilityOne API to rule them all, One spec to find them, One library to bring them all and in the TypeScript, bind them. When we were building the Infinite Chat API, initially, we only supported the OpenAI format. This was fine, until a lot of our customers started asking for more.
  24. The Wow Factor of Memory - How Flow Used Supermemory To Build Smarter, Stickier ProductsOverview: Flow is a note-taking app built around a bold vision: to create a more personal, context-aware writing experience powered by AI. At the heart of this mission is memory.
  25. The UX and technicalities of awesome MCPsLast month, we launched the Supermemory MCP, mostly to test our own infrastructure and get some initial traction. It blew up. To my absolute surprise, the initial launch itself got half a million impressions (!!!). Then, we launched and got #2 on ProductHunt too.
  26. Architecting a memory engine inspired by the human brainLanguage is at the heart of intelligence, but what truly powers meaningful interaction is memory — the ability to accumulate, recall, and contextualize information over time. Large Language Models (LLMs) have mastered language, but memory remains their Achilles’ heel.