Jev changes a lot in memory & context engineering. Here's exactly how.
We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short.

Everyone is talking about how Jev (a new system 1 model by TypeSafe AI that's really good at fast decisions) is changing computer use, browser use, labelling and many other things. It's exciting times, and it feels like a completely different paradigm of models.
I've been working on agent memory for the last 3 years, and I'm the founder of supermemory. Things like this completely change the dynamics of what we do at supermemory and how memory works in general (there's no right answer yet :)). So, we decided to explore how Jev can change, optimize, or improve our systems, or help us rethink how we do memory, and it's actually a lot of things throughout the pipeline!! Gonna be a fun and a long read — so buckle up!
How people do memory right now
Before we get into Jev-specific info, it helps to understand the general pipeline / methods people do memory, and then work backwards.
We looked at all popular memory systems right now (ChatGPT, Claude, Instinct, Openclaw, Hermes, Muse, etc.) and have been deconstructing / learning about them for a long time now. If you haven't yet, you should read my blog about Instinct.
All 'memory' pipelines that we've observed at scale, and even our customers, have a few things in common. It doesn't really matter if it's markdown-based, graph-based, fact-based etc., you can deduce it down to a few common points.
Raw data (messages, files, tool output) ────────────→ Current context
│ │
└→ Chunking / batching [fit into one model call] │
│ │
Observation / learning [off-loop] │
[background job, schedule, or trigger] │
│ │
Stored context │
[markdown files, vector DB, graph, KV] │
│ │ │
Summaries / profiles Search / reads ←── query ─────┤
[one-pager, profile] [grep, cosine, on demand] │
│ │ │
└────────────────────┴──────────────────────┤
Harness injection [hooks, tools, system prompt, recaps] ───────┤
↓
Agent's answer
- Finding relevant info / retrieval. It's not always retrieval as in RAG, but there almost always is a search step of sorts, which could be grep or cosine similarity.
- Chunking the data in order to do the observation (you can't fit the whole thing in another model, so you would batch / chunk it). This also applies to retrieval.
- Observation / learning happens outside of the 'main loop' (the main agent is not always deciding to learn things, it's at some schedule or trigger).
- Harness-specific logic to bring this context back into the model.
When I do these explorations, I try to forget about supermemory's specific architecture and think from first principles. And well, we found that —
Jev actually applies to all of the above
Thinking about it, we figured that Jev really does help in most of these — even if it's not practical in production in all of them yet. It decreases cost, increases accuracy, or makes things faster throughout the pipeline. It may or may not be the absolute best at everything, but this is more of an exploratory post, so we'll talk about them anyways. And I'll especially include the things that Jev is not good at.
The way to think about Jev is that it's really good at taking decisions really fast, and can output structured choices and probabilities. Not text generation.
Now, let's walk through them :)
Jev can make reranking better
We compared Jev as a reranker against other rerankers on public BEIR sets. We ran it through SciFact, NFCorpus, TREC-COVID, FiQA, SCIDOCS and some more benchmarks. Here's what we found.
-
Jev always helped the first-stage list. On SciFact, NFCorpus, TREC-COVID, FiQA, SCIDOCS it beat base BM25 search every time — which was expected (+0.05 to +0.17 nDCG@10).
We also tried different ways of using Jev, and found that Noul-as-a-delete-gate failed (it kept nothing, for some reason). Noul-as-a-sort and Score-10 both worked (almost the same); Score-10 was best among Jev-only methods (SciFact 0.751).
-
It is not the quality winner #1, and it is not cheap vs BGE. But the quality is much better than average. Versus bge-reranker-base, for example: Jev wins quality on the 3 overlapping sets (mean 0.612 vs 0.564) and costs ~$0.00037/q vs ~$0.00002/q (~20×). Versus jina-reranker-turbo: Jev loses the mean (0.612 vs 0.633). Versus monoT5 / RankGPT-4 on SciFact: 0.751 vs 0.766 / 0.756.
And it's certainly better in quality than using any standard model for reranking.
-
Cost sits next to Voyage, not next to BGE. Jev $0.042/MTok ≈ Voyage rerank-3 $0.05/MTok. Cohere is $0.002/q (5× Jev). RankGPT-4 is ~$0.04/q. The turbopuffer blog (Jev vs Voyage/Luna) is a different bakeoff; it never included BGE.
We always maintain the SOTA cost-quality ratio reranking at supermemory, so if you're a customer you don't really have to care about this, and we will keep running evals and remain frontier!
Turbopuffer also did some work benchmarking these, but they used GPT-5.6 sol Golden set to compare against, and we used public benchmarks.

My TL;DR here is that reranker models are pretty good! Jev does score really well here, but it's a bit more expensive and I'd stick to a reranker for now. But in the decision models world, I can see it become SOTA.
Umm wait, one more thought I'll put out there: could reranking models be used as we use Jev, if you reframe the question as a ranking? They are pretty good and pretty cheap! IDK, that's a discussion for another day.
Perfectly accurate chunking, even across languages and messy data
Right now, the popular ways of chunking are still a bit iffy. It's either too expensive (embedding-based), or too deterministic (markdown heading-based), or too vibes-based (sliding window chunking, fixed-length chunking). In all of these cases what we were seeing is that the chunks always had some issues.

With Jev, we can ask whether each sentence continues the previous thought or starts a new one, then use those answers to choose chunk boundaries near a target size. We built an internal benchmark with messy data, markdown data, and also multilingual data.

Jev is pretty clearly the SOTA at chunking.
Multilingual chunking is especially hard because the semantic boundaries are not very well defined. In many languages we put English words in the middle, or the other way around. Jev was able to perform pretty good even in these multilingual use cases :)

Shoutout to Chroma's chunking research for this way of thinking / benchmarking chunks — we took this as a reference to figure out "What is good chunking".
The caveat, again, is cost. Jev is much more expensive than other chunking methods (mainly because they are... essentially free). But Jev is the best chunker.
cost per 1k docs
rule-based $0.008 ██
embedding semantic $0.011 ██▌
jev continuation $0.080 ██████████████████▍
jev boundary $0.087 ████████████████████
Cleaning up context before observation
Right now, for pretty much all memory cases, memory generation can get a bit expensive because of two things:
- A model has to look through pretty much all context, so inference is run twice.
- In order to learn / register something, you need to know what's already learnt. You can't just remember a name every time you see it, you need to know if you already know the name, a different name, etc.
Most of the conversation or document is not even important for memory. It is headings, audio checks, "no action items," and talk that can stay searchable as source text.
The idea was to put Jev in front of that call. Split the document into sentences, ask a yes/no (Noul) on each: should this go to the extractor at all? Remove everything else. Basically a cheap filter before things go into the observer model.
On our internal benchmark, Jev was able to save us 58% of content tokens!! This means that we could literally spend half as much for theoretically just benefits!
The catch: however, obviously, this cut is not always free. Jev was good at removing a lot of this context, but removing some sentences independently seems to be the wrong direction, because of contextuality. One example of this is assistant turns — just trimming parts of it can damage the actual semantic meaning, which could be bad.

At the start of a legal document, or a conversation, a certain line/topic/word may not be useful, but it could be referenced sometime later on. We could not reliably figure out how to give Jev this full context because it's classifying sentences.
Even doing it in different ways had various problems of the same nature, so, as of right now, Jev is not a very good compactor for memory!
Because we have a small specialized model doing the learnings (learner-1), observation is insanely cheap anyways, so all roads lead to using supermemory for memory :) If you're building an agent you should use it.
Having decisions in the harness made by Jev
One of the product lines we work on at supermemory is plugins for all popular agents. The most popular and my favorite one is Claude Code — check it out. It automatically injects tokens to Claude Code to drive better outputs and improve personalization without damaging the context (average tokens injected is 250 tokens, and you'll always know when it does).

We tried using Jev to take decisions at a Claude Code UserPromptSubmit hook. Maybe it can figure out whether memory is needed or not for this query, but it is still a hook deciding the injection.
I have spoken a lot about my opinion on hooks vs tools for giving the model memory and context. But the TL;DR of it is that I believe models are very bad at deciding when some memory should be helpful, but are getting better at longer context.
We try to provide a magical experience with supermemory — especially in plugins where we can build them specific to the harness. You should try our Claude Code plugin!
So, we don't want the model to decide when to bring in memories. Having Jev decide feels really good! Some people mention that they sometimes don't want the model to use memories for certain questions. Because Jev is just taking the decision to search or not, the user can literally just say "without using memory, tell me..." And Jev would decide to not use the memory (without a tool call, in the UserPromptSubmit hook).
So, Jev is the perfect solution for making ad-hoc decisions on the harness, improving user love & satisfaction.
Closing it out
Jev is an incredible new model. There's a lot of hype and most of it is justified! It can completely change a lot of the parts of how people do memory. At the same time, it's not perfect for everything yet!
We have already started replacing a lot of our pipeline with decision models, and will continue doing work and research on them and writing about them! We also found ways to make supermemory much cheaper through Jev... more about that soon.
I'm biased, but if you're doing any context work (memory, retrieval, filesystems, even markdown), you should probably use supermemory!
Originally published by Dhravya Shah on X.