Blog·Learning

Memory and Retrieval for Large-Repository Coding Agents

Diagnose stale code, missing dependencies and lost decisions. Combine current source inspection with scoped memory and test the resulting changes.

By Shardul Mane·5 min read

Memory bottlenecks in large repositories

A coding agent working across a large repository needs current source code, relevant dependency information and the decisions that explain the design. Those are related but different inputs. A remembered architectural note cannot prove that a function still exists at the current commit.

When an agent produces a wrong change, locate the missing or stale evidence before blaming the context-window size. It may have retrieved an old revision, missed a caller, confused two services or treated an abandoned proposal as a current decision.

Separate source truth from remembered context

Input Useful for Required check
Current files and symbols Implementations, types and call sites Repository and commit match the task
Build and dependency metadata Package versions and integration boundaries Installed or deployed version is known
Decision records and issue history Why the team chose a design Decision is still current
Persistent notes or memory Relevant context from earlier work Source, scope and freshness are inspectable

Use memory to help locate and interpret evidence. Verify a proposed code change against the checkout and its tests. An explanation that sounds familiar is not a substitute for inspecting the relevant implementation.

Native persistence already exists in coding tools

Claude Code documents project instructions and auto memory. Cursor supports rules. GitHub documents Copilot Memory. Availability and configuration differ, so inspect the tool being used.

A maintained instruction file can contain a recent API deprecation or a hard-won debugging lesson. Its usefulness depends on maintenance, loading behavior and relevance, just as a retrieved memory's usefulness depends on capture and retrieval.

Before adding an external service, test whether the missing context was never saved, saved in the wrong scope or omitted from the request. Each failure calls for a different fix.

Diagnose stale retrieval at the source boundary

An index represents the source versions it has processed. If a rename or deletion has not propagated, an agent may retrieve obsolete code. Trace the affected revision through the indexing pipeline to find the synchronization delay.

Attach repository, branch or commit, path and source revision to indexed content. Define how updates and deletions are processed. When a retrieved implementation disagrees with the working tree, inspect the current file and report the mismatch instead of generating against the old signature.

Keep separate branches and releases distinct where their APIs differ. A correct description of the latest main branch can still be wrong for a service pinned to an earlier release.

Use syntax-aware chunks where they help

An abstract syntax tree can identify functions, classes and other structural boundaries. Chunking around these boundaries may make a retrieved passage easier to interpret than an arbitrary character split. It does not guarantee that every function fits the size limit or that every dependency is included.

Large functions may still need splitting. Imports, decorators, parent classes and call sites may require separate retrieval or added context. Parsing a file does not automatically build a complete cross-service dependency graph, especially with dynamic calls or generated code.

Compare chunking strategies on repository-specific questions. Check whether the result contains the full relevant logic and enough surrounding definitions.

Use relationship tools for relationship questions

A question about callers, imports or service contracts may need symbol search, a language server, static analysis, manifests, schemas or explicit dependency records. Semantic search can help locate candidates but does not guarantee an exhaustive call graph.

For example, changing a producer's event schema requires checking consumers and their deployed versions. Searching for the event name is a useful first step, but an absent text match does not prove no consumer exists. Generated clients and dynamic routing can hide the relationship.

An agent can combine those tools with retained design context. Keep the source of each dependency visible so a remembered note can be checked against current code and configuration.

Compare context strategies on repository tasks

The Lost in the Middle paper reports position-sensitive performance in the models and tasks it evaluated. Include relevant evidence at different positions when testing the models and tasks in your coding workflow.

A larger window can help when it contains necessary evidence. It can also add irrelevant material and cost. Compare a selected-context baseline with a larger-context path on the same tasks instead of declaring either universally superior.

When a failure occurs, inspect whether the evidence was absent, truncated, stale or simply used incorrectly. Retrieval and reasoning are separate stages.

Add external memory with explicit scope

For a team, define which notes belong to a person, repository or shared project. Resolve access before reading or writing. A tag supplied with a privileged API key is a retrieval parameter, not proof that the caller is authorized to use that scope.

Supermemory documents container tags and relationships between memories. Use them for scoped project context, alongside the symbol and dependency tools needed to inspect current code.

Store a decision with its source and relevant repository revision. Retrieve it for a related task, then verify that the current code still follows it. If a decision changes, retain enough history to explain the transition without presenting the old choice as current.

Evaluate the coding outcome

Build a small fixture containing a renamed symbol, an obsolete API, a cross-package caller, a changed event contract and a decision superseded in a later issue. Require the agent to identify the correct source revision and produce a change that passes the relevant checks.

Measure unsupported dependency claims, stale references, task completion, context tokens and elapsed time. Test whether memory improves those outcomes over the tool's native context features. A retrieval hit is useful evidence, but the final patch still needs review and validation.

Try Supermemory for project context when the baseline exposes a real continuity gap. Start with one source-backed decision across two tasks and keep current source inspection in the workflow.

Frequently asked questions

What context can coding tools retain natively?

Coding tools can retain rules, project instructions and memory. Check the current version and configuration to see which context is saved and loaded for a new task.

When does AST-aware chunking help?

It can preserve function and class boundaries so retrieved code is easier to interpret. Evaluate callers, imports and oversized functions separately to check whether the needed evidence is complete.

Is a memory graph automatically a complete code dependency graph?

No. Relationships between remembered facts are different from verified symbol, call and deployment dependencies.

  1. We're open sourcing the company brain. Here's how we designed the multiplayer harnessCompany Brain is now open source. A walkthrough of the multiplayer harness behind its Slack experience, from proactivity and memory boundaries to approvals and recovery.
  2. Jev changes a lot in memory & context engineering. Here's exactly how.We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short.
  3. I reverse-engineered Instinct's memory. Here's exactly how it worksInstinct keeps its memory as git-tracked markdown files, found with grep rather than vectors. Here is the whole system as far as black-box probing can reconstruct it, and how to rebuild it on supermemory in about 60 lines.
  4. An update to supermemoryWe've discontinued the supermemory company brain and Nova. Everyone who was charged has been refunded, our MCP and plugins continue to run, and we're going all in on the memory engine.
  5. Scaling Conversations: How Adapta Grew Usage Without Losing ContextAdapta added Supermemory as a persistent memory layer so every conversation keeps its context — letting the team scale usage without losing the thread.
  6. How Chatarmin Ditched RAG and Went Memory-Only with SupermemoryChatarmin replaced a heavy RAG pipeline with Supermemory's memory layer — cutting average AI response time from 40s to 12s and token usage by 40–50%.
  7. SMFS: making agentic retrieval 55% cheaper AND more accurateWe launched SMFS.ai (Supermemory Filesystem) a few weeks ago, with a simple bet: We can redesign the filesystem specifically for agents, with special files, structures, and commands that it can use for it's tasks. Today, SMFS is used by hundreds of companies to power their agents.
  8. Introducing Dynamic Dreaming: supermemory now connects the dots, for you.Dreaming is magical. TLDR: We're launching Dynamic Dreaming in supermemory today, which automatically works if you're using supermemory in any way - API, OpenClaw, Hermes agent, etc.
  9. Dear reader, we just made supermemory insanely cheap... the Context CloudWhen I first started building supermemory, I had one goal: To build the best memory system for AI. I would talk to customers, and find out that memory was not the only thing they needed - They were all setting up 7-8 different vendors at the same time.
  10. Introducing @supermemory/tools v2.0.0Today we're releasing v2.0.0. This release unifies the API across all agents sdk integrations from AI SDK to Mastra, makes conversation identity a first-class concept, and ships with memory saving on by default.
  11. Solving the Precision-Recall Tradeoff: Search Result AggregationWhen you're building memory for AI, search is your foundational layer. The way search generally works is straightforward: the user defines a query, and then sets a limit (top-K) on how many search results they want returned. Usually, this is set to 10 or 20.
  12. OpenClaw Memory Problems: Why It Forgets and How to Fix It (2026)TLDR: Today, we are releasing a new version of our openclaw plugin - https://github.com/supermemoryai/openclaw-supermemory. This post is going to be a bit technical, so bear with me (or bookmark for later!) In this post, I will talk about what we do about OpenClaw memory, and how we fix it.
  13. Stateful Coding Agents with Memory: Build Long-Running Agents (2026)We built a plugin for Claude Code and OpenCode that gives your coding agent persistent memory. It remembers your preferences, learns your codebase, and never loses context mid-conversation. The result is an agent you can run for months without starting over.
  14. Clawd / Molt bot's memory SUCKS. We gave it supermemory.I'm the founder of supermemory. Clawd/Molt bot is blowing up right now, with many, many use cases. I set it up, too, and have been using it through telegram. TLDR: just go to https://supermemory.ai/docs/integrations/clawdbot to set up supermemory for your clawd bot.
  15. Catch up with our UNFORGETTABLE Launch WeekOver the last year, one belief has guided almost everything we’ve built at Supermemory AI becomes meaningfully useful only when it remembers. Memory shouldn’t be something developers rebuild from scratch. It shouldn’t be fragile, expensive, or trapped inside a single tool.
  16. Empowering the Next Generation of Founders: Supermemory Startup ProgramIf there’s one thing we’ve learned while building Supermemory, it’s that most startups don’t fail because they didn't build features; they fail when infrastructure slows them down, or they built too slow.
  17. Building code-chunk: AST Aware Code ChunkingAt Supermemory, we're building context engineering infrastructure for AI. A huge part of that is dealing with code: ingesting repos, understanding structure, and making it searchable. The problem is that most code chunking solutions are terrible. We built code-chunk to fix this.
  18. Supermemory raises $3 million with the best memory engine for LLMsToday, I am excited to announce our first funding round to accelerate our mission of building an interoperable, scalable and reliable memory for LLMs and agents. Memory is one of the hardest challenges in AI right now.
  19. Mem0 vs Supermemory: Why Scira SwitchedScira AI moved its production memory layer from Mem0 to Supermemory. This is what failed, what improved, and how the team evaluated the two systems.
  20. Never Record Again: How Montra Uses Supermemory to Rethink Video CreationCampbell Baron, the founder of Montra, has been making videos since he was twelve. By thirteen, he was already doing brand work. Today, he’s betting on a very different future for creators: a world where recording is the exception, and most videos are generated from scratch.
  21. Unified Memory That Works Where You Work: Your Second Brain With SupermemoryHi everyone, I’m Dhravya, the founder of Supermemory. I want to start with a little story behind why this product means so much to me. You can also skip straight to what it is and how it works below.
  22. Supermemory just got faster on PlanetScaleWhat is Supermemory? Supermemory completes the missing part of the LLM puzzle: memory. Just as memory is crucial for human intelligence, it's essential for truly intelligent AI systems.
  23. Faster, smarter, reliable infinite chat: Supermemory IS context engineering.People are obsessed with prompts and prompt engineering. Sure, what you say is important, but what the model knows when you say it is the difference between a stateless text generator and an intelligent AI system. In short, context is the most crucial component.
  24. We solved AI API interoperabilityOne API to rule them all, One spec to find them, One library to bring them all and in the TypeScript, bind them. When we were building the Infinite Chat API, initially, we only supported the OpenAI format. This was fine, until a lot of our customers started asking for more.
  25. The Wow Factor of Memory - How Flow Used Supermemory To Build Smarter, Stickier ProductsOverview: Flow is a note-taking app built around a bold vision: to create a more personal, context-aware writing experience powered by AI. At the heart of this mission is memory.
  26. The UX and technicalities of awesome MCPsLast month, we launched the Supermemory MCP, mostly to test our own infrastructure and get some initial traction. It blew up. To my absolute surprise, the initial launch itself got half a million impressions (!!!). Then, we launched and got #2 on ProductHunt too.
  27. Architecting a memory engine inspired by the human brainLanguage is at the heart of intelligence, but what truly powers meaningful interaction is memory — the ability to accumulate, recall, and contextualize information over time. Large Language Models (LLMs) have mastered language, but memory remains their Achilles’ heel.
YOUR PRIVACY

Change or withdraw any time via Cookie settings in the footer. Read our cookie notice.