Blog·News

We're open sourcing the company brain. Here's how we designed the multiplayer harness

Company Brain is now open source. A walkthrough of the multiplayer harness behind its Slack experience, from proactivity and memory boundaries to approvals and recovery.

By Dhravya Shah·10 min read

Company Brain is now open source: an open workspace with three desks around a shared golden light.

It's been just two weeks since we discontinued the company brain harness. We had to offboard tons of people to other products, but many of our customers and others suggested we open source our work on the company brain. So, we have decided to open source it. You can use it here.

It's just one click to deploy on Cloudflare, and will take less than five minutes to set up. You'll own all the data.

In this blog, I wanted to talk through the architecture and some of the work we (mainly @maheshthedev, legend) did on the harness. It's honestly great engineering, built on top of Cloudflare Agents SDK, with many little cool things around permissioning, proactiveness, and more! Thanks @threepointone for Cloudflare Agents :)

As for why we shut it down, read our update. But the TL;DR is that we want to build the best memory system, and be the best in the world at it.

Unlike other blogs on X, this one is not gonna be AI slop, but kinda long because it's months of our work that I'll condense into simple architecture decisions that we made. I write these myself, so comment any feedback! This time, instead of boring you with words, this blog is full of diagrams and ASCII flows. Should be a fun read!

Designing a teammate, not a bot

When you think about a Slackbot, or an agent, only one thing comes to mind: "Send it a prompt, it will do the task, and reply with something."

We built the company brain around a different premise: instead of just getting the answer, it decides when to participate, gathers the right context, acts within the correct privacy boundaries, and can do extremely long-running tasks and sometimes just produce the most useful, smallest intervention.

The Company Brain harness is the runtime, policy layer, memory boundary, tool system, and conversation protocol surrounding the models. It turns raw model capability into the behavior of a dependable teammate: responsive when asked, selectively proactive, capable of investigation, quiet when appropriate, and recoverable when infrastructure fails.

The harness experience

A language model can classify a message, call a tool, or write a fluent answer. It does not inherently know: this is a multiplayer agent. So, there were a few things with multiplayer:

  • Proactivity with Slack events (even huddles, notes, etc. should be learned from).
  • Is the bot allowed to speak in the channel?
  • If a memory container is safe to search.
  • The right approval and access control policy. Approve before email send: yes or no? If you ask for too many approvals, that's bad too.
  • Whether a newer message has replaced the current request (steering and other stuff around it).
  • Whether a model-produced "working on it" is a final answer.
  • How to recover if the runtime disappears halfway through a tool call.
  • When silence is the correct outcome… when to shut up.
┌──────────────────────────────────────────────────────────────────────┐
│                     COMPANY BRAIN EXPERIENCE                         │
├──────────────────────────────────────────────────────────────────────┤
│ Slack UX                                                             │
│ mentions · DMs · threads · reactions · progress · approval cards       │
├──────────────────────────────────────────────────────────────────────┤
│ HARNESS                                                              │
│ routing · state · policy · memory scope · tools · concurrency          │
│ recovery · budgets · side effects · telemetry                         │
├──────────────────────────────────────────────────────────────────────┤
│ MODELS                                                               │
│ triage · main reasoning · classifiers · salvage · entity resolution   │
├──────────────────────────────────────────────────────────────────────┤
│ INFRASTRUCTURE                                                       │
│ Cloudflare · Postgres · AI Gateway · providers · connected apps        │
└──────────────────────────────────────────────────────────────────────┘

Our architecture

We built completely on Cloudflare, since the team has experience with it, it's the best DX for us and our agent, and scales really, really well. We also get a lot of features by using the Agents SDK out of the box.

Slack events pass through a verifying Cloudflare Worker to one CompanyBrainAgent Durable Object per organization, backed by local SQL, Postgres, authentication, models, memory, connected apps, a sandbox, and telemetry.

CompanyBrainAgent is the main thing, a Cloudflare Durable Object keyed by organization ID. It owns:

  • Fiber-backed turn lifecycles and recovery checkpoints.
  • Durable local SQL for active turns, inboxes, approvals, bot threads, cooldowns, rollout cursors, and traces.
  • Slack API operations after decrypting the bot token.
  • Triage, proactivity resolution, and action evaluation.
  • The main model/tool loop.
  • Approval suspension and resumption.
  • Billing and telemetry scheduling.

OK, so we have this. Now, let's walk through what happens when an event comes in from Slack. An event may be:

  • Explicit: an @mention, DM, assistant thread, or name wake.
  • Passive: a top-level message in a channel where proactive participation is possible.
  • Context-only: conversation that should be retained but not answered at the HTTP routing layer.
  • Ignored: bot/self traffic, unsupported subtypes, or ineligible events.

Based on these events, we classify them:

Slack message
    │
    ├─ bot or self message? ──────────────────────────► DROP
    ├─ unsupported subtype? ─────────────────────────► DROP
    ├─ clearly addressed to another person? ─────────► CONTEXT / DROP
    ├─ membership or entitlement failure? ───────────► NOTICE or DROP
    ├─ emoji-only passive message? ──────────────────► DROP
    ├─ quiet channel and not explicit? ──────────────► DROP
    ├─ proactive cooldown unavailable? ──────────────► DROP
    └─ active turn already running? ─────────────────► STEERING GATE
                                                         │
                                                         ▼
                                                    TRIAGE ELIGIBLE

Talking only when needed

One of the things that we really cared about is the impactfulness of the bot. A lot of the agents we tried out in public were proactive, and they honestly should not be. It gets annoying quite fast! A proactive agent must answer two questions:

  1. Could the agent do something useful?
  2. Should it interrupt this conversation now?

We initially tried to make the proactivity based on a yes/no decision, but a lot of these things are not binary. It slowly switched to a confidence model based on various different thresholds that a small model decides.

                     MESSAGE + CONTEXT
                            │
                            ▼
              ┌──────────────────────────┐
              │ Triage model             │
              │ estimates the situation  │
              │                          │
              │ usefulness               │
              │ confidence               │
              │ urgency                  │
              │ noise                    │
              │ interruption cost        │
              │ investigation value      │
              │ reaction fit             │
              └────────────┬─────────────┘
                           │ 0–100 scores
                           ▼
              ┌──────────────────────────┐
              │ Deterministic evaluator  │
              │ formulas + mode policy   │
              │ explicit-message rules   │
              └────────────┬─────────────┘
                           │
          ┌────────────────┼────────────────┬────────────────┐
          ▼                ▼                ▼                ▼
       ANSWER         INVESTIGATE      ACKNOWLEDGE         PASS
    full turn          tool check        reaction         silence

This lets the evaluator understand cases such as:

  • Useful but low-confidence → investigate instead of answer.
  • Socially appropriate but informationally empty → react.
  • Urgent but disruptive → investigate under a stricter mode.
  • Correct and useful but already answered → pass.
  • Explicit and lightweight → acknowledge rather than launch a full turn.

The output to this is an action + reaction (like a Slack emoji):

Primary action: ANSWER | INVESTIGATE | PASS
Reaction:       allowlisted emoji | none

BTW, for emojis we made this cool thing: emoji-resolve. So that the agent can just use whatever emoji it needs to, the resolver resolves it. This also means that the bot can use custom emojis that the organization has, once it learns that they exist.

That allows four natural behaviors:

  1. Reply without reacting.
  2. React without replying.
  3. React and reply.
  4. Do nothing.

Configuration of proactivity

Despite this, sometimes you just don't want the agent to be proactive. We built a way for users to control this behavior per channel.

Proactivity settings offer All channels or Only its own channel, with exceptions for individual Slack channels.

Proactivity modes: Listener responds only when asked; Investigator enters high-value situations; Balanced selectively answers, checks, and reacts; Teammate participates more often; Custom uses organization-defined thresholds.

An explicit @ tag is pretty much always set to answer.

Explicit message
      │
      ▼
Triage model
      │
      ├─ ACK ───────────────► allowlisted reaction
      ├─ ANSWER ────────────► full turn
      ├─ INVESTIGATE ───────► full turn with investigation framing
      ├─ PASS ──────────────► coerced to ANSWER
      └─ error / bad parse ─► fallback ANSWER

Doing the answering

When triage selects ANSWER or INVESTIGATE, we gotta start working on the actual answer. And here's where the actual harness comes in: a phased state machine, with awareness of time, etc., so that it can try to do the work in a certain way within expectations of the human.

┌────────────────────────────────────────────────────────────────────┐
│                         computeTurn                                │
├────────────────────────────────────────────────────────────────────┤
│ A. TOOL ASSEMBLY                                                   │
│    assemble tool set · determine active tools                      │
│                            │                                       │
│ B. CONTEXT LOADING                                                 │
│    workspace prompt · company context · memory profile             │
│    thread history · actor · directory · attachments                │
│                            │                                       │
│ C. MODEL GENERATION                                                │
│    runModelLoop · streamed text · tool calls · live updates         │
│                            │                                       │
│ D. PROGRESS CHECK                                                  │
│    final answer or merely “working on it”?                          │
│                            │                                       │
│ E. SALVAGE                                                         │
│    one tool-free pass if only progress text exists                 │
│                            │                                       │
│ F. APPROVAL HANDSHAKE                                              │
│    consequential action? → checkpoint and suspend                  │
│                            │                                       │
│ G. FINALIZATION                                                    │
│    inbox empty? claim current revision; otherwise continue         │
│                            │                                       │
│ H. COMPLETION                                                      │
│    publish · write memory · charge · telemetry · reflect           │
└────────────────────────────────────────────────────────────────────┘

The model will continue to update the humans on the work: "I found xyz. Next step, I'll work on abc."

This is just so that the human doesn't think the agent is stuck or something. We used AI SDK's streaming, and had a few pacing nudges that we inject using the onBeforeStep hooks.

runModelLoop
    │
    ├─ streamText({
    │    model: main profile,
    │    system: policies + workspace context,
    │    messages: conversation + turn state,
    │    tools: assembled ToolSet,
    │    activeTools: currently visible names,
    │    stopWhen: step budget reached,
    │    abortSignal: turn control
    │  })
    │
    └─ EACH STEP
         │
         ├─ onBeforeStep
         │    ├─ drain live thread inbox
         │    └─ send pacing/progress nudge if needed
         │
         ├─ prepareStep
         │    ├─ render latest turn state
         │    └─ hide tools during forced wrap-up
         │
         ├─ model output
         │    ├─ text
         │    ├─ tool calls
         │    └─ text + tool calls
         │
         ├─ execute tools and append results
         ├─ record generated step
         ├─ at limit − 3: warn about remaining budget
         └─ at limit − 1: remove tools and force final text

The step budget is an active control mechanism. Near the end, the harness tells the model how many steps remain. At the final boundary, it removes tool access so the model must synthesize an answer from what it already has. This prevents a common failure mode where an agent spends its final step launching one more search and never actually answers.

Tool assembly, progressive disclosure, and codemode

The model can access several classes of tools, but not all schemas need to be visible at once.

                         assembleTurnTools
                                │
     ┌──────────────┬───────────┼──────────────┬──────────────┐
     ▼              ▼           ▼              ▼              ▼
┌──────────┐  ┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐
│ Memory   │  │ Discovery │ │ Live apps │ │ Slack I/O │ │ Effects   │
│ search   │  │ people    │ │ Gmail     │ │ channel   │ │ save      │
│ inspect  │  │ entities  │ │ GitHub    │ │ search    │ │ connect   │
│ resolve  │  │ org tree  │ │ Linear    │ │ progress  │ │ forget    │
└──────────┘  └───────────┘ │ Notion    │ └───────────┘ └───────────┘
                            └───────────┘
                                  │
                                  ▼
                         ┌─────────────────┐
                         │ Lazy families   │
                         │ sandbox         │
                         │ scheduler       │
                         │ hidden until    │
                         │ enable_tools    │
                         └─────────────────┘

We used codemode and an activeTools selector to do that. activeTools controls which names are available at each step. Sandbox and scheduler schemas are hidden until the model explicitly unlocks those families. This provides three benefits:

  1. Lower context overhead: fewer schemas consume prompt space.
  2. Better tool selection: irrelevant tools do not compete for attention.
  3. Policy control: some capabilities only appear after an intentional transition.

The harness can also remove all tools during final wrap-up without rebuilding the entire model context.

Memory in a multiplayer setup

This was a bit tricky. Typically, a lot of the company brain startups are just ingesting everything into the vector store. We think that approach increases stale info and confuses the model a lot. We wanted to make sure there's a clear boundary between durable org knowledge and live operational state.

                 ┌────────────────────────┐
                 │       Agent turn       │
                 └───────────┬────────────┘
                             │
             ┌───────────────┴────────────────┐
             ▼                                ▼
┌────────────────────────┐       ┌────────────────────────┐
│ Company Brain memory   │       │ Connected applications │
│ decisions              │       │ current issue status   │
│ ownership              │       │ latest email           │
│ project context        │       │ live deployment        │
│ historical knowledge   │       │ recent Slack messages  │
└────────────────────────┘       └────────────────────────┘
        durable / semantic                live / transactional

A turn may start with memory to understand what "the migration" refers to, then query Linear or GitHub to verify the current status. Treating these sources as interchangeable would either make memory too volatile or live tools too context-poor.

As for the memory hierarchy, instead of everything being user-specific memory, we decided to take the channel-based approach.

Employee memory in a DM can read the employee’s private-channel memories and public-channel memory; private-channel memory can also read public-channel memory.

Inside Supermemory, the containerTags were arranged somewhat like this:

Public channel              sm_org_shared
Direct message              user_{userId}
Private channel or group DM slack_channel_{channelId}

Supermemory was all we needed to do memory, of course!

You need to know stuff before answering

One more interesting thing about our setup was that instead of just deciding whether or not to reply, the agent could decide to INVESTIGATE: investigate, and then decide whether you wanna talk.

INVESTIGATE
    │
    ▼
Gather evidence
memory · Slack · traces · connected apps
    │
    ▼
Is there something actionable or useful?
    │
    ├─ yes ─► synthesize and reply
    │
    └─ no ──► silentConclusion
              no Slack message
              release proactive slot

This is crucial for ambient intelligence. If every background check had to result in prose, the agent would pollute channels with messages like "I checked and found nothing." A silent conclusion lets the system optimize for useful outcomes rather than visible activity.

Approval flow for consequential tools

Model requests consequential tool
              │
              ▼
      Approval classifier
              │
      ┌───────┴────────┐
      ▼                ▼
 no approval       approval required
      │                │
 execute          store pending_approval
                       │
                       ▼
                checkpoint messages,
                tool state, and revision
                       │
                       ▼
                post Approve / Deny card
                       │
                 ┌─────┴─────┐
                 ▼           ▼
               Deny       Approve
                 │           │
              finish     resumeTurnAfterApproval
                             │
                             ▼
                      same model loop,
                      preserved context

Approval is not delegated to a separate agent. The current turn is suspended with a serializable resume state. If approved, the same reasoning process continues with the decision recorded. That preserves causality: the agent that proposed the action is the agent that sees the approval and completes the work.

Multiplayer and concurrency

Slack is not a request-response form. Threads continue while the agent works. A teammate can add evidence, correct a premise, replace the request, or ask the agent to stop. New messages entering an active thread go through an active-turn gate rather than normal triage.

When a new turn comes in, the small model decides whether to replace and supersede the original request, stop it, ignore the new message, or append and add context to the existing chain of thought.

New message arrives in active thread
                 │
                 ▼
       Active-turn gate model
                 │
    ┌────────────┼─────────────┬─────────────┐
    ▼            ▼             ▼             ▼
 IGNORE        APPEND        REPLACE       STOP
 unrelated     add context   supersede     cancel
                 │             request       │
                 ▼                          ▼
         durable inbox                abort signal
                 │
                 ▼
        drain before next step
                 │
                 ▼
    <live_thread_updates> injected

But that's too much complexity in a harness!

I know most of you would be thinking that's too much complexity in the harness, but it was needed. Here's why.

  1. Unlike Pi, OpenCode, Codex, etc. harnesses, we can't just YOLO and put everything in there. We needed determinism, because this would be an agent that would work in larger organizations. Permissioning, tool use, and a bunch of things are important. Also, we didn't want it to just be a question-answer thing; that's easy and boring. Making it truly multiplayer was a challenge.
  2. Product-specific things. We wanted to make sure that none of our users are scared of anything the agent could do that would be bad for the company. That's why we made sure that memory tool use, access control, and approval flow are something we have control over and can adjust based on our customers' needs.
  3. We were optimizing for different things. We weren't just concerned about response quality. For us, the larger objective was value provided as an employee.
              usefulness × correctness × timeliness
Agent value = ─────────────────────────────────────
                 noise × risk × interruption cost

A technically correct answer can still be a bad intervention if it arrives in the wrong channel, exposes private context, duplicates what a human already said, interrupts a sensitive discussion, or ignores a newer message. It gives the system a way to ask not only "What should I say?" but also:

  • Should I say anything?
  • Should I check first?
  • Is a reaction enough?
  • Which information boundary applies?
  • Has the request changed?
  • Do I need permission?
  • Can I finish reliably?

That is the difference between placing a model inside Slack and building an AI teammate!

Bored you enough. Now, try it!

Everything I mentioned is open source, so you can try it, change it, hack it, give us feedback, whatever works :)

Get Company Brain on GitHub.

I also tried to make the onboarding super simple, so it should literally just be one-click deploy with the Deploy to Cloudflare button. Thanks for reading, follow for more!


Originally published by Dhravya Shah on X.

  1. Jev changes a lot in memory & context engineering. Here's exactly how.We tested Jev across reranking, chunking, observation, and harness decisions. Here is where fast decision models help memory systems, where they cost more, and where they still fall short.
  2. I reverse-engineered Instinct's memory. Here's exactly how it worksInstinct keeps its memory as git-tracked markdown files, found with grep rather than vectors. Here is the whole system as far as black-box probing can reconstruct it, and how to rebuild it on supermemory in about 60 lines.
  3. An update to supermemoryWe've discontinued the supermemory company brain and Nova. Everyone who was charged has been refunded, our MCP and plugins continue to run, and we're going all in on the memory engine.
  4. Scaling Conversations: How Adapta Grew Usage Without Losing ContextAdapta added Supermemory as a persistent memory layer so every conversation keeps its context — letting the team scale usage without losing the thread.
  5. How Chatarmin Ditched RAG and Went Memory-Only with SupermemoryChatarmin replaced a heavy RAG pipeline with Supermemory's memory layer — cutting average AI response time from 40s to 12s and token usage by 40–50%.
  6. SMFS: making agentic retrieval 55% cheaper AND more accurateWe launched SMFS.ai (Supermemory Filesystem) a few weeks ago, with a simple bet: We can redesign the filesystem specifically for agents, with special files, structures, and commands that it can use for it's tasks. Today, SMFS is used by hundreds of companies to power their agents.
  7. Introducing Dynamic Dreaming: supermemory now connects the dots, for you.Dreaming is magical. TLDR: We're launching Dynamic Dreaming in supermemory today, which automatically works if you're using supermemory in any way - API, OpenClaw, Hermes agent, etc.
  8. Dear reader, we just made supermemory insanely cheap... the Context CloudWhen I first started building supermemory, I had one goal: To build the best memory system for AI. I would talk to customers, and find out that memory was not the only thing they needed - They were all setting up 7-8 different vendors at the same time.
  9. Introducing @supermemory/tools v2.0.0Today we're releasing v2.0.0. This release unifies the API across all agents sdk integrations from AI SDK to Mastra, makes conversation identity a first-class concept, and ships with memory saving on by default.
  10. Solving the Precision-Recall Tradeoff: Search Result AggregationWhen you're building memory for AI, search is your foundational layer. The way search generally works is straightforward: the user defines a query, and then sets a limit (top-K) on how many search results they want returned. Usually, this is set to 10 or 20.
  11. OpenClaw Memory Problems: Why It Forgets and How to Fix It (2026)TLDR: Today, we are releasing a new version of our openclaw plugin - https://github.com/supermemoryai/openclaw-supermemory. This post is going to be a bit technical, so bear with me (or bookmark for later!) In this post, I will talk about what we do about OpenClaw memory, and how we fix it.
  12. Stateful Coding Agents with Memory: Build Long-Running Agents (2026)We built a plugin for Claude Code and OpenCode that gives your coding agent persistent memory. It remembers your preferences, learns your codebase, and never loses context mid-conversation. The result is an agent you can run for months without starting over.
  13. Clawd / Molt bot's memory SUCKS. We gave it supermemory.I'm the founder of supermemory. Clawd/Molt bot is blowing up right now, with many, many use cases. I set it up, too, and have been using it through telegram. TLDR: just go to https://supermemory.ai/docs/integrations/clawdbot to set up supermemory for your clawd bot.
  14. Catch up with our UNFORGETTABLE Launch WeekOver the last year, one belief has guided almost everything we’ve built at Supermemory AI becomes meaningfully useful only when it remembers. Memory shouldn’t be something developers rebuild from scratch. It shouldn’t be fragile, expensive, or trapped inside a single tool.
  15. Empowering the Next Generation of Founders: Supermemory Startup ProgramIf there’s one thing we’ve learned while building Supermemory, it’s that most startups don’t fail because they didn't build features; they fail when infrastructure slows them down, or they built too slow.
  16. Building code-chunk: AST Aware Code ChunkingAt Supermemory, we're building context engineering infrastructure for AI. A huge part of that is dealing with code: ingesting repos, understanding structure, and making it searchable. The problem is that most code chunking solutions are terrible. We built code-chunk to fix this.
  17. Supermemory raises $3 million with the best memory engine for LLMsToday, I am excited to announce our first funding round to accelerate our mission of building an interoperable, scalable and reliable memory for LLMs and agents. Memory is one of the hardest challenges in AI right now.
  18. Mem0 vs Supermemory: Why Scira SwitchedScira AI moved its production memory layer from Mem0 to Supermemory. This is what failed, what improved, and how the team evaluated the two systems.
  19. Never Record Again: How Montra Uses Supermemory to Rethink Video CreationCampbell Baron, the founder of Montra, has been making videos since he was twelve. By thirteen, he was already doing brand work. Today, he’s betting on a very different future for creators: a world where recording is the exception, and most videos are generated from scratch.
  20. Unified Memory That Works Where You Work: Your Second Brain With SupermemoryHi everyone, I’m Dhravya, the founder of Supermemory. I want to start with a little story behind why this product means so much to me. You can also skip straight to what it is and how it works below.
  21. Supermemory just got faster on PlanetScaleWhat is Supermemory? Supermemory completes the missing part of the LLM puzzle: memory. Just as memory is crucial for human intelligence, it's essential for truly intelligent AI systems.
  22. Faster, smarter, reliable infinite chat: Supermemory IS context engineering.People are obsessed with prompts and prompt engineering. Sure, what you say is important, but what the model knows when you say it is the difference between a stateless text generator and an intelligent AI system. In short, context is the most crucial component.
  23. We solved AI API interoperabilityOne API to rule them all, One spec to find them, One library to bring them all and in the TypeScript, bind them. When we were building the Infinite Chat API, initially, we only supported the OpenAI format. This was fine, until a lot of our customers started asking for more.
  24. The Wow Factor of Memory - How Flow Used Supermemory To Build Smarter, Stickier ProductsOverview: Flow is a note-taking app built around a bold vision: to create a more personal, context-aware writing experience powered by AI. At the heart of this mission is memory.
  25. The UX and technicalities of awesome MCPsLast month, we launched the Supermemory MCP, mostly to test our own infrastructure and get some initial traction. It blew up. To my absolute surprise, the initial launch itself got half a million impressions (!!!). Then, we launched and got #2 on ProductHunt too.
  26. Architecting a memory engine inspired by the human brainLanguage is at the heart of intelligence, but what truly powers meaningful interaction is memory — the ability to accumulate, recall, and contextualize information over time. Large Language Models (LLMs) have mastered language, but memory remains their Achilles’ heel.