Coding harnesses
Every agent starts with the company's memory.
Connect Claude Code, Codex, Hermes, Cursor, OpenCode or any harness that supports MCP. Before an agent writes code, it asks the brain for team conventions, past decisions, who owns what, and what was tried and failed.
Works with your coding harness.
PlannedSupport comes in waves. Each harness stays marked planned until its integration test passes.
First wave planned
- Codex
- Hermes
Next planned
- Goose
- Aider
- Any MCP client
Planned integrations. Not affiliated with or endorsed by these companies; names and logos are trademarks of their owners.
Connect over MCP, or use the CLI.
PlannedAny MCP client works, including Claude Desktop. A Claude Code plugin will bundle the MCP server with skills and slash commands.
Built in Rust, and fast. The CLI is a single static binary that starts instantly. The same core runs the hosted service on Cloudflare's edge.
Context in, decisions out.
Agents → brain opt-in
The prompts your team runs and the decisions agents make along the way ("we switched to Postgres because…", "this module is fragile") become wiki pages that link to the session they came from. Secrets are redacted before anything is stored.
Brain → agents
Before an agent writes code, it asks for context: conventions, past decisions, owners, and what was tried and failed. Every agent on the team starts with the company's memory instead of a blank slate.
Every agent's logs, in one place.
Planned · Sample dataThe livingbrain CLI collects the session logs your coding harnesses already write: Claude Code, Codex, Cursor, OpenCode and Colonizer colonies. It redacts secrets on your machine and sends them to the brain only if you opt in. Search every session, see tokens and cost per person, repo and model, and spot where agents keep failing. Decisions found in the logs become wiki pages with a link back to the session.
| Agent | Sessions | Tokens | Cost | Ended in a merged PR |
|---|---|---|---|---|
| Claude Code | 214 | 38.2M | $61 | 64% |
| Codex | 97 | 14.9M | $22 | 58% |
| Colonizer colonies | 31 | 9.4M | $14 | 71% |
| Cursor | 122 | 6.1M | $9 | 49% |
Team prompt library
Sample dataThe team's best prompts are saved, versioned and suggested to everyone. The brain notices which ones work.
Write a migration the Kestrel way
Explain a stack trace using recent incidents
Draft a customer reply from this thread
It learns from every session.
The learning layer reasons over each person's prompts and agent sessions, not just chat. It learns how each developer works, which prompts lead to merged PRs and which lead to rework, and where agents keep getting stuck. Then it improves the prompt library, fills gaps in the wiki, and briefs each person's agent the way they like.
- W1Prefers small PRs, under 300 lines
- W2Writes tests first
- W3Gets stuck on auth middleware → auth notes now in her agent's brief
- W4Wants answers as diffs, not prose
Prompts ranked by merged PRs
Sample data- Review a PR against our conventions38
- Write a migration the Kestrel way31
- Fix a flaky test using incident history22
- Refactor this file · often reworked6
Fewer tokens. Less rework. Faster agents.
Planned · benchmark comingMost agents start every task cold: they reread threads, grep the repo and rediscover what the team already decided. Living Brain does that work once, at night, and hands every agent the result.
Fewer tokens
An agent asks for a short brief with citations instead of reading thirty threads and grepping the whole repo. The context is written once and reused by every agent on the team.
Less rework
Decisions and dead ends are on record, so agents stop retrying approaches the team has already ruled out.
Cheaper models do more
The hard reading happens when the wiki is written. Day to day, a small model with the right page can do work that a large model with no context gets wrong. Yes/no decisions, like whether to reply or what to remember, go to a fast judge model (Jev) instead of a large LLM.
Fast
One Rust binary, a local cache and Cloudflare's edge. Search answers from your machine before the network does.
Without the brain
Illustration- Read the chat threads about billing
- Grep the repo for every place that touches invoices
- Ask a teammate who owns it
- Try the approach the team rejected last month
With the brain
Illustration- brain_context_for("billing", task)
- One page: the owner, the conventions, the decision, the dead end, each with its source
- Start coding
The benchmark, in the open
We'll measure tokens, cost, time, success and rework, with and without the brain, against a plain retrieval baseline, across Claude Code, Codex and Colonizer colonies, on large and small models. The task set, the harness and the raw results will be public, and anyone can rerun them. Until then there are no numbers on this page. Follow the benchmark
Works with Colonizer
From chat to pull request.
When a job is bigger than an answer, the brain launches a colony on Colonizer, a sister Factory Zero venture: an isolated microVM with a coding agent inside that comes back with a pull request.
-
01 · ASK
Maya Okafor #routes
@livingbrain fix #142
#142 · ETA off by an hour after DST
-
02 · BRIEF
- repo: kestrel/routes
- owner: Maya Okafor
- tz logic: src/eta/tz.rs
- tried: offset patch, reverted
-
03 · COLONY c-7f3
microVM
$ opening pull request #219
- 04 · PULL REQUEST #219 Fix ETA drift across DST +42 −9 · 213 tests pass Posted back in the chat thread
-
05 · LEARNINGS
- +DST handling notes
- +tz.rs owner: Maya
- +Dead end: offset patch
Boots already briefed
Each colony starts with the brain's context for that repo and task.
Reports back to the wiki
When it finishes, its decisions, dead ends and the PR flow back into the brain.
Followed from the thread
Start a colony from Slack or Discord and watch it work live in the same thread.