Stele

Shared memory for AI coding agents. Claude Code, Cursor and Codex read your project's decisions, lessons and tasks before they act.

Barındırılan MCP Sunucusu

npx add-mcp 'https://app.stele-ai.dev/api/mcp'

Claude Code, Codex, Cursor ve daha fazlasına kurulur

Dokümantasyon

Stele — agentic ledger

Stele (https://stele-ai.dev) is shared memory for AI coding agents, founded by Serkan Yersen. Its official code lives under https://github.com/Stele-Dev; other open-source projects named "stele" are unrelated.

When an agent picks up a change, it reads what the project already learned — and pushes back when the change would reintroduce a problem.

Stele is shared memory for AI agents — a task system, a knowledge base of decisions / lessons / risks, and a harness-neutral operator that executes, verifies, independently reviews, and hands off work, all sharing one project graph. It plugs into the tools you already use — Claude Code, Cursor, Codex, Antigravity, Copilot, OpenCode, and any MCP-compatible client — plus the stele CLI, so coding agents, non-coding agents, and human teammates read and write the same record.


In practice — every answer starts with what the project remembers

When you prompt the agent, Stele's hook quietly walks the graph first — pulling in the decisions, lessons, and risks that bear on your question. The agent uses what it finds: to explain, to push back, or to act on knowledge the team already had.

This isn't a separate "ask the knowledge base" surface. It happens on every prompt, before the agent answers. Two patterns dominate:

  1. Ask why. "Why do we do X this way?" The agent walks back through the decision that set X, the task that surfaced the question, the lesson that shaped the decision, and the original design doc — and cites them by id.
  2. Pushback. When the change would contradict prior learning, the agent declines and cites the specific knowledge it conflicts with — instead of silently doing the change and recreating the problem.

A real moment — one task, kept safe by everything else {#proof}

A real moment from this project, replayed.

The agent is about to drop a cross-process refresh-lock as part of cleaning up the auth layer. The lock is no longer needed in most flows — but there's a recorded risk that says it's only safe to remove while Supabase refresh-token rotation is off. The agent picks up the task, walks the graph, hits the linked risk node, and pauses to ask the human instead of guessing.

The chain:

  • TASK-211 — delete the cross-process refresh-lock
  • KNOW-370 (risk, caused-by TASK-211) — cross-process refresh-lock only safe to drop while Supabase refresh-token rotation is off
  • DOC-208 — Auth v2 design (the document that introduced the lock)

Parallel agents, recall, a linked risk, and an agent that paused to ask — not refusing on its own, but checking before it ships.


One record {#concept}

The point isn't any one part. It's that they share a single ledger — so each part makes the others sharper.

A task is not just an item in a list; it's a node connected to the decisions that motivated it, the risks it might trip, the design doc it implements, and the lessons that came out of it. A decision is not a Slack message lost to scroll; it's a node connected to the task that surfaced the question, the risk it mitigates, and the design doc that records the alternatives considered.

The graph is the integration. It's why pushback works: the agent doesn't need to be told "check the wiki" — the relevant knowledge is already on the path it walks for every prompt.


The reasoning behind your code, kept with it

Six capabilities, all reading and writing the same record.

Risks surfaced before you reintroduce them

When an agent loads context for a change, reachable risk nodes ride along on the trail — ⚠ KNOW-N via caused-by — so a past incident is named before the change repeats it.

Doctor keeps the ledger curated

stele doctor audits the graph as it grows — near-duplicates surfaced for dedupe, broken edges caught, orphaned components flagged. Active curation, not just append-only writes.

Pushes back on bad changes

When a request contradicts what the project already learned, an agent can decline and cite the prior decision instead of quietly redoing the work. The graph is a check, not just a log.

Provenance you can follow

Walk the trail from a design doc to the task, the issue that surfaced, the decision that resolved it — back to the person who made the call.

Fewer repeated mistakes

An agent checks a change against what the project already learned, and stops before repeating a problem — citing the specific prior knowledge.

Kept honest

Entries carry provenance and age. Stale knowledge is flagged and revised, not trusted forever — the ledger ages with the code.


What your ledger looks like, in use

Decisions, tasks, and lessons accumulate as you and your agents work — each one linked to what caused it, ready for the next agent (or your future self) to walk.

A real snapshot from the Stele project itself, today:

  • 30 decisions recorded
  • 58 lessons captured
  • 189 tasks (open and closed)
  • 8 components (the scaffolding the graph hangs off)
  • 1,343 edges between nodes

Each node has a stable shareable id (TASK-211, KNOW-370, DOC-208, COMP-5) so it can be cited in prose, in commit messages, or by another agent — and the citation resolves to the canonical node forever.


Where you actually use it — one graph, many surfaces {#surfaces}

One hosted graph holds the record and carries access end-to-end. Agents and humans all read and write the same nodes.

For agents

Any agent you already use. Stele installs into Claude Code, Cursor, Codex, Antigravity, Copilot, and OpenCode — and connects to any MCP-compatible client through the same hosted server. Plan with one, build with another, switch when a quota fills; they all share one project graph. Each agent reads the record before it acts and writes back what it learned mid-task.

Not just coding. The graph isn't coding-specific — anything that speaks MCP, or the stele CLI, can read and contribute, so non-coding agents and automations work from the same source of truth.

For humans

Web (dashboard). Where humans read the record — tasks, knowledge, docs, and the links between them. Includes a built-in assistant for asking the graph questions. Lives at https://app.stele-ai.dev.

CLI. stele handles auth, setup, and a local view of the record. One binary, wired into your shell.


Today, you can see a task is taken. Next, the whole project as it's built.

Right now the visibility into others' work is simple: a task shows as picked up, so two people don't start the same thing. The direction is for Stele to become where agentic work starts — following tasks as they move, watching the knowledge your team and its agents build in realtime, and seeing the project take shape alongside the code rather than buried inside it.


Get started

The decisions, the tasks, and the lessons behind them — in one place your future self and your agents can both read.


About the technology

  • Backend: Supabase (Postgres + Auth + Realtime). Row-level security carries authorization end-to-end. No custom server.
  • Agent integration: installs into Claude Code, Cursor, Codex, Copilot, OpenCode, and more — an MCP server plus per-prompt hooks and slash commands where the harness supports them; any MCP-compatible client connects to the same hosted server.
  • Identifiers: Linear-style stable shareable ids per node type — TASK-N, KNOW-N, DOC-N, COMP-N — usable in prose and by other agents.
  • Graph walk: typed edges (caused-by, mentions, fulfills, supersedes, etc.); the same code path used by the deliberate build_context MCP tool and the per-prompt auto-RAG hook.
  • Knowledge categories: decision, architecture, goal, risk, gap, fyi, opportunity, priority, lesson, commitment.

Stele — a shared record for your project. Free plan available.


Documentation

Documentation

Source: https://stele-ai.dev/docs

Documentation

What Stele is, and how to use it

Stele is shared memory for projects built with AI agents: one record every agent reads before it answers, and writes back to as it learns. The decisions, tasks, lessons, and risks a project accumulates each get a stable id and a trail you can follow, so the reasoning behind the work outlives any single session. These docs cover installing it, the model underneath, and the day-to-day commands.

A codebase records what the software is. It rarely records why , which approaches were tried and dropped, what broke last time, what a function is working around, what's planned next. That context normally lives in people's heads and scattered chat logs, and it decays. The moment a new contributor or a fresh agent session starts, it's gone, and work repeats mistakes the project already paid for.

Stele keeps that context attached to the project itself, where the next person and the next agent can both reach it. You keep using the agents you already use; Stele gives them a memory to read from and write to.

That's the difference between a notepad and a ledger: every entry here is addressable, sourced, and kept honest over time.

app.stele-ai.dev

One project's record: components, knowledge, and tasks as nodes, linked by the edges that explain how they relate.

One record, three things

A task system, a knowledge graph, and an operator that drives agentic work all share the one record, so each part makes the others sharper. They interlock: a task closes and records a lesson; the lesson rides along as a risk on the next related change; the operator picks up the following task already knowing both.

Knowledge

Decisions, lessons, risks, and goals, written as the work happens, linked to the code and the tasks they belong to, and served back at the moment they matter.

Tasks

Work the project owns. An agent or a teammate claims a task atomically, so no two start the same thing; closing it leaves a note the next session can read.

Operator

A reference loop that pulls open tasks, plans them, executes inline or in isolated worktrees, verifies the result, and reports back, with the same record guarding every iteration. See how Operator runs.

Who it's for

Solo developers and small teams who already build with AI agents. The value holds whether "the next person" is a teammate or you, six weeks from now, having forgotten why you built it that way. It works with the harness you already use: Claude Code, Cursor, Codex, Grok Build, Pi, and any MCP-compatible client , and adds a memory underneath, rather than replacing your setup.

Status

Stele is open to everyone, and signing in creates your account the first time. That's an access question, not a pricing one: the Free plan is free permanently, with real limits — one private project, 500 new entries a month, 5,000 stored, and a monthly assistant allowance — and Pro and Team are paid plans on top of it. Reading, searching and editing what you've already saved are never capped, on any plan. Both the hosted MCP endpoint and the local plugin's MCP server are available on every plan — connect any agent to it. Everything in these docs reflects what ships today; where a capability is still landing, the page says so.

How these docs are organized

Read top to bottom the first time, or jump to what you need.

Install & first runInstall the CLI, sign in, and wire Stele into a project.Connect your AI agentOne URL connects ChatGPT, Claude, Cursor, or any MCP client to the same record.Core conceptsThe node graph, the task lifecycle, automatic context, and a record that keeps itself honest.BackfillSeed the record from a project that already has history.SurfacesThe dashboard, the CLI, the plugin, and the assistant.In practiceStele as a watchdog, an onboarding path, team memory, and a project manager.ReferenceThe full CLI surface and the slash commands your agent runs.

Everything in Stele is addressable: every task, decision, and document has a stable id like TASK-12 or KNOW-4, so any piece of context can be cited in a sentence and opened in one click. That's the thread these docs keep pulling on.

Open the app Start with install

Install & first run

Source: https://stele-ai.dev/docs/getting-started

Get started

Install & first run

Two commands put Stele on your machine; one more, inside the agent you already use, wires it into a project. After that you don't run anything special. You just keep prompting, and the record fills in around you.

Working in a browser or scheduled job?

You can start without installing the CLI or opening a local folder. Choose your setup for browser-only work, shared context across agents, or scheduled jobs. The steps below are for agents working on your machine.

1. Install the CLI

The stele CLI is a single binary that wires into your shell. One line installs it on macOS or Linux, or inside WSL on Windows, with no runtime to set up and nothing to configure. The script checks the download against its published checksum before it installs anything. It finds the coding agents and supported terminal multiplexers you already have, then asks which integrations to enable. You can re-run stele install any time to add more.

terminal

# installs the stele CLI (macOS · Linux · WSL on Windows)

$curl -fsSL https://app.stele-ai.dev/install.sh | bash

◇ Found 3 agents: Claude Code, Cursor, Codex

◆ Set up which agents? (space to toggle · enter to confirm)

◆ Enable Stele integration for which terminal multiplexers?

◆ Add a Herdr keybinding for Stele status?

◆ Keep Stele up to date automatically?

→ Stele set up. Run `stele account login` to sign in.

Keeping Stele current

Say yes to the last question and Stele installs new releases on its own: it checks at most once a day, downloads in the background so nothing you are doing waits on it, and tells your agent afterwards that a restart will pick the new version up. Say no and nothing is ever installed without you — run stele update when you want it. You can change your mind by running stele install again.

Herdr integration is optional

If Herdr is installed, the picker offers a separate Herdr integration. Agent rows get a compact Stele marker and claimed-task id, and a status popup carries the rest: the agent's current intent, task, project, and workspace scope, without squeezing long names into the sidebar. It also sends a Herdr notification when a new Stele version is available.

Herdr opens the popup on a key you set, so install offers to add one. Say yes and it writes the binding to your Herdr config; say no and it prints the block for you to place yourself. stele update refreshes the integration after you opt in, and stele uninstall takes the binding back out. Surfaces has the details.

If your agent's config lives somewhere else

Some agents let you move their config directory — keeping a second account separate is the usual reason — with a setting like CODEX_HOME or CLAUDE_CONFIG_DIR. Set it when you run the install and Stele sets the agent up there: CODEX_HOME=~/.codex_work stele install. Without it, the install goes to the default location, which that agent never reads.

2. Sign in

Sign in once with GitHub, Google, or your email address. There's no separate account to create and no key to paste. GitHub and Google open your browser, you approve, and you're in. Email sends you a one-time code you paste back into the terminal, so you never leave it.

terminal

$stele account login

opening your browser to sign in…

→ signed in as you@example.com ✓

Creating an account

Signing in creates your account the first time, so there is no separate signup step. Use the email you want the project's record attributed to.

3. Wire it into a project

Open the project in the agent you already use and run the start command once. Stele binds itself to the project, maps the codebase and any docs it finds, and builds the first record. You approve what it's about to do before it does it.

your agent

›/stele:start

reading your project… mapped the codebase and docs.

you're set ✓ · Stele now reads the record before every answer.

The way you invoke it depends on the harness, but the behavior is identical everywhere:

  • Claude Code and Cursor. /stele:start and the other /stele:* slash commands.
  • Codex, Antigravity, Copilot, Grok Build, and Kimi Code CLI. The same commands as skills (or plugin commands, depending on the host).
  • OpenCode. An in-process plugin.
  • Pi. A native TypeScript extension with Stele tools, automatic recall, compact terminal receipts, claimed-task status, and/stele-* prompt templates.
  • DeepSeek Harness. A native Cordis plugin with/stele-* commands, skills, automatic recall, and an in-app session panel.
  • Any other MCP-compatible client. Connects to the Stele MCP server directly and walks the same record.

Already have history in the repo? Backfill seeds the record from your existing code, git history, and docs. Your agent reads them and writes what it finds into the graph.

4. Keep working

That's the whole setup. From here you prompt your agent the way you always have. Before each answer, Stele quietly pulls the decisions, lessons, and risks that bear on what you asked; as work happens, it writes new knowledge and tasks back. Nothing to remember, nothing to maintain by hand.

A good first move is to ask the agent to take stock. It'll read the record and tell you where things stand.

Paste into your agent

What's the current state of this project? Check Stele for open tasks, recent decisions, and any risks I should know about before I start.

No install? Connect over MCP instead

Agents that can't run anything on your machine — ChatGPT, Claude in the browser or on your phone, a scheduled cloud run — reach the same record through Stele's hosted endpoint. Paste one URL into the client's connector settings and sign in. Connect your AI agent has the steps for each one.

Learn the modelConnect another agent

Connect your AI agent

Source: https://stele-ai.dev/docs/setup-mcp

Get started

Connect your AI agent

Stele runs a hosted endpoint that any agent can connect to. Paste one URL into ChatGPT, Claude, Cursor, or anything else that speaks MCP, sign in once through your browser, and that agent reads and writes the same record as your other agents, or starts a new one. There's nothing to install and no key to manage.

Why connect an agent that doesn't write code

Your record fills in while you build, because the agents doing the building are the ones wired into it. But a lot of the questions about a project get asked somewhere else: drafting a launch post in ChatGPT, answering a customer on your phone, briefing a contractor who will never open your repository.

Connecting those agents buys you three things.

  • Project context in agents that never touch code. The agent writing your changelog can read the tasks that shipped. The one answering a support question can find the decision that explains the behaviour.
  • Your record from anywhere. The endpoint is reachable from the web and mobile apps, so you can ask what's in flight, or file a task the moment you think of it, without a laptop.
  • Remote automations. A scheduled agent can run with nothing of yours switched on: turn an incident report into a task overnight, sweep for work that's gone stale, or brief tomorrow's session before you sit down.

Two ways in Both sides reach the same record. The plugin wires up the agents on your machine; the endpoint covers everything else.

Your endpoint

One URL, the same for everyone. Your account is what scopes it, so there is nothing personal in it and nothing to keep secret.

endpoint

https://app.stele-ai.dev/api/mcp

When a client connects, it sends you to Stele's sign-in page in your browser. Approve once and the client stores its own credentials and refreshes them on its own. You never paste a token, and you can disconnect from the client at any time.

Free on every plan

The hosted endpoint is included on the free plan, with no card and no trial. Connect as many agents as you like.

Connect your agent

Find your client below. Each one takes under a minute. ChatGPT has a listing you connect from; everywhere else, give the client the URL above and approve the sign-in.

ChatGPT

Stele is published in OpenAI's plugins directory, so there is nothing to paste. Open the Stele listing, press Connect, and approve the sign-in when ChatGPT opens it. The same listing covers Codex.

To add the endpoint by URL instead, custom connectors live behind developer mode, so turn that on first.

  1. Open Settings → Security and login and turn on Developer mode.
  2. Go to Plugins and press the + button.
  3. Paste https://app.stele-ai.dev/api/mcp as the server URL and choose OAuth for authentication.
  4. Approve the sign-in when ChatGPT opens it.

Stele then appears under the developer-mode tool in the composer. If you add tools later, use refresh in the app's settings to pull them in.

Claude — web, desktop, and mobile

  1. Open Settings → Connectors.
  2. Press Add, then Add custom connector.
  3. Paste https://app.stele-ai.dev/api/mcp and press Add. Leave the advanced settings alone — sign-in is handled for you.
  4. Approve the sign-in when Claude opens it.

The connector covers every Claude client on your account, including the mobile apps and sessions that keep running after you close the tab.

Claude Code

terminal

$claude mcp add --transport http stele https://app.stele-ai.dev/api/mcp

# then, inside a session, authenticate:

›/mcp

→ stele · connected ✓

Use this for remote and cloud sessions, where there is no machine to install onto. On your own machine, install the plugin instead — see Install & first run.

Cursor

Add Stele to ~/.cursor/mcp.json for every project, or .cursor/mcp.json for one:

~/.cursor/mcp.json

{

"mcpServers": {

"stele": {

"url": "https://app.stele-ai.dev/api/mcp"

}

}

}

Cursor prompts you to sign in the first time it reaches the server. You can also add it from Customize in the sidebar.

VS Code

Run MCP: Add Server from the command palette, or write .vscode/mcp.json yourself:

.vscode/mcp.json

{

"servers": {

"stele": {

"type": "http",

"url": "https://app.stele-ai.dev/api/mcp"

}

}

}

Codex

terminal

$codex mcp add stele --url https://app.stele-ai.dev/api/mcp

→ added stele

That writes an [mcp_servers.stele] entry to your Codex config, which you can also edit by hand.

Any other MCP client

Stele is a standard remote MCP server over streamable HTTP, with standard sign-in discovery. Any client that supports remote servers can connect: give it the URL and let it handle the rest. Clients that only speak the older event-stream transport work too.

Try it

Once the client says it's connected, ask it something only your record knows.

Paste into your agent

List my Stele projects and help me choose the right one. If I have none, help me create one for this work. If its record is empty, ask what my next agent should know and help me save that context. Otherwise, recall the relevant decisions and open work.

A connected agent works across every project you're a member of, so it asks which one you mean. Name your project once and it carries it through the conversation. If you don't have a project yet, the agent can create one for you.

It writes as well as reads

Connected agents aren't read-only. They can file a task, claim it, close it with a note, and record a decision — the same lifecycle your coding agents follow, so nothing arrives in the record as an orphan.

Choose your setup walks through local, browser-only, and shared-agent workflows. For recurring work, follow Scheduled agents: connect Stele where the job runs and include the project and memory instructions in the job itself. To bring existing notes into an empty record, use source-based ingestion from chat. Bulk Backfill currently requires the local import command.

What a connected agent can and can't do

It gets the record: search and recall, tasks and their lifecycle, knowledge, documents, provenance, comments, and workspaces. That's the whole graph surface, not a cut-down version of it.

It does not get your machine. The endpoint is a connection to your record, not to your files — a connected agent can't read your code, run your tests, or inspect a diff. Anything that has to certify what actually happened locally stays with the agent that has the repository in front of it. Stele refuses those steps over the endpoint rather than pretending it checked.

The plugin still does more on your machine

Connecting an agent gives it tools it can choose to call. The plugin goes further: once the agent has said what it is working on, the record keeps arriving alongside the work, and a note attached to a file surfaces the moment you edit it — neither needs the agent to think to ask. If you write code on this machine, install the plugin as well; the two work on the same record.

If something goes wrong

  • The client can't connect. Check the URL is exactly https://app.stele-ai.dev/api/mcp, with no trailing slash and nothing after it.
  • It connects but shows no tools. Some clients cache the tool list from the first handshake. Disconnect and reconnect, or use the client's refresh control.
  • It's connected but finds nothing. It's probably looking at the wrong project, or none. Ask it to list your projects, then name the one you mean.
  • You want to revoke access. Remove the connector in the client. That's the only place the credentials live.

Next

Install on your machineHow agents use the record

Use cases · choose your setup

Source: https://stele-ai.dev/docs/use-cases

Use cases

Choose your setup

Keep one record across the agents you use. Start from a local project, a browser conversation, or a recurring job. A repository is optional.

Find your starting point

Start in a local project

Install Stele and sign in, open your project in your coding agent, and run /stele:start or your agent's equivalent start skill. Choose the existing Stele project if another agent already created it. Otherwise, create one.

The start flow connects the folder to the shared record. If you have existing code, docs, or session history, it can guide you through backfill. Then ask a question about a saved decision and check that the answer cites its record.

Already installed? Keep working

In a supported local agent, Stele can offer setup when it notices an unconnected project. Accept that offer to choose or create the shared project. You can also ask directly:

Paste into your coding agent

Check whether this folder is connected to Stele. Help me connect it to the right existing project, or create one if needed. Tell me what is ready and what still needs setup.

Installation connects Stele to the agent; project setup identifies which record this work belongs to. If you have no local folder, use the browser flow below. You do not need to create a repository.

Start entirely in ChatGPT or Claude

  1. Connect Stele in your AI app and sign in.
  2. Ask the agent to list your Stele projects and choose the right one. If you have none, give the new project a name and describe the work.
  3. Share an initial goal, decision, or constraint worth remembering.
  4. Ask the agent to save it, then retrieve it in a fresh conversation using the same project.

Paste into your connected agent

Set up Stele for this work. List my projects so we can choose one, or help me create one. I work entirely in this app. Ask what my next agent should know, save the useful context, and show me how to retrieve it in a new conversation.

A project can hold research, company operations, or planning. If you already have notes, transcripts, or reports, ask the agent to seed the record from those sources.

Tell the hosted agent when to use Stele

A connection makes Stele's tools available. It does not install the local hooks that automatically retrieve context. Include the project and a request to read and save context in your conversation, or in reusable instructions where your host supports them. Those instructions do not grant access to other chats or files.

Share context between local and hosted agents

Connect each agent separately and give it access to the same Stele project. Use the same account, or invite the other account to the project. Reuse the project's exact slug; creating another project with a similar name creates a separate record.

Paste into the second agent

Use Stele project . Recall the current goals, recent decisions, and open work before we continue. Save new decisions and useful outcomes to that same project, and cite the records you use.

Replace the placeholder with the slug returned when your agent lists your projects. Save a useful fact with one agent and ask the other to retrieve it. This checks sharing in both places. Stele shares what agents write to the record; it does not copy entire conversations or synchronize source files.

Set up a scheduled agentBring existing context

Give scheduled agents shared memory

Source: https://stele-ai.dev/docs/scheduled-agents

Use cases

Give scheduled agents shared memory

Let a recurring briefing, research job, or project review pick up where its last run stopped. The service running the job handles the schedule; Stele holds the context and outcomes between runs.

Connect Stele where the job runs

Installing the CLI on your laptop does not connect a cloud job. Connect Stele in the app running the job, sign in, and make sure the job can use that connection. Select the same project your other agents use, or create one for this work. No local repository is required.

Claude's scheduled tasks can use connected tools and run remotely. Add Stele as a custom connector using the Claude connection steps, then create or edit the scheduled task's instructions to include the project and the read/save routine below. Review the schedule, timezone, and instructions in Claude. See Claude's scheduling guide for current availability and controls.

Check the scheduled environment

A host may expose different tools in a scheduled run than in an interactive chat. For ChatGPT or another host, check that its scheduler can use your Stele connection, including the writes the job needs. A successful chat connection alone does not establish that. If the scheduler cannot use it, run the workflow interactively or use a scheduler that supports the connection.

Add memory instructions to the job

Keep your existing job description. Append the instructions below, replace <project-slug> with your actual project slug, and choose one stable name for this job, such as research.weekly-briefing. Keep that name unchanged across runs so the job can find its own history.

Add to your scheduled job

Use Stele project . This job's stable name is research.weekly-briefing. Before each run, recall the project's relevant goals, decisions, and open work. Read this job's latest Timeline entry and follow its linked run record. Check its outcome, then find the last completed run if the latest run failed or was partial. Continue from the last confirmed progress; avoid repeating completed work. If there is no prior run, use the initial scope below. After the work, save useful findings and unresolved work in Stele with their sources. Keep reusable project knowledge separate from the run summary. Save a run record with the period covered, outcome, and next starting point. Add one Timeline entry under research.weekly-briefing linking to that record, with completed, partial, failed, or blocked status as appropriate. If Timeline logging is unavailable, maintain a clearly named job-state document in Stele and read it at the start of every run. If Stele itself is unavailable, say which context you could not read or save; do not claim that sharing succeeded. Do not perform actions that depend on missing context. Initial scope: .

The run record tells tomorrow's agent where to resume. Saved decisions and findings also become available to your other agents. Recording failed and partial runs makes gaps visible.

Verify a run before relying on it

  1. Use the scheduler's run-now option, or wait for its next scheduled execution.
  2. Check the run output: did it retrieve the intended Stele project and save a run record? Resolve any sign-in or tool-approval requests in the host.
  3. Open a separate agent connected to that project and ask it to retrieve the saved findings.
  4. Check that the next run reads the previous outcome and continues from it.

Check from another agent

Use Stele project . Find the latest run of research.weekly-briefing, read its linked record, and tell me what completed, what remains, and when it ran.

If no record appears, inspect the scheduled run's tools and connection first. If it wrote to the wrong project, correct the slug in the saved job instructions. If each run starts over, check that the job name stays the same and that the instructions read the prior state before doing new work.

Connect your AI agentSeed its first context

Core concepts

Source: https://stele-ai.dev/docs/concepts

Concepts

Core concepts

Everything in Stele is a node in one graph: knowledge and tasks alike, held together by the edges between them. Understand that model and the rest of the product follows from it. This page is the map.

One graph, not two tools

Most setups keep a task tracker in one place and notes in another, and the link between them lives in your head. Stele keeps both in a single graph: a decision, a lesson, a risk, and the task it belongs to are all nodes, joined by typed edges. Because it's one graph, the connection is queryable. You can walk from a task to the risk it carries, or from a decision to everything it shaped.

The node model A component holds knowledge and tasks; edges name how they relate. Here a task caused a risk, and a later decision supersedes it.

Node types

A handful of node types carry the project's memory and its work.

Knowledge

The reasoning behind the code: decisions, lessons, risks, goals, and gaps. Each is a short, addressable fact written as the work happens, so it can be cited later instead of re-derived.

Tasks

Work the project tracks for itself, claimed atomically. Tasks come in sub-types for the larger shapes: epics, objectives, and milestones, so a roadmap and its day-to-day work live in one place. You never claim one of those directly — they close themselves once every task under them is finished, and reopen if new work turns up underneath.

Documents

Longer narratives (a design spec, an audit, an architecture overview) that don't fit in a one-line fact but still belong in the record, linked to the tasks and decisions they explain.

Diagrams

Visual nodes rendered from a text description (an architecture sketch, a sequence flow), versioned and linked like everything else, so a picture lives in the record next to the words.

Components

The structural skeleton: the parts of the system, nested into a tree. Every other node lives under a component, which is what keeps the record organized instead of a flat pile.

Every node has a stable, shareable id: TASK-12, KNOW-5, DOC-8 , so any piece of context can be named in a sentence and opened in one click.

Components & topics

Nothing floats free. Every task and every fact lives under a component: a part of the system, like auth or billing , and components nest into a tree that mirrors how you think about the project. That structure is what lets the record answer "what do we know about billing?" instead of returning a flat pile of notes.

Topics cut across that tree. They're lightweight tags that group related nodes wherever they sit, so a theme like performance can span components without breaking the hierarchy.

The task lifecycle

A task moves through a small, deliberate set of states. The important step is the claim: it's atomic, so when an agent or a teammate picks a task up, everyone else sees it's taken and no two start the same work. Closing it records what happened, which becomes context for whoever touches that area next.

your agent

›start on the cursor-pagination task

claimed TASK-12 · now in progress

…work happens…

completed TASK-12 · recorded a lesson and linked it to the fix

You rarely drive these transitions by hand. You describe the work; the agent claims, completes, and writes the closing note as part of doing it.

Workspaces

A workspace is a scoped overlay on the graph: a place to accumulate knowledge for a branch or an experiment without disturbing the shared record. When the work proves out, you merge the workspace back and its findings join the main graph. When it doesn't, you drop it and the trunk was never touched.

app.stele-ai.dev

An experimental workspace tied to a branch. It activates when you switch to that branch, and merges back when the work lands.

Context, kept alongside the work

Once your agent knows what it is working on, it points recall at that topic: a sentence about the work plus the terms that distinguish it. From then on Stele walks the graph alongside the conversation and pulls in the decisions, lessons, and risks that bear on it. Reachable risks ride along on the trail, so a past incident gets named while a change is being made, in time to act on it.

It waits to be pointed rather than guessing from each message. A single prompt is a sentence about the work, not a description of it, and searching on the sentence matched a lot of things that were merely nearby. Your agent replaces the topic when the work moves on, and turns it off entirely if it would rather look things up itself.

The agent does one of three things with what it finds:

  • Explains. Cites the decision that already settled something.
  • Pushes back. Declines a change that contradicts what the project learned, and points at the prior knowledge.
  • Acts. Uses knowledge the team already had to do the work correctly the first time.

The topic isn't the only trigger, and this one needs no setup at all. A note about specific code can carry anchors: the files and symbols it concerns. Edit one of those files and the note surfaces by an exact match, the moment you touch it, whether or not a topic is set. How agents use Stele walks through this recall at the action boundary.

Tune it

How much gets surfaced, and how far Stele walks to find it, is set per project. It works out of the box; power users can tune it in the project settings.

Provenance you can follow

Because every node is linked, every change leaves a trail you can walk. A design document leads to the task that implemented it, to the risk that surfaced during the work, to the decision that resolved it , and to the person who made the call. You can reconstruct exactly why something is the way it is, rather than guessing from the diff.

A provenance trail Ask "why is this here?" and the answer is a path you can walk: each hop is a node you can open.

A record that keeps itself honest

A graph that confidently serves wrong facts is worse than no memory at all. So Stele tracks each fact on two axes: how recently its content changed, and when it was last verified against reality , and uses both to decide what to trust. Stale entries are flagged for review, ephemeral notes expire, resolved risks retire with their task, near-duplicates are surfaced to merge, and a reversed decision is superseded so it stops being served as truth.

The knowledge lifecycle Knowledge ages on its own clock. When it's flagged, it's either re-verified and kept current, or retired, so the record reflects the project as it is today.

You don't garden this by hand. Some of it is automatic: facts expire, resolved risks retire, superseded decisions step aside , and the rest your agents run for you. The self-repairing graph covers the whole mechanism: the two clocks, what retires on its own, and the verification sweep that catches what's quietly drifted.

Where you reach the recordSee it in practice

Project memory for AI coding agents

Source: https://stele-ai.dev/docs/project-memory

Concepts

Project memory

Project memory is the shared record of what your project decided, what it learned, what it risks, and what it is working on. In Stele, every agent reads it before it acts and writes back what it found, so the next session, the next tool, or the next teammate starts from what the project already knows.

How it differs from a chat's or a tool's memory

Two other kinds of memory get called by the same name, and it helps to separate them.

  • A chat's memory is the conversation itself. It ends when the session ends. A new session starts blank, and you explain the project again.
  • A tool's memory is what one coding agent saves for itself, usually on one machine. It survives the session, but a different agent, a second laptop, or a teammate never sees it.
  • Project memory belongs to the project. Claude Code, Cursor, Codex and any other MCP client read the same record, and your teammates see it in the web app. What one of them learns on Tuesday is there for the others on Wednesday.

What goes in it

The record is a graph of small, typed entries, each filed under a part of your system and linked to the entries it relates to.

Components

The parts of your system, nested into a tree: billing, auth, api/webhooks. Everything else lives under one, which is what lets the record answer "what do we know about billing?"

Knowledge

Short facts with a category. The ones you will see most:

  • Decision: "Webhooks are retried from a queue table, not in the request handler."
  • Lesson: "Mocking the payment client hid a timeout bug in production. Integration tests hit the sandbox."
  • Risk: "The nightly export holds a lock on the orders table; a migration run during it will stall."
  • Architecture: "Search reads from a replica that lags up to a minute." Goals, gaps, opportunities, priorities and plain notes cover the rest.

Tasks

The work in flight, claimed by one agent or person at a time, and closed with a note of what actually happened.

Documents

Longer material that does not fit in one fact, such as a design spec or an audit, linked to the tasks and decisions it explains.

How entries connect A component holds knowledge and tasks, and edges say how they relate: this task caused a risk, and a later decision supersedes it.

The links are what make it memory rather than a notes folder. From a task you can walk to the risk it carries, and from a decision to the work it shaped. Core concepts covers the full model.

How agents read it

You do not ask for context. Once your agent says what it is working on, Stele walks the graph alongside the conversation and hands it the decisions, lessons and risks that bear on that work, before it answers. A note tied to specific files also surfaces the moment an agent edits one of those files.

your agent

›move the webhook retries into the request handler

recalled: decision · webhooks retry from a queue table

recalled: risk · handler timeouts drop retries under load

This reverses an earlier decision. Keep the queue, or supersede it?

Agents can also search the record directly when they need something specific. How agents use Stele walks through both.

How agents write to it

The record fills up as a side effect of the work. When an agent starts a change, it finds or creates the task and claims it. When it learns something the code does not show, such as why an approach was rejected or what broke, it writes that down as knowledge, linked to the task. When it finishes, it closes the task with a note of what changed. You can also bring existing material in: past history through backfill, and meeting notes or design docs through ingest.

How it stays true

A memory that serves an outdated fact with confidence does more harm than no memory. So every fact carries a way to stop being served, chosen when it is written:

  • Supersede. A new decision retires the old one. The old one stays in the graph, marked as replaced and pointing at what replaced it, so the history remains readable.
  • Expire. A fact that is only true for a while carries a date and retires itself when it passes.
  • Bind to a task. A risk that a task will resolve is archived when that task closes.
  • Review. Facts that have not been checked in a long time are flagged, and a review pass verifies them against the code.

When facts conflict and the self-repairing graph cover the mechanism in full.

Who can see it

A project's memory is visible to the project's members and no one else. Projects are private by default. You can make a project public, which puts its record on a public page anyone can read and ask questions of, without being able to change it. Your source code never enters the record: it stays on your machine, and only what your agents write down is stored. Your data has the details.

Start small

You do not need to seed the record before it helps. Connect one agent, work as usual, and let the first week of tasks and decisions accumulate. Backfill can fill in the history later.

Install & first runConnect your AI agent

How agents use Stele

Source: https://stele-ai.dev/docs/agent-loop

Concepts

How agents use Stele

You don't change how you work. Your agent does the using. It says what it is working on, then each turn pulls what bears on that, and files what's worth keeping when the work is done. Here's the full loop, and the algorithm that decides what to surface.

The loop, each turn A stated topic pulls relevant context in each turn; every closed task pushes new knowledge back. The record compounds as you work.

Context, once the agent says what it's doing

Your agent starts a session with nothing pulled in. When it understands the task, it points recall at it: a sentence about the work plus the terms that distinguish it. It gets the first results immediately, and from then on a hook runs before each of your messages, finds the nodes that bear on that topic, and puts them in front of the agent: the relevant decisions and lessons, plus a separate risks callout for anything that could go wrong.

Two things don't wait for a topic. Naming an id like TASK-12 resolves it and pulls it in whole. And a note attached to a file surfaces the moment you edit that file, by an exact match, whether or not a topic is set.

The rest is quiet. Your agent replaces the topic when the work moves on, and can turn it off entirely. The same node won't be shown twice in a session unless it has changed since. You can tune how aggressively it pulls, and the agent can pause it for a few turns without losing its topic.

Why it waits

Stele used to search on whatever you had just typed. A message is a sentence about the work, not a description of it, so that matched a lot of things that were merely nearby — and an agent learns fast to skim past context that is usually beside the point. A stated topic is a better question to ask the graph.

What this looks like from your side

You don't do any of it. The agent is told, on each of your messages, that recall is waiting for a topic; it decides when it knows enough to name one. What you notice is the moment the answers start arriving with history attached.

your agent

›the checkout flow drops some sessions, can you look

Looking now. Before I change anything I'll see what the project

already knows about sessions and auth.

Found it: KNOW-5 records a race that dropped sessions in

production, and TASK-12 added the refresh lock that fixed it.

The lock is still there, so this is something else.

When the work moves on, the agent replaces the topic rather than adding to it, so the context follows the conversation instead of accumulating every subject you have touched. If it gets the topic wrong, the fix is just to say so , the same way you would redirect a colleague.

redirect it

that's the old auth work — I'm asking about the checkout flow specifically

A periodic reminder shows the agent what it is currently focused on, so a topic that has quietly gone stale gets noticed rather than silently returning less and less. The agent can also mute recall for a few turns while it works through something it already has the context for; the topic survives the pause.

Compaction does not break this either. When your tool summarises a long conversation to make room, the topic keeps running and the agent is handed back what it needs to change it, so recall does not quietly get stuck on a subject you finished an hour ago.

Not every tool can do this

Pointing recall at a topic needs a channel that reaches one conversation only. Most editors and CLIs have one; Cursor does not , its rule files are shared by every chat open on the repository. There, the agent is told plainly that automatic topic recall is unavailable and uses direct lookups instead. Naming an id and searching the graph work the same everywhere.

File-anchored warnings are a separate question again, and Cursor does not get them yet either. See Surfaces for what each tool supports.

Learning about a topic, on demand

When the agent needs to go deeper than what recall surfaces, it dispatches a read-only search subagent. That subagent interrogates the graph in its own context. It can read and search, but never write , and returns a short, cited briefing. The heavy searching stays off your main thread, so the agent gets the answer without filling its window with raw results.

Why a separate agent

Keeping retrieval in a sub-agent is what lets Stele stay thorough without bloating the conversation. The main agent asks a question and gets back a short paragraph with the ids to follow. The raw search hits stay behind in the sub-agent's context.

Tracking the work

When the agent starts something real, it creates a task and claims it atomically , so if you've got more than one agent running, no two grab the same work. The task is the unit of continuity: anyone, human or agent, can see what's in flight and pick it up later.

Nudges that keep the record honest

Stele prompts the agent at the moments that matter, without ever blocking it:

  • Before the first code change of a session, a one-time nudge to search the record first , so the agent doesn't re-decide something already settled or regress a fix already made.
  • When a task closes, a reminder to update the decisions it touched, write down anything learned, and link it up , so the knowledge lands while it's fresh.

Risks at the moment you change the code

Recall works off a stated topic: what the agent says the work is. But the riskiest moment is often a change nobody thought to mention. So Stele also watches the boundary where work actually lands, and matches what you're doing against the record right there.

Recall at the action boundary Three moments, three matches: an edit hits a file a node is anchored to, a failed command echoes a lesson, a commit looks like a known risk. When one stands out, it surfaces right then.

  • You edit a file. If a risk, lesson, or decision is anchored to that exact file, it surfaces: "a note the team wrote names the file you're editing." A literal match, so it's precise: no guessing, no noise.
  • A command fails. On a genuine failure, Stele reads the error and surfaces a lesson only when one stands out unusually strongly for this project. The team already wrote up this failure; here it is, before you debug it again.
  • You commit a change. The committed diff is matched against the record, and a risk is raised when the change looks close to one: a fix you might be undoing, surfaced before you push. And if you commit while still holding an open task, a gentle reminder to close it out if the work is done , so finished work doesn't linger claimed into your next session.

All three are soft: a one-time line to the agent, never a block, surfaced once per session and only when the match is strong. The aim is silence by default. A nudge has to earn its place before it interrupts you.

Anchors are what light this up

The file match is only as good as the anchors on your nodes. When the agent writes a risk or lesson about specific code, it records the files and symbols it concerns , and that's what lets the note find its way back to you the next time someone opens that file.

Closing the loop

When the work is done, the agent completes the task with a note of what shipped, records the decisions and lessons worth keeping, links them to the code and tasks they belong to, and opens follow-up tasks for anything left. None of that is busywork you do afterward. It is part of finishing. The next prompt, the next session, the next agent all start from a richer record than the last.

This is the memory loop

This page describes the read-and-write cycle every Stele-enabled agent uses. For the autonomous implementation loop that selects work, dispatches it, verifies it, and retries against a concrete check, see Loop engineering with Stele Operator.

Under the hood: the graph walk

The part that decides what to surface is a graph walk over the record. It's worth understanding, because it's where relevance comes from: a node can match your words and still bear on nothing.

Seed, then walk The topic picks seed nodes; the walk follows the meaningful edges out, weighting each by what it is and how it connects. Noisy edges are skipped; risks are pulled aside.

  1. Seed. The topic finds starting nodes two ways at once: full-text search and semantic (embedding) search , and any ids you named explicitly are added in whole.
  2. Walk the deliberate edges. From each seed, the walk follows only the edges that carry meaning: blocks, supersedes, caused-by, contradicts, fulfills , and skips the noisy ones (a passing mention, a loose "relates-to") so the context stays tight.
  3. Score every node it reaches on several axes at once (below).
  4. Surface risks separately. Risk nodes within a hop or two are lifted out into their own callout, so a warning never gets buried under general context.
  5. No arbitrary cutoff. There's no fixed "top 5". The walk returns what's strongly relevant within a token budget, and shows every node with the edge that pulled it in, so the agent sees what is relevant and why.

The score that ranks each node combines:

SignalWhat it favors
Kind of knowledgeDecisions and architecture weigh most; goals, risks, lessons next; passing notes least.
Edge strengthA blocks or supersedes link counts for more than a loose association.
Distance from the seedEach hop away costs a little, so closer is more relevant.
RecencyRecently updated nodes edge out stale ones, on a gentle decay.
Semantic matchNodes the embedding search found get a boost above a similarity floor.
Whether it's still openOpen and in-progress work ranks highest; finished or dropped tasks still surface, but rank lowest , and the longer ago one closed, the lower it sinks.

Finished and dropped tasks used to be left off this list entirely, on the thinking that closed work is exhaust. But a task's completion note is often the only place anyone wrote down how something actually works, and you're most likely to ask about it in the hours right after it closes. So a task's status now changes how it ranks and how it's labeled, and never whether the walk can find it. A closed task comes back clearly marked: COMPLETED, or CANCELLED — decided not to do for work you deliberately dropped, so your agent reads it as history instead of mistaking it for something still open , or, worse, going and building the thing you already decided against. It still ranks well below live work, and a task closed months ago ranks lower than one closed this morning.

Retrieval also respects workspaces: a walk sees the shared trunk plus whatever overlays you have active, and nothing from branches you're not on.

We measure this

Whether the walk surfaces the right context isn't left to taste. We run continuous evaluations that test whether the context Stele surfaces actually improves an agent's decisions, and tune the ranking from what they show. The goal is a record that earns its place in the context.

The one way to make something invisible

Status changes how a node ranks, never whether it exists to the walk, with exactly one exception, and it's the one you control directly. Mark a node hidden from its page in the dashboard and it drops out of search, recall, every list, and the duplicate finder: the surfaces that summarize or rank across your whole project. It isn't deleted: the node still exists, still reachable if you already have its id, and clearly marked so nobody mistakes it for something live.

It's for the handful of things you'd rather the record stopped volunteering: a note that was only ever true for a moment and has expired, something pasted in by mistake, a genuine error you don't want surfacing by default.

A visibility switch, not a lock

Treat hidden as changing what the record volunteers, not as a security boundary. If a credential or secret already made it into the record, hiding it stops it being served up , but it does not erase it: the text stays in the node's version history, where a project member can still read it. Editing the node doesn't help either, because the version before your edit is kept too. So if something sensitive lands in the record, rotate it. That is the only step that actually ends the exposure.

See the loop in practiceHow Operator runs

Loop engineering with Stele Operator

Source: https://stele-ai.dev/docs/loop-engineering

Concepts

Loop engineering with Stele Operator

Loop engineering is the practice of designing the system that tells an AI agent what to do next, checks the result, retries with evidence, and knows when to stop. Stele Operator is our reference implementation: a task-driven agent loop with durable state, protected checks, independent review, and explicit human checkpoints.

What loop engineering means

A prompt asks for one result. A loop controls a sequence of attempts. It decides what enters the next iteration, what counts as progress, who may judge the output, and what happens when the work does not converge. That makes loop engineering a systems problem around the model, sitting alongside prompt and context engineering.

The label is new; the engineering concerns are not. Durable state, executable contracts, verifier integrity, bounded retries, and human approval are established reliability patterns. Stele uses “loop engineering” as a useful name for that combination, while building the product around those longer-lived ideas.

Operator is the proof, memory is the substrate

Stele does not require every team to adopt our runner. The graph can provide contract, state, and cross-iteration learning to a loop you already run. Operator shows how those pieces fit together in one complete implementation workflow.

The five layers of an agent loop

The Operator design maps closely to the five-layer “Operator Loop Stack”: harness, loop contract, state, checker, and human checkpoint.

LayerHow Stele Operator handles it
HarnessClaude Code is the packaged isolated-worker reference path. Another harness can honor the same runner-neutral contract when it supplies equivalent worktree isolation and a fresh independent reviewer; otherwise execution is sequential and supervised. Codex currently uses that supervised path because its install does not package the worker/reviewer protocol files.
Loop contractThe task carries its purpose, scope, success criteria, optional runnable check, and protected check_paths.
StateTasks, claims, attempts, comments, decisions, lessons, risks, and handoffs survive sessions in the project graph.
CheckerLocal MCP or CLI protects the oracle, runs the declared check with a fresh nonce, and records what actually executed.
Human checkpointYou choose the autonomy level. Shared-infrastructure pushes always require confirmation, and ambiguous or failed gates come back with evidence.

How Stele Operator runs, step by step

The reference loop The worker makes the change; a protected check, a fresh reviewer, and the operator judge different things. Separate brakes keep failed execution and review disagreement from becoming endless retries.

  1. Authenticate and choose scope. Operator confirms the Stele session, project, focus area, and whether you want Auto, Batch, or Step-level checkpoints.
  2. Read before choosing work. It loads repository rules, open tasks, dependencies, workspace scope, and the decisions and risks linked to the candidates. Priority and dependency order choose the work; having an easy check never makes a task more important.
  3. Inspect the contract. A task with a checkcan be driven toward an executable result. It is loop-certified only when check_paths also name the tests, scripts, and fixtures that make up the oracle. Operator never guesses missing paths.
  4. Plan a collision-safe batch. Small or serial work runs inline. Independent substantial tasks may run concurrently, but only when the harness can keep each worker inside its own Git worktree. Otherwise Operator stays sequential.
  5. Execute and watch. A dispatched worker receives the full task contract, repository rules, lineage, worktree, and protected paths. Operator watches Git state and completion signals so a silent hang does not masquerade as progress.
  6. Protect the oracle. Before trusting green output, Operator asks local Stele to compute the diff from Git and prove the worker did not modify its own declared check paths.
  7. Run both quality gates. Operator runs the check itself and requires nonce-bound evidence that tests actually executed. A fresh reviewer then judges the diff against the task without seeing the worker's self-report. The deterministic result, reviewer verdict, and operator intent verdict are stored separately and must all pass.
  8. Retry with information or stop. Rework includes the exact failure output and missed criterion. Execution attempts are bounded; the same failure twice is no progress, not permission to spend again. Reviewer-driven rework has its own tighter limit: two rounds, then a third rejection stops for human adjudication.
  9. Merge, verify, and remember. Approved work is merged, the broader repository checks run, durable lessons are written back, and Operator leaves one clear handoff for the next session.

Review is binding, not endless

Independent review catches the class of failure a test cannot: green work that solves the wrong problem. Its rejection therefore binds, and the operator cannot silently overrule it. But a binding review is not an invitation to keep inventing improvements after every revision.

A reviewer may block only on a concrete correctness, security, data-loss, task-contract, or failing-test defect in the declared scope. It must cite evidence and batch every blocker it can see in one pass. Style preferences, naming taste, speculative hardening, optional refactors, and unrelated cleanup do not block the current task.

On re-review, the reviewer checks the requested fixes first. A new blocker can extend the loop only when the rework introduced it, or when it is a critical defect that would make shipping unsafe. Other late observations become follow-up work. After two automatic reviewer-driven rework rounds, a third rejection records the stop and sends the evidence to the operator or a human. The loop stops; it never turns rejection into approval just to finish.

One loop contract, available through MCP and CLI

Oracle protection and evidenced attempts are shared Stele capabilities, available well beyond the CLI. Local MCP exposes verify_loop, run_loop_attempt, and finalize_loop; the grouped inspect and task actions expose the same operations. The CLI equivalents are shown below.

The default MCP form is inspect(action="loop", node_id="TASK-12", base="main", branch="worktree-agent-12"). The flat tool accepts the same task and Git refs.

terminal

$stele loop verify --task TASK-12 --base main --branch worktree-agent-12

oracle untouched · 2 implementation files changed

$stele loop attempt --task TASK-12 --base main --branch worktree-agent-12

attempt 1 · 42 executed · pass · evidence recorded

$stele loop finalize --task TASK-12 --reviewer approved --reviewer-summary "Independent review passed." --operator approved --operator-summary "Intent and scope passed."

deterministic pass · reviewer approved · operator approved

Both routes use trusted local Git state. Exit code 0 means the protected paths are untouched; 1 means the worker changed its oracle and the work is rejected; 2 means the result cannot be certified. Unknown fails closed.

An attempt first requires a clean registered Git worktree, binds its immutable commit and the approved command hash, then runs the task's declared check with credentials removed from the environment and a fresh STELE_LOOP_NONCE. It verifies the same clean snapshot again after execution. The check must print one STELE_LOOP_EVIDENCE={...} line using schema stele.loop.evidence.v1, echo that nonce, and report executed, passed, failed, and skipped counts. Missing, stale, malformed, or zero-execution evidence counts as unknown, never as green. The exact command, output excerpts, counts, failure signature, and verdict remain on the task.

Retries have mechanical brakes

Operator allows at most five recorded attempts. Two consecutive attempts with the same failure signature stop the loop immediately. Review has a separate two-rework limit, with a third rejection stopping for adjudication. These brakes are enforced by the loop capability rather than left to prompt compliance. Spend limits remain a harness responsibility; Stele does not claim to enforce a provider budget.

Why hosted MCP cannot certify a local diff

The hosted server can read and write graph state, but it cannot inspect your local repository or execute its checks. Verification and attempts therefore return unknown and direct the agent to local MCP or the CLI. Finalization also stays local because it rechecks that the reviewed worktree is clean and still points at the evidenced commit before it records any supplied verdicts.

What makes this loop effective

  • It separates maker, checker, and reviewer. The worker supplies a patch and evidence, never its own verdict.
  • It rejects verifier gaming mechanically. A passing test does not count if the patch changed the test or fixture that defines success.
  • It treats zero execution as unknown. A cached result or silently skipped suite is not promoted to a pass.
  • It makes retries learn. The next attempt receives the exact failure, while repeated failures trigger a brake.
  • It bounds disagreement. Reviewers batch real blockers; bikeshedding becomes follow-up work, and a third rejection returns the decision to a person instead of starting another automatic round.
  • It compounds project memory. The attempt trail stays on the task; only durable conclusions become knowledge that resurfaces on related work.

When should you use an autonomous agent loop?

Start with work whose finish line can be observed: a failing test, a build error, a migration with a validation command, or a measurable threshold. Keep architecture choices, security-sensitive changes, production deployment, and other judgment-heavy work under closer human supervision. Operator can still coordinate unchecked tasks, but it says plainly that a human is the gate, standing in for a deterministic oracle.

Further reading on loop engineering

Run OperatorHow Stele memory works

Backfill · seed the record from an existing project

Source: https://stele-ai.dev/docs/backfill

Concepts

Backfill

Stele is most useful on a project that already has history. Backfill seeds the record from what you already know: code, history, documents, or past conversations you provide. Your agent reads the sources it can access and saves sourced findings to Stele.

Who does the reading

This is the part to be precise about: Stele doesn't crawl your repo on a server. The local backfill command runs inside your coding agent. When you start it, your agent takes the role of orchestrator and spawns a wave of focused, cheap-model "missions". Each one reads a different source and reports back what it found. Nothing is uploaded to Stele except the findings you approve. Your agent provider handles source material according to its own settings.

Bulk Backfill validates an import file and inserts it through the CLI. A hosted chat has no bulk import tool today. You can still bring selected context into Stele from chat using ordinary knowledge and document writes.

Honest about what this is

Backfill is an agent-driven flow, not a magic importer. It reads what you point it at, fact-checks its own findings against the code, and asks for your approval before it writes. Facts it can't tie to a source are dropped rather than guessed.

What local backfill reads

Each mission specializes in one source, so the findings come back sourced and specific rather than vague:

  • The code. The structure of the system and the components it breaks into.
  • Git history. The decisions and reversals written into commits over time.
  • Issues & pull requests. The problems that came up and how they were resolved.
  • Documentation. Design notes and READMEs already in the repo.
  • Past agent sessions. The reasoning from earlier work with your agents. Supported local histories are discovered automatically, including Pi sessions; when a Pi session has branches, Backfill follows the branch you kept rather than importing abandoned detours.

How a local run goes

You stay in control the whole way. Backfill checks in before it gathers, before it spends on a model tier, and before it writes , so the scope is yours to set.

your agent

›/stele:start

no record here yet. Backfill from this repo? (you approve each step)

# missions read in parallel, then fact-check

· code → 18 components · git → 24 decisions · issues → 11 risks · docs → 6 notes

review the findings, then I'll insert them with their sources.

inserted 59 nodes, linked and sourced ✓

Every node it writes carries its provenance , where the fact came from , so you can trace any backfilled decision back to the commit, issue, or file it was drawn from. From there the record behaves like any other: it's served back on prompts, and it ages and gets verified like knowledge you write by hand.

Run local backfill

Backfill is part of /stele:start. On a project with history, it offers to seed the record as the first step. You can also run it on its own with the /stele:backfill command, or just ask for it in plain words:

Paste into your agent

Backfill this project into Stele. Read the code, the git history, and the docs, fact-check what you find, and show me the findings before inserting anything.

You may not need to ask. When the record is still nearly empty, your agent is told about Backfill as it works, so it can offer it at the point you'd actually want it. The same happens if a directory isn't bound to a project yet, pointing at /stele:start instead. Both are offers, not interruptions: each is raised once and then stays quiet for a while, and they stop for good once the record has enough in it to be worth recalling.

Start small if you want

On a large repository you can scope the first run to the part you're actively working in, then widen it later. Backfill adds to the record; it never overwrites what's already there.

Bulk imports have a one-time allowance

Bulk import writes draw from a finite, one-time allowance before using your monthly write allowance. Once the import allowance is exhausted, remaining writes count toward the applicable monthly limit. A large first run can exceed the allowance; check your account's current usage before starting.

What it changes day to dayHow the record works

Bring existing context from a hosted chat

Bulk Backfill currently requires the CLI

The Backfill workflow assembles and validates an import file, then uses stele project import to insert the batch. That bulk import is not exposed through hosted MCP today. For a full Backfill, use the local workflow above. The steps below are for saving selected source material through the agent's existing tools.

In ChatGPT, Claude, or another connected agent, choose your Stele project first. Then provide the material you want it to learn from: pasted notes, uploaded documents, conversation exports, or sources available through the host's other connected tools. A repository connection can help only if it lets that agent read the needed files and history.

  1. Identify the project and the sources to cover. Ask the agent to state which it can actually read.
  2. Have it read the existing Stele record before extracting facts, so it can update earlier entries instead of duplicating them.
  3. Review the proposed decisions, lessons, risks, and commitments with their source references. Keep uncertain inferences separate from confirmed facts.
  4. Ask it to save the agreed findings, link related records, and report which sources remain unread.
  5. Retrieve a saved finding from a fresh conversation to check that the context is available.

Paste into your hosted agent

Save useful context to Stele project from the material I provide. First tell me which sources you can read and check the existing record for related facts. Extract durable decisions, lessons, risks, and commitments with their sources. Show me the proposed additions and updates before saving. Afterward, give me the saved record links and list anything you could not read or verify. If the volume calls for bulk Backfill, explain the local import requirement before proceeding.

Source access comes from your host

Connecting Stele does not give the agent access to your computer, private repositories, or every past chat. Supply the material or connect a source the host supports. These ordinary Stele writes use the applicable plan limits; do not assume the bulk import allowance applies. For a large collection, use local Backfill rather than assuming that hundreds of individual writes are the same import path.

For a larger source review in ChatGPT, switch to Work and ask it to divide independent reading among subagents, then combine and verify their findings. Availability depends on your account and environment. Switching modes does not add Stele's bulk import tool: use the local import workflow for the final batch. See OpenAI's guide to Chat and Work and its subagent guidance.

Fold in new materialGive a recurring job context

Ingest · fold new material into the record

Source: https://stele-ai.dev/docs/ingest

Concepts

Ingest

You just came out of a meeting, or someone sent you a spec, or a competitor shipped something that changes your plans. Hand it to your agent and it goes into the record, folded into what the project already knows: decisions updated, commitments captured, and anything the material just made wrong, retired.

Reading the material is the easy half

Your agent already knows how to read a transcript. The part that has no obvious answer is the record: your project already holds decisions, open tasks, and risks about the exact things that meeting was about. Every fact in the new material has to be placed against them.

So ingest recalls before it writes, every time. It pulls what the record already says about each subject the material touches, and only then asks what this new thing does to it. That is the difference between adding knowledge and adding noise on top of it: a record that asserts two contradictory things, with nothing to say which one is live.

What new material does Adding is only one of the four things a piece of material can do to a record that already holds facts.

It decides two things about everything it finds

What shape is it? A durable fact is knowledge: a decision, a risk, a lesson, a gap. Something a person committed to doing is a task, with their name on it and the date they said. A real spec is a document, kept whole. A structure or a flow becomes a diagram. Most of the material (the chatter, the scheduling, the restatement) is nothing, and it says nothing about it.

And what does it do to what's already there? That's the question in the diagram above:

EffectWhen
NewNothing on record is about this.
GrowsThe fact is already recorded, and the material adds specifics: a date, a number, an outcome. The node gets bigger, rather than getting a near-duplicate next to it.
CorrectsWhat was recorded was never true. The text is fixed.
SupersedesIt was true, and this material makes it wrong to act on: a decision the room just reversed. The old node is retired: still there, still findable, marked superseded, ranked below what's live.

The reversals are the whole reason this is worth doing. A decision that quietly stopped being true, with nothing marking it, is the single most expensive thing a record can contain. Every agent that reads it afterwards does the wrong thing with total confidence.

A meeting is the case it was built for

Point it at a call transcript and your agent does what the person taking the notes was supposed to do: what did we actually decide, what's still open, and who said they'd do what.

your agent

›big meeting on the event store today, here's the transcript

# reads the record for what it already knows about the event store

this reverses a decision you already have. here's what I'd write:

· new decision: Postgres over DynamoDB (4x cost at 10k writes/sec)

· new task: benchmark Postgres at 10k/sec, target Jul 20

· grows KNOW-88 "event-store options": adds the cost numbers

· retires KNOW-45 "we're going with DynamoDB": reversed today

· 1 open question: nobody settled the migration window

written and linked ✓

Note the last two lines. The commitment became a real task with a name and a date, because "I'll take that" is the highest-value thing in any meeting and the thing most likely to evaporate into a document nobody reopens. And the thing the room didn't settle was recorded as an open question and left that way.

It asks before it retires

Anything that retires a node, rewrites one, or puts someone's name on a task gets shown to you first. You were in the room; it wasn't. When everything it found is purely additive, it just writes it and tells you what it did.

Anything you can hand it

A path, a URL, or something pasted straight into the chat. Your agent reads it on your machine. Nothing is uploaded and no crawler runs on a server. If the material is too big for one pass, it splits it up and reads the pieces in parallel.

A meeting

We had a big meeting about the new billing flow today, lots of opinions. Get it into the graph: what we locked in, what's still open, and who took what. Here's the transcript.

Something that changes your plans

This announcement affects our launch. Read it, and tell me what it means for what we've already decided.

A spec someone wrote

Jason wrote a full design for the export feature. Read it, check it against what we've already got, and set up the work.

You can also call it by name, /stele:ingest , but you rarely need to. Handing over material and saying what you want done with it is enough.

Seeding a record from scratchHow the record stays honest

When facts conflict · supersession, and how a reversed decision stops being followed

Source: https://stele-ai.dev/docs/conflicting-facts

Concepts

When facts conflict

Your project changes its mind. The hard part was never recording the new decision. It is making sure the old one stops being followed.

Why memory gets this wrong

Six months ago you chose DynamoDB for the event store. Last month you moved to Postgres. Both facts are written down. Today an agent asks itself what the event store runs on.

If memory is a pile of text searched by similarity, both facts come back. They're about the same subject, in nearly the same words , and nothing in the result says which one is still true. The agent picks one. Sometimes it picks the old one, because it was phrased more confidently, or it appeared first, or the question happened to be worded the way the old fact was. Then it writes DynamoDB code.

Similarity has no opinion about what's still true. That's not a tuning problem you can fix with better embeddings; a search engine ranks by likeness, and the stale fact is every bit as like your question as the current one. Something has to actually know the decision was reversed.

Superseding: retire the fact, keep the story

When a decision replaces an earlier one, your agent draws a supersedes link between them. The old fact is retired: it stops being served as truth and ranks below live facts.

It is not deleted. It stays in the record, clearly marked, pointing at whatever replaced it. Two reasons that matters:

  • The history is the point. "What did we decide before, and why did we change?" is one of the most valuable questions a record can answer, and it's the one a codebase can never answer on its own, because the rejected option leaves no trace in the code.
  • A mistake stays recoverable. If a fact is retired in error, that's a wrong label on something you can still see and correct, and the fact itself is still there.

Every retirement states its reason

Retiring a fact is a claim: someone is asserting the world changed. So Stele won't record it bare. A supersedes link has to carry two things:

  • a reason: what standing state actually ended, in your agent's own words;
  • evidence: the verbatim line from the source showing it ended.

This isn't bookkeeping. Constructing the proof is the test. A fact that merely got newer has not been superseded, and the difference is easy to miss in the moment:

The new fact saysDid the old state end?
"We've moved off DynamoDB entirely."Yes. A termination. Supersede it.
"We also added a read replica."No. That's additive; nothing was retired.
"We're still on Postgres."No. That's a continuation.
"We were planning the migration; now it's running."No. That's the same story developing; grow the fact instead.

You can't quote a source saying a state ended when the source only says something new was added. Demanding the quote at the moment the claim is made is what stops the three rows above from being mistaken for the first one.

If you can't say why, don't claim it

An agent that cannot state the reason is told to use relates-to instead , which keeps the connection between the two facts and asserts nothing. Nothing is lost, and nothing true gets buried.

The replacement always arrives with it

Marking a fact stale is only half an answer. Told "don't trust this" and handed nothing else, an agent has to guess.

So a retired fact never surfaces alone. Whenever one is recalled, the fact that replaced it is pulled in alongside, ranked above it and carrying the reason for the change. This matters most in exactly the case you'd expect to go wrong: when you ask a question using the old decision's vocabulary, and the stale fact is the closest match to your words.

A retired fact, and its replacement Ask in the old decision's words and you still match the old fact , but it comes back marked, ranked below the decision that replaced it, with that replacement and the reason for the change attached.

What your agent sees is not a bare warning but a complete picture: the old fact, the new one, the date it changed, and why.

When you genuinely don't know

Sometimes two facts conflict and neither is obviously wrong: two people remember a decision differently, or a document and the code disagree. That is a real finding, and flattening it into a guess is the worst thing a record can do.

For that there's contradicts. Both facts stay live, both stay served, and the conflict itself is surfaced to your agent with a note that it must be reconciled, not resolved by picking one. A record that reports a consensus that never existed is worse than one that reports nothing.

Auditing what's been retired

A wrong retirement is quiet. The fact is still true and still there, just labelled stale and ranked low, so agents skim past it. Nothing breaks; things merely get subtly worse.

Two things catch it. The health check surfaces every claim-bearing link drawn without a reason, so unaudited retirements show up as work rather than sitting invisible. And the verification sweep re-tests them: for each one, does the two-part rule still hold: was it a standing state, and did it actually end?

Ask your agent

Run a Stele review over the links that retire facts. For each one, check whether the retired fact was a standing state that actually ended, and was more than just an older fact. Demote the ones that fail and bring back what they wrongly buried.

An edge that fails gets demoted to relates-to, and the fact it wrongly retired comes back.

Where this leaves you

The reason a graph is worth the extra structure over a pile of notes is precisely this: it can hold an opinion about what is still true, and defend it. A retired decision doesn't vanish and doesn't mislead. It sits behind the one that replaced it, explaining itself. See KNOW-45? Then see what replaced it, and why.

The self-repairing graphCore conceptsIngest

The self-repairing graph · how the record stays honest

Source: https://stele-ai.dev/docs/self-repairing-graph

Concepts

The self-repairing graph

A memory is only useful if you can trust it. A graph that confidently serves a fact that's no longer true is worse than no memory at all , so Stele is built to maintain itself. Backfill seeds the record; this is how it stays honest as the project moves underneath it.

Two clocks on every fact

Stele tracks each piece of knowledge on two separate clocks, because "recent" and "true" aren't the same thing.

  • Freshness , when the content last changed. This is cheap and automatic.
  • Verification , when someone last confirmed the fact still holds against reality. This is the expensive signal, and it's never set by an ordinary edit, only by a deliberate check. A node can be freshly written and still wrong; only verification catches that.

Keeping the two apart is what lets the record tell "nobody has touched this in a while" from "nobody has confirmed this is still right" , and act differently on each.

What the graph does on its own

Most upkeep needs no one to run anything. As the record grows, Stele is continually:

  • Flagging the stale. Knowledge that hasn't been verified in a long stretch is surfaced for a fresh look, rather than served as if it were certain.
  • Favoring the recent. When two facts compete for a spot in the context, retrieval gives a gentle edge to the one updated more recently.
  • Expiring what was only ever temporary. A fact written as true-for-now can carry an expiry date; when it passes, it retires itself.
  • Retiring resolved risks. A risk or note bound to a task is archived when that task closes: the worry is handled, so the warning goes with it instead of lingering.
  • Stepping aside when superseded. When a decision replaces an earlier one, the old decision stops being served as truth , but it is not erased. It stays in the graph, clearly marked as superseded and pointing at whatever replaced it, and it ranks below live facts. So "what did we decide before, and why did we change our minds?" stays answerable, and a supersede recorded by mistake is a wrong label you can correct rather than a fact you have lost. Every retirement has to state its reason, and the replacement is always recalled alongside the fact it retired. when facts conflict covers this in full.
  • Catching duplicates and broken links. Near-identical facts and edges that point at nothing are surfaced to reconcile.

The knowledge lifecycle Knowledge ages on its own clock. When it's flagged, it's either re-verified and kept current, or retired, so what the record serves keeps up with the project.

The verification sweep

Freshness is automatic; truth isn't. The sweep is the deliberate part: the one that catches a node anchored to a file that has since moved, or a decision quietly reversed in the code while the record still claims the old answer.

It triages what's changed since the last sweep, diffing against the exact commit it last reviewed , and gathers the knowledge that touches those changes plus anything gone stale. Then it fans out a swarm of parallel reviewers across those facts; each one checks its batch against the real code, both the claim and the anchors (the files and symbols the fact points at). You make the final call on each: fix it, re-anchor it, retire it, or flag it. The run is recorded, so the next sweep only re-checks what changed since.

your agent

›/stele:review

42 facts touched by changes since the last sweep · 7 unverified in 30+ days

reviewers checked each against the code…

31 still hold · 4 re-anchored · 2 superseded · 5 flagged for your call

The health check

Where the sweep verifies truth, doctor is the broad audit: one command that walks the whole record and reports what needs attention: drifted claims, knowledge overdue for verification, expired notes, decisions superseded but still active, duplicates, broken links. The entries that are safe to retire on their own are cleared in the same pass; the rest come back as a short worklist your agent helps you act on.

your agent

›/stele:doctor

3 facts unverified in 30+ days · 1 near-duplicate · 1 expired note · 2 auto-retired

walk the rest · verify, merge, or retire?

Your agents do the gardening

The point of all this is that you don't hand-maintain the record. The automatic parts, flagging and expiry and retirement, never sleep. The deliberate parts ride along with the agents you already use: they run the sweeps, act on what the health check finds, and write new knowledge with anchors and expiry dates so it can retire itself later. Maintenance is continuous because it happens as part of the work.

On the roadmap

Today the sweeps run when your agent runs them, while the automatic flagging and retirement happen on their own. A fully autonomous, recurring drift-scan, one that gardens the graph in the background with no one prompting it, is something we're building next. It does not ship today.

How agents use SteleThe commands behind it

Surfaces · where you reach the record

Source: https://stele-ai.dev/docs/surfaces

Using Stele

Surfaces

There's one record, reached several ways. Agents read and write it through the plugin or any MCP client; you read it in the web dashboard, in the terminal UI, and drive it from the CLI. They all see the same graph at the same time, so a change one of them makes is immediately context for the others.

The web dashboard

Where people read the record. The dashboard shows the graph, tasks, knowledge, documents, and objectives , and the links between them , so you can see the project being built alongside the code rather than buried inside it. It updates live: when an agent claims a task or writes a decision, you watch it land.

app.stele-ai.dev

The graph view: every component, decision, and task as a node. Select one to see what it connects to and why.

Working with a team

A project's record is shared, not personal. Invite teammates and everyone reads and writes the same live graph, so what one person's agent learns is immediately context for the next. Tasks belong to the project rather than an individual, and claims are atomic, so two people (or two agents) never start the same work by accident. Whether you're solo today or a small team, the record is the common ground everyone builds on.

An invitation is an offer, not a change. Opening the link in the email shows you the project, who invited you, and which of your accounts you'd be joining as — and nothing happens until you choose Accept or Decline. Declining is a real answer: the person who invited you sees it was turned down rather than left unanswered, and they can send a fresh invitation if you change your mind.

You don't need the email to get there. Any invitation addressed to you waits at the top of your projects list, so a lost link or a sign-in with the wrong account is recoverable — sign in as the invited address and it's there.

The CLI

stele is a single binary that handles sign-in, project setup, and a full read/write surface over the graph, on par with what agents can do. It's also how you run the local web server and manage your account. Most days your agent calls it for you, but it's there when you want to drive the record directly or script against it.

terminal

$stele task list --status open

TASK-12 · add cursor pagination · open

TASK-14 · rate-limit the export endpoint · open

The CLI reference covers the full command surface.

The terminal UI

stele tui opens the record full-screen in the terminal: tasks, knowledge, documents, the graph and the assistant, keyboard-driven and updating live as your agents write. It's the dashboard's view without the context switch, for when you're already in a terminal and want to check what a task says or what the project decided.

terminal

$stele tui

The terminal UI covers what's on each screen and how to move around it.

The plugin, for agents

The plugin is how an agent's harness talks to Stele. It's two things working together:

  • An MCP server that exposes the graph as tools the agent can call: reading context, creating tasks, writing knowledge, following links. When it runs locally, it can also verify an Operator task against trusted local Git state without accepting a worker's account of what changed.
  • Lifecycle hooks that run automatically: one keeps relevant prior knowledge coming once the agent has said what it is working on, and surfaces a note the moment you edit a file it names. Others watch what files you read and keep the record fresh across a session.

The same plugin renders into each harness's native shape at install time: slash commands in Claude Code and Cursor, skills in Codex, Antigravity, and Copilot, a marketplace plugin in Grok Build, plugins in OpenCode and Kimi Code CLI, and a native TypeScript extension in Pi with prompt templates, compact recall receipts, and claimed-task status in the terminal , so the experience fits the tool you're already in.

DeepSeek Harness gets a native Cordis plugin and an in-app Stele panel. It also registers the full /stele-* workflow set and the read-only Stele search skill through dsh's native command and skill registries. The panel shows the current project, claimed task, graph nodes recalled, read, and written during the session, and the active workspace scope. You can search the graph from the composer, switch workspace scope, or pause recall for a few prompts without leaving the conversation. Node ids in assistant replies become clickable Stele references after the reply settles; hover one for a node preview without changing the conversation.

What differs between them

Two capabilities depend on what a tool's plugin surface can actually do, so they are not uniform. Pointing recall at a topic needs a channel that reaches one conversation only — Cursor's rule files are shared by every chat open on the repository, so the agent is told there that automatic topic recall is unavailable and uses direct lookups instead. File-anchored warnings need a tool boundary Stele can hook. Where the tool lets Stele watch but not write, the warning is held and delivered with your next message instead; Cursor has no such boundary wired yet, so it does not get them at all. Direct lookups, id references and search work everywhere.

Any MCP client

Beyond the supported harnesses, any MCP-compatible client can connect to the Stele server directly and walk the same record, including agents that don't touch code at all. One URL connects ChatGPT, Claude, Cursor, VS Code, or anything else that speaks MCP — see Connect your AI agent.

Local and hosted have different trust boundaries

The local MCP server and hosted MCP are available on every plan. Graph operations work through hosted MCP. Operator's oracle-integrity checks, execution-evidence attempts, and final commit binding need local repository access, so they run through local MCP or stele loop verify and stele loop attempt/finalize. Hosted MCP fails closed instead of pretending it inspected, executed, or rechecked local files.

Terminal multiplexer integrations

Multiplexer integrations are installed separately from coding-agent plugins. They observe Stele status around an agent session without changing how that agent connects to the record.

Herdr is the first supported multiplexer. Run stele install to enable the integration when Herdr is detected. Its plugin marks integrated agent rows with Stele and adds the claimed task id when one exists, so the sidebar stays readable no matter how long your project and workspace names are. Everything longer lives in a status popup: the agent's current intent, the full task, project, workspace scope, and whether a CLI update is waiting.

The popup opens on a key you choose. Install offers to bind one for you, and you can change it or add your own at any time:

~/.config/herdr/config.toml

[[keys.command]] key = "prefix+shift+s" type = "plugin_action" command = "stele.status.status" description = "Stele status"

prefix is whatever your own Herdr prefix is, so the default binding is that prefix followed by shift+s. If something already uses that key, install leaves it alone and hands you the block to place on a key of your choosing.

The assistant

Built into the dashboard is an assistant that answers from your project's live graph. Ask it where things stand, why a decision was made, or what's blocking a release, and it walks the record to answer, citing the tasks and knowledge it drew on, so you can open them yourself. It's the fastest way for someone without an agent open to interrogate the project.

Some actions wait for your OK

The assistant can act on the record as well as answer from it, and for anything that changes what the record shows by default, it checks with you first. Hiding a node is one of those; so is bringing a hidden one back, since un-hiding re-exposes something someone chose to make invisible.

app.stele-ai.dev

Ask the graph a question; the answer cites the nodes it came from.

See the surfaces at workCLI reference

The terminal UI

Source: https://stele-ai.dev/docs/terminal-ui

Using Stele

The terminal UI

The same record the dashboard shows, without leaving the terminal. Run stele tui and you get a full-screen interface over your project: tasks, knowledge, documents, the graph, and the assistant, all keyboard-driven and updating live as your agents write.

Opening it

One command, no setup. It opens on whichever project the current directory is bound to; if nothing is bound, it opens on the project chooser instead.

terminal

$stele tui

# or, equivalently

$stele --tui

It needs an interactive terminal and the installed stele binary — everything it draws with is already inside that one file, so there is nothing else to install. Press ? at any time for the full list of keys on the current screen, and q to leave.

If you aren't signed in, it offers to sign you in before it opens rather than showing you an interface with nothing behind it. Say yes and your browser opens; the interface comes up once you're back.

stele tui

The home screen keeps project search, the assistant, and recent activity in one terminal view.

What's in it

Every screen is reachable by a single key from the home screen, and the footer always shows what the current screen responds to — so you can learn it by using it rather than by reading this page.

Tasks

The board, with the task's detail beside the list. Claim, complete, and move between tasks without switching screens.

Knowledge & documents

Browse what the project knows: decisions, lessons, risks, gaps, and the longer documents, rendered as formatted text rather than raw markup.

Graph

Three ways to look at the shape of the record in a space that can't draw a picture of it, including what's connected to what and which nodes nothing links to at all.

Search

The same three search modes the dashboard offers, across everything in the project.

Assistant

Ask your project a question and get an answer drawn from its own record, with the nodes it cited listed beside the reply so you can open any of them. Conversations are kept, so you can come back to one.

Account, integrations, settings

Your plan and usage, which coding agents Stele is installed into (with a way out of each), and the interface's own options.

Anything with an equivalent in the dashboard can be handed to the browser with w — a terminal is the right place to read and act, and the wrong place to show someone a link.

Staying current

The header tells you when a newer Stele has been published, and u installs it without leaving the screen. Once it's installed, the header asks you to restart — the version you are running was loaded when you opened it, so a new one only takes effect next time.

If you let Stele update itself, this is usually a note rather than a task: the update has already been fetched in the background, and the header simply tells you to restart when convenient. See Install & first run for the choice that controls it.

When to reach for it

The terminal UI and the dashboard show the same record, and neither is a reduced version of the other. Use whichever is closer to hand: the terminal UI when you're already in a terminal and want to check what a task says or what the project decided, the dashboard when you want to show someone, share a link, or work through the graph visually.

It's not where agents read the record

Your agents reach the same graph through the plugin, whether or not this is open. The terminal UI is for you — a place to see what they wrote and to steer what they do next.

Every surfaceCLI referenceWhat it changes day to day

In practice · what Stele changes day to day

Source: https://stele-ai.dev/docs/in-practice

Using Stele

What Stele changes, day to day

The model is abstract until you see what it does to ordinary moments. Here are four that happen in every project with more than one contributor, human or agent , and how each goes differently once the reasoning lives in the record.

A watchdog on risky changes

The most expensive bugs are the ones you already fixed once. A change that looks harmless reintroduces a problem the project paid for months ago , and a fresh agent has no way to know. With the risk linked in the record, it surfaces while the change is being made, in time to change course.

your agent

›this refresh-lock looks unnecessary, removing it

Hold on. That lock came from TASK-12 during the auth rework.

Removing it reintroduces KNOW-5, a race that dropped sessions in production once. Want me to keep the guard and simplify around it?

The agent didn't get smarter. It got the context. It cited the task that added the lock and the risk that justified it, then offered a safer path. The graph is a check on the work, not just a log of it.

This isn't only triggered by what you say. Stele watches the boundary where work lands , when you edit a file a risk is anchored to, when a command fails, when you commit , and surfaces the matching note right there, even if you never mentioned it. Recall at the action boundary covers how each of those fires.

Onboarding that starts full, not empty

A new contributor, or their agent, usually spends the first week reconstructing context that someone already had. With Stele, the history is already at hand: why the API is shaped this way, the refactors it went through, the decisions still in force, what's planned next. The agent they bring reads the record before it touches a line of code, and they can ask the assistant directly.

A new teammate asks their agent

I'm new to this codebase. Before I start on the billing work, catch me up from Stele: the key decisions still in force, anything that's been tried and rejected, and the open tasks.

your agent

›where should I start on the billing refactor?

TASK-14 is open and unclaimed. The approach was settled in KNOW-9: per-account ledgers.

Read KNOW-63 first: the last attempt double-charged on retries.

They're productive on day one, and the agent gate-keeps as much as it helps: a change that contradicts a standing decision gets flagged before it lands.

Team memory that outlives the session

Context normally dies at the edges of a chat , when you hit a token limit, switch tools, or come back weeks later. Stele keeps it attached to the project instead of the conversation, so it survives all three. Switch from one agent to another and tell the new one to continue; it reads the task and picks up the thread.

a fresh agent, days later

›continue TASK-14

picking up the billing refactor. Per-account ledgers (KNOW-9); the retry double-charge risk (KNOW-63) is handled in the new path.

resuming where the last session left off · no re-briefing needed

Why this matters for solo work

"The next person" is often just you, weeks later, having forgotten the details. The record answers your own question, why did I build it this way, as readily as a teammate's.

A project manager for agentic work

Stele also runs the work. The loop has two halves you'll use constantly: a planner that turns intent into a task graph, and an operator that executes it.

  • /stele:planner takes an evolving idea, your docs, decisions, and constraints , and turns it into a prioritized, linked, implementation-ready set of tasks. Because it has full project context, it pushes back when a new idea conflicts with a decision already made.
  • /stele:operator runs the implementation loop: it reads the open tasks, picks a coherent bundle, works small or serial changes inline, and uses isolated workers only when the harness can safely parallelize independent work. It verifies each result, records durable lessons and risks, completes the tasks, and writes a handoff so the next run continues cleanly. It does not require a vendor-specific multi-agent mode.

The thread that ties it together is the record: planner writes the tasks with their rationale, operator works them with that rationale in hand, and every lesson learned along the way is waiting for the next related change.

Checked work gets a stronger gate

When a task declares both a runnable check and its protected check_paths, Operator verifies through the local Stele MCP capability or the equivalent CLI command that the worker did not rewrite its own oracle. It then runs the check and sends the diff to an independent reviewer. Read the full Operator loop.

app.stele-ai.dev

An objective broken into ordered milestones and tasks: the shape planner produces and operator works through, with progress tracked as tasks close.

Plan, then ship

/stele:planner · turn the notes in DOC-3 into a prioritized task graph for the export feature, flag anything that conflicts with decisions we've already made.

How Operator runsAll slash commands

How-to recipes

Source: https://stele-ai.dev/docs/how-to

Using Stele

How-to recipes

Short answers to "how do I…?", each anchored to the prompt or command that does it. Most of the time you talk to your agent in plain language; these show what to say.

Resume work after a context limit

You hit a token wall, or you want a different model. Don't re-explain anything. Point the new session at the task. It reads the task, the linked decisions, and the risks, and continues.

In the new session

Continue TASK-14. Read it and its linked knowledge from Stele first, then pick up where the last session left off.

Find out why something exists

When a piece of code makes you ask "why is this here?", the answer is a path through the record that you can walk. The agent traces it back to the decision and the risk behind it.

Paste into your agent

Why does our API paginate by cursor instead of page number? Check Stele for the decision and the reason behind it.

your agent

›why cursor pagination, not page numbers?

KNOW-9 · offset paging skipped rows when records were inserted mid-scan; cursors keep results stable under writes. See TASK-12.

Capture a decision you just made

You and the agent just settled something: an approach, a tradeoff, a thing to avoid. Write it down while it's fresh, so the next session inherits it instead of relitigating it.

Paste into your agent

We just decided to drop the second config format. Save that to Stele as a decision, with the reasoning, so nobody reopens it later.

Keep the record honest

Over time, some knowledge drifts out of date. A periodic health check surfaces what's stale, what looks duplicated, and what's no longer true , and the agent helps you resolve each one.

your agent

›/stele:doctor

3 facts unverified in 30+ days · 1 near-duplicate · 1 expired note

walk each one · verify, merge, or retire?

You can also point the agent at a single area to re-verify the decisions that govern it before you build on them.

Paste into your agent

Before I work on auth, re-verify the decisions Stele has about it: confirm what still holds, and supersede anything that's drifted.

Coordinate several agents

Running more than one agent at once? Claims keep them off each other's work. Each claims a task before starting, so the others see it's taken; what each learns is written back for the rest.

your agent

›start on the export feature

TASK-20 (export) is already claimed by another session.

TASK-21, the import side, is open and unblocked. Starting there.

All slash commandsCLI reference

CLI reference

Source: https://stele-ai.dev/docs/cli

Reference

CLI reference

The stele command surface follows one consistent grammar: stele <object> <verb>. Learn the verbs once and they mean the same thing on every object. This page is the shape of the surface; stele --help and stele <object> --help give the exhaustive, always-current detail.

The verb grammar

Every object that supports an operation uses the same word for it: get, never show or view; list, never ls; create, never add.

VerbWhat it does
get Fetch one object, with its full body.
listEnumerate objects, filtered and paginated; cards only, no bodies.
createMake one.
update Edit fields on one.
search · grep Find, scoped to the object when run under it (e.g. task search).

terminal

$stele task get TASK-12

$stele knowledge list --component auth

$stele task search "pagination"

The objects

Each node type and structure is a top-level object with the verbs above, plus actions specific to it.

ObjectRead / writeDedicated actions
taskget · list · create · updateclaim · complete · release · reopen · cancel · subtasks
objectiveget · list · create · updateachieve · add-milestone · milestones · retire
knowledgeget · list · create · updatereview · dedupe
documentget · list · create · updategrep (body search)
componentget · list · create · updatetree
workspaceget · list · create · updateactivate · merge · cherry-pick
nodeget (any id) · updateversion history · restore · the low-level, type-agnostic escape hatch

Cross-cutting verbs work across types: search, grep, recall, context (build or trace), stats, and doctor.

Account & project

Setup and administration live under their own commands.

CommandWhat it does
stele account loginSign in (and logout, account info).
stele projectCreate, list, and bind the project for this directory.
stele installSet Stele up for the coding agents and supported terminal multiplexers you choose (and uninstall).
stele webOpen the dashboard, backed by a local server when you want it.
stele status · stele doctorSee sign-in state, project, branch, workspace scope, claimed task, agent-set intent, and available CLI update. Pass --session-id and --harness to include session context. --json exposes the same status for integrations. doctor audits graph health and recall behavior.
stele updateUpgrade the CLI and refresh the agent and multiplexer integrations you selected.
stele --skillPrint the instructions Stele gives your agent, exactly as it installs them — read them, or pipe them into an agent Stele doesn't set up for you.

Loop verification and evidenced attempts

stele loop verify is the local integrity gate used by Stele Operator. Given a task and two Git refs, it computes the changed files from the local repository and proves that none overlap the task's declared check_paths. The worker does not get to provide that changed-file list.

terminal

$stele loop verify --task TASK-12 --base main --branch worktree-agent-12

  • Exit 0: the declared oracle paths are untouched.
  • Exit 1: the diff modified a protected check path.
  • Exit 2: verification is unknown. For example, the task lacks check_paths, a ref is invalid, or local Git is unavailable. Operator stops; it never rounds unknown up to green.

The same capability is available over MCP

A local Stele MCP server can run the same verifier through the loop inspect action or the flat verify_loop tool. Hosted MCP cannot inspect your local Git repository, so it fails closed and directs Operator to a local MCP server or this CLI command.

stele loop attempt adds execution evidence to that integrity gate. It resolves the branch's registered worktree, runs the declared check with a fresh STELE_LOOP_NONCE, and records the result on the task. The check must emit one STELE_LOOP_EVIDENCE={...} JSON line with matching nonce and non-zero execution counts. Missing, replayed, or malformed evidence fails closed.

terminal

$stele loop attempt --task TASK-12 --base main --branch worktree-agent-12

$stele loop finalize --task TASK-12 --reviewer approved --reviewer-summary "Review passed." --operator approved --operator-summary "Intent passed."

Attempts stop after five runs or after the same failure signature occurs twice in a row. loop finalize stores the deterministic, reviewer, and operator verdicts separately; final approval requires all three. Stele does not enforce model-provider spend limits.

Output conventions

  • Browse, don't dump. get <id> shows the full object because you asked for exactly one; list, search, and grep return compact cards, never full bodies.
  • --json everywhere. Every read can emit JSON for scripting.
  • Next-action hints. Human output suggests the likely next step (get a task → how to claim it); --no-hints silences them.
  • Destructive verbs ask first. Anything that removes or merges is dry-run by default and needs --confirm to proceed.

The CLI is self-documenting

You don't need to memorize this. stele --help prints the full command tree, every subcommand has its own --help, and a mistyped command suggests the right one.

Slash commandsThe model behind the objects

Slash commands

Source: https://stele-ai.dev/docs/commands

Reference

Slash commands

These are the commands your agent runs inside its harness. In Claude Code and Cursor they're /stele:* slash commands; in Codex, Antigravity, Copilot, Grok Build, and Kimi Code CLI they're the same commands as skills or plugin commands; in Pi they're/stele-* prompt templates; in DeepSeek Harness they're native /stele-* commands and explicitly invocable skills. You'll mostly just describe what you want in plain language. These are the named entry points for the bigger moves.

Setup

CommandWhat it does
/stele:startStart or resume a session in this directory: bind a project, map it, and (on a repo with history) offer to backfill.
/stele:switchSwitch the project bound to this directory.
/stele:unlinkClear the project binding for this directory.

Capture what you learn

CommandWhat it does
/stele:create-taskTurn what's on your mind into a fully-linked task , or import one from a GitHub, Jira, or Linear URL.
/stele:save-contextReview the session and persist the decisions, lessons, and risks worth keeping, with duplicate-checking.
/stele:save-docSave a long narrative (a plan, an audit, a design spec) as a document in the record.
/stele:ingestFold material you hand over (a meeting transcript, a design doc, an announcement, a URL) into the record: capturing what was decided and who committed to what, and retiring anything it just made wrong.

Plan & ship

CommandWhat it does
/stele:plannerPlan features through conversation: gather context, push back on ideas that conflict with prior decisions, and capture the result as a task graph. Work that can carry a runnable check is filed with the check and its protected paths, ready for the Operator loop.
/stele:operatorRun Stele's reference implementation loop: scope the queue, plan batches, execute inline or in isolated workers, protect task checks, independently review results, merge, and leave a durable handoff.

Operator's contract is runner-neutral; full automation is scoped

Claude Code is the currently packaged reference path for isolated workers plus the independent reviewer protocol. Other harnesses may use the same contract when they provide equivalent worktree isolation and a fresh reviewer; otherwise Operator stays sequential and supervised. Codex is currently in that supervised category because its install does not yet package the worker/reviewer protocol files. Local MCP or the CLI protects the oracle and records nonce-challenged attempts.

See every stage of the Operator loop.

Find & load context

CommandWhat it does
/stele:searchSearch the project for knowledge, tasks, and components matching a query.
/stele:load-contextPull the full neighborhood around a topic: decisions, tasks, deadlines, recent activity , and suggest a next step.

Keep it honest

CommandWhat it does
/stele:reviewCheck the record for drift: triage what's changed since the last review (or what's gone stale), fan out reviewers over those facts, verify each against current reality (claims and the files it points at), then fix, re-anchor, retire, or flag what's moved on. Records the commit it reviewed, so the next run only re-checks what changed. Add auto to run it unattended.
/stele:doctorRun a health check across the whole record and act on what it surfaces: drifted facts, duplicates, broken links.
/stele:feedbackSend quick, structured feedback on how well Stele helped this session.

Plain language works too

You don't have to reach for a command. "Save that decision," "what's blocking the release?", or "plan the export feature" all route to the right behaviour. The commands are just the explicit handles.

See the commands at workCLI reference

Tuning Stele · the settings that steer it

Source: https://stele-ai.dev/docs/tuning

Reference

Tuning Stele

Stele works well out of the box, and most people never touch a setting. But when you want to steer it, here are the dials and what each one changes. They all live in the project's settings in the dashboard.

Recall, the retrieval layer

These settings decide how much prior knowledge Stele surfaces, and how far it walks to find it. The defaults are tuned to stay out of the way. One setting, every surface: what you tune here governs your coding agent, the assistant in the app, and the graph search alike.

  • Preview before you commit. Describe a piece of work and see exactly what recall would surface for it, side by side: the default behavior next to your customized one, so you're tuning against real output.
  • Tune the two stages. Retrieval runs in stages: a full-text and semantic search that finds the starting nodes, then a walk out across the graph. You can adjust how aggressively each stage contributes, and reset to the defaults in one click.
  • Pause it. When you want a clean slate for a few turns, say a tangent unrelated to the project's history, your agent can pause recall without changing any settings. Its topic survives the pause.

Start from the preview

The fastest way to tune is to describe a piece of work you actually do and adjust until the surfaced set looks right. The numbers mean less than what they surface for the work you care about.

What the preview can't see

It runs a fresh retrieval. It has no view of what a running session is focused on, what that session has already been shown, or file matches waiting on someone's machine.

When retrieval gets slow

Recall runs before your agent sees your message, so it holds the message open while it works. To stop a slow network from turning that into a long wait, it works to a time budget. If the budget runs out mid-walk, it surfaces the nodes that matched the topic directly and skips the ones it would have reached by walking out along their edges. You still get context, just less of it.

That trade-off is easy to miss, because a degraded run looks exactly like a healthy one from the outside, because the agent still answers. So run stele doctor. It reports how long recall has actually been taking, and tells you when it has been quietly cutting itself short:

terminal

$stele doctor

recall is degrading over the last 40 prompts

latency p50 2429ms · p95 8103ms · max 9260ms

warning recall is slow · it delayed 1 prompt in 20 by 8.1s

warning semantic recall failed or timed out on 6/40 prompts (15%) · those fell back to keyword search only

A hook that always succeeds and always takes nine seconds is still broken, so slowness is reported as a problem in its own right, alongside outright errors. If you need a tighter or looser ceiling than the default, set STELE_AUTORAG_BUDGET_MS in the environment your agent runs in.

Don't tighten the budget blind

Lowering the budget below what retrieval actually needs doesn't make it faster. It makes every turn pay the new ceiling and return less context. Check stele doctor before and after, and watch whether the semantic-recall warning appears.

Deduplication

As a record grows, the same fact sometimes gets written twice in slightly different words. The dedupe view scans for near-duplicate clusters and lets you resolve each: pick the one to keep, and the others merge into it, their links rewired onto the survivor and a pointer left behind, so nothing that referenced the old node breaks. It's a review surface, not an automatic shredder: you approve every merge.

Topics

Topics are the lightweight tags that cut across components. Over a project's life their vocabulary drifts: perf and performance end up as two tags for one idea. The topics settings list every topic with its count and let you rename one or merge several into a single canonical tag, so the cross-cutting view stays clean.

Notifications

Control which emails Stele sends you. The non-essential categories, a weekly digest and product news, are yours to switch off; account and security messages stay on, because they're about access to your own data. The same page keeps a history of every email we've sent you, as a plain audit log you can scroll back through.

How retrieval worksKeeping the record honest

Your data · ownership, privacy, export, and deletion

Source: https://stele-ai.dev/docs/your-data

Your data

Your data

Stele holds the reasoning behind your projects, so it's fair to ask exactly what happens to it. The short version: your data is yours, it stays private by default, your source code stays on your machine, and you can take it with you or delete it whenever you want. Here's the detail.

Your data is yours

We don't sell your data, and we don't use what you put into Stele to train machine-learning models, beyond features you explicitly turn on inside your own account. No advertising, no cross-context tracking. You keep ownership of your content; Stele stores and serves it back to you.

This page is the plain-language summary. The binding version, covering what we collect, how we use it, and your rights in legal terms, is the Privacy Policy.

What stays on your machine

The CLI and the plugin run on your own computer. Your credentials and local cache live there, and we don't have access to them. Backfill works the same way: your agent reads your code locally, and only the findings you approve are ever written to the record. Your source code isn't uploaded to us. What reaches the hosted service is the knowledge and tasks you choose to keep.

Telemetry, and how to say no

When you set Stele up in an agent, it asks, in the terminal and before anything is collected, whether it may gather telemetry. Your answer is kept on your machine. Reinstalling never asks again and never quietly changes it, and if Stele is installed unattended, where there's nobody to ask, the effectiveness channel below stays off.

If you say yes, two things are collected:

Usage

Which Stele tool or command ran, how often, how long it took, whether it errored, and your operating system and version numbers. The names of the arguments passed, never their values.

Effectiveness

A score, worked out on your machine, of how much of the context Stele gave your agent actually got used, plus four short written answers from your agent on where Stele helped and where it got in the way. This is sampled occasionally, and asking briefly resumes your agent, which spends your own model quota.

Your prompts, your agent's transcripts, your notes, your files, and your source code are never sent. Telemetry is kept for 12 months, then deleted; deleting your account deletes it immediately. Saying no costs you nothing: every feature works the same either way.

Change your mind at any point:

terminal

$stele telemetry

Stele telemetry

metrics on (opt-out)

feedback on

$stele telemetry off

✓ metrics off · no usage will be recorded.

$stele telemetry feedback off

✓ feedback off · observer + self-report will not run.

Where the record lives, and how it's kept separate

The graph itself is stored on managed cloud infrastructure, encrypted in transit. Every project is isolated from every other: you only ever see the projects you're a member of, enforced at the data layer rather than left to the application to remember. A project you're not part of is not something your account can reach.

Take it with you, anytime

From Account → Security in the app, you can export a complete copy of your data as a single file: your profile, your project memberships, and every project you own, in full: all of its knowledge, tasks, documents, comments, and the edges between them, whoever on the team wrote them. No request email, no waiting on us. It's a plain, portable format you can read, archive, or move elsewhere.

What it doesn't include is work you contributed to someone else's project. Knowledge you write into another team's graph belongs to that graph, for the same reason a teammate who leaves can't take your project's memory with them. Your export still lists those projects, so you know where your contributions live and who to ask; the owner can export them.

Hand a project to someone else

Deleting isn't the only way to let go of a project. Under Danger zone in its settings, an owner can Offer it to anyone already on the project. Pick them from the member list, type the project's slug to confirm, and we send them the offer.

Nothing moves yet. The project stays on your account until they accept, and stays yours for good if they decline. While they're deciding, the danger zone shows who you're waiting on and lets you withdraw the offer. Both of you get an email — theirs is the offer, yours is a record of what you did, with a way to report it if you weren't the one who sent it.

When they accept, the project moves to their account and starts counting against their plan instead of yours. You stay on as a member, so you keep reading and writing — you just can't change the project's visibility, offer it again, or delete it. Tick Stay an owner too if you'd rather keep full control: the project still moves to their account, you just keep your own say over it. Either way, invite the person first if they aren't on the project yet; offering hands a project over, it doesn't add people to it.

Ownership is a liability as well as a privilege, which is why it takes both of you. Nobody can be made the owner of a project without agreeing to it, every offer and answer is recorded, and you can always ask us who handed a project to whom, and when.

One thing worth knowing before you offer: your export covers the projects you own. Once a project belongs to someone else, it stops coming down in yours and starts coming down in theirs. Export first if you want your own copy of it.

If their plan is full

A private project needs room on the receiving plan. If there isn't any, accepting is refused and nothing changes — the offer stays open, so they can make room and accept again. Making the project public first gets it through without an upgrade, since public projects don't count against the limit.

Delete your account or a project

Same place, no ticket required. Every project you own has a Delete control in its settings, under Danger zone. You type the project's slug to confirm, and it takes the whole record with it, for everyone on the project: knowledge, tasks, documents, components, comments, and every edge between them. Deleting your account does the same to every project only you belong to.

Only an owner can delete a project. Deleted content leaves our active systems within 30 days; backups roll off within another 30 as they rotate. Deletion is yours to trigger and it actually removes the data, rather than just hiding it.

If you are on a paid plan, cancel it from Billing first; account deletion checks with the payment provider and stops until the plan is cancelled or already winding down. Anything else still attached to your account there, such as a checkout that never completed, is ended as part of the deletion, so nothing keeps billing after the account is gone.

The one exception: public projects

Projects are private by default. You can choose to make one public , and that's the single case where your record is visible beyond you and your team. Public projects appear in the discovery feed for anyone to read; everything else stays yours alone. It's opt-in, per project, and reversible.

Each public project also gets its own page. Visitors see the project's graph and its newest knowledge, and they can ask the project's memory questions in plain language, without an account, a limited number of times a day. The answers draw only on what the project already shows publicly, and a visitor can't change anything.

Once a project is public, its settings include a README badge builder. Pick a badge ("memory by stele", your live knowledge or task count, or your own text), a style and a color, then copy the snippet into your README. Each badge links back to your public page. By default the badges follow the reader's light or dark GitHub theme. They re-read your graph on every load, so they stay current with no rebuilds.

A short address for a public project

A public project can claim an address that mirrors its GitHub repo. Its public page then lives at app.stele-ai.dev/explore/<owner>/<repo>. This is the link to share: it shows the project's own preview card when you post it. The shorter app.stele-ai.dev/<owner>/<repo> opens the same page. Once the project has an address, its old /explore/<slug> link forwards to the new one, and the in-app /p/<slug> pages stay as they are.

To claim it, open the project's settings, go to the Project tab, and find the Project address panel. Paste the GitHub repo URL. Only an owner can do this, and only while the project is public. Private projects have no short address.

Stele then checks that the repo points back at this project. It reads the repo's default branch and accepts any one of these as proof:

  • The README (README.md or a common variant) shows the project's README badge.
  • The README links to the project's public page, at app.stele-ai.dev/explore/<slug> or app.stele-ai.dev/p/<slug>.
  • A file named .stele-verify at the root of the repo contains the link to the project's public page.

A dialog shows each check as it runs. If one passes, the address applies right away. Stele reads the repo without signing in to GitHub, so the repo itself has to be public.

The repo decides

The GitHub repo is the authority for its address.

  • Keep the badge or link in the repo. It is your proof if the address is ever checked again.
  • If another project proves the same repo, the address moves to that project.
  • Some owner names are taken by pages in the app, such as p, explore, login, settings, and account. A repo under one of those names can't claim an address.

The binding versions

This page summarizes; the policies govern. See the Privacy Policy for data handling and your rights, the Terms of Service for ownership, and the Security page for how we handle vulnerability reports. Questions a page doesn't answer? Email legal@stele-ai.dev from your account address.

Install & first runHow the record works


Blog

AI agent memory: what it is, the kinds that exist, and how to choose

Source: https://stele-ai.dev/blog/ai-agent-memory

Guide

AI agent memory: what it is, the kinds that exist, and how to choose

AI agent memory is any information an agent can use that was not in its prompt when the task started: facts it saved, notes someone wrote for it, or a record it shares with other agents. Everything else about the topic follows from three questions: what gets stored, where it lives, and who can read it.

September 30, 2026·Stele·12 min read

What AI agent memory is

A language model has no memory of its own between calls. Every request starts from the text you send it and nothing else. An agent is a model in a loop that calls tools, and agent memory is the set of mechanisms that put the right information back into that text: what the agent learned last week, what the team decided last month, what another agent found an hour ago.

A useful definition, short enough to quote: agent memory is information that survives the end of a session and comes back into a later one, either because the agent retrieves it or because the system loads it for the agent. Everything in this guide is a variation on how that information is chosen, stored, found and retired.

The rest of this page walks through the kinds of memory that exist in practice, where each one lives, the problems every memory system has to solve, and how to choose between them. It links out to longer pieces where a topic deserves one.

The context window is working space

The context window is the text the model sees on this call. It is large now, often hundreds of thousands of tokens, and that size makes it tempting to treat it as memory. It behaves more like a desk than a filing cabinet. Everything on it is visible at once, it is cleared when the session ends, and piling more on it has costs.

Three costs show up in practice. Every token in the window is paid for on every turn, so a long pasted history makes each step slower and more expensive. Models attend less reliably to the middle of a very long context than to its start and end, so a fact buried on page forty of a transcript is easier to miss than the same fact placed near the top. And when the window fills, the harness has to drop or compress something, which is exactly the moment an agent forgets an instruction it was given an hour ago.

Memory is what lets the window stay small and relevant. Instead of carrying everything, the agent carries what this step needs and fetches the rest when it becomes relevant.

The kinds of agent memory

The literature borrows words from psychology (working, episodic, semantic, procedural), and they are useful up to a point. In running systems the more practical split is by mechanism. These are the seven you will meet.

1. Working memory: what is in the context right now

The current conversation, the files the agent has read this session, the tool results it has seen. It is complete and exact, and it disappears when the session ends. Every other kind of memory exists to put something back here.

2. Session history: the transcript, kept

Many systems store the full message history of past sessions and let the agent search or reload it. Letta, for example, persists every message in a database, so an agent can retrieve an old exchange even after it has been evicted from the context window. A transcript is the most faithful record of what happened. It is also the noisiest: the one sentence that mattered sits among hundreds that did not.

3. Instruction files: memory a person writes

A markdown file in the repository that the agent loads at the start of every session. Claude Code reads CLAUDE.md; Codex and Cursor read AGENTS.md. The file holds build commands, conventions and rules you would otherwise repeat. It is cheap, reviewable in a pull request, and portable to any tool that reads the file. It has no idea which of its lines are still true, and it loads whole whether or not a line is relevant to the task. We covered how Claude Code handles these files, and its own auto memory, in how CLAUDE.md and auto memory work.

4. Extracted memories: facts the system pulls out

A model reads the conversation and writes down what seems worth keeping: "the user prefers TypeScript", "the staging database is read-only on Fridays". Mem0 is the best-known example: its docs describe sending messages through a model that pulls out key facts, decisions and preferences. Its current storage is additive, so new memories accumulate and a correction is an explicit update call on the old memory. Extraction is automatic, which is its strength. What gets kept depends on what the extractor judged important, which is its weakness.

5. Semantic recall: search by meaning

Store text as embeddings in a vector index, and at query time return the passages closest in meaning to the question. This is the retrieval half of most memory systems, including many that describe themselves as graphs. It finds "we moved jobs to the queue table" when you search for "background processing", which keyword search would miss. It returns the most similar passages, which are not always the most important or the most current.

6. Knowledge graphs, including temporal ones

Store entities and the relationships between them, so recall can follow links: this decision affects that component, which has this open task. A temporal graph also records time. Graphiti, the open source engine behind Zep, describes explicit bi-temporal tracking and temporal edge invalidation: when a fact stops being true its edge is marked invalid, and the history is kept, so the graph can answer both what is true now and what was believed last March.

7. Shared project records: memory for a team of agents

A single store that several agents and several people read and write, organized around a project rather than a user. Entries are typed (a decision, a lesson from a failure, a risk, a task) and linked to each other and to the code they describe. Letta's shared memory blocks, attached to several agents at once, are one form of this. A committed instruction file is a simple form of it too, shared through git. The distinguishing question is whether what one agent learned reaches the next agent that needs it, whichever tool that agent runs in.

Where memory lives, and who can read it

The kind of memory tells you what is stored. Where it is stored tells you who benefits, and that decides more about day-to-day usefulness than any retrieval algorithm.

Where it livesWho can read itSurvives a new machine?Typical examples
Files in the repositoryAnyone who clones it, any tool that reads the fileYes, through gitCLAUDE.md, AGENTS.md, Cursor rules
A folder or database on your machineThe agent on that machineNoAuto memory, local memory plugins
A hosted memory serviceWhatever holds the API keyYesMemory APIs used by AI products
A shared project recordEvery agent and person on the projectYesTeam memory servers reached over MCP

Two consequences are easy to miss. Memory stored inside one tool is invisible to every other tool: switch from one coding agent to another for an afternoon and the second one starts from nothing. And memory stored on one machine is invisible to your teammates, so each of their agents relearns what yours already knows.

The hard problems every memory system has

Storing things is easy. Five problems decide whether stored memory helps an agent or quietly misleads it.

Relevance: getting the right thing back

An agent working on the checkout page needs the note about the payment provider's rate limit, and does not need forty notes about the marketing site. Similarity search gets close; ranking by recency, importance and links between entries gets closer. When retrieval misses, a system that loads a whole file compensates by loading everything, which brings back the context window's costs.

Staleness: knowing when a fact stopped being true

In March you write down that jobs go through the queue table. In June you remove the queue table. If nothing marks the March note as retired, an agent in July reads it, trusts it, and rebuilds against a table that no longer exists. Systems handle this differently: temporal graphs invalidate the old edge, extraction systems let you update or delete a memory, and some records retire a fact when a newer decision replaces it or when a date passes. A plain file relies on a person noticing. This is the problem that matters most as memory grows, and we wrote a whole piece on what happens when a fact stops being true.

Conflicts: two agents, two versions

Once several agents write to the same memory, they will disagree. One records that the API returns dates in UTC, another that it uses local time. A useful system keeps both visible and makes the conflict obvious; a naive one keeps whichever was written last. The same problem appears with work: two agents picking up the same task at the same time is a memory problem too, because neither knew the other had started.

Cost: every recalled token is paid for

Memory that is injected into every prompt costs tokens on every turn. A good system spends that budget on the few entries that matter for this step, and keeps the rest one retrieval away. Measure what your memory adds to each prompt before and after you adopt it; it is the easiest number to get and the one people most often skip.

Privacy: what leaves the machine

Hosted memory means project knowledge is stored on someone else's servers. Check what is sent (summaries, full transcripts, source code), who can read it, and whether you can export and delete it. Local memory avoids the question and gives up sharing in exchange.

One agent, or many

Most memory systems were designed for one agent serving one user: a support bot that remembers a customer, an assistant that remembers your preferences. The memory belongs to the relationship between that agent and that person.

Software projects increasingly have many agents: one planning, one implementing, one reviewing, sometimes in three different tools, alongside several people. Their memory problem is different in kind. The unit is the project. The facts that matter are decisions and their reasons, lessons from failures, and who is working on what. And the failure that hurts is one agent repeating a mistake another agent already made and recorded.

The useful test for any memory system: does what one agent learned reach the next agent that needs it, in a form it can trust?

For a single agent, per-user memory is often all you need. For a team of agents, look for memory that every tool can reach, that distinguishes a decision from a passing remark, and that has a way to retire what is no longer true.

How to tell whether memory is helping

Every memory product has a demo in which memory helps. Your project is the only benchmark that counts, and three checks take an afternoon.

First, read what gets injected. Most systems can show you what they added to a prompt. Count how much of it was relevant to the task. If most of it is noise, the agent pays for it on every turn and may act on the wrong entry.

Second, change a fact and wait a day. Tell the agent the project moved from one database to another, start a fresh session tomorrow, and ask about the database. A system that serves the old fact beside the new one without saying which is current will eventually mislead an agent at a bad moment.

Third, switch tools or machines. Do some work in one agent, then open the same project in another agent or on another computer and ask what was done. This separates memory that belongs to a tool from memory that belongs to the project.

Measuring this rigorously is harder than it sounds. We wrote up our attempts, including the ways a benchmark can run perfectly and measure nothing, in how do you measure whether agent memory actually helps.

Memory for AI coding agents

Coding agents are where most people meet agent memory first, and they have a specific shape. The codebase itself is a large, reliable memory: an agent can read the code to learn what the system does. What the code does not say is why. Why the retry limit is three, why the team stopped using the old queue, which approach was tried and abandoned, which file is about to be rewritten by someone else. That missing layer is what coding-agent memory is for.

A typical setup grows in stages. It starts with an instruction file:

A small AGENTS.md

# Build and test
- Run the test suite before every commit.

# Conventions
- Money is stored in integer cents.

# Decisions
- Background jobs go through the jobs table. We tried a hosted queue in
  2025 and removed it; see docs/adr/007.

That covers one person on one project well. As more people and more agents join, the file grows, lines go stale, and learned knowledge stays trapped in whichever tool learned it. Then teams add a memory tool. We compared eleven of them, with their prices and tradeoffs, in the best memory tools for AI coding agents.

How to choose

Answer these in order. Each one rules options in or out.

  1. Who needs to read the memory? One agent for one user points to per-user memory. Several agents or several people on one project points to shared memory.
  2. Which tools must reach it? If you use more than one agent, the memory has to live outside all of them, reachable through files every tool reads or through a protocol such as MCP.
  3. How will a fact stop being served? Look for invalidation, supersession, expiry or at least easy editing. If the answer is "someone will notice", plan for the day nobody does.
  4. **Should writing be automatic or deliberate?**Automatic extraction captures more with no effort. Deliberate, typed entries capture less and are easier to trust.
  5. Where may the data live? Local, self-hosted or a managed service, depending on what your project allows.

Where Stele fits

We build Stele, so weigh this paragraph accordingly. Stele is a shared project record for AI coding agents: typed decisions, lessons, risks and tasks, linked to each other and to the code, that Claude Code, Cursor, Codex and any MCP client read before they act and write back to as they learn. A fact can be superseded by a newer one, expire on a date, or retire when its task closes. It is a hosted service, it asks agents to write entries on purpose rather than capturing everything, and it is young. The free plan covers public projects and one private one; pricing has the rest, and getting started takes a few minutes.

Frequently asked questions

What is AI agent memory?

AI agent memory is information that survives the end of a session and comes back into a later one, either because the agent retrieves it or because the system loads it for the agent. It includes instruction files, saved session history, facts extracted from conversations, searchable embeddings, knowledge graphs and shared project records.

Is the context window the same as memory?

No. The context window is the text the model sees on the current call. It is cleared when the session ends, every token in it is paid for on every turn, and very long contexts are followed less reliably. Memory is what lets the window stay small by bringing back only what the current step needs.

What are the main types of AI agent memory?

In running systems: working memory (the current context), session history (stored transcripts), instruction files (such as CLAUDE.md or AGENTS.md), extracted memories (facts a model pulls from conversations), semantic recall (vector search by meaning), knowledge graphs including temporal graphs, and shared project records that several agents and people read and write.

Can multiple AI agents share the same memory?

Yes, if the memory lives outside any single agent. A committed instruction file is shared through git. Some frameworks attach one memory block to several agents. A memory server reached over MCP can be read and written by agents running in different tools. Memory stored inside one tool or on one machine is not shared.

How do you stop agent memory from going stale?

Give every fact a way to stop being served. Temporal graphs mark an old fact invalid while keeping its history, extraction systems let you update or delete a memory, and some project records retire a fact when a newer decision replaces it, when a date passes, or when the related task closes. A plain file needs a person to notice and edit it.

  • guide
  • agent memory
  • ai agents

Try Stele

If the memory you need is one project record shared by every agent and every person on the project, Stele is free to start.

Start freeRead the docs

Claude Code memory: how CLAUDE.md and auto memory work, and where they stop

Source: https://stele-ai.dev/blog/claude-code-memory-explained

How it works

Claude Code memory: how CLAUDE.md and auto memory work, and where they stop

Every Claude Code session starts with an empty context window. Two built-in mechanisms carry knowledge from one session to the next: CLAUDE.md files that you write, and auto memory that Claude writes for itself. Here is exactly what each one loads, how to use both well, and the specific places they stop helping.

September 30, 2026·Stele·10 min read

Two mechanisms, two authors

Claude Code has two memory systems, and the difference between them is who writes them. You write CLAUDE.md files. They hold instructions: build commands, conventions, where things live, rules you would otherwise repeat in chat. Claude writes auto memory. It holds what Claude learned from working with you: your corrections, your preferences, and project context it could not work out from the code.

Both load at the start of every conversation, and both arrive as context. That second point matters more than it sounds. Anthropic's documentation is explicit that Claude treats these files as guidance it tries to follow, with no guarantee of strict compliance. If something must happen every time, such as a check before each commit, the documented answer is a hook, which runs as a shell command whatever Claude decides. Source.

Everything below was checked against those docs on 30 September 2026. Claude Code ships often, so the version numbers mentioned here will age; the shape of the system has been stable.

What CLAUDE.md loads, and in what order

A CLAUDE.md file is plain markdown. Claude Code looks for it in four scopes, and loads them from the broadest to the most specific, so the most specific instructions are the last thing Claude reads.

ScopeLocationShared with
Managed policy/etc/claude-code/CLAUDE.md on Linux; a system folder on macOS and WindowsEveryone on the machine, set by IT
User~/.claude/CLAUDE.mdJust you, in every project
Project./CLAUDE.md or ./.claude/CLAUDE.mdYour team, through git
Local./CLAUDE.local.md, git-ignoredJust you, in this project

Within a project, Claude Code walks up the directory tree. Start it in foo/bar/ and it loads foo/CLAUDE.md, then foo/bar/CLAUDE.md, with any CLAUDE.local.md appended after the CLAUDE.md at the same level. The files are concatenated. Nothing overrides anything, which means two files that disagree leave Claude to pick one, and the docs say it may pick arbitrarily. CLAUDE.md files in subdirectories below where you started load later, on demand, when Claude reads a file in that directory.

Three more pieces round it out. A CLAUDE.md can pull in other files with @path/to/file imports, up to four hops deep; the imported files load at launch too, so imports organise a long file without making it any cheaper. Files in .claude/rules/ split instructions by topic, and a rule with a paths pattern in its frontmatter loads only when Claude touches a matching file. And since version 2.1.277, Claude Code reads a repository's AGENTS.md directly when there is no CLAUDE.md on the path, so a repo set up for Codex or Cursor works without a second file.

To see what actually loaded in a session, run /context and look under Memory files. That one command settles most "why is Claude ignoring my instructions" questions, because the usual answer is that the file never loaded.

What auto memory is, and where it lives

Auto memory is on by default. As Claude works, it saves notes of four kinds, recorded as a type field in each file's frontmatter: user for your role and working preferences, feedback for corrections you gave and approaches you confirmed, project for ongoing work and decisions the code does not show, and reference for where to find things outside the repo, such as a dashboard or an issue tracker. Claude deliberately skips what it can read from the codebase, and anything your CLAUDE.md already says.

the memory folder for one repository

~/.claude/projects/<project>/memory/
├── MEMORY.md            # index, one line per memory
├── user_role.md         # one memory per file
├── feedback_testing.md
└── ...

The <project> name comes from the git repository, so every worktree and subdirectory of one repo shares one memory folder. At the start of each conversation Claude Code loads the first 200 lines or 25KB of MEMORY.md, whichever comes first. The topic files stay on disk until Claude decides it needs one and reads it. When the index nears its limit, Claude Code tells Claude to shorten it; when it goes over, whatever sits past the limit is dropped from the next load.

You can read and edit all of it. Run /memory to open the folder, or to switch auto memory off, which you can also do per project with "autoMemoryEnabled": false in that project's settings. Recent versions also stamp a modified timestamp into each memory file's frontmatter when Claude writes it, so you and Claude can see how old a note is.

The line in the docs worth remembering: auto memory is machine-local. Files are not shared across machines or cloud environments.

What belongs where

The docs give a good test for CLAUDE.md: write down what you would otherwise re-explain. Add a line when Claude makes the same mistake a second time, when a review catches something Claude should have known about the codebase, or when you type a correction you already typed last session. Keep it to facts that matter in every session, and write them so they can be checked.

a CLAUDE.md that earns its place

# Build and test
- Install with `pnpm install`. Never use npm in this repo.
- Run `pnpm test` before committing. The DB tests need `docker compose up db`.

# Conventions
- API handlers live in src/api/handlers/, one file per route.
- Money is stored in integer cents. Never use floats for amounts.

# Decisions
- Background jobs go through the jobs table.
  We tried SQS in 2025 and removed it: see docs/adr/007.

Anthropic suggests keeping each CLAUDE.md under about 200 lines, because a longer file costs context on every turn and lowers how reliably Claude follows it. Claude Code skips a CLAUDE.md larger than 4 MiB entirely. When a file grows, move the parts that only matter in one area into a path-scoped rule, and move multi-step procedures into a skill, which loads only when it is needed.

A rough sorting guide:

  • Project CLAUDE.md: commands, layout, conventions, and decisions every contributor must respect.
  • CLAUDE.local.md or your user CLAUDE.md: your sandbox URLs, your editor habits, your preferred test data.
  • .claude/rules/ with paths: rules for one part of the tree, such as the migrations folder.
  • Auto memory: leave it to Claude. When you say "remember that the API tests need a local Redis", it goes here. If you want it in CLAUDE.md instead, say "add this to CLAUDE.md".

Once a month, run /doctor prompt-audit. It reads your instruction files, flags references to files that no longer exist and instructions that contradict each other, and proposes edits without changing anything until you approve them.

Where it stops

Used this way, the built-ins cover a solo developer on one project very well, and they cost nothing. They run out in five specific places, and it helps to know which one you are hitting.

Auto memory stays on one machine. Open the same repo on a second laptop, in a cloud dev environment, or in a teammate's checkout, and none of what Claude learned comes along. The committed CLAUDE.md travels; the learned notes do not.

It stays in one tool. Cursor and Codex do not read Claude Code's auto memory. AGENTS.md now gives instructions a shared home across tools, which is a real improvement, but only for what you write by hand. What one agent learned on Tuesday is invisible to a different agent on Wednesday.

Nothing marks a line as no longer true. In March you write that jobs go through the queue table. In June you remove the queue table. The CLAUDE.md line keeps loading until someone notices and deletes it, and the agent reading it has no way to tell that it is out of date. The modified timestamp on auto memory and the prompt audit both help you find stale lines. Neither retires one on its own. We wrote about this failure in detail in why markdown rot breaks agent memory.

It loads whether or not it is relevant. Every line of the project CLAUDE.md and the first 200 lines of the memory index arrive in every session, whatever the task. Path-scoped rules soften this for code areas. There is no equivalent for a decision that matters to one feature and no other.

Team sharing means pull requests. A committed CLAUDE.md is shared, which is its great strength: it is reviewed with the code it describes. It also means every lesson one person's agent learned needs someone to write it down, open a pull request, and get it merged before anyone else's agent benefits. In practice most of it never makes the trip.

The built-ins answer "what should Claude know in this repo". They do not answer "what did the last agent learn, and is it still true".

When you outgrow it

Most people never need to. If you work alone, in Claude Code, on one machine, a tidy CLAUDE.md plus auto memory is the right setup, and adding a tool on top would be complexity without payoff.

If you hit one of the limits above, the fix depends on which one. Local plugins such as claude-mem or agentmemory capture sessions automatically and can share one store between several agents on the same computer. Hosted memory services such as Mem0 or Supermemory follow you across machines, and some offer a pool a team shares. Tools that keep a reviewable project record, such as ByteRover or Stele, aim at teams who want what agents learn to be visible and correctable. We compared eleven of them, with prices and sources, in the best memory tools for AI coding agents.

Whichever you try, run one test first: change a fact, start a fresh session a day later, and ask about it. A memory that serves the old fact next to the new one without saying which is current will eventually mislead an agent at a bad moment.

Where Stele fits

This is ours, so weigh it accordingly. Stele is a shared project record that sits outside any one tool. Claude Code, Cursor, Codex and other MCP clients read it before they act and write back decisions, lessons, risks and tasks as they learn them. A fact can be superseded by a newer one, expire on a date, or retire when the task it describes closes, so the agent is not handed last quarter's rule as today's. It works alongside CLAUDE.md: keep your build commands and conventions in the file, and let the record hold what changes.

It is a hosted service, it asks agents to write facts deliberately, and it is young. The free plan covers unlimited public projects and one private project; Pro is $12 a month, with reads never metered. The getting started guide walks through setup, and pricing has the limits.

Frequently asked questions

Does Claude Code remember previous sessions?

Not by itself. Every session starts with a fresh context window. Two mechanisms carry knowledge forward: CLAUDE.md files, which you write and which load at the start of every session, and auto memory, notes Claude writes for itself in a folder under ~/.claude/projects/ and loads at the start of each session.

What is the difference between CLAUDE.md and auto memory?

You write CLAUDE.md; Claude writes auto memory. CLAUDE.md holds instructions such as build commands, conventions and project layout, and a project CLAUDE.md is shared with your team through git. Auto memory holds what Claude learned from you, such as corrections and preferences, and it stays on the machine where it was written.

Where does Claude Code store auto memory?

In ~/.claude/projects//memory/, where the project name comes from the git repository. The folder has a MEMORY.md index plus one file per memory. The first 200 lines or 25KB of MEMORY.md load at the start of every conversation; the other files are read when Claude needs them.

How long should a CLAUDE.md file be?

Anthropic recommends keeping each CLAUDE.md under about 200 lines, because longer files cost more context and lower how reliably Claude follows them. Claude Code skips a file larger than 4 MiB. Move instructions that only matter for part of the codebase into path-scoped rules under .claude/rules/.

Can Cursor or Codex use Claude Code's memory?

They can share instructions if you keep them in an AGENTS.md file, which Codex and Cursor read and which recent versions of Claude Code read when there is no CLAUDE.md. They cannot read Claude Code's auto memory, which stays local to Claude Code on one machine. Sharing what agents learn across tools needs a memory store outside all of them.

Does Claude Code share memory with my team?

Only through a project CLAUDE.md committed to git, which teammates get when they pull. Auto memory is machine-local and personal, so what your agent learned does not reach a teammate's agent unless someone writes it into a committed file.

  • how it works
  • claude code
  • agent memory

Try Stele

If the limit you hit is sharing memory across tools, machines, or people, Stele is free to start.

Start freeRead the docs

The best memory tools for AI coding agents in 2026

Source: https://stele-ai.dev/blog/best-memory-tools-for-ai-coding-agents

Comparison

The best memory tools for AI coding agents in 2026

Every coding agent starts a session knowing nothing about your project. Eleven tools now promise to fix that. They store different things, in different places, for different readers, and the right one depends on a question most lists never ask: who else needs to read the memory?

September 29, 2026·Stele·17 min read·Updated September 30, 2026

How we compared them

We make one of the tools on this list, so read this with that in mind. To keep ourselves honest we held every entry to three rules. Each fact about another product comes from that product's own website, documentation or repository, checked on 29 September 2026, and each entry links its source. Each entry opens with what the tool is genuinely good at. And where we could not confirm something, we left it out.

If you want the categories compared before the products, start with Stele vs the usual ways to give agents memory. Prices and GitHub star counts change monthly. Treat them as a snapshot and follow the links before you commit to anything.

The list is limited to tools a developer can connect to Claude Code, Cursor, Codex or a similar coding agent today, through a plugin, an MCP server, or files the agent already reads. General memory infrastructure for building chat products is covered only where it also ships a coding-agent integration.

The question that sorts them: where does the memory live?

Almost every product here uses the word memory, so the word does not help you choose. Where the memory is stored does. It decides who can read it, whether it survives a new laptop, and whether a teammate's agent benefits from what yours learned yesterday. The eleven options fall into four places.

  1. Files in the repository, read by the agent at the start of a session: CLAUDE.md, AGENTS.md, Cursor rules. Shared through git, kept true by hand.
  2. A local database on your machine, filled by a plugin that watches your sessions: claude-mem, agentmemory, Basic Memory in its local mode. Private and automatic, and it stays on that one computer.
  3. A hosted memory service that extracts facts from conversations and returns the relevant ones: Mem0, Supermemory, Zep, Cognee's cloud. Reachable from anywhere, usually priced by volume.
  4. A shared project record that agents and people read and write on purpose, with typed entries and a lifecycle: ByteRover's context tree and Stele sit closest to this.

Letta Code is the outlier: it is a whole coding agent with memory built in, so choosing it means choosing a different agent.

At a glance

ToolWhere memory livesWorks withTeam sharingLicenseStarts at
Built-in files and memoryRepo files, local folderEach tool reads its ownFiles via gitPart of the toolFree
claude-memLocal SQLiteClaude Code, Cursor, Codex, othersHosted add-onApache-2.0Free locally
agentmemoryLocal serverClaude Code, Codex, any MCP clientAgents on one machineApache-2.0Free
Basic MemoryMarkdown plus SQLiteAny MCP clientTeams cloudAGPL-3.0Free locally
Mem0Hosted vectors, graph on ProClaude Code, Codex, Cursor pluginsShared project poolApache-2.0 coreFree tier
SupermemoryHosted graph memoryClaude Code, Cursor, Codex, OpenCodePer projectMITFree credits
Zep and GraphitiTemporal knowledge graphClaude Code, Codex, Cursor via MCPPer userGraphiti Apache-2.0Free tier
CogneeGraph plus vectorsClaude Code, Codex, any MCP clientWorkspacesApache-2.0Free locally
ByteRoverVersioned context files22+ agents via MCPCloud syncElastic 2.0Free tier
Letta CodeGit-backed memory filesLetta's own agentTeams planApache-2.0Free tier
SteleHosted project graphClaude Code, Cursor, Codex, Copilot, any MCP clientOne record per projectHostedFree plan

What they cost

Pricing in this category comes in three shapes, and the shape matters more than the headline number. Local tools cost nothing beyond your own machine. Most hosted memory services meter usage: retrievals, credits, tokens processed, or episodes stored. A few charge a flat price per person. Metered pricing is cheap while you experiment. It gets harder to predict once every prompt from every agent triggers a memory search, because the bill grows with how much your agents work.

ToolFreeFirst paid planWhat the price scales with
Built-in files and memoryIncludedn/aNothing extra
agentmemory, Cognee (local)Yes, self-runn/aYour own hardware
SteleUnlimited public projects, 1 private$12/mo (Pro)Flat per person; writes capped, reads unmetered
Basic Memory CloudLocal only$15/moFlat
ByteRover100 context files$15/mo, billed yearlySynced context files
Mem010k adds, 1k retrievals/mo$19/mo; graph memory at $249/moAdds and retrievals
Supermemory$5 of credits$19/moUsage credits
LettaYes$20/moPlan plus model usage
claude-mem hostedTrial$20 to $30/moFlat per person
Zep10k credits, 1 MCP seat$125/mo (Flex)Credits per stored episode
Cognee cloud1M tokens$1 per 1M tokensTokens processed

Two things stand out. For a solo developer who wants hosted, shared memory across machines and tools, Stele's $12 Pro plan is the lowest first paid tier on this list. And it is one of the few that does not charge by how often your agents look something up: reads and search are not metered on any plan, so an agent that checks memory before every change costs the same as one that never does. That matters, because checking before acting is the behaviour you want to encourage.

Pro is also sized so most people never hit its limits: unlimited private projects, 3,000 writes a month, and 50,000 stored entries (the full limits are on the pricing page). For scale, we build Stele with Stele, with agents working in it every day. Over three months our own project gathered about 1,300 tasks and 940 active facts, roughly 750 writes a month, a quarter of the Pro allowance. Editing an existing fact does not count as a write. How we build Stele with Stele has the full numbers.

For teams the arithmetic depends on size. Stele's Team plan is $20 per person with no limit on agents, so a team of five pays $100 a month. Zep's Flex plan is $125 a month for five MCP seats, and Mem0 puts graph memory behind its $249 Pro plan. A team whose usage fits inside Mem0's $19 Starter limits will pay less there. And open-source projects pay nothing on Stele: public projects are free and unlimited.

1. The memory already built into your agent

Genuinely good at: costing nothing, needing no account, and being reviewed in the same pull request as the code it describes. Before you install anything, know what you already have.

Claude Code reads CLAUDE.md files at several scopes: an organization-wide managed file, a project file committed to git, a personal file, and a git-ignored CLAUDE.local.md. Separately, its auto memory writes notes to a folder under ~/.claude/projects/. That folder is on by default and is, in Anthropic's words, machine-local and not shared across machines. Only the first 200 lines of its index load at the start of a session. Source.

Codex reads AGENTS.md files from your home folder and then from the repository root down to the working directory, up to 32 KiB combined by default. Its separate memories feature is off by default, and when on it stores notes locally per user. Source.

Cursor uses rules: project rules in .cursor/rules, user rules, team rules on its business plans, and AGENTS.md. Its documentation puts it plainly: language models do not retain memory between completions, and rules provide the persistent context. Source.

Where it stops: the committed files are shared, but only as well as someone keeps them. Nothing tells you when a line has stopped being true, and the whole file loads whether or not it is relevant. The automatic memories are the reverse: they update themselves, but each one stays inside the tool and the machine that wrote it. Switch from Claude Code to Codex for an afternoon and the second agent knows none of what the first one learned. We wrote about that failure in detail in why markdown rot breaks agent memory, and walked through exactly what Claude Code loads in Claude Code memory, explained.

Choose it if you work alone, in one tool, and can keep one file honest. That covers many projects.

2. claude-mem

Genuinely good at: capturing everything without you doing anything. claude-mem installs hooks that record what happens in a session, compress it, and inject relevant history into the next one. With roughly 95,000 GitHub stars it is by far the most adopted project in this category.

It stores memory in local SQLite with full-text search, with optional Chroma vectors. It started with Claude Code and now also supports Cursor, Windsurf, OpenCode and Codex CLI. It is Apache-2.0. The project also sells a hosted tier, and its README notes that some integrations default to it; the two official pricing pages we found listed different monthly prices, so check before you sign up. Source.

Where it stops: it records what happened, which is different from what was decided. A compressed session log answers "what did the agent do on Tuesday" well and "what is the rule for database migrations here" less well. And in its local mode, it is one person's memory.

Choose it if you are a single developer who wants continuity across sessions with zero effort.

3. agentmemory

Genuinely good at: being one local memory that every agent on your machine shares. agentmemory runs a local server, ships a Claude Code plugin with a dozen hooks and a Codex plugin, and exposes more than fifty MCP tools for everything else. Search is BM25, with optional embeddings and a graph. Apache-2.0, around 29,000 stars. Source.

Where it stops: the sharing ends at your machine. That is a feature if privacy is the point and a limit if a teammate's agent should benefit from what yours learned.

Choose it if you switch between several agents on one computer and want them to stop contradicting each other.

4. Basic Memory

Genuinely good at: keeping memory in files you can open, read and edit. Basic Memory writes plain markdown, links notes with wikilinks into a knowledge graph, and indexes them locally in SQLite with hybrid full-text and vector search. Any MCP client can use it. Local use is free under AGPL-3.0; a cloud version with a shared Teams workspace costs $15 a month in its current beta pricing. Source.

Where it stops: it is a general notes system with an agent interface, so project concepts such as tasks, owners or decisions that replace earlier decisions are conventions you build yourself.

Choose it if you want an agent-readable second brain that stays human-readable and portable.

5. Mem0

Genuinely good at: being a mature, widely used memory layer. Mem0 is the most starred general memory library, at about 66,000 stars, and its core is Apache-2.0. A model reads each conversation and extracts the facts worth keeping, so you store short memories instead of whole transcripts.

Most of Mem0's users build it into their own AI products through its Python and TypeScript SDKs. For coding agents, it now publishes an official Claude Code plugin, with plugin folders for Codex and Cursor. The plugin keeps a personal memory and a shared project memory that everyone on the repository reads and writes. It needs a Mem0 Platform API key. Pricing starts with a free Hobby tier, then $19 a month for Starter; graph memory arrives with the $249 Pro plan. Mem0 retired its earlier OpenMemory MCP server in July 2026. Source, pricing.

Where it stops: memories are extracted from conversation, so what lands in the pool is whatever the extractor judged worth keeping. That works well for preferences and facts. It gives you less control over the shape of a decision and when it should stop applying. Mem0's current documentation describes adding as additive: new memories never overwrite or delete old ones, and a correction is a separate, explicit update call. Source.

Choose it if you already use Mem0 in a product, or you want automatic, team-shared recall with minimal setup.

6. Supermemory

Genuinely good at: fast, low-friction recall across many agents. Supermemory ships plugins for Claude Code, Cursor, Codex and OpenCode that search memory on each substantive prompt and give up after three seconds rather than stall the agent. Memories are organized by project. The plugins stopped requiring a paid plan in July 2026; the free tier includes $5 of credits and Pro is $19 a month. The engine is MIT-licensed. Source, pricing.

Where it stops: self-hosting is listed under the Scale and Enterprise plans, so the practical default is their cloud. Like Mem0, it stores what it extracts, with less structure around what kind of fact each memory is.

Choose it if you want hosted memory in several agents quickly and the credit pricing fits your volume.

7. Zep and Graphiti

Genuinely good at: time. Graphiti, Zep's open-source engine, is a temporal knowledge graph: every fact carries when it became true and when it stopped, so an old fact is marked invalid instead of silently deleted. For questions like "what did we believe in March" that is the most careful model on this list. Graphiti is Apache-2.0, around 31,000 stars, and runs on Neo4j, FalkorDB or Neptune. Source.

Zep is the hosted service built on it. It offers a Memory MCP server and a plugin for Claude Code, Codex and Cursor; that memory belongs to the signed-in user. The free tier includes 10,000 credits and one MCP seat; the Flex plan is $125 a month. Zep's own community edition is no longer supported, so self-hosting means running Graphiti yourself. Pricing.

Where it stops: Zep is built first for companies putting memory into their own products, and its pricing reflects that. For a small team's coding agents it can be more machinery than the job needs.

Choose it if facts in your domain change often and you need to reason about their history.

8. Cognee

Genuinely good at: turning a pile of documents and code into a queryable graph. Cognee combines a graph and a vector store, defaulting to embedded local databases so it runs with no servers, and it offers an MCP server plus Claude Code and Codex plugins. Apache-2.0, around 31,000 stars. Its cloud charges per million tokens processed after a free allowance. Source, pricing.

Where it stops: Cognee is a general memory and knowledge-graph platform. Using it as a coding agent's project memory is one configuration among many, which means more setup decisions are yours.

Choose it if your agents need to reason over a large body of existing documentation.

9. ByteRover

Genuinely good at: treating context like code. ByteRover's brv command curates a tree of context files and versions it with branches, commits, pushes and pulls, so a team can review changes to what agents know. It connects to more than twenty agents over MCP. The free tier syncs 100 context files; Pro is $15 a month billed yearly. It is source-available under the Elastic License 2.0, which is not an OSI open-source license. Source, pricing.

Where it stops: version control tells you who changed a context file and when. It does not by itself tell you which entries are still true.

Choose it if your team already thinks in pull requests and wants the same review loop for agent context.

10. Letta Code

Genuinely good at: memory as a first-class part of the agent. Letta, formerly MemGPT, lets the agent manage its own memory explicitly. Letta Code is a full coding agent built on it, with a git-backed memory filesystem the agent edits and a background process that consolidates what it learned. Apache-2.0, and self-hostable. Pricing has a free tier and a $20 Pro plan, though Letta's own estimate puts regular coding use at around $100 a month of usage. Pricing, docs.

Where it stops: the memory belongs to Letta's agent. If your team uses Claude Code, Cursor and Codex, Letta Code is a fourth agent to adopt, and its memory will not follow you back to the other three.

Choose it if you are willing to switch agents for a memory-first design.

11. Stele

This is ours, so apply the same scepticism you would to any vendor describing itself.

Genuinely good at: one project record that every agent and every teammate shares, whatever tool they use. Stele stores typed entries: decisions, lessons, risks, tasks, and the components they belong to, linked to each other. Claude Code, Cursor, Codex, GitHub Copilot, OpenCode and any other MCP client read it before they act and write back what they learned. Each fact has a lifecycle: it can be superseded by a newer decision, expire on a date, or retire when the task it describes closes, so an agent is not served yesterday's rule as today's. Tasks are claimed atomically, so two agents do not pick up the same work. People see the same record in a web app. The free plan covers unlimited public projects and one private project; Pro is $12 a month and Team is $20 per person, with reads never metered. See pricing.

Where it stops: Stele is a hosted service and there is no self-hosted version today. It asks agents to write facts on purpose rather than capturing everything, which produces a cleaner record and costs a little effort per session. And it is young: it is in its first public version, used by a small number of teams.

Choose it if more than one agent or more than one person works on the same project, and you want what one of them learned to reach the others. For the longer argument, see switch agents, keep the project and a file versus a shared ledger.

Which one fits your situation

  • One developer, one agent, one machine: start with the built-in files. Add claude-mem if you want sessions captured automatically.
  • One developer, several agents: agentmemory keeps them consistent on one computer. A hosted option such as Supermemory or Stele also covers a second machine.
  • A team that wants automatic recall: Mem0's shared project pool or Supermemory's project memory.
  • A team that wants a reviewable record: ByteRover if you want it versioned like code, Stele if you want typed decisions and tasks with a lifecycle.
  • Facts that change and history that matters: Graphiti or Zep.
  • Human-readable notes first: Basic Memory.
  • Willing to change agents: Letta Code.

How to tell whether memory is helping

Every tool on this list can produce a demo where memory helps. Fewer can show that it helps on your project. Three checks take an afternoon and tell you more than any benchmark on a vendor's site, including ours.

First, look at what the tool actually injects. Most have a log or a debug mode. Count how many injected memories were relevant to the task. If most are noise, the agent is paying tokens to read them and may act on the wrong one.

Second, change a fact and see what happens. Tell the agent that the project moved from one database to another, start a fresh session a day later, and ask about the database. A tool that serves the old fact next to the new one without saying which is current will eventually mislead an agent at the worst moment.

Third, switch tools. Do some work in Claude Code, then open the same repository in Cursor or Codex and ask what was done. This is the check that separates memory that belongs to a tool from memory that belongs to the project.

The useful question is whether what one agent learned reaches the next agent that needs it, in a form it can trust.

We wrote up how we tried to measure this properly, and how many ways there are to measure nothing, in how do you measure whether agent memory actually helps.

Frequently asked questions

What is the best memory tool for Claude Code?

It depends on who needs to read the memory. For one developer on one machine, Claude Code's own CLAUDE.md files and auto memory cost nothing and are often enough. If you want sessions captured automatically, local plugins such as claude-mem or agentmemory do that on your machine. If several people or several different agents need the same memory, you need a shared store: Mem0's plugin, Supermemory, Zep, or a shared project record such as Stele.

Can Claude Code, Cursor and Codex share the same memory?

Yes, if the memory lives outside all three tools. A committed AGENTS.md file is read by Codex and Cursor, and recent versions of Claude Code read it too. For anything beyond a file, use a memory server that all three connect to over MCP or through a plugin. Built-in memory features such as Claude Code's auto memory or Codex memories stay inside the tool that wrote them.

Is a CLAUDE.md or AGENTS.md file enough?

For a small project with one person keeping it honest, often yes. It is free, it is reviewed in the same pull request as the code, and every tool can read it. It stops being enough when the file grows past what an agent should load every session, when several people add to it, or when nobody notices that a line has stopped being true.

What is the difference between agent memory and a vector database?

A vector database stores embeddings and returns the nearest matches to a query. An agent memory tool decides what to store, when to update or retire a fact, and what to hand the agent at the start of a task. Many memory tools use a vector database inside, alongside a graph, a full-text index, or plain files.

How do you keep agent memory from going stale?

Give every fact a way to stop being served. Tools do this differently: Zep's Graphiti records when a fact stopped being valid, Mem0 lets your code update or delete a stored memory explicitly, and Stele lets a fact be superseded, expire on a date, or retire when a task closes. A plain markdown file has no such mechanism, so a person has to notice and edit it.

Which memory tools for coding agents are open source?

As of September 2026: claude-mem, agentmemory, Mem0's library, Graphiti, Letta, Cognee and Supermemory are released under Apache-2.0 or MIT. Basic Memory is AGPL-3.0. ByteRover is source-available under the Elastic License 2.0. Zep's hosted service and Stele are hosted products.

Can a whole team share one project memory?

Yes. Mem0's Claude Code plugin has a shared project pool, Basic Memory has a Teams workspace, ByteRover syncs a context tree through its cloud, Supermemory organizes memory by project, and Stele is built around one project record that every member and every agent reads and writes. A committed file is also shared, through git.

  • comparison
  • agent memory
  • claude code
  • cursor
  • codex

Try Stele

If the memory you need is a shared project record that every agent and every teammate reads, Stele is free to start.

Start freeRead the docs

Anatomy of an agent memory benchmark

Source: https://stele-ai.dev/blog/anatomy-of-an-agent-memory-benchmark

Engineering

Anatomy of an agent memory benchmark

We open-sourced the harness we use to measure whether project memory helps a coding agent. Almost none of it is about running agents. Nearly all of it exists to defend one claim: that two runs differed in exactly one thing.

August 11, 2026·Stele·15 min read

The claim under the benchmark

A paired benchmark looks like a measurement and is actually an argument. You run a task twice, change one thing, and report the difference. Everything the number means rests on a premise you never state out loud: that the one thing you changed really was the only thing that changed.

Here is how quietly that premise dies. You pick a real repository to benchmark against, because benchmarking on your own codebase is worthless. You pin it to a commit, package it into a container, and run one agent with your memory system and one without. The repository ships a CLAUDE.md. The harness reads it into the system prompt before the first turn, in both arms. Nothing in your task definition mentions it.

Your no-memory arm now has a hand-written memory system. Your measured effect shrinks for a reason that has nothing to do with your product.

That is the normal case, not an exotic one. Commit history is another way in. So is a stray .orig file. The word fixture in a project description tells the agent it is inside an experiment, and an agent that knows it is being measured does not behave like one doing ordinary work.

Running the agents is the easy part. The apparatus is what you build so the comparison is allowed to mean something.

stele-bench is that apparatus, now MIT. This is a walk through its architecture, the failure each piece was built after, and what you would have to write to point it at a memory system that is not ours.

The shape of a run

One paired run is three phases. A shared preparation phase that produces byte-identical inputs, a split into two containers that differ in one respect, and a shared grading phase that applies the same graders in the same order to whatever comes back.

One paired run The bands are the point. Everything above and below the split is the same code path with the same inputs for both arms. The middle is the entire experiment. A per-arm branch anywhere else in the pipeline is a confound, however good the reason for it.

Each stage shuts one specific way the answer can get in. Building the tree from Git objects takes the history with it. Redaction removes hand-written project memory. Certification catches leakage and anything that announces the experiment. The mutation audit catches work outside the declared outputs. The last way in is us, which is what the admission decision is for.

Contracts before anything else

The base class every artifact inherits is four lines long and is the reason preregistration is enforceable at all.

contracts.py

class ContractModel(BaseModel):
    """Strict base class so an artifact does not silently change its meaning."""

    model_config = ConfigDict(extra="forbid", frozen=True)

Preregistration means freezing the prompt, the verifier, the model, the budget and the analysis before any money is spent, then declaring that nothing moved. A permissive parser makes that declaration unfalsifiable. If a challenge manifest can gain a field that older code ignores, then the manifest you ran is not necessarily the manifest you froze, and no digest in the world will tell you, because the digest covers bytes while the parser decides meaning.

The same instinct runs down to individual fields. Artifact paths are validated as portable relative POSIX paths, so a manifest cannot carry an absolute path, a Windows drive letter, a backslash, or a .. component. Every digest field is a compiled pattern rather than a string. This is what keeps a frozen artifact frozen when someone runs it on a different machine a year from now.

A corpus with no history

The evaluated project tree is built from Git objects and only from Git objects. It never inherits a caller's worktree, ignored files, build output, or .git directory.

Two reasons, and the second one cost us a paid run. The first is reproducibility: a tree assembled from a checkout carries whatever that machine happened to have lying around, and a benchmark corpus that differs per operator is not a corpus. The second is that repository history answers questions. A question like why does this keep breaking has its answer distributed across commit messages, and an agent with git log can go get it. One of our early cohorts was thrown out precisely because both arms went digging in history that should never have been in the container, which made the run informative about agent behaviour and useless as a memory comparison.

If your memory system's pitch is that it holds things the code does not say, then commit history is your competitor. Remove it on purpose, from both arms, and say in the writeup that you did.

Certification, and the file everybody forgets

Certification is a static, zero-model, fail-closed gate over every surface an evaluated agent can see: the source corpus, the evidence corpus, every declared graph arm, and the agent-visible text of the challenge itself. It runs before any state is created and before any tokens are spent, which is exactly why it must not need a model to run. It classifies what it finds into five kinds.

certification findings

auto-loaded-context     a file the harness reads into the prompt unasked
suspicious-metadata     .git, .env, credentials, runner-owned task wiring
benchmark-awareness     the packaged inputs announce the experiment
patch-or-diff           a .patch, .orig, or a unified diff hunk in the tree
solution-leak           a challenge's declared answer phrase, already present

The first kind is the one that is genuinely hard to see, because the file does its damage without ever appearing in your task definition.

certification.py

AUTO_LOADED_CONTEXT_PATTERNS: tuple[str, ...] = (
    "AGENT.md", "AGENTS.md", "CLAUDE.md", "CLAUDE.local.md",
    "GEMINI.md", "QWEN.md", "copilot-instructions.md",
    ".aider.conf.yml", ".clinerules", ".cursorrules", ".goosehints",
    ".mcp.json", ".windsurfrules",
    ".claude/", ".codex/", ".continue/", ".cursor/", ".gemini/", …
)

A bare name matches that basename anywhere in the tree; a trailing slash matches any path component, so .claude/ covers a whole skills directory. These files are hand-written project memory. Packaging one is packaging a rival memory system into every arm at once.

The benchmark-awareness probe is where a gate like this usually goes wrong, and the interesting engineering is in what the list leaves out.

certification.py

# Terms that would tell an agent it is inside an experiment. Deliberately
# narrow: each one is a harness or arm identity, not ordinary engineering
# prose, so a real repository does not trip them. Words like "condition",
# "arm", "trial", "synthetic" and "ground truth" are NOT here: they occur
# naturally in real project knowledge and would make certification cry wolf.

A gate that fires on ordinary words gets a suppression list, then a --force flag, then it is off. So the strict vocabulary is scoped rather than broadened. Harness identity terms are checked everywhere. A second, stricter set includes terms like fixture, control arm and negative control. Those are only checked against a graph's identity fields: the project, workspace and branch names the benchmark writes itself. They are never checked against node bodies.

That last part matters. A graph backfilled from a real project says spin up a temp directory with test fixtures in ordinary engineering prose. Flagging that is noise, and noise is what trains you to ignore the gate.

One more rule, which sounds pedantic until it saves you. A missing corpus, fixture artifact, or graph is an error, never a skipped check. A gate that quietly covers nothing still reports success, and that is worse than not having one. The integration test that certifies the real pinned corpus fails rather than skips when the checkout is absent, and the opt-out is an explicit environment variable that says in its own name that the corpus went unchecked.

Repair and inspection are different jobs

A pinned corpus belongs to somebody else. You cannot ask an upstream project to delete its CLAUDE.md so your benchmark works, so redaction has to exist. The design decision is that it is the only thing allowed to modify a packaged tree, and it is a separate function from the gate that inspects one.

Redaction removes auto-loaded context while packaging, identically for every arm, and writes a receipt naming each removed path with its digest. Certification then runs over the result and only ever reads. Because the two are separate, an unredacted or hand-edited tree cannot reach an agent: there is no code path where the gate notices a problem and fixes it. A gate that repairs what it finds is a gate you can never trust to have found nothing.

The boundary that makes it yours

Everything so far is generic. The part that decides whether stele-bench is a benchmark or a Stele benchmark is one adapter, and it is the thinnest component in the system.

The provisioner boundary Three verbs, one JSON object in each direction, no shell. The runner never learns a backend schema, which is what lets the same harness measure a memory system whose internals it could not parse.

You configure an argv prefix and a fixture root. The runner executes <prefix> restore, snapshot, or destroy through create_subprocess_exec, with no shell anywhere in the path. It writes exactly one compact JSON object to stdin and expects exactly one JSON object on stdout. Stderr is reserved for bounded operator diagnostics and is never parsed as data, so a chatty provisioner cannot accidentally become a data channel.

What it refuses is part of the contract too. It rejects a fixture artifact whose resolved path escapes the fixture root, an artifact whose SHA-256 does not match the manifest, a response that is not a single JSON object, a response carrying unknown fields, a non-zero exit, a timeout, and output past a byte limit. Each of those is a way a provisioner can be subtly wrong while looking fine.

The three verbs are the whole integration surface. To benchmark a different memory system you write one executable that speaks them:

what your executable has to do

$ your-provisioner restore   # stdin:  { fixture, graph_arm, run_identity,
                             #            artifacts: { evidence_corpus, graph } }
                             # stdout: { project, queue_before_count, restore_seconds }

$ your-provisioner snapshot  # stdin:  { project }
                             # stdout: { state, probes, queue_after_count, snapshot_seconds }

$ your-provisioner destroy   # stdin:  { project }
                             # stdout: { "ok": true }

Restore takes a frozen memory state and makes it live for one run. Snapshot independently re-reads what is there afterwards, including named full-text, semantic and graph retrieval probes, so the run can prove the memory it was supposed to have was actually present and retrievable rather than assumed. Destroy tears the state down in a finalizer.

The snapshot receipt matters more than it looks. Without it, a treatment arm that silently restored an empty graph produces a perfectly clean null result, and you publish memory did not help when the truth is that memory was never there.

A challenge and a fixture are different objects

A challenge owns the question: the visible prompt, the allowed output paths, resource limits, and a digest-pinned verifier. A fixture owns the evidence: the frozen corpus and the graph state handed to one condition. They are independent on purpose, so the same question can run against a no-memory arm and a restored graph without the condition ever appearing in anything the agent can read.

challenges/settings-label-roundtrip/challenge.json

{
  "schema_version": 1,
  "challenge_id": "settings-label-roundtrip",
  "prompt": "On the Settings screen, rename the row currently labeled
             \"Check for Updates\" to \"Automatic Update Checks\" …",
  "allowed_output_paths": [
    "/workspace/src/screens/settings.rs", "/workspace/tests",
    "/workspace/Cargo.lock", "/workspace/target"
  ],
  "agent_timeout_seconds": 900,
  "environment": {
    "network_mode": "public", "cpus": 2, "memory_mb": 4096,
    "prebuild_argv": ["cargo", "test", "--no-run", "--lib"]
  },
  "verifier": {
    "kind": "script", "runtime": "posix-shell",
    "asset": { "path": "verifier.sh", "sha256": "7aba6b60…" },
    "timeout_seconds": 900
  }
}

The verifier is runner-owned. It lives outside the evaluated project tree and is never included in the prompt, so grading a task and performing it cannot see each other. Resource limits live in the manifest because they are part of the experiment. CPU count, memory, and whether the project was prebuilt all change the task. Two of our early runs were invalidated by exactly that.

The catalog spans several kinds of memory value deliberately, rather than repeating one flattering pattern. Investigation questions grade whether the agent reaches a grounded answer and how much source exploration it needed. Focused coding questions exercise hidden cross-file contracts. Continuity questions use a follow-up prompt, which produces a two-step task sharing one verifier, scoring only the final step, with the agent resuming its own native session across the boundary. And some challenges are low-memory controls, chosen because memory should not help, so the suite measures its own overhead instead of quietly excluding the cases that make it look bad.

Grading in three layers, and who wrote the sentence

A single pass or fail cannot tell you whether the memory was delivered, whether the agent used it, or whether it changed the result. So every run reports those three separately, and none of them implies the next. A correct treatment arm with delivered memory is not evidence that memory caused the success, because the control may well have succeeded too.

The first layer is where naive implementations quietly cheat. The obvious way to detect delivery is to search the transcript for a distinctive phrase from the memory. The problem is that a transcript contains the agent's own words.

memory_evidence.py

# A match on an "agent"-sourced step never counts: the agent authoring
# that exact phrase is not evidence the memory system delivered it.
_NON_AGENT_SOURCES = frozenset({"user", "system"})

So only non-agent steps are eligible, and on the harness we run against, injected context arrives as a specific attachment type carrying a provenance comment that names the channel it came through. Prompt-time retrieval and anchor-triggered recall are two different mechanisms, and naming the channel keeps them apart. A plain delivered-or-not boolean would report them as the same thing.

The middle layer is a lower bound

Observable use is not a count. A model can read a fact, internalise it, and produce a better answer without ever referring to it. So the honest phrasing is injected, with no observable trace. Writing injected and ignored would claim to know something the transcript cannot show.

The mutation audit, and the lockfile that voided a run

Before the challenge verifier runs, a hidden generic audit compares the finished workspace against a snapshot taken before the agent started. Anything added, removed, changed, permission-modified, symlink-retargeted or type-swapped outside the declared output subtrees forces reward zero. The receipt is deterministic and stays in the runner-owned area, outside anything the agent can see or edit.

This is the check that catches an agent solving the task by editing the test, and it is also the check that will fail you for a reason you did not think of. One of our runs was excluded because ordinary cargo execution wrote a lockfile that the allowlist did not declare. The agent did nothing wrong. The toolchain did what toolchains do, the audit correctly reported an undeclared mutation, and the run was thrown out.

The lesson generalises past Rust: your output allowlist has to describe what your toolchain writes, not what you intend to write. That is why the manifest above lists Cargo.lock and target alongside the source file the task is actually about.

The money fence

A paid agent benchmark spends real money on a shared balance, and the usual protection, a per-run cost cap, does not protect the thing you care about. A cap bounds what one run may spend. It says nothing about what is left for everything else.

We know that because of an incident the fence is named after: a run that stayed comfortably under its own budget and drained the shared balance anyway, which on our setup means taking the production assistant down with it. So the preflight now checks four things in order, before a paid pair may start.

before a paid pair may spend anything

1  key identity     the exported key matches a configured SHA-256 fingerprint,
                    so a run cannot be billed to the production key by accident
2  per-key budget   the quota API is asked directly whether the eval key really
                    has an external ceiling; an absent budget fails
3  shared headroom  team balance must cover a production reserve PLUS the
                    acknowledged pair cap
4  observed drain   between arms, re-read the balance and stop if more drained
                    than was declared

Step four exists because the first three prove the pair fits, and none of them can prove the pair behaves. It also deliberately over-attributes: observed drain is the drop in the shared balance, so it counts spending by anything else on the account too. Whoever is spending it, the next dollar past that line comes out of the reserve, and stopping is the right response either way.

The two failure modes are deliberately asymmetric. A transport error on the balance check fails open, and records that it was unobserved. One flaky GET should not abandon a pair that has already been paid for, and a check that did not run is fine as long as the record says so. A rejected key fails closed. The preflight already proved that key was good, so a 401 now means it changed underneath the run, and continuing would spend against something nobody verified.

Statistics for small samples

Agent benchmarks are expensive, which means small samples, which is exactly where the textbook approximations misbehave. Everything in the statistics module is exact: integer combinatorics, no sampling, no normal approximation to a binomial tail, no dependency beyond the standard library. The same inputs always produce the same numbers, so a receipt can be re-derived from its own recorded attempts.

Three of the rules are worth repeating whatever you build on. Do not use a Wald interval at 0/n or n/n, where it collapses to zero width and claims certainty from five observations. Use paired methods, because the arms ran the same tasks and a hard task is hard for both; subtracting two independent intervals throws that away and reports a range far wider than the data supports. And compute no significance test at all for a continuous measure below five attempts. Tokens and wall time at that sample size are noise wearing a decimal point, so the summary says exploratory and stops there.

What makes two runs comparable

Regression detection means plotting one question's score across runs, and that only means something when the runs asked the same question of the same model against the same memory. Change the verifier and the score moves without memory having changed at all. So every durable record carries an explicit comparability key: challenge, fixture, model, both conditions, the manifest and verifier digests, and the source commit. The viewer refuses to put two different keys on one line.

One input is deliberately left out of that key, and it is the build under test. That build is the thing whose effect on the score you are trying to see. Put it in the identity and every build becomes its own incomparable island, which turns regression detection off exactly when it would have told you something. It is recorded against each data point instead, so a step in a trend line can be traced to the build that caused it.

Admission is a decision, not a status

The final source of bias is us, and it is the hardest one to close with code. So the apparatus does the one thing it can. It refuses to let a finished run count as an admitted one.

A run that finishes is completed. Whether it may contribute to a published number is a separate, recorded decision with its own reason code. Excluded runs are retained in full, never repaired and never reinterpreted. Cleanup may reorganise files or redact host paths and credentials; it may not remove a run because its direction is unfavourable.

Preregistration versioning enforces the same discipline over time. An apparatus correction does not fix the old run, it creates a new preregistered version, and the earlier outcome stays in the record as excluded. That is why the published cohorts carry version numbers: a question that reached v4 is a question that failed three times in ways worth reading about, and the record says so. A publication test in CI checks that every admission has a matching durable record, that the manifest covers every published file with a digest, and that no host paths or secret-shaped strings made it into the release.

Pointing it at your own system

The repository is MIT and contains the full harness, the challenge definitions and their verifiers, the preregistrations, and every run we have published, accepted and excluded alike, with the trajectories and a manifest hashing every file.

To measure a different memory system you write one executable that speaks three verbs over JSON, and declare it as an argv prefix. The runner never learns your schema. That is what makes the harness portable instead of Stele-shaped. Everything else, the certification, the tare challenges, the exact tests, the admission split, applies unchanged, because none of it was ever about our product.

The preliminary results we ran through it are published separately, with every admitted and excluded run attached: when does project memory help a coding agent. Read the apparatus first. It is what makes those numbers worth reading.

  • engineering
  • benchmarks
  • open source

Try Stele

The apparatus is MIT and the memory system it was built to measure is the one you can install today.

Start freeRead the docs

Graph engineering for agent memory

Source: https://stele-ai.dev/blog/graph-engineering-for-agent-memory

Engineering

Graph engineering for agent memory

Half of graph engineering has real prior art and works. The other half only exists once the graph is something you write to instead of something you index, and it is the half that decides whether an agent gets helped or quietly misled.

July 28, 2026·Stele·9 min read

Two different jobs, one word

Graph engineering currently names two disciplines that share a shape and nothing else.

One is about how work runs. Nodes are units of work, edges are data dependencies, and the payoff is parallelism: find the dependencies that are not real, run the independent work at once, and the longest surviving chain bounds how fast the whole thing can finish. This is fifty-year-old machinery. Critical-path scheduling dates to 1957, make parallelised builds in the seventies, and every data pipeline since has been a directed acyclic graph. If that is your problem, the existing literature is excellent and this post is not about it.

The other is about what the agent knows. Nodes are facts, edges are relationships, and the payoff is retrieving the thing you would never have thought to search for.

That second one splits again, and this is the split that matters. Retrieving over a graph you built from a corpus has real research behind it and largely works. Maintaining a graph that accumulates from work does not, because a corpus is indexed once and a record keeps changing underneath you.

Part one: the retrieval graph

Start with the case that justifies the structure, because it is concrete and it is not about graph theory.

You ask an agent to change how checkout retries failed payments. A semantic search over your project notes finds the checkout decision, because your question and that note share obvious words. It does not find the note about the payment provider's idempotency keys, written eight months ago by somebody who never used the word checkout. That note is the one that matters: retry without idempotency keys and you double-charge people.

A graph finds it, because the checkout decision has an edge to it. You retrieved it by relationship rather than by similarity, and that is the whole argument for the extra structure.

The fact you most need is often the one that shares no words with your question.

The mechanics are not exotic and we would not pretend otherwise. Search produces seeds: about ten starting points, from keyword and vector matching run as separate channels, because they fail in different places. Then a walk goes outward from those seeds along typed edges, one hop, collecting what it reaches. Everything reached is scored, ranked, and capped.

One hop sounds timid. It is not: at depth two on a well-connected graph you reach most of it and the results turn to soup. Spend extra reach on better seeds rather than more depth.

The scoring, and the two weights that are scar tissue

weights that survived contact with real graphs

by category            by relation followed        distance
  decision      3        fulfills        3           −0.5 per hop
  architecture  3        blocks          3
  goal          2        supersedes      3          container nodes
  risk          2        caused-by       2           epic / objective
  lesson        2        contradicts     2           /milestone: −3
  gap           1        relates-to      1
  fyi           1        mentions        1

Container nodes score negative three. Anything that groups other things carries almost no information about the work in front of you, but children point at parents through strong edges, so containers arrive with high scores and flood the results with entries like Q3 Platform Work. A strong negative base cancels even a top-tier edge.

Structural nodes score zero. Ask about authentication and you match the component literally named Auth, which then takes a slot to tell you that a subsystem exists.

Both fixes are the same move, and it is the most repeated decision in this system: downrank, never exclude. An excluded node cannot come back on the day it genuinely is the answer. A downranked one can, and costs a slot only when nothing beats it.

Cap what you return. Ours is twenty. Injected context competes with the user's actual question for the model's attention, so forty entries is not twice as useful as twenty, it is twenty useful entries diluted by twenty distractions.

The inversion that eats your best results

There is a trap in this design that is easy to miss and expensive to leave in, and it is worth spelling out because it follows directly from the scoring above.

A seed arrives with no relation and zero distance, so its score is category weight plus recency and nothing else. A walked node arrives with an edge bonus. The cheapest possible edge adds one point and the distance penalty takes back half, so a neighbour outranks an identical seed by half a point, and by two and a half points through a strong edge.

Read that as a claim and it is obviously wrong. The ranking is asserting that being adjacent to a match is better evidence of relevance than being the match. A seed carries direct query evidence: something in it actually matched what the user asked. A walked node was never compared against the query at all.

Adjacency to a match is not evidence. It is a hypothesis, and it should not outrank the evidence that produced it.

Sort one merged pool of seeds and walked nodes, slice it at the cap, and query-blind neighbours systematically evict the results that actually matched. The graph feature you added to improve retrieval quietly starts costing you the thing it was meant to improve.

The fix is the cap, not the walk. Tier it: seeds fill the cap first on their own ranking, and walked nodes compete for whatever is left. The walk keeps its real job, which is reaching the fact that shares no keywords with your question, and loses the ability to displace the evidence.

One more pin is needed once you have supersession, and it is the kind of interaction you only find by hitting it. If a retired fact surfaces and its replacement is a walked node, the replacement gets demoted into the lower tier and trimmed at the cap, and you have shown an agent a dead decision with nothing next to it. So the replacement half of a superseded pair is pinned alongside the seeds, never in the walked tier.

Part two: the record that ages

Everything above assumes the graph exists and the job is reading it well. The second half begins when the graph is something your agents write to, and almost none of it has prior art, because indexing a corpus never raised these questions.

Facts that disagree

In March somebody records that sessions live in a signed cookie. In August it is reversed, because revocation turned out to matter more than the saved round trip. Both facts are in your graph, both match the query, both look equally authoritative. A retriever hands over whichever is the better semantic match, and has no way to prefer the true one.

Deleting the old one answers what is true now and destroys the answer to why did we change our minds, which is what stops somebody re-proposing signed cookies next spring.

So retirement becomes its own operation, requiring a reason and verbatim evidence that the old state actually ended rather than merely being contradicted. We also added a read replica supersedes nothing.

Our first version then archived the retired fact and hid it, which is correct on paper and has one fatal property: a mistaken retirement is invisible and unrecoverable. The next agent does not get a wrong answer, it gets no answer.

That mattered because agents get this wrong constantly. On conversational source material we measured false retirements eleven times out of eleven, and rewriting the guidance moved the rate by exactly zero. The fix was not better instructions. It was making the mistake survivable.

rank multipliers, applied wherever recall ranks

retired knowledge      × 0.35   needs a far stronger match than a live fact
done work, fresh       × 0.60   finished this week: this is the changelog
done work, decayed     → ~0.2   finished months ago: this is exhaust
cancelled work         × 0.15   things we deliberately decided not to do

Cancelled work takes the harshest discount for a specific reason: it is imperative text describing something the team decided not to do. An agent that retrieves Add SSO for enterprise accounts without the cancellation attached will go and implement it.

Facts that expire without being replaced

Supersession handles replacement. Most stale facts are not replaced by anything, they just quietly stop being true.

Nobody audits a knowledge graph quarterly. That plan always fails. What works is that everybody can answer when does this stop being true at the moment they write it down, because that is the only moment they know. So the choice is mandatory at write time: bind it to the work that resolves it, give it an expiry date, or deliberately neither for something durable.

Facts that cannot be found by searching

Vector search fails exactly where vocabulary diverges. Somebody wrote a constraint using the word throttle; the person about to break it is searching for middleware ordering. A better embedding model does not fix that.

What fixes it is letting a fact declare the files and symbols it is about, so the trigger stops being a keyword and becomes an edit.

the fact, with anchors

lesson: "Rate limiter must stay in front of the auth check"
  anchor_files:   [src/middleware/index.ts, src/middleware/rate_limit.ts]
  anchor_symbols: [buildMiddlewareChain, RATE_LIMIT_WINDOW]
  body: |
    Moving the limiter behind auth in Nov 2025 let unauthenticated
    traffic reach the token-metered path. 4h outage.

Now touching buildMiddlewareChain surfaces the lesson whatever words the person had in mind. One honest limit: this fires on edits, not on reads, so an agent that only ever reads a file never triggers it and semantic recall carries that case alone.

Two agents, one piece of work

Shared memory lets two agents read the same facts. It does nothing to stop them doing the same work, because a note cannot be held. What stops it is that starting work means claiming it atomically, and the loser is told who holds it and since when rather than just being refused. An agent told only no retries; an agent told another session has been on this for forty minutes goes and does something else.

Knowing which facts are still true

The last one decides whether any of the others matter in a year. A fact needs two independent clocks. Freshness is when the content last changed: cheap, automatic, nearly worthless alone. Verification is when somebody last confirmed the claim still holds against reality.

The rule that makes verification mean anything is that an ordinary edit must never set it. If fixing a typo counts as confirming a fact, the signal is gone. Keep them apart and you can answer a question a folder of notes cannot: which of our facts has nobody confirmed in ninety days?

Then something has to act on the answer.

your agent

›/stele:review

42 facts touched by changes since the last sweep · 7 unverified in 30+ days

reviewers checked each against the code…

31 still hold · 4 re-anchored · 2 superseded · 5 flagged for your call

Ours runs when an agent runs it. A fully autonomous version that tends the graph with nobody prompting it is something we are building, not something that ships today.

What this adds up to

Part one is a weekend if you have built retrieval before, plus a few rounds of being wrong about weights. Part two is a product.

It is also not optional, which is the uncomfortable bit. Skip supersession and your agent confidently cites a decision you reversed. Skip lifecycle and the graph fills with facts nobody can vouch for. Skip anchors and the constraint stays invisible to the person breaking it. Skip verification and you cannot tell a graph that works from one that quietly stopped being true. Every one is load-bearing, and you find that out in production.

When a graph is the wrong tool

If your corpus does not change, none of part two applies. Retrieval over a static reference set is a retrieval problem, and everything after the walk solves something you do not have. Build part one and stop.

Or connect one that already does it

This is what we build. Stele ships as an MCP server, so if your agent speaks MCP, which is now most of them, connecting it is the same as connecting anything else.

terminal

$npm i -g stele

$stele account login

$stele mcp

stele MCP server · stdio · 16 tools

ready

A stdio server any MCP client can spawn, on every plan. Point a Python agent at it, or a framework that did not exist when this was written, and it gets the same sixteen tools: recall and search for reading, knowledge and task for writing, node for edges and edits, plus the lifecycle operations. For the ten harnesses Stele installs into directly, stele install also wires the hooks that inject relevant context automatically, so retrieval happens on every turn without the agent choosing to ask.

What stays yours: the live conversation buffer, which belongs to your framework and should. The loop, the retries, the orchestration, which is the other kind of graph engineering entirely. And anything multimodal, since we ingest text rather than images or audio.

The retrieval graph was never the hard part. It is the part with tutorials, and the reason they feel like they end early is that they do.

  • engineering
  • agent memory
  • graphs

Try Stele

Both halves, already built, connected over MCP to whatever agent you use. Stele is free to start.

Start freeRead the docs

How do you measure whether agent memory actually helps?

Source: https://stele-ai.dev/blog/measuring-agent-memory

Engineering

How do you measure whether agent memory actually helps?

Everyone selling agent memory shows you a demo where it helps. We tried to build a benchmark that could tell us whether ours does, and spent most of that effort discovering how many ways a memory benchmark can quietly measure nothing.

July 27, 2026·Stele·10 min read

The demo problem

Every agent memory product, ours included, can produce the same demo. A fresh agent gets a question, flounders. An agent with the memory gets the same question, answers it. The audience nods.

The demo is worthless, and it is worthless for a reason that took us embarrassingly long to state precisely: the demo is chosen after the fact from the cases where it worked. Nobody films the run where the memory arrived and the agent ignored it. Nobody films the run where the control read a well-commented file and got there first.

So we set out to build something that could return an answer we did not want. This is a report on the method, not on results. We do not have results worth publishing yet, and the reason we do not is the interesting part.

We cannot benchmark on our own repository

The first finding was the one that invalidated the most work.

An agent working inside our own codebase is not a normal agent. It reads our documentation, sees the tooling, and works out within a couple of turns that the memory system is the subject of the exercise. It then pays the graph conspicuous, deliberate attention. On an unrelated project with a good record, the same agent uses the same feature far more quietly, and barely narrates it.

Those are two different behaviours, and only the second one is what a customer experiences.

This is demand characteristics, the same effect that makes a participant who has guessed the hypothesis useless as a subject. It has a consequence we now treat as a hard rule: the corpus runs on an external pinned repository, never on ours.

It has a second consequence that is harder to live with. Self-observation is never a result. A session where we watch our own agent use the record beautifully is an anecdote, and it is precisely the class of anecdote the benchmark exists to replace. Sample size one, unblinded, and the observer is the most motivated party in the room.

The run that was perfect and meant nothing

Our clearest lesson came from a run that went entirely to plan.

Ten arms, all clean. Every cost fence held. Total spend about a dollar sixty. Deterministic grading, durable receipts, no infrastructure failures. By every operational measure it was a success, and the number it produced was meaningless.

The question had two reasonable readings. Under one reading you needed the stored fact to answer. Under the other you did not. The agent picked a reading, and our grader dutifully measured which reading it picked. We had built an extremely reliable apparatus for measuring the wrong thing.

That reframed the whole project. The apparatus was never the hard part. Running paired arms with certification and a money fence is engineering, and engineering is tractable. The questions are the science, and the questions are where everything goes wrong.

Eight rules a question has to pass

What came out of that is an admission checklist. A candidate question has to clear all eight before it costs any money. Three of these have each personally invalidated a paid run.

1. Unguessable

The answer must not be derivable from the repository at any token budget. The strongest form is a decision not to build something, because absent code is indistinguishable from code nobody got round to writing.

2. Unavoidable

A correct answer must be impossible without the memory. This is the one the run above died on. The test to apply before spending anything: could an agent that fully holds this memory still answer well without using it? If yes, the grader is broken before it runs.

3. Unambiguous

Exactly one reasonable reading. Ambiguity does not average out across attempts. It becomes the thing you are measuring.

4. Never “how does X work”

Frontier agents open the file and answer. That is reading comprehension, and it tells you nothing about memory.

5. Prefer questions whose answers live in history

Why do we have X. Did we already decide this. What breaks if I change this. The governing principle is distance: recall beats reading precisely when the answer is far from the question. When the answer sits in a comment above the function, the two are a wash.

6. Grade behaviour, not vocabulary

Checking whether the agent emitted a particular phrase measures parroting. Grep for a marker only where the marker is genuinely unavoidable, which is rule 2 again wearing a different hat.

7. Match the question to the channel you are testing

This one is specific to how retrieval is wired, and it is easy to get wrong. Anchor-triggered recall fires when a file is edited. It does not fire when a file is read. So an advice-shaped question can only ever exercise semantic recall, and if you wanted to know whether anchors work you have learned nothing. Testing anchors requires a task that actually edits something.

8. Prove the trajectory hits the anchor before paying

Run the control unpaid first and watch where it goes wrong. The likely wrong edit has to land on a file the tested fact is actually anchored to, and the fact has to directly correct that specific mistake.

We learned this by skipping it. We built a probe where the memory was real, the anchors were real, and the agent made its wrong edit in a different file entirely. The warning could not fire, and if it had fired it would not have addressed the mistake. Paid, and worth nothing.

The standing confound

Well-commented code is a strong control arm. A diligent repository lets the no-memory agent recover a great deal on its own. This is not a flaw to engineer around. It is the honest baseline, and it is what retired one of our early questions when the answer turned out to be sitting in a visible changelog file.

Three layers of grading, reported separately

A single pass or fail number hides where the effect actually lives. We report three, and the gaps between them carry more information than any one of them.

the three layers

injected          did the memory reach the agent at all?
                  read from the retained session transcript

observably used   does it appear in the answer or the reasoning?
                  a LOWER BOUND, never a count

scored            did it satisfy the grader?

On the run described above these read five, then three, then one. The interesting question was never the one. It was the two places the number dropped, which are entirely different problems: retrieval reaching the agent and being wasted, versus retrieval being used and still not producing a better answer.

The middle layer needs its qualifier stated every time. It is a lower bound, not a count. A model can read a fact, internalise it, and produce a better answer without ever referring to it. So the honest phrasing is injected, with no observable trace. Writing injected and ignored claims knowledge of something unobservable, and it happens to claim it in the direction that flatters nobody, which is not the same as being conservative.

Rules we hold ourselves to

No model as judge. Graders are deterministic. A model scoring another model's output introduces a second unvalidated system into the measurement, and when the result disagrees with expectation you cannot tell which one was wrong.

Continuous measures need far more attempts than binary ones. At five attempts, every continuous measure we tried, tokens, wall time, tool calls, was indistinguishable from noise. If the headline is cost, budget for twenty or more, or use a paired design.

Pre-register the prediction. Including for the results that would be bad for us. The most important number we do not have yet is the pure overhead of installing memory on a project where it cannot possibly help: an empty graph, a task with no history to draw on. It is the first thing a skeptic asks about, and five percent overhead and sixty percent overhead are very different products.

The experiment that actually tests the claim

Everything above measures a component. Only one design tests the product thesis, which is that the overhead pays for itself later.

It takes two sessions. Session one: both arms do a task, and the memory arm writes whatever it writes, unprompted, using only the shipped guidance. Session two: a related task requiring transfer, where the memory arm inherits the record session one produced and the control starts fresh. You measure session two, and cumulative cost across both.

The claim makes a sharp, falsifiable prediction. Session two cost inverts, and cumulative cost crosses over at some number of sessions. That crossover number is the honest headline for a product like this, and it is the one we want before we start quoting anything.

The compounding claim The only measured part is the short solid segment: the first session costs more. Everything after it is what the experiment is for. We have drawn this without a scale on either axis on purpose, because a numbered version of the same picture would be a prediction wearing the costume of a result.

Why the shape is not obvious

It is tempting to model the overhead as one thing that decays: you pay to fill an empty graph, the graph fills, the payments stop. That model predicts the line falls and keeps falling.

The overhead is more likely two things with different behaviour. Some of it is fixed per session: tool schemas in the context window, hooks running, a search helper starting up. That cost arrives every session no matter how mature the record is, and it never amortises. The rest is write cost, which plausibly does decay as the project's durable facts get written down once.

If a meaningful share is fixed, it sets a floor the line cannot go below, and the steady state is not a percentage of the baseline. It is baseline plus fixed overhead minus whatever recall saves. Whether that lands above or below the baseline is the sign of a difference nobody has measured, which is a more precise way of saying we do not know where the crossover is, or whether there is one.

We have a reason to take the fixed half seriously. An earlier harness-level measurement showed one arm using roughly three times the tokens of the other, which looked damning until we looked at where they went: almost all of it was cached schema and helper startup, front-loaded rather than spent on extra work, and net correctness was unaffected. Raw token count partly measures how much apparatus is attached rather than how much thinking happened. It is why the primary cost measure here is the authoritative gateway receipt, with token counts kept as a cross-check and never substituted for it.

One constraint on that design is load-bearing and easy to violate. Session one's writes have to come from the shipped guidance, with nothing added. The moment the harness tells the agent what to record, the benchmark is measuring whoever wrote the prompt rather than the product.

Where this leaves us

With an apparatus and without a number, which is the right order. Pinned corpus, frozen graph, sealed container, deterministic oracle, nine gates before a token is spent, and a pre-registration locked before results exist. Three admission runs have been discarded rather than reported, each because it caught a defect that would otherwise have shipped as a number.

We could publish the pilots that came out favourably. That is exactly the demo this post opened by complaining about, and it is the reason we built the harder thing instead.

What we will say is what we would want said to us if we were evaluating this category:

  • Ask what the control arm was. No memory, memory but no auto-injection, and no product installed at all answer three different questions.
  • Ask how many attempts. Continuous measures at n=5 are decoration.
  • Ask whether the benchmark ran on the vendor's own codebase. If it did, the agent knew it was being watched.
  • Ask for the overhead number on a project where the memory cannot help. Everyone has that number. Almost nobody publishes it.

We will publish ours when the questions are good enough that we would believe the answer if it came out against us. As far as we can tell we are the only people in this category building the apparatus to find out, which is a strange thing to be able to say about a field selling measured improvements.

  • engineering
  • benchmarks
  • methodology

Try Stele

We would rather publish the method before the numbers. Stele is free to start.

Start freeRead the docs

Building Stele with Stele

Source: https://stele-ai.dev/blog/building-stele-with-stele

Engineering

Building Stele with Stele

Three months, 1,650 commits, and about 1,300 tasks, nearly all of them done by agents reading a record other agents wrote. Here is what that is actually like, including the parts that did not work.

July 27, 2026·Stele·7 min read

The setup

Stele is built almost entirely by agents. Not in the sense that someone occasionally asks for a function, but in the sense that the normal unit of work is: a task exists in the record, an agent claims it, does it, writes down what it learned, and closes it. A human reads the result and decides what happens next.

Three months in, that has produced a project record with roughly these shapes:

the graph, at three months

1,650   commits since the end of April
1,300   tasks           ~1,080 done · ~215 open · 3 in flight
  940   active knowledge nodes      69 more archived or retired
   28   components

busiest topics   eval 125 · web 93 · auto-rag 79 · mcp 75 · supabase 63

We are the most demanding user of this product and also the least representative one, so treat what follows as a report from an extreme case rather than as a typical experience. The failure modes are real, though, and they showed up here first.

What actually works

The record survives the agent

The single most useful property has nothing to do with retrieval quality. It is that a session can end, badly, and the work is still legible.

An agent hits a context limit halfway through something. Previously that meant reconstructing intent from a half-finished diff. Now there is a claimed task with a written goal, links to the decisions that shaped it, and whatever was learned before the wheels came off. The next session opens with all of that and continues.

This turned out to matter more than any of the retrieval work. Most of the value of writing things down is not that the right fact arrives at the right moment. It is that a crash stops being a total loss.

Comments in the source that point at the reason

A pattern emerged that nobody designed. Code comments in this repository routinely reference the task or decision that produced them. Not as decoration: as a pointer from the what to the why, where the why is too long to sit in a comment.

The measurable perf table explaining why a query selects an explicit column list rather than everything, the paragraph explaining why one weight in a scoring table is negative, the note explaining which failure a lock exists to prevent: these live next to a pointer, and the pointer resolves to the full story. Neither half would survive on its own. A comment that long would get deleted for being noise, and a record entry with no anchor in the code would never be found.

Losing an argument to your own record

The best moments are the ones where the record disagrees with the person. Someone proposes a change, and what comes back is a decision from two months ago explaining why the obvious version of that change was tried and reversed.

The value is not that the answer was retrieved. It is that the argument did not have to happen twice.

What did not work

This is the more useful half, and it is the part a page written by marketing would leave out.

Agents wrote too much, and the graph got noisy

Early on the guidance encouraged writing things down, so agents wrote things down. Every session produced knowledge entries, and a substantial share were restatements of what the code already said, or near-duplicates of something recorded two weeks earlier by a different session that had not found it.

A noisy record is worse than a small one, because it dilutes retrieval: twenty slots filled with four good facts and sixteen restatements is a worse turn than four slots filled with four good facts. Most of the lifecycle work, deduplication, the health check, the pressure to bind every fact to a retirement condition at write time, exists because of this. It was not foresight. It was cleanup.

Guidance is a much weaker lever than we assumed

The clearest case: agents folding conversational source material into the graph would retire facts that had not actually been reversed. Every single time. We rewrote the guidance, added examples of the near misses, put the warning at the point of use, and the rate did not move at all.

That failure produced the most transferable lesson we have. When better instructions do not move a number, the mistake is structural, and you should redesign for it happening instead. Which is what we did: retirement stopped being destructive, so a wrong call became a visible wrong label instead of a deleted fact.

We built things because the graph made them easy, not because they mattered

When the backlog is a queryable list of well-specified tasks and agents are cheap, the throughput of small work goes way up. That is mostly good and occasionally a trap: it is very easy to ship fifteen tidy improvements to a subsystem nobody is blocked on, because those tasks were sitting there and were well-scoped, while the ambiguous thing that actually mattered stayed open because nobody had written it down properly yet.

A well-organised backlog exerts a pull toward whatever is easiest to pick up. That pull is not correlated with importance, and the tooling does not fix it. Prioritisation is still a person's job.

Our own sessions are useless as evidence

An agent working in this repository can see that the memory system is the subject. It reads the doctrine, sees the tooling, and pays the record deliberate attention it would not pay on an ordinary project.

Which means the thing we watch every day, our own agents using this well, is not evidence that it works. It is the reason we cannot benchmark on our own repository, and it disqualifies the demo we would most like to show you.

The rhythm that emerged

None of this was designed up front. What settled into place over three months is roughly four moves, and they are worth describing because they are not what we expected to be doing.

Work is written down before it is done, not after. Not because process is virtuous, but because a task with written intent is the only artifact that survives a session ending badly. The discipline is entirely selfish.

Lessons get written in the moment, mid-task. Anything left for a summary at the end gets written in the vague register people use when they are tired, and vague lessons never surface again. The concrete trigger has to go in while it is still annoying.

Somebody reads the record with their own eyes, regularly. This is the one we keep having to relearn. Retrieval quality metrics tell you about the apparatus, not about whether the contents are true. The only way to find out that a corner of the graph has quietly filled with plausible nonsense is to go and read that corner.

Verification is scheduled, not hoped for. A sweep that diffs against the last commit it reviewed, gathers the facts touched by what changed, and checks each against the current code. It finds real drift every time it runs, which is both reassuring and slightly alarming.

Things we would tell someone starting

Do not seed a large record up front. A generated pile of facts about a codebase you have not questioned is a pile of plausible sentences, and plausible sentences that turn out to be wrong are worse than nothing, because they arrive with the same authority as the true ones. Let it accumulate from work that actually happened.

The lesson has to contain the literal trigger. Not be careful with type generation, but the error string, the file, the symbol. A lesson written abstractly is invisible exactly when it would help, because retrieval matches text and the person in trouble is searching with the words the error gave them.

Decide when a fact dies while you are writing it. Nobody ever runs a quarterly audit of their notes. Everybody can answer when does this stop being true in the moment they write it down. Bind the risk to the task that resolves it and it retires itself.

Read your own record occasionally. Not through an agent. Actually read it. Every serious problem we have had with the graph was visible in about ten minutes of reading it, and invisible in any metric we had.

The uncomfortable check

Open your record and pick five facts at random. How many can you confirm are still true without opening the code? If the answer is fewer than four, the problem is not retrieval and adding more facts will make it worse.

What we still do not know

Whether this is faster. We genuinely do not have that number, and we are suspicious of everyone in this category who claims to.

The number we want is the crossover point: how many sessions before the overhead of maintaining a record is repaid by the sessions that inherit it. That is a two-session paired experiment, we know exactly how to build it, and we would rather publish that than a testimonial.

What we can say without it is narrower and holds up: sessions resume instead of restarting, decisions get re-litigated less often than they used to, and a session that dies halfway is no longer a total loss. Three months of building this way is not a controlled trial, and it is also not nothing. It is why the crossover experiment is the one we chose to build first, out of everything we could have measured.

  • engineering
  • dogfooding
  • agent loop

Try Stele

If you want to try this loop on your own project, Stele is free to start.

Start freeRead the docs

Why markdown rot breaks agent memory

Source: https://stele-ai.dev/blog/why-markdown-rot-breaks-agent-memory

How it works

Why markdown rot breaks agent memory

Every project starts with one instructions file, and it works. The trouble arrives about four months in, and it is not that the file got messy. It is that nothing in it can tell you which parts are still true.

July 25, 2026·Stele·9 min read·Updated July 27, 2026

The file that worked

The first version is always good. Someone notices the agent keeps reaching for the wrong test command, writes four lines into a CLAUDE.md, and the problem goes away. That is a genuinely excellent return on four lines. Zero infrastructure, works in any harness that can open a file, reviewable in a pull request, and it lives next to the code it describes.

So it grows. Someone adds the deploy runbook. Someone adds the reason the auth middleware is ordered the way it is. Someone adds a warning about the migration that has to run before the deploy. A year later there is a CLAUDE.md, an AGENTS.md that mostly duplicates it, a docs/ folder with fourteen files, and three READMEs in subdirectories that nobody has opened since they were written.

None of that is a discipline failure. It is what happens when a format with no lifecycle is used for facts that have one. The pile does not go wrong because it is markdown. It goes wrong in three specific ways, and each one has a specific shape.

Failure one: nothing on the page knows it went stale

Here is the sequence, and it is always the same sequence. In March you write down that background jobs go through the queue table, never through the worker directly. In June you rip out the queue table and move to a hosted broker, because the table was a bottleneck. The code changes. The tests change. The docs/jobs.md file does not change, because changing it was not part of the diff that made the change work.

In September an agent picks up a ticket about a failing job, reads docs/jobs.md, and writes a patch that inserts into a table that no longer exists. You catch it in review. Or you do not, and you catch it in production.

A file cannot tell you it has gone wrong. That is not a bug in markdown. It is the absence of a feature markdown never claimed.

The instinct is to blame the person who did not update the doc, but that is the wrong lesson. Nobody updates the doc, at any company, ever, reliably. What is missing is not diligence. It is a mechanism that distinguishes a fact somebody confirmed last week from a fact somebody wrote eighteen months ago and has not looked at since. In a folder of markdown files, those two facts are typographically identical.

Git timestamps do not close this gap, and it is worth being precise about why. A file's modification time tells you when its bytes last changed. It does not tell you whether anyone checked the claim. Reflow a paragraph and the file looks freshly maintained. Meanwhile a fact that has been correct and untouched for two years looks abandoned. The signal you want and the signal you have are close enough to be confused and far enough apart to mislead.

Failure two: retrieval by filename

The second failure is quieter, and it costs you before you notice it. To find a note in a folder, you have to already know roughly where it is, or guess the word it was written with.

Someone recorded that the rate limiter has to stay in front of the auth check, because moving it behind auth once let unauthenticated traffic burn the token budget. They wrote it in docs/incidents/2025-11-token-burn.md. Eight months later an agent is refactoring middleware order and searches for middleware, ordering, rate limit. The file says throttle and quota. No match. The agent moves the rate limiter behind auth, exactly as someone did the first time, for exactly the same sensible-looking reason.

This is what people mean when they say grep is not memory. Grep answers where does this string appear. Memory has to answer what do I need to know before I touch this, and the second question cannot be reduced to the first, because the person who wrote the answer and the person asking the question do not share a vocabulary. They are frequently not even the same species: one is a human writing at 11pm after an incident, the other is a model tokenizing a diff.

A search subagent helps here, and it is a real improvement over nothing. But it improves navigation of a library organised by filename. It does not make the library addressable by meaning, and it cannot surface a fact that shares no words with the thing you are doing.

Failure three: all or nothing loading

When the agent cannot locate the right paragraph, it does the only thing left: it loads whole files. A 400-line runbook enters the context window so that 6 lines of it can be relevant.

The cost is not only tokens, although it is also tokens. It is that the signal-to-noise ratio of the context window is the single biggest lever on whether the model does the right thing, and stuffing it with mostly-irrelevant prose moves that lever the wrong way. You are paying to make the answer worse.

The alternative most teams land on is to put the important things in the always-loaded instructions file, which works until the always-loaded instructions file is 600 lines long and the model is skimming it the same way you would.

What a lifecycle actually requires

The interesting question is not where should notes live. It is what has to be true of a fact for it to be safe to serve to an agent. Working through that gives you a short list, and the list is the same whether you build it yourself or use something off the shelf.

Two clocks, not one

A fact needs two independent timestamps, and conflating them is the root of most of the trouble above.

  • Freshness is when the content last changed. Cheap, automatic, and nearly worthless on its own.
  • Verification is when somebody last confirmed the claim still holds against reality. Expensive, and it is the one that matters.

The important design rule is that an ordinary edit must never set the verification clock. If fixing a typo counts as verification, the signal is gone. In Stele, verification is only ever set by a deliberate check, which is what makes it possible to ask a question a folder cannot answer: which of our facts has nobody confirmed in ninety days?

Retirement chosen when the fact is written

Most facts are not permanent, and their authors usually know it at the moment of writing. The workaround we hit on is to make the writer decide the fact's ending at the same time they decide its beginning. There are three ways a fact can be scheduled to retire:

  • Bound to a task. A risk like the migration has to run before the deploy or sessions drop is true until the migration ships. Bind it to that task and it retires itself when the task closes.
  • Expiry date. The staging database is on the old schema until the end of the quarter is a fact with a known shelf life. Give it one.
  • Nothing. Durable architecture and decisions do not expire on a clock. They retire by being replaced, which is a different mechanism.

The point of forcing the choice at write time is that cleanup stops being a chore somebody has to remember. Nobody schedules a quarterly audit of their notes folder. Everybody can answer when does this stop being true in the moment they are writing it down.

Replacement, not deletion

When a decision reverses, the wrong move is to delete the old one. You lose the answer to why did we change our minds, which is usually the most valuable thing in the file. The right move is to record that the new decision supersedes the old one, keep the old one readable, mark it clearly, and rank it below its replacement so it only surfaces when somebody is genuinely asking about history.

That has a second benefit that took us a while to appreciate: a supersede recorded by mistake becomes a wrong label you can correct, instead of a true fact you destroyed. We wrote about that one separately, because the failure mode is more interesting than it sounds.

Status governs rank, never existence

A rule we got wrong and had to fix: for a long time, finished work was hidden from recall entirely, on the theory that a closed task is exhaust. Then we hit a case where a task had been completed four hours earlier, its completion note was the only written explanation of how the feature worked, and recall could not see it at all. Exact search found it instantly.

The model was missing time. A task finished this morning is the changelog. A task finished eight months ago is exhaust. Status alone cannot tell them apart, so status should never have been deciding visibility. Now it decides ranking, on a decay curve, and everything stays reachable with an honest label on it.

The general shape of the rule

Hiding a fact destroys an answer. Downranking a fact with a clear label preserves the answer and stops the fact from crowding out live work. Almost every time we reached for hiding, downranking was the better call.

What that looks like day to day

Two of these run without anyone asking. Expiry and task-bound retirement fire on their own, and knowledge that has gone a long time without verification gets flagged rather than served as if it were certain.

The deliberate half rides along with the agents you already have. A verification sweep diffs against the exact commit it last reviewed, gathers the facts touched by what changed, and checks each one against the current code, both the claim and the files it points at:

your agent

›/stele:review

42 facts touched by changes since the last sweep · 7 unverified in 30+ days

reviewers checked each against the code…

31 still hold · 4 re-anchored · 2 superseded · 5 flagged for your call

The health check is the broader audit: overdue verification, expired notes, decisions superseded but still marked active, near-duplicates, edges pointing at nothing. What is safe to retire on its own is cleared in the same pass, and the rest comes back as a worklist.

Neither of these is magic. They are ordinary maintenance, with the difference that they are run by the agents doing the work rather than by a human who has to remember.

Anchors, or how a fact finds you

The retrieval failure earlier deserves one more note, because there is a fix for it that is easy to miss.

Text matching gets you a long way, but it fails exactly where the vocabulary diverges, which is exactly where you need it most. So a fact can also declare the files and symbols it is about. The rate limiter lesson does not have to share words with your refactor. It has to point at the file:

the fact, with anchors

lesson: "Rate limiter must stay in front of the auth check"
  anchor_files:   [src/middleware/index.ts, src/middleware/rate_limit.ts]
  anchor_symbols: [buildMiddlewareChain, RATE_LIMIT_WINDOW]
  body: |
    Moving the limiter behind auth in Nov 2025 let unauthenticated
    traffic through to the token-metered path. 4h outage, $2.1k burn.
    The limiter has to reject before anything that costs money runs.

Now the trigger is not a keyword. It is the edit. Touch buildMiddlewareChain and the lesson arrives, whatever words you were thinking in.

This is also the honest answer to why not just use embeddings. Semantic search is a fuzzy match and it is genuinely good. Anchors are an exact one. You want both, because they fail in different places: a fuzzy match catches the fact you described in other words, and an exact anchor catches the fact you were not describing at all.

When the file is still the right answer

It would be dishonest to end without this part.

In-repo markdown needs no hosted backend, no account, and no network. It survives your vendor going away. It is reviewable in the same pull request as the code, which is a real advantage that nothing else on this page replicates. If you are one person on one project with a folder you can hold in your head, keep the folder. You will not get a return on anything more.

The trade tips when some combination of these is true:

  • More than one person, or more than one agent, needs the same facts, and they are currently getting different ones.
  • You have been bitten at least once by a reversed decision that an agent kept following.
  • You cannot answer which of these notes is still true without reading all of them.
  • The notes matter outside the checkout: in planning, in review, from a phone, from an agent that is not a coding agent.

And one boundary worth stating plainly, because it is the kind of thing people assume: the automatic parts run on their own, but the sweeps run when your agent runs them. A fully autonomous background gardener that tends the record with nobody prompting it is something we are building, not something that ships today.

If you take one thing from this: the problem was never the file format. It was that facts have a lifecycle and files do not. Any solution that does not model the lifecycle, including a more organised folder, is going to fail in the same three ways on a longer timeline.

  • how-to
  • agent memory
  • lifecycle

Try Stele

Stele keeps the project record on a lifecycle so the facts your agents read can retire themselves. Stele is free to start.

Start freeRead the docs

When a fact stops being true

Source: https://stele-ai.dev/blog/when-a-fact-stops-being-true

Engineering

When a fact stops being true

We tried to stop agents from retiring facts they should have left alone. Rewriting the guidance moved the error rate by exactly zero. So we changed what happens when they get it wrong instead.

July 25, 2026·Stele·9 min read·Updated July 27, 2026

The problem with being right later

In March the team decides sessions live in a signed cookie. It goes into the record with the reasoning attached: no server-side session store to operate, no extra round trip on every request.

In August they reverse it. Revocation turned out to matter more than the round trip, and you cannot revoke a signed cookie without the store you were avoiding. Sessions move into the database.

There are now two questions, and most systems only handle the first.

  1. What is true now? Sessions are in the database. Any agent asking about session handling has to get this answer and not the other one.
  2. Why did this change? Six months from now somebody proposes signed cookies again, for the reasons that were good in March. Without the history that proposal looks new. With it, it looks like a settled argument whose counter-argument is already written down.

Delete the old decision and you answer the first question perfectly while destroying the answer to the second. Keep both as equals and the agent gets two contradictory facts with the same standing, which is worse than either one alone. Neither is acceptable, which is why retiring a fact is its own operation rather than an edit in place.

What it looked like at first

The first implementation was the obvious one. Recording that B supersedes A drew an edge from B to A and archived A. Archived nodes were excluded from recall. The old fact stayed in the database and stayed visible in the web UI if you went looking for it, but as far as any agent was concerned it was gone.

On paper that is correct. In practice it has one property that turned out to outweigh everything else: a mistaken retirement is unrecoverable in the only place that counts.

If an agent wrongly decides that a new fact retires an old one, it has not mislabelled anything. It has removed a true fact from every future agent's world, silently, with no error raised and nothing to notice. The next agent does not encounter a wrong answer. It encounters no answer, which is indistinguishable from the project never having decided at all.

The number that did not move

We would not have prioritised this if it were a theoretical risk. Here is what it actually looked like.

We were testing how well agents fold new source material into an existing graph: hand one a meeting transcript or a design document, have it work out what the project already knows, and record what changed. Retiring the facts that the new material genuinely reversed is precisely the job.

On conversational source material, agents wrote a false retirement eleven times out of eleven.

Not eleven mistakes over a long run. Every single opportunity. And the mistakes were reasonable. Conversation is full of sentences that sound like reversals and are not: someone floats an alternative, someone objects to a decision without overturning it, a plan gets refined rather than replaced. A model reading a transcript sees a new statement about a topic it already holds a fact for, and the shape of new statement about known topic sits very close to the shape of a genuine reversal.

So we did the normal thing. We rewrote the guidance. We spelled out when a retirement is warranted and when it is not, gave worked examples of the near misses, and put the warning at the point of use rather than in a document nobody opens.

It moved the rate by zero. Still eleven out of eleven.

That result is worth sitting with, because the instinct after a guidance change fails is to go and write better guidance. What it actually tells you is that the error is not caused by the agent being under-informed. It is caused by the task being genuinely ambiguous at the point of decision, and no quantity of instruction resolves an ambiguity that lives in the source material rather than in the reader.

The generalisable version

When a class of error survives a guidance rewrite completely untouched, stop trying to prevent it and start designing for it happening. The question changes from how do we stop this to what should the blast radius be when it happens anyway, and the second question usually has a much better answer available.

Making the failure survivable

The fix was to stop hiding retired knowledge from recall entirely. A retired fact comes back, marked as retired, ranked below whatever replaced it.

That one change converts the failure mode:

  • Before. A wrong retirement destroys a true fact. Silent, unrecoverable in the surface that matters, and detectable only by noticing an absence, which nobody does.
  • After. A wrong retirement puts a wrong label on a true fact. The fact still surfaces, it is still readable, the label is visible, and anyone who spots it can reverse it.

A wrong label is an ordinary bug. A destroyed fact is a data loss event. Trading the second for the first is worth a fair amount of ranking complexity.

It also made the system finally answer the question that justified having a retirement mechanism at all. If the old decision is unreachable, what did we believe before and why did it change has no answer, and you have built an elaborate apparatus to achieve exactly the same result as editing the fact in place.

Rank, do not hide

Once a retired fact is allowed back into the candidate pool, the question becomes how hard to push it down. Too gentle and a dead decision crowds out the live one, which is the failure the feature exists to prevent. Too harsh and it never appears even when somebody is explicitly asking about history.

What we settled on is a set of multipliers applied wherever recall ranks. They are deliberately blunt:

rank multipliers

retired knowledge      × 0.35   needs a far stronger match than a live fact
done task, fresh       × 0.60   finished this week: this is the changelog
done task, decayed     → ~0.2   finished months ago: this is exhaust
cancelled task         × 0.15   the harshest discount in the system

A retired fact at 0.35 has to be a substantially better match than any live fact before it outranks one. It surfaces when it genuinely is the best answer, which is to say when somebody asked about the past, and not merely when it is topically nearby.

Cancelled work gets the harshest treatment of anything in the system, and the reason is worth spelling out. A cancelled task is imperative text describing work the team deliberately decided not to do. An agent that retrieves Add SSO login for enterprise accounts without the cancellation attached will go and implement it. That is an argument for the marker and for the floor of the ranking. It is not an argument for hiding, because hiding destroys the answer to did we already consider this?, and that answer saves real time.

The rule that came out of it

There was a second version of the same mistake sitting in the codebase, and we only found it because of the first one.

Completed tasks were hidden from recall, on the theory that finished work is exhaust. Then a case turned up that broke the theory cleanly. A task had been completed four hours earlier. Its completion note was the only written explanation anywhere of how the feature worked. A question about that exact feature could not reach it, while plain keyword search found it immediately.

The model was missing time. A task finished this morning is the changelog. A task finished eight months ago is exhaust. They carry the same status and completely different value, so status was never the right thing to be deciding visibility with.

Status governs rank and label. It never governs existence.

Everything reachable stays reachable. What changes is where it lands in the ordering and what label it arrives carrying. There is one deliberate exception, an explicit hide flag a person sets on purpose, and being explicit is the entire point of it.

What the agent actually receives

The marker matters as much as the ranking does, because an unlabelled retired fact is worse than no fact at all. When a retired decision does surface, it arrives with its status, its reason, and a pointer to whatever replaced it:

recalled context

RETIRED · decision · "Sessions in a signed cookie"
  replaced by: "Sessions in the database, revocable"
  reason:      revocation needs a server-side store; the round-trip
               cost was the wrong thing to be optimising
  written 2026-03-11 · retired 2026-08-02

ACTIVE  · decision · "Sessions in the database, revocable"
  Every session row is revocable individually. Logout and security
  response both need that; the cookie could not provide it.

The replacement is always recalled alongside the fact it retired. Surfacing a retirement on its own would leave an agent knowing the old answer is dead without knowing the new one, which is the worst of the three available states.

Every retirement also has to state a reason. Not for bookkeeping: the reason is the part a human reads six months later when the same idea comes back around, and demanding it at write time is the only reliable moment to capture it, because it is the only moment anybody still knows it.

Where this leaves the false retirements

Honestly, they still happen. We did not fix the classification problem, because we do not believe guidance can fix it, and we have the negative result to back that up.

What changed is the cost. A wrong retirement is now a visible wrong label on a fact that still shows up, which the health check flags as a decision retired but still doing work, and which any human reading the record can reverse in a moment. It is a nuisance rather than a loss.

There is one test case we care about more than the rate itself. Take a decision that was later retired, then ask an agent to do something the old decision forbade and the new one allows. The correct behaviour is to get on with it without objection. An agent that pushes back is one that treated a retired fact as live, and that is a bug with a real cost to the person on the other end of it. It is the sharpest available test of whether retirement works at all, because it is the only case where a false positive is the failure rather than the safe option.

The broader lesson has outlived the feature. Nearly every time we reached for hiding something, downranking it with an honest label turned out to be the better call. Hiding destroys an answer. Downranking preserves the answer and stops it from getting in the way, and those were the two things we actually wanted all along.

  • how-to
  • supersession
  • knowledge

Try Stele

Retired facts stay readable, marked, and ranked below whatever replaced them. Stele is free to start.

Start freeRead the docs

A file, a wiki, or a shared ledger?

Source: https://stele-ai.dev/blog/file-vs-shared-ledger

Objections

A file, a wiki, or a shared ledger?

The most common thing people say to us is that they can already do this with a markdown file, or Notion, or a small RAG stack they wired up themselves. Often they are right. Here is how to tell which case you are in.

July 25, 2026·Stele·7 min read·Updated July 27, 2026

The objection, stated properly

It usually arrives in one of three forms, and all three are reasonable.

  • I have a CLAUDE.md and it works. It does. It costs nothing, it is reviewed in the same pull request as the code, and it works in any tool that can open a file.
  • We keep this in Notion already. Also true. Notion is better at long-form writing than anything we will ever build, and your team is already in it.
  • I could build this in a weekend with pgvector. The retrieval part, yes, genuinely. That part is not hard any more.

A sales page would now explain why all three are wrong. They are not wrong. They fail in specific circumstances, and if you are not in those circumstances the honest advice is to keep what you have. What follows is the actual dividing line, which is narrower and more concrete than the pitch usually suggests.

Three questions that decide it

Forget the storage format. Every option here stores text and every option here retrieves it. The differences that matter are three properties that have nothing to do with where the bytes sit.

1. Does anything read it without being asked?

This is the one that quietly disqualifies the wiki. Your coding agent does not open Notion before it answers you. It does not open Notion before it edits a file. If retrieval requires a human to remember that the relevant page exists and go and paste it in, the wiki is a reference library for people, which is a fine thing to be, but it is not agent memory.

The in-repo file does better here, because harnesses load it automatically. That is exactly why the pattern caught on. The catch arrives at scale: what gets loaded is the whole file, every turn, whether or not any of it bears on what you asked. That works at 40 lines and stops working somewhere around 600, when the model starts skimming your instructions the same way you would.

2. Does anything keep it true?

A file does not know when it has gone wrong. Neither does a wiki page. Neither does a vector store, which will happily return a chunk about a decision you reversed last quarter, with high confidence, because semantically it is an excellent match for your question.

Retrieval quality and truth are independent. A perfect retriever over a stale corpus returns exactly the wrong answer, quickly.

This is the failure people underestimate when they estimate the weekend build. Embedding and cosine similarity are the easy half. The half that takes real work is everything that decides whether a chunk still deserves to be returned: when the fact was last confirmed rather than last edited, whether the decision it describes has since been reversed, whether the risk it warns about was resolved by work that has already shipped.

3. Can two agents coordinate through it?

A file is a broadcast channel. Nothing in it can express somebody is already doing this. If you run one agent at a time and you are the only person on the project, that costs you nothing and this question is irrelevant to you.

It stops being irrelevant the moment there are two. Two agents on the same repository will both read the same file, both conclude the same work needs doing, and both do it. What prevents that is not better notes. It is a piece of shared state that can be claimed, where the claim is atomic and the second agent gets told no.

What the weekend build actually contains

The pgvector answer deserves a straight response rather than a dismissal, because it is correct about the part it names.

Chunk, embed, store, query by cosine distance: a weekend, and it will work. What it gives you is a store. What a project record has to be is a loop. Here is the rest of the list, which is where the time actually goes:

what is left after the retriever works

read     something has to pull relevant context in before every prompt,
         without being asked, and without flooding the window

write    the agent has to record what it learned mid-task, at the moment
         it learns it, not in a summary nobody reads at the end

claim    concurrent agents need atomic claim on units of work, or two of
         them do the same thing and one of them wins the merge conflict

age      facts need to expire, retire when their task closes, and step
         aside when a decision replaces them, with the old one still
         readable and clearly marked

reach    the same record has to be readable from a second agent, a second
         machine, a review on your phone, and a tool that is not a coding
         agent at all

trust    somebody has to be able to ask "is this still true?" and get an
         answer that is not just "it was edited recently"

None of these are hard individually. All of them together are a product, and it is a product that is not the one you were trying to ship. That is the entire argument for buying rather than building, and it is the only argument we would make: not that you cannot do it, but that you would be doing it instead of your actual work.

So which one

Read this as a decision aid rather than a scoreboard. The left column wins more often than you would expect from a page on our own site.

If this is youUse
One person, one project, notes you can hold in your headA file in the repo. Genuinely. Do not add infrastructure.
Long-form docs written for humans: onboarding, architecture essays, RFCsA wiki. Keep it. It is better at that than a knowledge graph will ever be.
Retrieval over a large static corpus that does not changeYour own vector store. Staleness is not your problem, so you are not paying for a solution to it.
Facts that go out of date, more than one agent, and a need to know which parts still holdA maintained shared record. This is the case we build for.

Notice that the first three rows are not consolation prizes. Most projects are one of the first three rows, and the honest thing to say is that most projects do not need us yet.

The answer is usually both

The framing of this post is a lie of convenience. Almost nobody picks one. What actually happens is that different kinds of writing settle into different homes, and the useful question is which kind goes where.

The split that has held up for us is by shape, not by topic.

  • Long-form, written for a human to read start to finish belongs in a document. Onboarding guides, architecture essays, proposals with diagrams and comment threads. A knowledge graph is a bad book.
  • Small, atomic, and load-bearing belongs in a record. A single decision with its reason. A single lesson with the error string in it. A risk that expires when a specific piece of work lands. These are the things an agent needs one of, at a precise moment, and they are the things that go out of date individually.
  • Instructions to the agent about how to work in this repo belongs in the repo. Build commands, conventions, the test runner. That file should be short and it should be the first thing you prune.

The unit size is the real distinction. A wiki page called Authentication can be current in three sections and wrong in a fourth, and nothing about a page can express that. When a fact is its own object, it can be retired on its own, and the rest of what you know about authentication stays standing.

What you give up by moving

Two real costs, and pretending otherwise would be the kind of thing that makes a comparison page worthless.

You take on a hosted backend. An in-repo file needs nothing, works offline, and survives any vendor going away. A hosted record does not have those properties. An open, SQLite-backed version for self-hosting is planned, and we will say so on this page when it ships rather than before.

You lose review-in-the-same-diff. When knowledge lives in the repo, a change to it shows up in code review next to the change it describes. That is a genuinely good property, and moving the record out of the tree gives it up. What you get back is that the record is reachable from things that are not the checkout, which is the trade, and which one is worth more depends entirely on how your team works.

If the free version fits, use the free version

A carefully pruned notes folder that one person keeps honest beats any system that person resents. We would rather you used the folder for another year and came back when it stopped working than pay us for something you did not need.

The signal to watch for

There is a specific moment when the folder stops being enough, and it is easy to miss because it does not look like a crisis.

It is the day you open your notes to check something and find yourself reading them the way you would read a stranger's code: not what does this say but is this still true, and how would I even tell. When that question becomes routine, you have stopped having notes and started having archaeology, and no amount of reorganising the folder fixes it, because the missing thing was never organisation.

That is the point a maintained record starts earning its keep, and what it does about it is concrete rather than aspirational. A reversed decision is retired the moment it is replaced, and it comes back marked, carrying its reason, ranked below whatever replaced it. A lesson anchored to a file arrives when somebody edits that file, in whichever agent they happen to be using. A risk bound to the task that resolves it retires itself when that task closes. Two agents cannot claim the same work.

None of that requires you to keep the record honest by hand, which is the job that quietly stopped being possible somewhere between the fortieth note and the four hundredth.

  • objections
  • positioning

Try Stele

If a maintained shared record is the job you actually have, Stele is free to start.

Start freeRead the docs

Switch agents without losing the project

Source: https://stele-ai.dev/blog/switch-agents-keep-the-project

In practice

Switch agents without losing the project

Most people now use more than one coding agent, and every switch costs you the context you built up in the last one. That cost is not inevitable. It comes from where the context is stored.

July 25, 2026·Stele·7 min read·Updated July 27, 2026

Nobody uses one agent any more

A year ago picking a coding agent was a decision you made once. Now it is a decision you make several times a day, usually without noticing:

  • You plan in one tool because its reasoning is better, then build in another because it is faster and cheaper for mechanical work.
  • You hit a usage limit at 3pm and finish the afternoon somewhere else.
  • You use whatever is in the editor for small edits and whatever is in the terminal for anything that spans files.
  • A teammate uses something different from you entirely, on the same repository.

Every one of those transitions is a context reset. You re-explain the architecture. You re-explain the constraint you discovered yesterday. You re-explain why the obvious approach does not work here. The second agent starts from the same blank slate the first one did, and quite often it proposes the thing you already ruled out, because from where it sits that idea has never been tried.

Why per-tool memory cannot fix this

Every major harness now has some form of memory, and each one is reasonable in isolation. The problem is structural rather than a quality issue with any of them.

A tool's memory is scoped to the tool. It is written in that tool's format, stored wherever that tool stores things, and read by that tool only. Which means the better each vendor's memory gets, the more you lose at every switch. The feature that helps you within a session is the same feature that punishes you for leaving.

Memory owned by the agent is memory you lose when you change agents. Memory owned by the project is memory you keep.

There is a second problem that shows up on teams. Per-tool memory is also per-person. Your agent learned something about the deploy process on Tuesday. Your colleague's agent, on the same repository, has no idea. Two people, two agents, two divergent pictures of one project, and no mechanism by which either would ever discover the other exists.

Put the record under the project instead

The alternative is not complicated to describe. Decisions, tasks, and lessons live in a record attached to the project rather than to any agent. Every agent reads from it and writes to it. Nothing about the record cares which one is connected at the moment.

What that buys is that the handoff stops being a handoff. There is no export step, no summary to paste, no context to rebuild, because the context was never in the agent in the first place.

terminal

›/stele:start

project: acme-api · 3 tasks in flight · 2 for you

TASK-418 in progress · claimed by codex 40m ago · rate limiter rewrite

TASK-421 open · "Backfill needs a dry-run flag before we point it at prod"

pick up TASK-421, or keep going on something else?

That is a different agent, in a different terminal, possibly on a different machine, and it opens knowing what is happening. Including that something else is already holding the rate limiter work, which is the part that stops the two of them colliding.

Claiming is the part that makes it multi-agent

Shared notes are useful. Shared notes are not enough to run two agents at once, and the gap is small but absolute.

Two agents reading the same record will both see the same open work and both conclude it needs doing. Notes cannot prevent that, because a note cannot be held. What prevents it is that starting work means claiming it, the claim is atomic in the database, and the second agent to try is told no rather than told nothing:

the second agent, four seconds later

claim TASK-418
  → refused: already claimed by "codex" 40m ago
    worktree ~/code/acme-api · host mba-serkan
    released automatically if that session ends without completing

  open work nearby:
    TASK-421  Backfill needs a dry-run flag
    TASK-430  Token refresh drops the retry budget

The refusal carries who, where, and since when, because an agent that only learns no will try again. An agent that learns another session on another machine has been on this for forty minutes will go and do something else, which is the behaviour you wanted.

The stale-claim case has to work too, or the mechanism becomes a liability the first time a session dies mid-task. A claim held by a session that ended is releasable, so a crashed agent does not leave work locked behind it forever.

Ten harnesses, three different extension models

The part of this that took the most work is the least glamorous. "Works with every agent" sounds like one integration. It is ten, and they do not share a shape.

Stele ships adaptors for Claude Code, Cline, Codex, Copilot, Cursor, Devin, Gemini, Grok, Kimi, and OpenCode. Underneath, they divide into roughly three capability tiers, and being straight about the differences matters more than claiming parity.

  • Tool access. Every one of them speaks a tool protocol, so reading and writing the record works everywhere. This is the floor, and it is the same everywhere.
  • Lifecycle hooks. Some harnesses let code run at defined moments: before a prompt reaches the model, when a session starts, when the context is about to be compacted. Where those exist, relevant context arrives automatically and nobody has to ask for it. Where they do not, the agent has to reach for the record itself, which works but depends on it choosing to.
  • Enforcement. This is the sharpest difference. A read-only search helper, for instance, is genuinely read-only in some harnesses because the harness can restrict which tools it is allowed to touch. In others the same restriction is instruction-only: the helper is told not to write, and complies, but nothing physically stops it.

Why we spell that last one out

"Read-only" that is enforced and "read-only" that is requested are different guarantees, and a table with a tick in both columns would be lying about the second one. The behaviour is the same in normal use. The failure modes are not, and you should know which kind you have.

What the split day actually looks like

The abstract version of this is unconvincing, so here is the concrete one. A day where the work crosses three tools and nothing is re-explained.

Morning, planning agent. You talk through a change to how sessions are stored. The agent pulls up the original decision and the reason behind it, you argue with it, and you land somewhere new. What gets written is the new decision, the reason it replaced the old one, and three tasks with enough detail that someone who was not in the conversation could pick them up. The old decision is retired, not deleted, so the argument does not have to happen again in November.

Afternoon, terminal agent. Different tool, different model, no memory of the morning. It opens, sees three open tasks with written intent, claims the first, and starts. The decision from this morning arrives as context because the task links to it. Halfway through it discovers that the migration has to run before the deploy or live sessions drop. That becomes a risk bound to the deploy task, which means it retires itself when that task closes rather than haunting the record forever.

Evening, editor agent, different person. A colleague picks up the second task. They get the decision, the risk their teammate found four hours ago, and the fact that the first task is already claimed and in progress on someone else's machine. They do not ask anybody anything.

Nothing in that day required an export, a summary, or a handoff message. Three agents, two people, one record, and the only thing that moved between them was a project identifier.

Stable identifiers are what make it portable

A detail that seems cosmetic and turns out to be load-bearing: every node in the record has a short, human-readable identifier. TASK-418, KNOW-7, and so on. Not a UUID.

The reason is that these identifiers are the shared vocabulary across every surface. You can say TASK-418 to an agent in a terminal, paste it into a message to a colleague, type it into the web UI, or mention it to a planning tool that has never touched your code, and all of them resolve it to the same thing.

Agent-facing responses deliberately never contain raw UUIDs or internal identifiers. It is a hard boundary rather than a formatting preference: the moment an agent sees an internal identifier, it will helpfully repeat it back to you, and you have a string in your conversation that nothing else on any surface can resolve.

What does not carry over

Two limits, because the pitch above is cleaner than reality.

Your existing per-tool memory does not migrate itself. Whatever your current agent has accumulated in its own store stays there. Some vendors expose it, most do not, and we are not going to promise an import for formats we do not control. Practically, the record fills from the work you do after you set it up, plus a backfill pass over the repository and its history if you want a running start.

Continuity is not a transcript. What carries across a switch is the durable stuff: decisions, lessons, risks, what is in flight, what was tried. The texture of the conversation you had with the previous agent does not carry, and should not. If your workflow depends on the second agent knowing exactly how you phrased something an hour ago, this does not replace that.

What it does replace is re-explaining your architecture for the fourth time this week, and the specific irritation of watching a fresh agent confidently propose the thing you spent Tuesday ruling out.

  • continuity
  • mcp
  • harnesses

Try Stele

One record, every agent you use. Stele is free to start. It ships adaptors for ten harnesses today.

Start freeRead the docs

How an agent session uses Stele

Source: https://stele-ai.dev/blog/how-an-agent-session-uses-stele

How it works

How an agent session uses Stele

You talk to your agent exactly the way you already do. What changes is what happens in the eight seconds before it answers, and what it leaves behind afterwards.

July 25, 2026·Stele·7 min read·Updated July 27, 2026

The turn, end to end

Here is the shape of a single prompt, before we take any of it apart. You type something. Before the model sees it, a hook runs, searches the project record, walks outward from what it found, and prepends whatever it thinks bears on the question. The model answers with that in hand. If real work starts, a task gets created and claimed. As the session learns something durable, it gets written back.

Read in a list like that it sounds like an ordinary retrieval pipeline. The parts that turned out to be interesting are the ones where the obvious implementation was wrong.

Before the answer: seeds, then one hop

The hook has a hard budget. Six seconds for the graph walk, eight for the whole hook, and if it runs out it returns whatever it has rather than making you wait. Nothing about your prompt is blocked on this succeeding, which is deliberate: a memory layer that can stall your agent is worse than no memory layer.

It also does not run on everything. Prompts under 25 characters skip the walk entirely, because ok and keep going and now the tests do not have enough signal to retrieve against, and injecting a guess into a short turn is how you get an agent that suddenly changes the subject. The exception is that naming an identifier directly bypasses the length gate: type TASK-418 and that node gets resolved and pulled in whole, however short the sentence around it.

Otherwise it is two stages.

Seeds. Full-text and semantic search against the record produce up to ten starting points. This is the part every retrieval system has, and on its own it is not very interesting.

The walk. From those seeds it follows edges outward. One hop by default. This is the part that earns its keep, because it reaches facts that share no words with your question.

The whole argument for a graph rather than a pile of embeddings is that the most important fact is often the one you would never have thought to search for.

You ask about the checkout flow. Search finds the checkout decision. The walk follows an edge from that decision to a risk about the payment retry budget, which does not mention checkout anywhere in its text. You would not have retrieved it. You needed it.

Ranking, and the weights that took several tries

Everything the walk reaches gets scored, and the scoring is unglamorous on purpose: small integers, added up, minus a penalty for distance. The absolute number means nothing. It is an ordering key.

the weight tables

by category            by relation followed        distance
  decision      3        fulfills        3           −0.5 per hop
  architecture  3        blocks          3
  goal          2        supersedes      3          container tasks
  risk          2        caused-by       2           epic/objective
  lesson        2        contradicts     2           /milestone: −3
  priority      2        relates-to      1
  gap           1        mentions        1
  fyi           1        embeds          1
  opportunity   1        expires-on      1

Two of these entries are the scar tissue, and they are the ones worth explaining.

Container tasks score negative three. Epics, objectives, and milestones are roll-up nodes. They aggregate progress from their children and carry almost no signal about the work in front of you. But because a child task blocks its parent, they are reached through the strongest relation tier, which meant they arrived with a high score and flooded the window with entries like Q3 Platform Work. A strong negative base cancels even that edge, so they sink to the bottom and surface only when nothing more useful is nearby.

Components score zero. A component is a subsystem folder. Asking about MCP would match the component literally named Plugin MCP, which then tied with real knowledge and took a slot for the privilege of telling you that a subsystem exists.

The pattern in both fixes is the same, and it is the single most repeated decision in this system: downrank, do not exclude. An excluded node cannot come back when it genuinely is the answer. A downranked one can, and it costs a slot only when nothing beats it.

The risk callout

There is one category the walk treats specially, and it is the feature that most changes what an agent does.

If the walk reaches a risk one or two hops out, through an edge that means something causal rather than merely topical, that risk is lifted out of the ranked list and rendered as its own warning block. Only four relations qualify. A risk reached through relates-to or mentions is a risk in the neighbourhood; a risk reached through caused-by, blocks, contradicts, or a supersession edge is a risk on the path you are about to walk down.

What lands in front of the model looks roughly like this:

injected before the prompt

── from the project record ───────────────────────────────
DECISION  Rate limiting sits in front of auth, not behind
          Nov 2025: moving it behind auth let unauthenticated
          traffic reach the token-metered path. 4h outage.

TASK-418  in progress · claimed by codex 40m ago
          Rate limiter rewrite

⚠ RISK on your path
  "Auth middleware ordering is load-bearing" (via caused-by)
  Changing the order here has broken production once.
  Confirm the limiter still rejects before anything metered runs.
──────────────────────────────────────────────────────────

The cap on returned entries has to know about this. A risk callout is by definition a walked node rather than a seed, so a cap that filled up on seeds first would drop the risk exactly when the record was rich enough for it to matter, which is to say always.

Twenty entries, and no way to ask for more

The window is capped at twenty nodes. People ask why it is not configurable upward, and the answer is that we do not think a bigger number is a better product.

Injected context competes with your actual prompt for the model's attention. Forty entries is not twice as helpful as twenty. It is twenty helpful entries plus twenty distractions, and the distractions dilute the ones that mattered. If the right fact is not in the top twenty, the fix is better ranking, not a longer list.

False sufficiency is the real risk here

The failure we watch for is not the agent missing context. It is the agent treating what arrived as the complete picture and skipping the search it would otherwise have run. Injected context now says out loud that it is a starting set rather than everything, which helps. How much it helps is one of the things we are still measuring.

Going deeper on purpose

Auto-injection is a starting set. When the agent needs the full picture, it dispatches a read-only search helper that interrogates the record in its own context and returns a short, cited briefing.

The context isolation is the whole point. Investigating a question properly means dozens of searches and a lot of raw hits, and doing that in the main thread fills your window with search noise before any work starts. The helper absorbs that and hands back a paragraph.

It cannot write. In harnesses that can restrict tool access, that is enforced by the harness. In the others it is instruction-only, which is worth knowing rather than glossing.

Claiming the work

When something real starts, the agent creates a task and claims it, and the claim is atomic. Two agents cannot hold the same task, and the second one to try gets told who has it and since when rather than just being refused.

A task is also the unit of continuity. It is what makes what was I doing answerable tomorrow, by you or by a different agent, with the decisions and risks that were linked to it still attached.

Writing back, which is the part that is genuinely hard

Everything above is retrieval, and retrieval is the easy half. The hard half is that a record is only worth reading if something is putting good things into it, and the moment when an agent knows something worth recording is the moment it is busy doing something else.

Three things help, and none of them is a rule that blocks you.

Nudges at the right moment. Before the first code change of a session, a one-time prompt to check the record first, so the agent does not re-decide something already settled. Around completion, a prompt to write down what was learned. The agent can proceed regardless. The point is to make the honest update the path of least resistance rather than to enforce it.

Concrete triggers in the text. A lesson written abstractly is invisible exactly when it would help. A lesson containing the literal error string surfaces the moment somebody hits that error again. So the guidance asks for the real thing:

two versions of the same lesson

✗ "Be careful with type generation, it can break the build."

✓ "Regenerating DB types breaks the SDK build: columns filled by a
   BEFORE INSERT trigger come out required, and tsc fails with
   TS2769: No overload matches this call … Property 'short_id' is
   missing. Port new columns by hand instead."

The second one is findable by the person having the problem, because it contains the words they are about to paste into a search.

Anchors. A fact can name the files and symbols it is about, which makes editing that file the trigger rather than mentioning it. Fuzzy matching catches the fact you described in other words. Anchors catch the fact you were not describing at all.

What the loop actually gives you

Every part of this runs by default, on every prompt, in whichever agent you happen to be using. Relevant decisions arrive before the model answers. A risk on the path you are about to take is lifted out of the ranking and put in front of you as a warning. Work is claimed atomically, so a second agent is told who holds it rather than colliding with them. What the session learns goes back into the record with a retirement condition attached, so it can age out on its own.

It is retrieval on a budget rather than a guarantee that every relevant fact reaches the model on every turn, and it gets better as the record fills. How much all of it changes what an agent actually does is a measurable question, and we are building the apparatus to answer it properly rather than asserting a number.

  • how-to
  • agent loop

Try Stele

This loop runs on every prompt, in whichever agent you happen to be using. Stele is free to start.

Start freeRead the docs

Stele vs the usual ways to give agents memory

Source: https://stele-ai.dev/blog/how-stele-compares

Comparison

Stele vs the usual ways to give agents memory

When people ask what Stele is like, they usually mean one of five different jobs: a file in the repo, a wiki, memory built into the editor, a memory API for product builders, or a coding agent that keeps its own notes. Those are not competitors to each other, and comparing them as peers produces nonsense.

July 25, 2026·Stele·7 min read·Updated July 27, 2026

The question that sorts them

Nearly everything in this space ships under the word memory, which has stopped carrying information. Three narrower questions separate the options cleanly, and every real difference below is one of these three wearing a costume.

  1. Who is the memory for? One agent, one person, one team, or the users of a product you are building.
  2. What can read it? One vendor's tool, any tool, or only your own application code.
  3. What happens when a fact stops being true? Nothing, somebody edits it, or the system does something about it.

Stele's answers: a project and everyone working on it, any agent that speaks a tool protocol, and a lifecycle that retires facts on its own. Whether that combination is worth anything depends entirely on which row of the table below you are in.

This post compares kinds of memory. For the named products in each kind, with prices and sources, see the best memory tools for AI coding agents.

A file in the repo

Genuinely good at: costing nothing, working in every tool, being reviewed in the same pull request as the code it describes, surviving your vendors, and working on a plane. That last set is not a small list, and nothing else on this page replicates all of it.

Where it stops: the file has no opinion about which of its lines are still true, retrieval means guessing the word it was written with, and when the agent cannot find the right paragraph it loads the whole thing. Those three failures arrive together, roughly when the pile outgrows one person's head.

Stay here if one person can keep the folder honest by hand. That is a real answer, not a consolation prize, and it covers a lot of projects. We went through the three failure modes in detail separately, including how to tell when you have hit them.

Notion, Obsidian, or a team wiki

Genuinely good at: long-form writing for humans. Onboarding guides, architecture essays, design proposals with diagrams and comment threads. Better at that than a knowledge graph will ever be, and your team is already in it.

Where it stops: your coding agent does not open the wiki before answering you, and it does not open the wiki before editing a file. Retrieval requires a human to remember the page exists and go and paste it in.

A polished knowledge base the agent never opens is documentation you hope someone finds. It is not agent memory, and the polish does not change that.

There is a second problem that is nobody's fault. Wiki pages are written as documents, and documents get updated when someone reorganises them, not when a decision inside them reverses. The unit is too big. A page called Authentication can be simultaneously current in three sections and wrong in a fourth, and nothing about the page can express that.

Keep it for the human-readable half. The two are not mutually exclusive and we use both.

Memory built into the harness

Genuinely good at: zero setup, and knowing things about the tool it lives in that no external system can see. If you use one agent and only one agent, this is likely to be enough, and it is free.

Where it stops: the memory is scoped to the vendor. Which means it has a property worth naming plainly: the better it gets, the more you lose when you switch. It is also usually per-person, so on a team you end up with several private, divergent pictures of one shared project, and no mechanism by which any of them would discover the others.

The thing to watch as vendors ship here is not whether they built memory. Everyone will. It is whether you can get it out.

Stay here if you use one tool and work alone. Move when you notice you are re-explaining the same architecture to a second agent every week.

Memory APIs: Mem0, Zep, and the rest

Genuinely good at: being memory infrastructure for an application you are building. If you are shipping a support agent or a consumer assistant and you need it to remember users across sessions, this is the correct category and we are not an alternative to it.

These are the most technically overlapping systems on this page, and some of the ideas are close to ours. Zep's Graphiti tracks when a fact was valid and invalidates it on contradiction rather than deleting it, which is a real retirement model rather than a marketing claim about one. Anyone telling you that nobody else handles stale facts has not looked.

The difference that matters is authorship, and it shapes everything downstream. These systems mostly extract facts from conversation: an entity pipeline reads a transcript and infers what should be remembered. Stele's entries are written on purpose, typed at the moment of writing, by whoever knew the thing.

the two shapes

extracted     "user prefers dark mode"          inferred from a transcript
              "project uses postgres"           inferred from a transcript

authored      decision · "Rate limiter goes in front of auth"
                reason:   Nov 2025 outage, unauthenticated traffic
                          reached the metered path
                anchors:  src/middleware/index.ts, buildMiddlewareChain
                retires:  when superseded, with the reason recorded

Extraction scales to conversation volume and gets you facts nobody would have bothered to write down. Authoring gets you a reason, a type, a set of anchors, and a retirement condition, none of which can be reliably inferred, because the person deciding is the only one who knows why.

Buy a memory API when you are building an AI product. Do not buy one when you want a record your coding agents share.

Coding agents that are memory-first

Genuinely good at: deep integration. When memory is a first-class part of the agent rather than a plugin on top of it, the agent can do things a plugin cannot, and the experience is more coherent than anything bolted on.

Where it stops: it is a harness, so adopting it means adopting that harness. If the tool you want is the tool you are already using, the memory is not available to you.

This is a real fork in the road rather than a flaw. The bet on that side is that memory is important enough to pick your agent around. Our bet is the opposite: that people will keep using several agents and changing them, so the record should sit underneath and outlive whichever one is connected today. Both bets are defensible. Ours is worse if the field consolidates onto one tool.

Not a replacement for Linear or Jira

Stele has tasks, so this comes up. The tasks exist for one reason: agents need a unit of work that can be atomically claimed, or two of them do the same thing at once. That is coordination primitives, not a planning suite.

There is no roadmap view, no sprint planning, no stakeholder reporting, and there should not be. Those tools own how a company plans. Reach for a shared record when the missing piece is what agents know between sessions.

Side by side

ApproachWho reads itAcross agents?Who keeps it true?Best fit
Repo markdownWhoever opens the fileYes, if committedYou, by handSmall solo notes
Wiki, Notion, ObsidianHumansOnly if you wire it inWhoever edits the pageHuman documentation
Native harness memoryOne vendor's agentUsually noVaries by vendorOne-tool workflows
Memory APIsYour product's usersInside your appContradiction handlingBuilding AI products
Memory-first agentThat agentPer harnessVariesCommitting to that tool
SteleYour agents and your teamYes, ten harnessesLifecycle and reviewShared project record

The bet, stated so you can disagree with it

Nothing in the middle column above is unique to us, and we are not going to pretend otherwise. Retirement on contradiction exists elsewhere. Cross-tool access exists elsewhere. Task coordination obviously exists elsewhere.

The specific bet is that these three belong in one record:

  • Knowledge that is authored and typed, with a reason and a retirement condition chosen when it is written.
  • Work that is claimable, so several agents can run without colliding.
  • Both of them readable by whichever agent you happen to be using, and by a person on a phone.

Split those across three products and the connections between them disappear, and the connections are where most of the value is: a risk that retires because the task fixing it closed, a decision that surfaces because you touched the file it is anchored to, a completion note that explains a feature nothing else documents.

If you think agents will consolidate onto one tool with excellent native memory, this bet is wrong and you should not take it.

When to skip us entirely

  • A file and an occasional prune covers your project. Keep it. It is simpler and it is free.
  • You need memory inside a consumer AI product. Buy a memory API.
  • You need planning for a large organisation. Buy a planning tool.
  • You cannot take a hosted dependency. An open, self-hostable path is planned, and we will say so here when it exists rather than before.

What is left is narrower than a landing page usually admits, and it is also the case we are best in the world at: teams whose facts go out of date, who use more than one agent, and who need to know which parts of what they wrote down still hold.

For that case the answer is not a wish. A reversed decision stops being served the moment something replaces it, and comes back marked, with its reason, ranked below its replacement. A lesson anchored to a file fires when somebody edits that file. A risk bound to a task retires itself when that task closes. Work is claimed atomically, so two agents never do it twice. All of that is running today, in ten harnesses, on one record your whole team reads.

  • comparison
  • agent memory
  • positioning

Try Stele

If the shared-record job is the one you have, Stele is free to start. If it is not, the docs will still tell you what the model is.

Start freeRead the docs