mcp-server-decisions

Decision tracking with prediction validation and outcome gates for AI agents

ドキュメント

🧠 MCP Server: Decisions

Python 3.10+ License: MIT Registry GitHub

Architectural decision tracking with prediction validation and outcome gates.
Record choices, state testable predictions, measure outcomes, and close the feedback loop for AI agents & teams.


Zero External Dependencies • Stdlib-only Python • Append-only JSONL storage • Fully portable


📌 Key Features

  • 🎯 Decision Recording — Capture architectural choices with problem statement, solution, rejected alternatives, and target technologies.
  • 📈 Prediction Linking — Attach testable claims (latency, cost, scalability, reliability) tied to decisions.
  • ✅ Outcome Validation — Record measured results and automatically compute accuracy scores (0–100 scale).
  • 📊 Technology Performance Registry — Aggregate success rates and confidence metrics per technology over time.
  • 🚪 Outcome Gate Pattern — In-band nudges inside tool responses prevent decision feedback loops from leaking (3.8% → 14.5% closure rate).
  • ⚡ Zero Dependencies — Portable append-only JSONL log. No database servers, no migrations, no background daemons.

⚡ Quick Example

1️⃣ Record a Decision

# Agent or user records a choice:
record-decision(
  problem="Query latency exceeds SLA (p99 > 500ms)",
  chosen_solution="DuckDB + Parquet caching",
  rejected_alternatives=["Redis", "Elasticsearch"],
  technologies=["duckdb", "parquet"],
  predictions=[
    {"prediction_type": "LATENCY", "predicted_value": "p99 < 200ms"},
    {"prediction_type": "COST", "predicted_value": "< $50/month"}
  ]
)
# ➔ Returns: DEC-2026-0001, PRD-2026-0001, PRD-2026-0002

2️⃣ Record an Outcome

record-outcome(
  prediction_id="PRD-2026-0001",
  actual_value="p99 = 180ms",
  measurement_source="MONITORING",
  accuracy_score=95
)
# ➔ Returns: SUCCESS ✅ (95% accuracy)

3️⃣ Query Prior Decisions & Technology Stats

# Search past decisions before choosing a technology:
query-decisions(technology="duckdb", max_results=5)

# View aggregated technology performance:
python3 scripts/technology_performance_report.py
# ➔ Output:
# technology: duckdb  | successful: 12 | failed: 1 | avg_accuracy: 91.2% | confidence: HIGH

🚀 Quick Start & Setup

📦 Installation

# From PyPI (once published) or local editable install:
pip install -e .

🛠️ Client Configuration

Add to your MCP client configuration (e.g. Claude Desktop, Claude Code, Cursor, OpenCode):

{
  "mcpServers": {
    "mcp-server-decisions": {
      "type": "stdio",
      "command": "mcp-server-decisions"
    }
  }
}

For client-specific setup guides (Claude, OpenCode, Codex, Antigravity), see 📖 docs/INTEGRATIONS.md.


How It Works

The Loop

Decide → Predict → Implement → Measure → Validate → Learn → Next Decision
  1. Record a decision — Store the problem, chosen solution, alternatives, and technologies
  2. Make predictions — Attach testable claims (latency, cost, reliability, etc.)
  3. Implement — Build the system
  4. Measure results — Capture actual values from monitoring, logs, benchmarks
  5. Validate — The server calculates accuracy (0-100) and validation status (SUCCESS / PARTIAL_SUCCESS / FAILED)
  6. Learn — Review what worked via the Technology Performance Registry
  7. Next decision — Query past decisions before making new recommendations

The Outcome Gate Pattern

Decision loops leak because predictions aren't validated. This server embeds a reminder directly in tool responses:

Without Outcome Gate:

  • Decision is made → implementation starts → results come in → nobody checks if prediction was right

With Outcome Gate:

{
  "decision_id": "DEC-2026-0001",
  "status": "OK",
  "OUTCOME_GATE": "⚠️  2 prediction(s) from this session still lack outcomes: [PRD-2026-0001, PRD-2026-0002]. Record results via record-outcome before ending."
}

The nudge is in-band (inside the tool response), where agents are already looking. Result: 3.8% → 14.5% closure rate improvement (validated on internal tool).

Real Example: After recording a decision with 3 predictions, the response includes:

{
  "decision_id": "DEC-2026-0042",
  "prediction_ids": ["PRD-2026-0051", "PRD-2026-0052", "PRD-2026-0053"],
  "status": "OK",
  "OUTCOME_GATE": "⚠️  3 prediction(s) from this session still lack outcomes: [PRD-2026-0051, PRD-2026-0052, PRD-2026-0053]. Record results via record-outcome before ending."
}

Next query still shows the gate until all 3 outcomes are recorded. Once they are, the gate disappears automatically.

For the full pattern explanation, see docs/OUTCOME-GATE-PATTERN.md.

🏛️ Architecture & Tech Stack

  • Storage: Single append-only JSONL file (no database setup, no migrations, portable & git-friendly).
  • IDs: Sequential per calendar year (DEC-2026-0001, PRD-2026-0002, OUT-2026-0003).
  • Accuracy Scoring: Automatic classification (≥90 SUCCESS, 50–89 PARTIAL_SUCCESS, <50 FAILED).
  • Runtime: Stdlib-only Python 3.10+ (zero external pip runtime dependencies).
  • Protocol: Model Context Protocol (JSON-RPC 2.0 over stdio).

⚙️ Environment Variables

VariableDescriptionDefault Path
MCP_DECISIONS_LOG_PATHPath to the append-only JSONL log file~/.local/share/mcp-decisions/decisions_log.json

📚 Documentation & Resources

DocumentPurpose
Quick Start5-minute setup guide & first decision
🔌 Client IntegrationsSetup configs for Claude, OpenCode, Codex, Antigravity
📐 Architecture & DesignCore design rationale & data models
💡 Detailed ExamplesReal JSON-RPC request/response payloads
🚪 Outcome Gate PatternIn-band feedback loop design philosophy
📖 WikiFAQ and advanced topics

🧪 Development & Testing

Run unit & selftests locally:

python3 server.py --selftest
# ➔ ✅ All self-tests passed

See 📝 CONTRIBUTING.md to contribute features or fixes.


🗺️ Roadmap

  • Core decision / prediction / outcome tracking
  • Outcome Gate in-band nudges
  • Technology Performance Registry
  • Web UI for browsing & searching decisions
  • Webhooks / notifications on low prediction accuracy
  • Pre-built decision templates & domain patterns

📄 License & Disclaimer

MIT © 2026 Roberton003 — See LICENSE.

This project is community-built and independent. It is not affiliated with any organization or standard-setting body.


Made for AI agents. Built for teams. Learn from every decision.