caveman-manage

作者: juliusbrussee

检查 Caveman Cloud 的评估门控实验生命周期,并阻止不安全的执行。当用户要求启动、批准、取消、提升或回滚 Caveman 实验,或询问实验证据支持什么操作时使用。先读取证据;在服务器权威的转换和证据门控上线之前,不要执行生命周期变更。

npx skills add https://github.com/juliusbrussee/caveman --skill caveman-manage

Manage eval-gated experiments

Treat every lifecycle change as a production control action. Read current state and results, then report one supported recommendation or block. Current agent MCP is intentionally read-only: control-api does not yet enforce a complete lifecycle transition table and evidence gate atomically.

Non-negotiable gates

  1. A request to review, inspect, explain, or recommend authorizes reads only.
  2. Never approve an experiment whose results are pending, whose required guardrails are absent, or whose evidence reports a breach.
  3. Never convert experiment lift into verified_savings. Only active real traffic plus provider-causal, provider-complete ledger evidence can do that.
  4. Never supply an organization id. Project and tenant scope come from the logged-in Caveman identity and server RBAC.
  5. Never execute a lifecycle mutation, even after user approval. Exact <action>:<experiment_id> strings are agent-generatable and are not proof of human intent.
  6. Unknown states and server errors fail closed. Report exact cave_snake_code.

Step 1 — Load project and experiment

Prefer MCP:

caveman_context {}
caveman_experiment_get {"action":"get","experiment_id":"<id>"}
caveman_experiment_get {"action":"results","experiment_id":"<id>"}

Use {"action":"list"} when the user has not named an id.

CLI fallback:

caveman cloud experiments list
caveman cloud experiments show <id>
caveman cloud experiments results <id>

Stop if login, project, experiment, or results are unavailable.

Step 2 — Evaluate evidence

Report:

  • current lifecycle state and safety class;
  • control and candidate sample sizes;
  • quality or eval result;
  • latency, error, cost, retry, drop, and escalation guardrails when present;
  • evidence cost;
  • rollback or hold reason;
  • whether result is pending, failed, promotable, or active.

Absence is not a pass. If a required field is absent, state evidence incomplete and do not propose approval.

Step 3 — Propose one action

Allowed actions:

  • start — only from a startable draft or queued state with configured graders;
  • approve — only with complete passing evidence and a safety class the current role may approve;
  • cancel — stop a non-active experiment the user no longer wants;
  • rollback — revert an active or harmful change through the server's linked policy path. Current deployments may reject this honestly with cave_not_implemented; never describe that response as a rollback.

Show recommendation and id:

Proposed action: approve experiment 7f...
Reason: candidate passed quality and every configured guardrail.
Execution: blocked until server-authoritative lifecycle and evidence gates ship.

Do not treat earlier generic statements such as "manage it" or "do what is best" as mutation approval.

Step 4 — Block unsafe execution

Do not emit or run an executable lifecycle command. Explain that current server does not yet enforce every evidence/state transition atomically. CLI and MCP agent surfaces therefore expose experiment reads only.

Step 5 — Re-read after external operator action

If operator says they executed command, read detail and results again. Report server-observed post-state, audit or result response, and any policy-delivery status returned. Never infer success from operator intent alone.

Use this close:

Action: <action> <experiment-id>
Before: <state>
Server response: <status and cave_snake_code if any>
After: <re-read state>
Basis: experiment evidence only. Verified savings unchanged unless the signed
ledger independently records active, provider-causal real-traffic savings.

来自 juliusbrussee 的更多技能

caveman
juliusbrussee
超压缩沟通模式。通过像原始人一样说话,将令牌使用量削减约75%,同时保持完整的技术准确性。支持强度级别:lite、full(默认)、ultra、wenyan-lite、wenyan-full、wenyan-ultra。当用户说“caveman mode”、“talk like caveman”、“use caveman”、“less tokens”、“be brief”或调用/caveman时使用。在请求令牌效率时也会自动触发。
communicationproductivity
caveman-commit
juliusbrussee
超精简提交信息生成器。去除提交信息中的冗余内容,同时保留意图和理由。采用常规提交格式。主题不超过50个字符,仅在“原因”不明确时添加正文。当用户说“写提交”、“提交信息”、“生成提交”、“/commit”或调用/caveman-commit时使用。暂存更改时自动触发。
developmentcode-review
caveman-compress
juliusbrussee
将自然语言记忆文件(CLAUDE.md、待办事项、偏好设置)压缩为穴居人格式以节省输入令牌。保留所有技术内容、代码、URL和结构。压缩版本覆盖原文件。人类可读备份保存为FILE.original.md。触发方式:/caveman-compress 文件路径 或 "压缩记忆文件
developmentdocument
caveman-help
juliusbrussee
所有穴居人模式、技能和命令的快速参考卡。一次性显示,非持久模式。触发词:/caveman-help、"caveman help"、"what caveman commands"、"how do I use caveman"。
developmentdocumentproductivity
caveman-review
juliusbrussee
超精简代码审查评论。减少PR反馈中的噪音,同时保留可操作的关键信息。每条评论仅一行:位置、问题、修复。当用户说“审查此PR”、“代码审查”、“审查差异”、“/review”或调用/caveman-review时使用。审查拉取请求时自动触发。
developmentcode-review
caveman-stats
juliusbrussee
显示当前会话的实际令牌使用量和预估节省量。直接从Claude Code会话日志读取——无AI估算。通过/caveman-stats触发。输出由mode-tracker钩子注入;模型本身不计算这些数字。
developmentdata-analysis
cavecrew
juliusbrussee
我们要求翻译一段文本,目标语言是简体中文。文本内容是关于一个名为"cavecrew"的代理技能的描述。需要保留名称"cavecrew"以及其中的子代理名称如"cavecrew-investigator"、"cavecrew-builder"、"cavecrew-reviewer"等。同时要保留技术术语如"Explore"、"diff review"等。不要添加任何额外内容,只翻译<text>内的文本。 翻译时注意:保持原意,简洁。原文中有一些英文术语和代码风格,需要保留。例如"cavecrew-investigator"等子代理名称不翻译。"Explore"可能是一个命令或功能,保留不译。"caveman-compressed"可以翻译为"穴居人压缩"或类似,但为了保持风格,可以译为"穴居人式压缩"。"~60% smaller"译为"约小60%"。"main context"译为"主上下文"。"Trigger"译为"触发词"或"触发条件"。 整体翻译要流畅,符合中文表达习惯
developmentcode-reviewapi
caveman-explore
juliusbrussee
只读仓库浏览器。在冷启动探索、广泛的跨文件定位,或直接搜索失败且需要定位目标所在位置时,应主动使用。当问题已明确指向具体文件或符号,或上一轮已返回可用的 file:line 证据时,跳过此工具。仅返回紧凑的 path:line 引用;其读取和 grep 操作不会进入主对话。