caveman-evidence-review

作者: juliusbrussee

审查Caveman Cloud只读证据:成本、Cave Score、Cave Plan、工作流、追踪、延迟、错误、压缩、路由及已验证的节省。当用户询问Caveman发现了什么、LLM支出去向、成本或质量变化原因、哪些工作流需要关注,或要求进行追踪或分析审查时使用。优先使用Caveman MCP工具;回退到CLI JSON。

npx skills add https://github.com/juliusbrussee/caveman --skill caveman-evidence-review

Review Caveman evidence

Act as a read-only operator. Build conclusions from current Caveman data, not from repository guesses. Never start, approve, cancel, or roll back an experiment from this skill.

Hard rules

  1. Keep these buckets separate:
    • measured provider-complete list-price cost;
    • inferred daily headroom;
    • verified ledger savings;
    • evidence cost. Never add or relabel them.
  2. Do not fetch prompt, completion, tool, or artifact payloads unless the user explicitly asks for payload review. Metadata, spans, timing, models, token counts, status, and optimizer attribution are enough for the default review.
  3. Scope every read to the project selected by Caveman context. Never supply an organization id.
  4. Empty results are evidence of no current signal, not zero cost or zero risk.
  5. Cite trace ids and exact time windows used. Do not claim a cause from an aggregate alone.

Step 1 — Load context

Prefer MCP:

caveman_context {}

CLI fallback:

caveman cloud whoami
caveman cloud projects list

Stop if login or project selection is missing. Ask the user to run caveman login or select a project; never guess.

Step 2 — Establish baseline

Use caveman_report for:

  • overview
  • costs
  • score
  • workflows
  • verified_savings

Then use caveman_plan for ranked daily headroom. If question is narrow, skip unrelated reports. Read shortest set that can answer it.

CLI fallback:

caveman cloud costs
caveman cloud score
caveman cloud plan --json

State report window and basis before interpreting direction.

Step 3 — Test the leading explanation with traces

Use caveman_trace_search. Choose a bounded window and closed filters: workflow, agent, model, provider, error code, runtime mode, cache status, optimization id, status class, token/cost/latency bounds, compression, or monitor verdict.

Useful groupings:

  • workflow — find jobs driving cost or failures;
  • model — compare model mix;
  • session — isolate retry or loop behavior;
  • ungrouped — identify exact traces.

Compare a suspect cohort with a control cohort or earlier bounded window. Do not infer causality from one expensive trace.

CLI fallback:

caveman cloud traces search \
  --workflow <slug> \
  --from <RFC3339> \
  --to <RFC3339> \
  --sort total_cost_usd \
  --dir desc \
  --limit 25

Step 4 — Inspect representative traces

Call caveman_trace_get for a small number of high-signal trace ids. Inspect request and span metadata, latency, status, token counts, cache state, applied optimizers, and model route. Keep payload retrieval off.

CLI fallback:

caveman cloud traces show <trace-id> --spans

Step 5 — Report

Use this shape:

## Caveman evidence review

Scope: <project> · <from> to <to>
Measured cost: <value and basis>
Verified savings: <ledger value, kept separate>
Inferred headroom: <per-day band, kept separate>

Findings:
1. <finding> — <aggregate evidence> — traces <ids>
2. <finding> — <aggregate evidence> — traces <ids>

Unproven:
- <plausible explanation lacking a control, trace, or eval>

Next read-only check:
- <one bounded query>

Possible action:
- <proposal only; use caveman-manage for read-only lifecycle review and safety gate>

If data is missing, name missing signal and stop at strongest supported statement. Never turn a catalog subtotal into an invoice or an experiment result into verified savings.

来自 juliusbrussee 的更多技能

caveman
juliusbrussee
超压缩沟通模式。通过像原始人一样说话,将令牌使用量削减约75%,同时保持完整的技术准确性。支持强度级别:lite、full(默认)、ultra、wenyan-lite、wenyan-full、wenyan-ultra。当用户说“caveman mode”、“talk like caveman”、“use caveman”、“less tokens”、“be brief”或调用/caveman时使用。在请求令牌效率时也会自动触发。
communicationproductivity
caveman-commit
juliusbrussee
超精简提交信息生成器。去除提交信息中的冗余内容,同时保留意图和理由。采用常规提交格式。主题不超过50个字符,仅在“原因”不明确时添加正文。当用户说“写提交”、“提交信息”、“生成提交”、“/commit”或调用/caveman-commit时使用。暂存更改时自动触发。
developmentcode-review
caveman-compress
juliusbrussee
将自然语言记忆文件(CLAUDE.md、待办事项、偏好设置)压缩为穴居人格式以节省输入令牌。保留所有技术内容、代码、URL和结构。压缩版本覆盖原文件。人类可读备份保存为FILE.original.md。触发方式:/caveman-compress 文件路径 或 "压缩记忆文件
developmentdocument
caveman-help
juliusbrussee
所有穴居人模式、技能和命令的快速参考卡。一次性显示,非持久模式。触发词:/caveman-help、"caveman help"、"what caveman commands"、"how do I use caveman"。
developmentdocumentproductivity
caveman-review
juliusbrussee
超精简代码审查评论。减少PR反馈中的噪音,同时保留可操作的关键信息。每条评论仅一行:位置、问题、修复。当用户说“审查此PR”、“代码审查”、“审查差异”、“/review”或调用/caveman-review时使用。审查拉取请求时自动触发。
developmentcode-review
caveman-stats
juliusbrussee
显示当前会话的实际令牌使用量和预估节省量。直接从Claude Code会话日志读取——无AI估算。通过/caveman-stats触发。输出由mode-tracker钩子注入;模型本身不计算这些数字。
developmentdata-analysis
cavecrew
juliusbrussee
我们要求翻译一段文本,目标语言是简体中文。文本内容是关于一个名为"cavecrew"的代理技能的描述。需要保留名称"cavecrew"以及其中的子代理名称如"cavecrew-investigator"、"cavecrew-builder"、"cavecrew-reviewer"等。同时要保留技术术语如"Explore"、"diff review"等。不要添加任何额外内容,只翻译<text>内的文本。 翻译时注意:保持原意,简洁。原文中有一些英文术语和代码风格,需要保留。例如"cavecrew-investigator"等子代理名称不翻译。"Explore"可能是一个命令或功能,保留不译。"caveman-compressed"可以翻译为"穴居人压缩"或类似,但为了保持风格,可以译为"穴居人式压缩"。"~60% smaller"译为"约小60%"。"main context"译为"主上下文"。"Trigger"译为"触发词"或"触发条件"。 整体翻译要流畅,符合中文表达习惯
developmentcode-reviewapi
caveman-explore
juliusbrussee
只读仓库浏览器。在冷启动探索、广泛的跨文件定位,或直接搜索失败且需要定位目标所在位置时,应主动使用。当问题已明确指向具体文件或符号,或上一轮已返回可用的 file:line 证据时,跳过此工具。仅返回紧凑的 path:line 引用;其读取和 grep 操作不会进入主对话。