caveman-manage

โดย juliusbrussee

ตรวจสอบวงจรชีวิตการทดลองแบบ eval-gated ของ Caveman Cloud และบล็อกการดำเนินการที่ไม่ปลอดภัย ใช้เมื่อผู้ใช้ขอให้เริ่ม อนุมัติ ยกเลิก เลื่อนระดับ หรือย้อนกลับการทดลอง Caveman หรือถามว่าหลักฐานของการทดลองสนับสนุนการดำเนินการใด อ่านหลักฐานก่อน อย่าดำเนินการเปลี่ยนแปลงวงจรชีวิตจนกว่าการเปลี่ยนผ่านที่ได้รับการยืนยันจากเซิร์ฟเวอร์และเกตหลักฐานจะพร้อมใช้งาน

npx skills add https://github.com/juliusbrussee/caveman --skill caveman-manage

Manage eval-gated experiments

Treat every lifecycle change as a production control action. Read current state and results, then report one supported recommendation or block. Current agent MCP is intentionally read-only: control-api does not yet enforce a complete lifecycle transition table and evidence gate atomically.

Non-negotiable gates

  1. A request to review, inspect, explain, or recommend authorizes reads only.
  2. Never approve an experiment whose results are pending, whose required guardrails are absent, or whose evidence reports a breach.
  3. Never convert experiment lift into verified_savings. Only active real traffic plus provider-causal, provider-complete ledger evidence can do that.
  4. Never supply an organization id. Project and tenant scope come from the logged-in Caveman identity and server RBAC.
  5. Never execute a lifecycle mutation, even after user approval. Exact <action>:<experiment_id> strings are agent-generatable and are not proof of human intent.
  6. Unknown states and server errors fail closed. Report exact cave_snake_code.

Step 1 — Load project and experiment

Prefer MCP:

caveman_context {}
caveman_experiment_get {"action":"get","experiment_id":"<id>"}
caveman_experiment_get {"action":"results","experiment_id":"<id>"}

Use {"action":"list"} when the user has not named an id.

CLI fallback:

caveman cloud experiments list
caveman cloud experiments show <id>
caveman cloud experiments results <id>

Stop if login, project, experiment, or results are unavailable.

Step 2 — Evaluate evidence

Report:

  • current lifecycle state and safety class;
  • control and candidate sample sizes;
  • quality or eval result;
  • latency, error, cost, retry, drop, and escalation guardrails when present;
  • evidence cost;
  • rollback or hold reason;
  • whether result is pending, failed, promotable, or active.

Absence is not a pass. If a required field is absent, state evidence incomplete and do not propose approval.

Step 3 — Propose one action

Allowed actions:

  • start — only from a startable draft or queued state with configured graders;
  • approve — only with complete passing evidence and a safety class the current role may approve;
  • cancel — stop a non-active experiment the user no longer wants;
  • rollback — revert an active or harmful change through the server's linked policy path. Current deployments may reject this honestly with cave_not_implemented; never describe that response as a rollback.

Show recommendation and id:

Proposed action: approve experiment 7f...
Reason: candidate passed quality and every configured guardrail.
Execution: blocked until server-authoritative lifecycle and evidence gates ship.

Do not treat earlier generic statements such as "manage it" or "do what is best" as mutation approval.

Step 4 — Block unsafe execution

Do not emit or run an executable lifecycle command. Explain that current server does not yet enforce every evidence/state transition atomically. CLI and MCP agent surfaces therefore expose experiment reads only.

Step 5 — Re-read after external operator action

If operator says they executed command, read detail and results again. Report server-observed post-state, audit or result response, and any policy-delivery status returned. Never infer success from operator intent alone.

Use this close:

Action: <action> <experiment-id>
Before: <state>
Server response: <status and cave_snake_code if any>
After: <re-read state>
Basis: experiment evidence only. Verified savings unchanged unless the signed
ledger independently records active, provider-causal real-traffic savings.

Skills เพิ่มเติมจาก juliusbrussee

caveman
juliusbrussee
โหมดสื่อสารแบบบีบอัดสูง ลดการใช้โทเค็นประมาณ 75% โดยพูดแบบมนุษย์ถ้ำ แต่ยังคงความถูกต้องทางเทคนิคครบถ้วน รองรับระดับความเข้มข้น: lite, full (ค่าเริ่มต้น), ultra, wenyan-lite, wenyan-full, wenyan-ultra ใช้เมื่อผู้ใช้พูดว่า "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief" หรือเรียกใช้ /caveman และจะทำงานอัตโนมัติเมื่อมีการขอประสิทธิภาพการใช้โทเค็น
communicationproductivity
caveman-commit
juliusbrussee
เครื่องมือสร้างข้อความคอมมิตที่บีบอัดอย่างยิ่ง ลดความรกในข้อความคอมมิตขณะคงเจตนาและเหตุผลไว้ รูปแบบ Conventional Commits หัวข้อไม่เกิน 50 ตัวอักษร เนื้อหาเฉพาะเมื่อ "เหตุผล" ไม่ชัดเจน ใช้เมื่อผู้ใช้พูดว่า "write a commit", "commit message", "generate commit", "/commit" หรือเรียก /caveman-commit ทำงานอัตโนมัติเมื่อมีการจัดเตรียมการเปลี่ยนแปลง
developmentcode-review
caveman-compress
juliusbrussee
บีบอัดไฟล์หน่วยความจำภาษาธรรมชาติ (CLAUDE.md, todos, preferences) ให้เป็นรูปแบบ caveman เพื่อประหยัดโทเค็นอินพุต คงไว้ซึ่งเนื้อหาทางเทคนิค โค้ด URL และโครงสร้างทั้งหมด เวอร์ชันที่บีบอัดจะเขียนทับไฟล์ต้นฉบับ สำเนาสำรองที่มนุษย์อ่านได้จะถูกบันทึกเป็น FILE.original.md ทริกเกอร์: /caveman-compress FILEPATH หรือ "compress memory file
developmentdocument
caveman-help
juliusbrussee
บัตรอ้างอิงด่วนสำหรับโหมด ทักษะ และคำสั่งทั้งหมดของ caveman แสดงผลครั้งเดียว ไม่ใช่โหมดถาวร คำสั่งเรียก: /caveman-help, "caveman help", "what caveman commands", "how do I use caveman
developmentdocumentproductivity
caveman-review
juliusbrussee
ความคิดเห็นรีวิวโค้ดแบบบีบอัดสูง ลดสัญญาณรบกวนจากฟีดแบ็ก PR ขณะคงสัญญาณที่นำไปปฏิบัติได้ไว้ แต่ละความคิดเห็นคือหนึ่งบรรทัด: ตำแหน่ง ปัญหา วิธีแก้ไข ใช้เมื่อผู้ใช้พูดว่า "review this PR", "code review", "review the diff", "/review" หรือเรียกใช้ /caveman-review ทำงานอัตโนมัติเมื่อรีวิว pull requests
developmentcode-review
caveman-stats
juliusbrussee
แสดงการใช้งานโทเค็นจริงและประมาณการประหยัดสำหรับเซสชันปัจจุบัน อ่านโดยตรงจากบันทึกเซสชันของ Claude Code — ไม่มีการประมาณด้วย AI เรียกใช้ด้วย /caveman-stats ผลลัพธ์ถูกแทรกโดย mode-tracker hook; ตัวโมเดลไม่ได้คำนวณตัวเลขเอง
developmentdata-analysis
cavecrew
juliusbrussee
We need to translate the given English text into Thai. The text describes a decision guide for delegating to caveman-style subagents. It mentions specific names: cavecrew-investigator, cavecrew-builder, cavecrew-reviewer, and cavecrew. These names should be preserved as is. Also preserve "Explore" (capitalized) and the trigger phrases. The translation should be natural Thai, keeping technical terms like "subagent", "tool-result", "context", "inline", "diff review", etc. The instruction says to preserve URLs, numbers, technical terms. No extra commentary. Just output the translation. Let me translate paragraph by paragraph. First sentence: "Decision guide for delegating to caveman-style subagents." -> "คู่มือการตัดสินใจสำหรับการมอบหมายงานให้กับซับเอเจนต์สไตล์ถ้ำมนุษย์" (caveman-style might be "สไตล์ถ้ำมนุษย์" but careful: "caveman" is a term, but we preserve "caveman-style" as is? The
developmentcode-reviewapi
caveman-explore
juliusbrussee
เครื่องมือสำรวจพื้นที่เก็บข้อมูลแบบอ่านอย่างเดียว ใช้เชิงรุกสำหรับการสำรวจเบื้องต้น การค้นหาข้ามไฟล์ในวงกว้าง หรือเมื่อการค้นหาตรงล้มเหลวและจำเป็นต้องระบุตำแหน่งของสิ่งนั้น ข้ามไปเมื่อประเด็นระบุไฟล์หรือสัญลักษณ์ที่แน่นอนอยู่แล้ว หรือรอบก่อนหน้าให้หลักฐาน file:line ที่ใช้งานได้แล้ว ส่งคืนเฉพาะการอ้างอิง path:line แบบกระชับ การอ่านและการ grep ไม่เคยเข้าไปในบทสนทนาหลัก