cultivar

Sử dụng CLI cultivar để kiểm tra xem kỹ năng tác nhân có cải thiện hành vi hay không — tạo khung tác vụ, chạy có/không có kỹ năng trên Claude/Copilot/Gemini (cục bộ hoặc…

npx skills add https://github.com/pinecone-io/cultivar --skill cultivar

cultivar

cultivar is a CLI that measures whether an agent skill actually improves an agent's behavior. For each task it runs the agent with the skill and without it (and optionally with the source docs), then an LLM grader scores each run against a natural-language rubric. Use this skill when the user wants to create, run, or interpret cultivar evals.

The loop

  1. cultivar init <skill> — scaffold tasks/<skill>.yaml + a SKILL.md stub.
  2. Edit the task file (intent + PASS/FAIL criteria) and the skill.
  3. cultivar run -s <skill> -r <runner> --grade — run all variants and grade.
  4. cultivar report / cultivar show latest -t <task> — read the outcome.
  5. Iterate on the skill; re-run; compare.

Always confirm the install first with cultivar hello (or cultivar hello --no-grade when no ANTHROPIC_API_KEY is available) — it runs a packaged smoke task end-to-end.

Commands

  • cultivar init <skill> [--skills-dir DIR] — scaffold task YAML + SKILL.md stub.
  • cultivar run -s <skill> -r <claude|copilot|gemini> — run. Key flags:
    • -t <task> one task · -v <with-skill|without-skill|with-docs> one variant
    • --remote run in isolated Modal sandboxes · -n N repeat · -p N parallelism
    • --grade grade after running · --title NAME label the run · --dry-run print the prompt + command without calling anything · --timeout S per-call budget (default 90)
  • cultivar grade <run|latest> -s <skill> [--report] — (re)grade an existing run.
  • cultivar report [run] — summary table across runners/variants.
  • cultivar show <run|latest> -t <task> [--grader|--conversation-only|--workdir] — inspect one run.

--dry-run is the safe way to preview exactly what will be sent before spending tokens.

Variants (the controls)

  • with-skill — skill loaded; prompt prefixed Use the /<skill>.
  • without-skill — no skill; identical otherwise. The baseline.
  • with-docs — no skill, but the task's ground_truth.context_refs files are prepended. Only runs for tasks that declare context_refs.

Read two deltas: with-skill vs without-skill ("does the skill do anything?") and with-skill vs with-docs ("is the distilled skill better than dumping the raw docs?").

Tasks

tasks/<skill>.yaml holds one or more tasks. Each task:

tasks:
  - id: a-short-id
    intent: "what you'd ask the agent to do"
    category: cli            # or: code-gen
    # setup / teardown / verify: optional shell hooks
    # env: ["SOME_KEY"]      # required env vars, checked upfront
    ground_truth:
      criteria: |
        PASS requires <2-3 concrete, checkable things>.
        FAIL if <a common failure mode>.
      # context_refs: [docs/ref.md]   # activates the with-docs variant

Guidance:

  • For code-gen tasks, the intent must say "write a file … in the current directory." Anything the agent writes to its cwd is captured and shown to the grader. A code-gen task that produces no file auto-fails.
  • Write criteria as crisp PASS conditions + at least one concrete FAIL mode — vague criteria produce vague grades.

Where skills live

cultivar tests exactly one skill per run (the -s one). It resolves the skills root as: --skills-dir flag → CULTIVAR_SKILLS_DIR env → ./.claude/skills. Keep skills-under-test outside .claude/ (e.g. ./skills, via CULTIVAR_SKILLS_DIR=skills) if you don't want your interactive coding agent to auto-load them.

Local vs remote

  • Local (default) — uses the runner CLI installed on your machine + its auth.
  • --remote — each (task, variant, repeat) runs in its own Modal sandbox: clean isolation, parallelism, reproducibility. Requires a Modal account (modal token new) and a secret holding the agent's ANTHROPIC_API_KEY (default secret name eval-sandbox-secrets; override with CULTIVAR_MODAL_SECRET). Prefer --remote for rigorous comparisons. The grader always runs locally and needs ANTHROPIC_API_KEY.

Reading results

results/<timestamp>[__title]/ holds per-run .json (stats), .md (readable trace), .jsonl (raw events), and .workdir/ (files the agent wrote). grades.json holds the verdicts. Use cultivar report for the table and cultivar show … --grader for the grader's reasoning + suggestions on a failure.

Gotchas

  • Grading needs ANTHROPIC_API_KEY (loaded from a .env in the cwd). hello --no-grade and run --dry-run need no key.
  • tasks/, examples/, and results/ are cwd-relative and user-owned.
  • One run is a sample, not a signal — use -n 3 (or more) for anything you'll act on.

Thêm skills từ pinecone-io

workdir-smoke
pinecone-io
Kỹ năng giữ chỗ chỉ được sử dụng để kiểm tra thử cơ chế thu thập thư mục làm việc của khung đánh giá. Không phải kỹ năng thực tế — hãy xóa khi có kỹ năng tạo mã thực tế.
official
pinecone:assistant
pinecone-io
Tạo, quản lý và trò chuyện với Pinecone Assistant để hỏi đáp tài liệu có trích dẫn. Xử lý tất cả các thao tác với assistant - tạo, tải lên, đồng bộ, trò chuyện, ngữ cảnh…
official
pinecone:cli
pinecone-io
Hướng dẫn sử dụng Pinecone CLI (pc) để quản lý tài nguyên Pinecone từ terminal. CLI hỗ trợ TẤT CẢ các loại index (standard, integrated, sparse) và tất cả…
official
pinecone:docs
pinecone-io
Tài liệu tham khảo được tuyển chọn dành cho nhà phát triển xây dựng với Pinecone. Chứa các liên kết đến tài liệu chính thức được sắp xếp theo chủ đề và tham chiếu định dạng dữ liệu. Sử dụng khi…
official
pinecone:full-text-search
pinecone-io
Tạo, nạp dữ liệu vào và truy vấn chỉ mục tìm kiếm toàn văn (FTS) Pinecone bằng API xem trước (2026-01.alpha, bản xem trước công khai). Sử dụng khi người dùng hoặc tác nhân yêu cầu…
official
pinecone:help
pinecone-io
Tổng quan về tất cả các kỹ năng Pinecone có sẵn và những gì người dùng cần để bắt đầu. Gọi khi người dùng hỏi về các kỹ năng có sẵn, cách bắt đầu với…
official
pinecone:mcp
pinecone-io
Tài liệu tham khảo cho các công cụ máy chủ Pinecone MCP. Ghi lại tất cả các công cụ có sẵn - list-indexes, describe-index, describe-index-stats, create-index-for-model,…
official
pinecone:n8n
pinecone-io
Xây dựng quy trình làm việc n8n bằng cách sử dụng nút Pinecone Assistant hoặc nút Pinecone Vector Store. Sử dụng khi xây dựng các pipeline RAG, quy trình làm việc chat-with-docs, cấu hình…
official