behavioral-evals

作者: google-gemini

创建、运行、修复和推广行为评估的指南。用于验证代理决策逻辑、调试故障、调试提示…

npx skills add https://github.com/google-gemini/gemini-cli --skill behavioral-evals

Behavioral Evals

Overview

Behavioral evaluations (evals) are tests that validate the agent's decision-making (e.g., tool choice) rather than pure functionality. They are critical for verifying prompt changes, debugging steerability, and preventing regressions.

[!NOTE] Single Source of Truth: For core concepts, policies, running tests, and general best practices, always refer to evals/README.md.


🔄 Workflow Decision Tree

  1. Does a prompt/tool change need validation?
    • No -> Normal integration tests.
    • Yes -> Continue below.
  2. Is it UI/Interaction heavy?
  3. Is it a new test?
    • Yes -> Set policy to USUALLY_PASSES.
    • No -> ALWAYS_PASSES (locks in regression).
  4. Are you fixing a failure or promoting a test?

📋 Quick Checklist

1. Setup Workspace

Seed the workspace with necessary files using the files object to simulate a realistic scenario (e.g., NodeJS project with package.json).

2. Write Assertions

Audit agent decisions using rig.setBreakpoint() (AppRig only) or index verification on rig.readToolLogs().

3. Verify

Run single tests locally with Vitest. Confirm stability locally before relying on CI workflows.


📦 Bundled Resources

Detailed procedural guides:

  • creating.md: Assertion strategies, Rig selection, Mock MCPs.
  • fixing.md: Step-by-step automated investigation, architecture diagnosis guidelines.
  • promoting.md: Candidate identification criteria and threshold guidelines.

来自 google-gemini 的更多技能

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
使用Gemini API CLI工具的指南。当你需要通过命令行与Gemini API交互、管理代理或生成媒体(图像、……)时使用。
gemini-live-api-dev
google-gemini
通过WebSocket与Gemini进行实时双向流式传输,支持音频、视频和文本对话。具备音频输入/输出(16 kHz PCM)、视频帧、文本以及带语音活动检测的自动转录功能,可实现中断处理。包含原生音频特性:情感对话、主动音频和思考模式;支持同步和异步工具调用的函数调用;以及Google搜索接地功能。提供会话管理,包括上下文压缩、恢复等。
gemini-omni-flash-api
google-gemini
使用此技能进行生成式视频编辑、文本转视频、图像参考视频生成以及首帧到视频的过渡动画,基于…
gemini-api-dev
google-gemini
使用Google的Gemini模型构建应用程序,支持多模态内容、函数调用和结构化输出,覆盖Python、JavaScript、Go和Java语言。可访问当前Gemini 3模型(Pro、Flash、Pro Image),具备100万token上下文;旧版Gemini 2.x和1.5模型已弃用。支持文本生成、图像/音频/视频理解、函数调用、结构化JSON输出、代码执行、上下文缓存和嵌入功能。提供官方SDK:google-genai(Python)...
gemini-interactions-api
google-gemini
Gemini模型与代理的统一接口,支持服务端状态、流式传输和工具编排。兼容当前多款模型(gemini-3-flash-preview、gemini-3-pro-preview、gemini-2.5-flash/pro)及Deep Research代理;自动将已弃用的模型ID替换为当前替代方案。通过previous_interaction_id将对话历史卸载至服务端,实现有状态的多轮交互,无需手动管理历史记录。内置工具编排功能包括...
deliver
google-gemini
将简报的浓缩版发布到 Google Chat 或 Slack 的传入 webhook,使每日运行自行投递——当没有 webhook 时静默跳过……
fetch-news
google-gemini
获取读者感兴趣的所有主题和类别的Google News和Hacker News最新条目,并针对之前运行中显示过的所有条目进行去重。