behavioral-evals

작성자: google-gemini

행동 평가를 생성, 실행, 수정 및 홍보하기 위한 지침입니다. 에이전트 결정 로직 검증, 오류 디버깅, 프롬프트 디버깅 등에 사용하세요.

npx skills add https://github.com/google-gemini/gemini-cli --skill behavioral-evals

Behavioral Evals

Overview

Behavioral evaluations (evals) are tests that validate the agent's decision-making (e.g., tool choice) rather than pure functionality. They are critical for verifying prompt changes, debugging steerability, and preventing regressions.

[!NOTE] Single Source of Truth: For core concepts, policies, running tests, and general best practices, always refer to evals/README.md.


🔄 Workflow Decision Tree

  1. Does a prompt/tool change need validation?
    • No -> Normal integration tests.
    • Yes -> Continue below.
  2. Is it UI/Interaction heavy?
  3. Is it a new test?
    • Yes -> Set policy to USUALLY_PASSES.
    • No -> ALWAYS_PASSES (locks in regression).
  4. Are you fixing a failure or promoting a test?

📋 Quick Checklist

1. Setup Workspace

Seed the workspace with necessary files using the files object to simulate a realistic scenario (e.g., NodeJS project with package.json).

2. Write Assertions

Audit agent decisions using rig.setBreakpoint() (AppRig only) or index verification on rig.readToolLogs().

3. Verify

Run single tests locally with Vitest. Confirm stability locally before relying on CI workflows.


📦 Bundled Resources

Detailed procedural guides:

  • creating.md: Assertion strategies, Rig selection, Mock MCPs.
  • fixing.md: Step-by-step automated investigation, architecture diagnosis guidelines.
  • promoting.md: Candidate identification criteria and threshold guidelines.

google-gemini의 다른 스킬

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLI 도구 사용 가이드입니다. 명령줄을 통해 Gemini API와 상호작용하거나, 에이전트를 관리하거나, 미디어(이미지 등)를 생성해야 할 때 사용하세요.
gemini-live-api-dev
google-gemini
WebSocket을 통해 Gemini와 실시간 양방향 스트리밍을 지원하여 오디오, 비디오, 텍스트 대화를 처리합니다. 오디오 입력/출력(16kHz PCM), 비디오 프레임, 텍스트, 음성 활동 감지를 통한 자동 전사 및 인터럽트 처리를 지원합니다. 네이티브 오디오 기능(감정 대화, 능동적 오디오, 사고 모드), 동기 및 비동기 도구 사용을 위한 함수 호출, Google Search 접지 기능을 포함합니다. 컨텍스트 압축, 재개 등을 통한 세션 관리를 제공합니다.
gemini-omni-flash-api
google-gemini
이 스킬을 사용하여 생성형 비디오 편집, 텍스트-투-비디오, 이미지 참조 비디오 생성, 첫 프레임-투-비디오 전환 애니메이션 등을 수행할 수 있습니다…
gemini-api-dev
google-gemini
Google의 Gemini 모델로 애플리케이션을 구축하며, 멀티모달 콘텐츠, 함수 호출, 구조화된 출력을 Python, JavaScript, Go, Java에서 지원합니다. 최신 Gemini 3 모델(Pro, Flash, Pro Image)에 1M 토큰 컨텍스트로 접근 가능하며, 레거시 Gemini 2.x 및 1.5 모델은 지원 중단되었습니다. 텍스트 생성, 이미지/오디오/비디오 이해, 함수 호출, 구조화된 JSON 출력, 코드 실행, 컨텍스트 캐싱, 임베딩을 지원합니다. 공식 SDK 제공: google-genai (Python),...
gemini-interactions-api
google-gemini
Gemini 모델 및 에이전트를 위한 통합 인터페이스로, 서버 측 상태, 스트리밍 및 도구 오케스트레이션을 제공합니다. 여러 현재 모델(gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro)과 Deep Research 에이전트를 지원하며, 더 이상 사용되지 않는 모델 ID를 현재 대안으로 자동 대체합니다. previous_interaction_id를 통해 대화 기록을 서버에 오프로드하여 수동 기록 관리 없이 상태 저장 다중 턴 상호작용을 가능하게 합니다. 내장된 도구 오케스트레이션을 포함합니다...
deliver
google-gemini
브리핑의 축약 버전을 Google Chat 또는 Slack 인커밍 웹훅에 게시하여 일일 실행이 자동으로 전달되게 합니다 — 웹훅이 없으면 조용히 건너뜁니다…
fetch-news
google-gemini
구독자의 관심 분야에 해당하는 모든 토픽과 장르의 최신 Google News 및 Hacker News 항목을 가져오며, 이전 실행에서 이미 표시된 항목은 중복 제거합니다.