code-change-verification

작성자: openai

OpenAI Agents Python 저장소에서 변경 사항이 런타임 코드, 테스트, 또는 빌드/테스트 동작에 영향을 미칠 때 필수 검증 스택을 실행합니다.

npx skills add https://github.com/openai/openai-agents-python --skill code-change-verification

Code Change Verification

Overview

Ensure work is only marked complete after formatting, linting, type checking, and tests pass. Use this skill when changes affect runtime code, tests, or build/test configuration. You can skip it for docs-only or repository metadata unless a user asks for the full stack. This is a post-review final gate: when $implementation-final-review applies, do not invoke the broad stack until its clean-review condition applies to the stable task diff.

Quick start

  1. Keep this skill at ./.agents/skills/code-change-verification so it loads automatically for the repository.
  2. Codex on macOS/Linux: /usr/bin/env -u OPENAI_API_KEY OPENAI_AGENTS_TEST_IN_CODEX_SANDBOX=1 UV_DEFAULT_INDEX=https://pypi.org/simple bash .agents/skills/code-change-verification/scripts/run.sh.
  3. Other macOS/Linux environments: env UV_DEFAULT_INDEX=https://pypi.org/simple bash .agents/skills/code-change-verification/scripts/run.sh.
  4. Windows: powershell -ExecutionPolicy Bypass -File .agents/skills/code-change-verification/scripts/run.ps1.
  5. On macOS/Linux, the script runs make format, make lint, make typecheck, and make tests sequentially and stops at the first failure. Parallelism inside each Make target, including pytest workers, is unchanged.
  6. The Bash script streams each command's output directly. The Windows wrapper retains parallel lint, typecheck, and test steps with periodic heartbeat updates.
  7. If any command fails, fix the issue, rerun the script, and report the failing output.
  8. Confirm completion only when all commands succeed with no remaining issues.

Start condition and host capacity

  • During iterative review, use only focused tests and a narrowly targeted static check when the changed typing boundary requires one. Defer repository-wide make typecheck and the rest of this complete stack until review is clean.
  • Immediately before starting the complete stack, use available read-only task or process evidence to check whether another repository-wide test, typecheck, build, examples runner, or integration command is already active on the same host.
  • When concrete contention is visible, continue useful non-heavy work such as review, remediation, evidence preparation, or focused checks, then check again later. Do not create or wait on a repository lock, host-wide mutex, or sentinel file.
  • Start automatically once review is clean, the diff is stable, and observable host capacity is available. Do not require a user-triggered finalize message. If host telemetry is unavailable, do not block solely because capacity cannot be measured.

Codex execution policy

Repository verification and all child processes must remain in the normal Codex workspace sandbox. Never request elevated sandbox permissions for the verification wrapper, and never retry the wrapper with broader host access after a failure.

On macOS, tests marked requires_native_macos_sandbox need to start their own sandbox-exec process. The Codex command sets OPENAI_AGENTS_TEST_IN_CODEX_SANDBOX=1, which skips only that marker before nested sandbox creation. All other tests remain enabled. Ordinary local and CI runs do not set this variable and therefore keep the marked tests enabled.

The marked tests run separately on a disposable GitHub-hosted macOS runner. If that trusted runner is unavailable, report the missing native-macOS coverage; do not compensate by weakening the Codex sandbox boundary.

Environment setup

The verification scripts assume repository dependencies are already installed. Do not run make sync as part of every verification pass; use it for a fresh checkout, after dependency files change, or when dependency resolution fails before the checks start.

On Linux, some Python packages with native extensions may require system packages such as libffi-dev, Python development headers, or build tools. If verification cannot start because one of these packages is missing, treat it as a local environment setup issue. Install the missing dependency when possible, or report the failing command and missing dependency in the PR test plan before rerunning verification in a prepared environment.

Manual workflow

  • For a fresh checkout, or if dependencies are not installed or have changed, run make sync first to install dev requirements via uv.
  • Run from the repository root with make format first, then make lint, make typecheck, and make tests.
  • Do not skip steps; stop and fix issues immediately when a command fails.
  • Run the manual steps sequentially and stop at the first failure. Keep the parallelism provided by each Make target.
  • Re-run the full stack after applying fixes so the commands execute in the required order.

Resources

scripts/run.sh

  • Runs make format, make lint, make typecheck, and make tests sequentially from the repository root. It streams output, preserves the first failure or cancellation status, and cleans up the active step's process group before continuing or exiting.

scripts/run.ps1

  • Windows-friendly wrapper that runs the same sequence with make format first and the remaining steps in parallel with fail-fast semantics, plus periodic heartbeat updates while work is still running. Use from PowerShell with execution policy bypass if required by your environment.

openai의 다른 스킬

release
openai
커밋된 버전을 올리고, 이를 반영하고, 병합된 커밋에 태그를 단 후, Burrito 릴리스 워크플로우를 검증하여 Symphony 릴리스를 진행합니다. 다음과 같이 요청받았을 때 사용합니다…
signing-entitlements
openai
macOS 앱의 서명, 자격, 강화된 런타임 및 Gatekeeper 문제를 검사합니다. 코드 서명 실패, 누락된 자격 등을 진단하라는 요청을 받을 때 사용하세요.
building-ai-agent-on-cloudflare
openai
Cloudflare에서 Agents SDK를 사용하여 상태 관리, 실시간 WebSockets, 예약 작업, 도구 통합, 채팅을 통해 AI 에이전트를 구축합니다…
epigraphdb-skill
openai
온톨로지, 문헌, MR, 유전자-약물 및 지원 경로 증거에 대한 간결한 EpiGraphDB API 요청을 제출합니다. 사용자가 간결한 EpiGraphDB 요약을 원할 때 사용하세요.
runtime-behavior-probe
openai
런타임 동작 조사를 계획하고 실행하며, 임시 프로브 스크립트, 검증 매트릭스, 상태 제어, 결과 우선 보고서를 사용합니다. 다음 경우에만 사용하세요…
deep-security-scan
openai
사용자가 심층적이고, 철저하며, 다중 패스 또는 변동성을 줄이는 저장소 전체 또는 범위가 지정된 경로의 Codex Security 스캔을 요청할 때 사용합니다. 반복적으로 독립적인…
define-security-policy
openai
저장소 또는 구성 요소에 대한 SECURITY.md 지침을 정의, 검토 또는 업데이트합니다. 사용자가 Codex Security가 검토해야 할 대상과 범위를 벗어나는 항목을 명확히 하려 할 때 사용합니다…
validation
openai
Codex가 보안 스캔의 검증 단계에 이미 있거나 사용자가 하나 이상의 후보 보안 결과를 판별하도록 명시적으로 요청할 때 사용합니다…