bootstrap-realtime-eval

bởi openai

Khởi tạo một thư mục đánh giá thời gian thực mới trong kho cookbook này bằng cách chọn harness phù hợp từ examples/evals/realtime_evals, tạo scaffold prompt/tools/data…

npx skills add https://github.com/openai/openai-cookbook --skill bootstrap-realtime-eval

Bootstrap Realtime Eval

Use this skill when the user wants a new realtime eval scaffold under examples/evals/realtime_evals/.

This skill is repo-specific. Do not copy harness code into the generated folder. The generated eval should point at the shared harnesses already in:

  • examples/evals/realtime_evals/crawl_harness
  • examples/evals/realtime_evals/walk_harness
  • examples/evals/realtime_evals/run_harness

Inputs To Collect

Always ask the user for the minimum set needed to choose and scaffold the eval before you create files, run the scaffold script, or author starter data. Do not skip this just because you can infer a default.

Ask for:

  • Eval name
  • Goal or scenario
  • Harness choice, or enough context to recommend one
  • System prompt path or inline text
  • Tools JSON path or tool descriptions
  • Data path or source materials
  • Desired graders

If the user does not know which harness they want, explain the options briefly and recommend one. See references/harness-selection.md.

When the user asks for synthetic audio but does not specify a harness, default to crawl text-to-TTS unless they need the generated audio to carry particular noise, telephony artifacts, speaker characteristics, or other replay-specific properties. Use walk for those cases.

Keep the questions concise and grouped into one short batch whenever possible.

If the user only provides user_text or a short task description, still ask the questions above first. If they answer only partially, then infer the remaining low-risk details, call out the assumptions, and make the scaffold easy to revise later.

Workflow

  1. Ask the user for the required inputs first.

    • Do this before making files or selecting a final harness.
    • If the user already supplied some of the inputs, ask only for the missing ones.
    • If you recommend a harness, wait for the user response before scaffolding.
  2. Pick the harness.

    • crawl: single-turn text-to-TTS.
    • walk: replay saved audio or generate audio from text rows.
    • run: multi-turn simulation with tool mocks and judge criteria.
    • If the user wants synthetic audio but does not care about replay-specific audio characteristics, prefer crawl over walk.
  3. Normalize the inputs.

    • If the user gives inline prompt or tool content, write it into the generated folder.
    • If the user gives CSV data, inspect it with pandas before wiring it in.
    • If the data is not yet harness-ready, scaffold the files and then have you author the starter dataset directly in those files.
    • If the user only gives user_text, infer example_id values and leave optional grading fields blank unless you have enough signal to fill them.
    • Ground the starter data in the use case, prompts, tools, and any other material the user provided.
    • Make the starter data realistic. Put yourself in the shoes of the end user in that use case and craft datapoints that a real user would plausibly say or do.
  4. Be proactive when data is missing.

    • crawl: you should author 3 starter rows covering one happy path and a couple of nearby variants or edge cases.
    • walk: you should author 3 source CSV rows and prepare the audio-generation step so the user can create audio immediately.
    • run: you should author 2 starter simulations, not 1.
    • Do not rely on deterministic script-generated samples for use-case-specific starter data.
    • Show the generated starter samples to the user, ask for a quick greenlight or correction, then expand only after feedback when the task calls for more coverage.
  5. Run the scaffold script:

python examples/evals/realtime_evals/skills/bootstrap-realtime-eval/scripts/bootstrap_realtime_eval.py --name "<eval_name>" --harness "<crawl|walk|run>"

Add flags for prompt, tools, data, graders, and run-specific fields as needed. Read the script help if you need the exact flag names.

  1. Review the generated folder.

    • Confirm the README is accurate for the selected harness.
    • Confirm the system prompt, tools, and data files point at the generated folder.
    • If the user provided real data, make sure it is wired in instead of leaving starter placeholders.
    • If you authored inferred starter data, show the user the generated rows or simulations before scaling up.
  2. Enrich the scaffold.

    • For crawl and walk, make the CSV realistic and ensure the expected tool columns are present.
    • For walk, if the dataset lacks audio_path, use the shared walk_harness/generate_audio.py flow described in the README.
    • For run, make sure simulations.csv and the starter sim_*.json file reflect the user’s scenario, tool mocks, and graders.
  3. Validate before returning.

    • Run a smoke eval for the generated folder.
    • If the smoke eval succeeds, automatically run the full eval for the generated folder in the same turn.
    • If the smoke eval fails, stop and fix or report the blocker before attempting the full eval.
    • Run pytest examples/evals/realtime_evals/tests -q.
    • Inspect the generated README and at least one generated data file.

Hard Rules

  • Keep the generated folder inside examples/evals/realtime_evals/.
  • Ask the user for the missing setup information before scaffolding. Do not jump straight to building a default eval when the request is to create a new eval.
  • Do not duplicate the main harness scripts into the generated folder.
  • Do not write generic placeholder starter data when the use case provides enough context to do better.
  • Use the prompt, tools, and scenario details to make the initial datapoints feel like a real user interaction.
  • The README must include:
    • why this harness was chosen
    • which files to edit first
    • smoke and full run commands
    • data contract notes
    • troubleshooting notes
  • Prefer assistant-specific flags for run harness commands:
    • --assistant-system-prompt-file
    • --assistant-tools-file

Completion Criteria

Treat the task as complete only when:

  • the folder exists under examples/evals/realtime_evals/<name>_realtime_eval/
  • README.md, system_prompt.txt, and tools.json exist
  • harness-specific starter data exists, authored by you when the user did not provide it
  • the smoke command has been run or is blocked for a clear reason
  • if the smoke command succeeded, the full eval command has also been run or is blocked for a clear reason
  • pytest examples/evals/realtime_evals/tests -q has been run

Learnings

When this skill uncovers a reusable workflow or harness constraint that should guide future bootstrap work, add a short note here.

Keep learnings concise and action-oriented:

  • Problem -> Fix -> Why

Only add items that are likely to help future realtime-eval scaffolding in this repo. Remove stale items when they no longer apply.

  • gpt-realtime temperature is unsupported -> Do not add a temperature field or CLI flag when scaffolding gpt-realtime evals -> Avoids invalid config and keeps runs aligned with the realtime harness constraints.
  • Starter samples should come from the model, not the scaffold script -> Use the script to create the folder structure and template files, then have you author the initial rows or simulations from the user’s use case -> Produces more relevant starter data and makes iteration easier with the user.
  • Synthetic audio can be overfit to the wrong harness -> Default unspecified synthetic-audio requests to crawl text-to-TTS and reserve walk for replay-specific audio characteristics like noise or telephony artifacts -> Keeps the bootstrap path simpler unless audio realism is the actual target.
  • Smoke-only validation can leave scaffolds half-proven -> After a successful smoke run, automatically run the full eval before declaring the bootstrap complete -> Catches dataset-wide or late-row failures that a one-example smoke test misses.

Thêm skills từ openai

user-context
openai
Tải hoặc quản lý các tùy chọn định tuyến nguồn bền vững, logic giới thiệu, tiến trình thiết lập và sổ đăng ký lớp ngữ nghĩa của plugin Phân tích Dữ liệu.
official
notion-research-documentation
openai
Nghiên cứu nội dung Notion và tổng hợp thành các bản tóm tắt có cấu trúc, báo cáo hoặc so sánh kèm trích dẫn. Tìm kiếm và truy xuất các trang Notion bằng truy vấn mục tiêu, sau đó sắp xếp kết quả theo chủ đề với trích dẫn nguồn trong văn bản và phần tài liệu tham khảo. Chọn từ bốn định dạng đầu ra (tóm tắt nhanh, tổng hợp nghiên cứu, so sánh, báo cáo toàn diện) dựa trên phạm vi và mục tiêu của người dùng. Tạo và cập nhật các trang Notion bằng mẫu có sẵn; liên kết trực tiếp nguồn và theo dõi thay đổi khi
official
rcsb-pdb-skill
openai
Gửi yêu cầu RCSB PDB nhỏ gọn để lấy siêu dữ liệu cốt lõi, truy vấn API Tìm kiếm và tải xuống FASTA. Sử dụng khi người dùng muốn tóm tắt RCSB ngắn gọn; lưu JSON thô hoặc…
official
pdf
openai
Đọc, tạo và xác thực PDF với kết xuất trực quan và tạo theo chương trình. Kết xuất các trang PDF sang PNG để kiểm tra trực quan bố cục, khoảng cách và kiểu chữ trước khi bàn giao bằng Poppler (pdftoppm). Tạo PDF theo chương trình với reportlab để định dạng đáng tin cậy; trích xuất văn bản và siêu dữ liệu bằng pdfplumber hoặc pypdf. Thực thi các tiêu chuẩn chất lượng: không có văn bản bị cắt, phần tử chồng lấn, bảng bị hỏng hoặc hiện vật kết xuất; chỉ sử dụng dấu gạch nối ASCII, trích dẫn dễ đọc cho con người. Sử dụng...
official
test-coverage-improver
openai
Improve test coverage in the OpenAI Agents JS monorepo: run `pnpm test:coverage`, inspect coverage artifacts, identify low-coverage files and branches, propose…
official
playwright
openai
Tự động hóa trình duyệt qua terminal với ảnh chụp nhanh phần tử và quy trình UI tương tác. Hoạt động thông qua script wrapper playwright-cli (yêu cầu npx); hỗ trợ chế độ headless và headed để gỡ lỗi trực quan. Quy trình cốt lõi: mở trang, chụp nhanh để tham chiếu phần tử ổn định, tương tác bằng refs, chụp lại sau khi điều hướng hoặc thay đổi DOM. Bao gồm điền biểu mẫu, nhấp chuột, gõ văn bản, quản lý nhiều tab, chụp ảnh màn hình/PDF và ghi lại trace để gỡ lỗi luồng. Tham chiếu phần tử (ví dụ: e3, e15)...
official
ukb-topmed-phewas-skill
openai
Lấy các bản tóm tắt PheWAS UKB-TOPMed nhỏ gọn cho các biến thể đơn lẻ bằng cách chấp nhận đầu vào rsID, GRCh37 hoặc GRCh38 và phân giải thành truy vấn GRCh38 cần thiết. Sử dụng khi một…
official
code-review-context
openai
Ngữ cảnh hiển thị của mô hình
official