transcribe

por openai

We need to translate the given text from English to Brazilian Portuguese. The text describes a skill that transcribes audio files to text with optional speaker diarization and known-speaker hints. It mentions specific model names: gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize. These should be preserved as is. Also preserve "OPENAI_API_KEY". The output formats: plain text, JSON, diarized JSON. The text says "up to 4 speakers". Also "configurable output directories". The instruction says to preserve product names, protocol names, URLs, numbers, technical terms. So we keep "gpt-4o-mini-transcribe", "gpt-4o-transcribe-diarize", "OPENAI_API_KEY", numbers like 4, and terms like "JSON", "plain text", "diarized JSON". Also "speaker diarization" and "known-speaker hints" are technical terms, but we should translate them appropriately? The instruction says "preserve technical terms" but also "translate

npx skills add https://github.com/openai/skills --skill transcribe

Audio Transcribe

Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.

Workflow

  1. Collect inputs: audio file path(s), desired response format (text/json/diarized_json), optional language hint, and any known speaker references.
  2. Verify OPENAI_API_KEY is set. If missing, ask the user to set it locally (do not ask them to paste the key).
  3. Run the bundled transcribe_diarize.py CLI with sensible defaults (fast text transcription).
  4. Validate the output: transcription quality, speaker labels, and segment boundaries; iterate with a single targeted change if needed.
  5. Save outputs under output/transcribe/ when working in this repo.

Decision rules

  • Default to gpt-4o-mini-transcribe with --response-format text for fast transcription.
  • If the user wants speaker labels or diarization, use --model gpt-4o-transcribe-diarize --response-format diarized_json.
  • If audio is longer than ~30 seconds, keep --chunking-strategy auto.
  • Prompting is not supported for gpt-4o-transcribe-diarize.

Output conventions

  • Use output/transcribe/<job-id>/ for evaluation runs.
  • Use --out-dir for multiple files to avoid overwriting.

Dependencies (install if missing)

Prefer uv for dependency management.

uv pip install openai

If uv is unavailable:

python3 -m pip install openai

Environment

  • OPENAI_API_KEY must be set for live API calls.
  • If the key is missing, instruct the user to create one in the OpenAI platform UI and export it in their shell.
  • Never ask the user to paste the full key in chat.

Skill path (set once)

export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"

User-scoped skills install under $CODEX_HOME/skills (default: ~/.codex/skills).

CLI quick start

Single file (fast text default):

python3 "$TRANSCRIBE_CLI" \
  path/to/audio.wav \
  --out transcript.txt

Diarization with known speakers (up to 4):

python3 "$TRANSCRIBE_CLI" \
  meeting.m4a \
  --model gpt-4o-transcribe-diarize \
  --known-speaker "Alice=refs/alice.wav" \
  --known-speaker "Bob=refs/bob.wav" \
  --response-format diarized_json \
  --out-dir output/transcribe/meeting

Plain text output (explicit):

python3 "$TRANSCRIBE_CLI" \
  interview.mp3 \
  --response-format text \
  --out interview.txt

Reference map

  • references/api.md: supported formats, limits, response formats, and known-speaker notes.

Mais skills de openai

release
openai
Crie um release da Symphony atualizando a versão commitada, fazendo o merge, criando a tag no commit mesclado e verificando o workflow de release do Burrito. Use quando for solicitado a...
signing-entitlements
openai
Inspecione problemas de assinatura, entitlements, runtime protegido e Gatekeeper em aplicativos macOS. Use quando solicitado a diagnosticar falhas de assinatura de código, entitlements ausentes,…
building-ai-agent-on-cloudflare
openai
Constrói agentes de IA na Cloudflare usando o Agents SDK com gerenciamento de estado, WebSockets em tempo real, tarefas agendadas, integração de ferramentas e chat…
epigraphdb-skill
openai
Envie requisições compactas da API EpiGraphDB para ontologia, literatura, MR, gene-droga e evidências de suporte a caminhos. Use quando um usuário desejar resumos concisos do EpiGraphDB.
runtime-behavior-probe
openai
Planeje e execute investigações de comportamento em tempo de execução com scripts de sonda temporários, matrizes de validação, controles de estado e relatórios focados em descobertas. Use apenas quando…
deep-security-scan
openai
Use quando o usuário solicitar uma varredura de segurança Codex Security scan profunda, exaustiva, de múltiplas passagens ou com redução de variância, em todo o repositório ou em caminho específico. Execute repetidas verificações independentes…
define-security-policy
openai
Defina, revise ou atualize as orientações do SECURITY.md para um repositório ou componente. Use quando o usuário quiser esclarecer o que o Codex Security deve revisar, o que está fora…
validation
openai
Use quando o Codex já está na fase de validação de uma varredura de segurança ou quando o usuário pede explicitamente para determinar se uma ou mais descobertas de segurança candidatas…