tools

Use ao escrever, editar ou revisar ferramentas de linha de comando e scripts auxiliares neste repositório (itens em bin/). Abrange onde as ferramentas estão localizadas, análise de argumentos com…

npx skills add https://github.com/astronomer/astronomer --skill tools

Writing tools in bin/

Repo tooling (setup scripts, one-off utilities, anything a human or CI invokes directly) lives in bin/ and follows a few rules so tools are discoverable, self-documenting, and safe to run by accident.


Critical Rules

  1. Tools live in bin/ as executable scripts: a shebang (#!/usr/bin/env python3 for Python) plus chmod +x. Python tools run via uv run bin/<tool>.py.
  2. Every tool parses arguments with argparse (or the language equivalent) so --help works and every argument is self-documenting. Parse arguments as the first thing main() does.
  3. --help and insufficient/invalid arguments must do no work. They print usage and exit before any side effect. argparse gives this for free as long as parsing happens before any side-effecting code.
  4. A tool must not perform a destructive or state-mutating operation by default. Merely running it (or running it to read --help) must not create/delete Kubernetes objects, write/delete files, call external services, or change the active context.

Non-destructive by default

The failure mode to design against: someone runs bin/some-tool.py (or bin/some-tool.py --help) expecting it to be inert or to print help, and instead it mutates whatever ambient context it finds — the current kube context, the current directory, a live cluster.

The rule that prevents it: do not give a safe-looking default to any argument that determines where a mutation lands (a namespace, a cluster, a path, a target host). Make those arguments required with no default, so a bare or accidental invocation aborts before doing anything.

With argparse, a required=True argument with no default means:

  • tool (no args) → prints usage to stderr and exits non-zero, before main() reaches any side effect.
  • tool --help → prints help and exits 0.
  • tool --namespace foo ... → runs, because the caller was explicit about the target.

Worked example: bin/setup-forgejo-ca.py

This script creates and deletes Kubernetes Secrets in a cluster. It originally defaulted its namespaces (astronomer, git-forgejo). Running bin/setup-forgejo-ca.py --help to read the help text would instead have run the whole thing against the reader's current kube context — a potentially destructive surprise.

The fix was to make the namespaces required, with no defaults:

def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument("--platform-namespace", required=True, help="...")
    parser.add_argument("--forgejo-namespace", required=True, help="...")
    return parser.parse_args()

def main() -> None:
    args = parse_args()          # aborts here on a bare run or --help, before any kubectl
    ...                          # cluster mutations only happen after this line

Now --help and a bare run both abort before touching the cluster, and any real run has to name its target namespaces on purpose.

Automation still works

Making the target arguments required does not break automated callers — it just moves the intent to the caller, where it belongs. The automated invocation passes the values explicitly. For example, the git-sync-private-ca scenario's pre_helm_scripts entry names the namespaces:

pre_helm_scripts:
  - bin/setup-forgejo-ca.py --platform-namespace astronomer --forgejo-namespace git-forgejo

Checklist for a new or edited tool

  • Lives in bin/, is executable, has the right shebang.
  • Uses argparse; --help works and does nothing else.
  • Arguments that decide where a mutation lands are required with no default.
  • Side effects run only after arguments parse successfully.
  • Fails loudly on error (non-zero exit), and is idempotent (safe to re-run) where practical.
  • Callers (CI, scenario manifests, other scripts) pass the required arguments explicitly.

Mais skills de astronomer

airflow
astronomer
Consulte, gerencie e solucione problemas de DAGs, execuções, tarefas e configuração de sistema do Apache Airflow. Suporta mais de 30 comandos para inspeção de DAGs, gerenciamento de execuções, registro de tarefas, consultas de configuração e acesso direto à API REST. Gerencie múltiplas instâncias do Airflow com configuração persistente; descubra automaticamente implantações locais e Astro. Dispare execuções de DAG de forma síncrona (aguardando conclusão) ou assíncrona, diagnostique falhas, limpe execuções para repetição e acesse logs de tarefas com filtragem por repetição/índice de mapa. Saída...
official
airflow-hitl
astronomer
Portões de aprovação humana, entradas de formulário e ramificações em DAGs do Airflow usando operadores adiáveis. Quatro tipos de operadores: ApprovalOperator para decisões de aprovar/rejeitar, HITLOperator para seleção de múltiplas opções com formulários, HITLBranchOperator para roteamento de tarefas orientado por humanos e HITLEntryOperator para coleta de dados de formulário. Todos os operadores são adiáveis, liberando slots de worker enquanto aguardam resposta humana via a aba Ações Necessárias da interface do Airflow ou API REST. Suporta recursos opcionais incluindo personalização...
official
airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
official
analyzing-data
astronomer
Consulte seu data warehouse para responder perguntas de negócios com padrões em cache e mapeamentos de conceitos. Suporta busca de padrões e cache para tipos de perguntas repetidas, com registro de resultados para melhorar consultas futuras. Inclui cache de mapeamento conceito-tabela e descoberta de esquemas de tabela via INFORMATION_SCHEMA ou grep no código-fonte. Fornece funções de kernel run_sql() e run_sql_pandas() que retornam DataFrames Polars ou Pandas para análise. Comandos CLI para gerenciar caches de conceitos, padrões e tabelas, além de...
official
annotating-task-lineage
astronomer
Anotar tarefas do Airflow com linhagem de dados usando inlets e outlets. Suporta objetos OpenLineage Dataset, Assets do Airflow e Datasets do Airflow para definir entradas e saídas em bancos de dados, data warehouses e armazenamento em nuvem. Use como fallback quando operadores não possuem extratores OpenLineage integrados; segue um sistema de precedência de quatro níveis onde extratores personalizados e métodos OpenLineage têm prioridade. Inclui auxiliares de nomenclatura de datasets para Snowflake, BigQuery, S3 e PostgreSQL para garantir consistência...
official
authoring-dags
astronomer
Fluxo de trabalho guiado para criação de DAGs do Apache Airflow com integração de validação e testes. Abordagem estruturada em seis fases: descobrir o ambiente e padrões existentes, planejar a estrutura da DAG, implementar seguindo as melhores práticas, validar com comandos da CLI af, testar com consentimento do usuário e iterar em correções. Comandos da CLI para descoberta (af config connections, af config providers, af dags list) e validação (af dags errors, af dags get, af dags explore) fornecem feedback imediato sobre a DAG...
official
authoring-go-sdk-tasks
astronomer
Writes Airflow task logic in Go using the Airflow Go SDK. Use when the user wants to implement Airflow tasks in Go, asks about `BundleProvider`/`RegisterDags`,…
official
authoring-java-sdk-tasks
astronomer
Escreve a lógica de tarefas do Airflow em Java, Kotlin ou qualquer linguagem JVM usando o Airflow Java SDK. Use quando o usuário quiser implementar tarefas do Airflow em Java/JVM, pedir…
official