tools

작성자: astronomer

이 저장소의 bin/에 있는 명령줄 도구와 헬퍼 스크립트를 작성, 편집, 검토할 때 사용합니다. 도구의 위치, 인자 파싱 등을 다룹니다…

npx skills add https://github.com/astronomer/astronomer --skill tools

Writing tools in bin/

Repo tooling (setup scripts, one-off utilities, anything a human or CI invokes directly) lives in bin/ and follows a few rules so tools are discoverable, self-documenting, and safe to run by accident.


Critical Rules

  1. Tools live in bin/ as executable scripts: a shebang (#!/usr/bin/env python3 for Python) plus chmod +x. Python tools run via uv run bin/<tool>.py.
  2. Every tool parses arguments with argparse (or the language equivalent) so --help works and every argument is self-documenting. Parse arguments as the first thing main() does.
  3. --help and insufficient/invalid arguments must do no work. They print usage and exit before any side effect. argparse gives this for free as long as parsing happens before any side-effecting code.
  4. A tool must not perform a destructive or state-mutating operation by default. Merely running it (or running it to read --help) must not create/delete Kubernetes objects, write/delete files, call external services, or change the active context.

Non-destructive by default

The failure mode to design against: someone runs bin/some-tool.py (or bin/some-tool.py --help) expecting it to be inert or to print help, and instead it mutates whatever ambient context it finds — the current kube context, the current directory, a live cluster.

The rule that prevents it: do not give a safe-looking default to any argument that determines where a mutation lands (a namespace, a cluster, a path, a target host). Make those arguments required with no default, so a bare or accidental invocation aborts before doing anything.

With argparse, a required=True argument with no default means:

  • tool (no args) → prints usage to stderr and exits non-zero, before main() reaches any side effect.
  • tool --help → prints help and exits 0.
  • tool --namespace foo ... → runs, because the caller was explicit about the target.

Worked example: bin/setup-forgejo-ca.py

This script creates and deletes Kubernetes Secrets in a cluster. It originally defaulted its namespaces (astronomer, git-forgejo). Running bin/setup-forgejo-ca.py --help to read the help text would instead have run the whole thing against the reader's current kube context — a potentially destructive surprise.

The fix was to make the namespaces required, with no defaults:

def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument("--platform-namespace", required=True, help="...")
    parser.add_argument("--forgejo-namespace", required=True, help="...")
    return parser.parse_args()


def main() -> None:
    args = parse_args()  # aborts here on a bare run or --help, before any kubectl
    ...  # cluster mutations only happen after this line

Now --help and a bare run both abort before touching the cluster, and any real run has to name its target namespaces on purpose.

Automation still works

Making the target arguments required does not break automated callers — it just moves the intent to the caller, where it belongs. The automated invocation passes the values explicitly. For example, the git-sync-private-ca scenario's pre_helm_scripts entry names the namespaces:

pre_helm_scripts:
  - bin/setup-forgejo-ca.py --platform-namespace astronomer --forgejo-namespace git-forgejo

Checklist for a new or edited tool

  • Lives in bin/, is executable, has the right shebang.
  • Uses argparse; --help works and does nothing else.
  • Arguments that decide where a mutation lands are required with no default.
  • Side effects run only after arguments parse successfully.
  • Fails loudly on error (non-zero exit), and is idempotent (safe to re-run) where practical.
  • Callers (CI, scenario manifests, other scripts) pass the required arguments explicitly.

astronomer의 다른 스킬

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
지원되지 않는 Airflow 연산자와 복잡한 계보 시나리오를 위한 맞춤형 OpenLineage 추출기. 두 가지 접근 방식: 소유한 연산자에 직접 OpenLineage 메서드를 추가(권장)하거나, 수정할 수 없는 타사 연산자를 위한 맞춤형 추출기를 생성합니다. 추출기는 세 지점에서 연산자 실행을 가로챕니다: 정적 계보를 위한 실행 전, 런타임에 결정된 출력을 위한 성공 후, 그리고 선택적으로 부분 계보를 위한 실패 후. airflow.cfg 또는 환경을 통해 추출기를 등록합니다...
debugging-dags
astronomer
체계적인 근본 원인 분석 및 구조화된 조사 워크플로를 통한 실패한 Airflow DAG의 문제 해결. 4단계 진단 프로세스를 안내합니다: 실패 식별, 오류 세부 정보 추출, 컨텍스트 정보 수집, 실행 가능한 수정 단계 제공. 실패를 네 가지 유형(데이터, 코드, 인프라, 종속성)으로 분류하여 조사에 집중하고 적절한 수정을 제안합니다. 로그 검색, 실행 비교, 작업 정리, DAG...을 위한 즉시 사용 가능한 CLI 명령을 제공합니다.
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
Airflow DAG 및 프로젝트를 배포합니다. 사용자가 코드를 배포하거나, DAG를 푸시하거나, CI/CD를 설정하거나, 프로덕션에 배포하거나, 배포 전략에 대해 질문할 때 사용하세요.
deploying-go-sdk-bundles
astronomer
컴파일된 Airflow Go SDK 번들을 빌드, 패킹 및 배포하여 ExecutableCoordinator가 실행할 수 있도록 합니다. 사용자가 Go 태스크 번들을 컴파일하려고 하거나 요청할 때 사용합니다.
testing-dags
astronomer
포괄적인 실패 진단 기능을 갖춘 Airflow DAG의 반복적인 테스트-디버그-수정 주기. af runs trigger-wait <dag_id>로 시작하여 DAG를 실행하고 완료를 기다립니다. 사전 점검은 필요하지 않습니다. 실패 시 af runs diagnose를 사용하여 포괄적인 실패 요약을 확인하고, af tasks logs를 사용하여 특정 태스크의 오류 세부 정보를 검사합니다. 사용자 정의 구성, 시간 제한 및 재시도 횟수를 지원하며, 명확한 응답 해석으로 성공, 실패 및 시간 초과 시나리오를 처리합니다. 빠른 검증 가능...
tracing-downstream-lineage
astronomer
테이블이나 DAG를 수정하기 전에 다운스트림 데이터 계보를 추적하여 변경 영향을 평가합니다. 소스 코드 검색, 뷰 종속성, BI 도구 연결을 통해 대상 테이블 또는 DAG의 직접적인 소비자를 식별합니다. 테이블에서 대시보드, ML 모델에 이르기까지 모든 다운스트림 영향을 매핑하는 전체 종속성 트리를 구축합니다. 종속성을 중요도(심각, 높음, 중간, 낮음)별로 분류하여 이해관계자 커뮤니케이션 및 테스트의 우선순위를 지정합니다. 위험 평가, 영향을 받는 항목이 포함된 영향 보고서를 생성합니다...