tools

Verwenden beim Schreiben, Bearbeiten oder Überprüfen von Befehlszeilen-Tools und Hilfsskripten in diesem Repository (Dinge in bin/). Deckt ab, wo Tools leben, Argument-Parsing mit…

npx skills add https://github.com/astronomer/astronomer --skill tools

Writing tools in bin/

Repo tooling (setup scripts, one-off utilities, anything a human or CI invokes directly) lives in bin/ and follows a few rules so tools are discoverable, self-documenting, and safe to run by accident.


Critical Rules

  1. Tools live in bin/ as executable scripts: a shebang (#!/usr/bin/env python3 for Python) plus chmod +x. Python tools run via uv run bin/<tool>.py.
  2. Every tool parses arguments with argparse (or the language equivalent) so --help works and every argument is self-documenting. Parse arguments as the first thing main() does.
  3. --help and insufficient/invalid arguments must do no work. They print usage and exit before any side effect. argparse gives this for free as long as parsing happens before any side-effecting code.
  4. A tool must not perform a destructive or state-mutating operation by default. Merely running it (or running it to read --help) must not create/delete Kubernetes objects, write/delete files, call external services, or change the active context.

Non-destructive by default

The failure mode to design against: someone runs bin/some-tool.py (or bin/some-tool.py --help) expecting it to be inert or to print help, and instead it mutates whatever ambient context it finds — the current kube context, the current directory, a live cluster.

The rule that prevents it: do not give a safe-looking default to any argument that determines where a mutation lands (a namespace, a cluster, a path, a target host). Make those arguments required with no default, so a bare or accidental invocation aborts before doing anything.

With argparse, a required=True argument with no default means:

  • tool (no args) → prints usage to stderr and exits non-zero, before main() reaches any side effect.
  • tool --help → prints help and exits 0.
  • tool --namespace foo ... → runs, because the caller was explicit about the target.

Worked example: bin/setup-forgejo-ca.py

This script creates and deletes Kubernetes Secrets in a cluster. It originally defaulted its namespaces (astronomer, git-forgejo). Running bin/setup-forgejo-ca.py --help to read the help text would instead have run the whole thing against the reader's current kube context — a potentially destructive surprise.

The fix was to make the namespaces required, with no defaults:

def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument("--platform-namespace", required=True, help="...")
    parser.add_argument("--forgejo-namespace", required=True, help="...")
    return parser.parse_args()


def main() -> None:
    args = parse_args()  # aborts here on a bare run or --help, before any kubectl
    ...  # cluster mutations only happen after this line

Now --help and a bare run both abort before touching the cluster, and any real run has to name its target namespaces on purpose.

Automation still works

Making the target arguments required does not break automated callers — it just moves the intent to the caller, where it belongs. The automated invocation passes the values explicitly. For example, the git-sync-private-ca scenario's pre_helm_scripts entry names the namespaces:

pre_helm_scripts:
  - bin/setup-forgejo-ca.py --platform-namespace astronomer --forgejo-namespace git-forgejo

Checklist for a new or edited tool

  • Lives in bin/, is executable, has the right shebang.
  • Uses argparse; --help works and does nothing else.
  • Arguments that decide where a mutation lands are required with no default.
  • Side effects run only after arguments parse successfully.
  • Fails loudly on error (non-zero exit), and is idempotent (safe to re-run) where practical.
  • Callers (CI, scenario manifests, other scripts) pass the required arguments explicitly.

Mehr Skills von astronomer

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
Benutzerdefinierte OpenLineage-Extraktoren für nicht unterstützte Airflow-Operatoren und komplexe Lineage-Szenarien. Zwei Ansätze: Fügen Sie OpenLineage-Methoden direkt zu Operatoren hinzu, die Sie besitzen (empfohlen), oder erstellen Sie benutzerdefinierte Extraktoren für Drittanbieter-Operatoren, die Sie nicht ändern können. Extraktoren greifen an drei Punkten in die Operatorausführung ein: vor der Ausführung für statisches Lineage, nach Erfolg für zur Laufzeit bestimmte Ausgaben und optional nach Fehlschlag für partielles Lineage. Registrieren Sie Extraktoren über airflow.cfg oder Umgebungsvariablen...
debugging-dags
astronomer
Systematische Ursachenanalyse und Behebung fehlgeschlagener Airflow-DAGs mit strukturierten Untersuchungsabläufen. Führt durch einen vierstufigen Diagnoseprozess: Fehler identifizieren, Fehlerdetails extrahieren, Kontextinformationen sammeln und umsetzbare Abhilfeschritte liefern. Kategorisiert Fehler in vier Typen (Daten, Code, Infrastruktur, Abhängigkeiten), um die Untersuchung zu fokussieren und geeignete Korrekturen vorzuschlagen. Stellt einsatzbereite CLI-Befehle für Logabruf, Ausführungsvergleich, Task-Löschung und DAG... bereit.
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
Airflow-DAGs und -Projekte bereitstellen. Verwenden, wenn der Benutzer Code bereitstellen, DAGs pushen, CI/CD einrichten, in die Produktion bereitstellen oder nach Bereitstellungsstrategien fragt…
deploying-go-sdk-bundles
astronomer
Erstellt, packt und stellt kompilierte Airflow Go SDK-Bundles bereit, damit der ExecutableCoordinator sie ausführen kann. Verwenden Sie dies, wenn der Benutzer ein Go-Task-Bundle kompilieren möchte, fragt…
testing-dags
astronomer
Iterative Test-Debug-Fix-Zyklen für Airflow-DAGs mit umfassender Fehlerdiagnose. Starten Sie mit af runs trigger-wait <dag_id>, um einen DAG auszuführen und auf dessen Abschluss zu warten; keine Pre-Flight-Checks erforderlich. Bei Fehlern verwenden Sie af runs diagnose für eine umfassende Fehlerzusammenfassung und af tasks logs, um Fehlerdetails von bestimmten Tasks zu überprüfen. Unterstützt benutzerdefinierte Konfiguration, Timeouts und Wiederholungsversuche; behandelt Erfolgs-, Fehler- und Timeout-Szenarien mit klarer Antwortinterpretation. Schnelle Validierung verfügbar...
tracing-downstream-lineage
astronomer
Verfolgen Sie die nachgelagerte Datenherkunft, um die Auswirkungen von Änderungen vor der Modifikation von Tabellen oder DAGs zu bewerten. Identifiziert direkte Konsumenten einer Ziel-Tabelle oder eines Ziel-DAGs durch Quellcode-Suche, View-Abhängigkeiten und BI-Tool-Verbindungen. Erstellt einen vollständigen Abhängigkeitsbaum, der alle nachgelagerten Auswirkungen abbildet – von Tabellen über Dashboards bis hin zu ML-Modellen. Kategorisiert Abhängigkeiten nach Kritikalität (kritisch, hoch, mittel, niedrig), um die Kommunikation mit Stakeholdern und Tests zu priorisieren. Generiert einen Auswirkungsbericht mit Risikobewertung, betroffenen...