analyzing-data

作成者: astronomer

データウェアハウスにクエリを実行し、キャッシュされたパターンと概念マッピングを使用してビジネス上の質問に回答します。繰り返し発生する質問タイプのパターン検索とキャッシュをサポートし、結果を記録して将来のクエリを改善します。概念からテーブルへのマッピングキャッシュと、INFORMATION_SCHEMAまたはコードベースのgrepによるテーブルスキーマ検出を含みます。分析用にPolarsまたはPandas DataFrameを返すrun_sql()およびrun_sql_pandas()カーネル関数を提供します。概念、パターン、テーブルキャッシュを管理するCLIコマンド、さらに...

npx skills add https://github.com/astronomer/agents --skill analyzing-data

Data Analysis

Answer business questions by querying the data warehouse. The kernel auto-starts on first exec call.

All CLI commands below are relative to this skill's directory. Before running any scripts/cli.py command, cd to the directory containing this file.

Workflow

  1. Pattern lookup — Check for a cached query strategy:

    uv run scripts/cli.py pattern lookup "<user's question>"
    

    If a pattern exists, follow its strategy. Record the outcome after executing:

    uv run scripts/cli.py pattern record <name> --success  # or --failure
    
  2. Concept lookup — Find known table mappings:

    uv run scripts/cli.py concept lookup <concept>
    
  3. Table discovery — If cache misses, search the codebase (Grep pattern="<concept>" glob="**/*.sql") or query INFORMATION_SCHEMA. See reference/discovery-warehouse.md.

  4. Execute query:

    uv run scripts/cli.py exec "df = run_sql('SELECT ...')"
    uv run scripts/cli.py exec "print(df)"
    
  5. Cache learnings — Always cache before presenting results:

    # Cache concept → table mapping
    uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>
    # Cache query strategy (if discovery was needed)
    uv run scripts/cli.py pattern learn <name> -q "question" -s "step" -t "TABLE" -g "gotcha"
    
  6. Present findings to user.

Kernel Functions

FunctionReturns
run_sql(query, limit=100)Polars DataFrame
run_sql_pandas(query, limit=100)Pandas DataFrame
run_sql_many(queries, limit=100)List of Polars DataFrames (one per query)

pl (Polars) and pd (Pandas) are pre-imported.

Run independent queries together with run_sql_many — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:

uv run scripts/cli.py exec "dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])"

run_sql_many is fail-fast: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate run_sql calls if you need partial results.

Timeouts: exec waits up to 120s by default, then interrupts the query and returns a "client stopped waiting" message (the query may still finish server-side). Raise it for known long-running queries: uv run scripts/cli.py exec "..." -t 600.

Idle kernel: the kernel self-terminates after 2h idle (preserving state until then). Override with ASTRO_KERNEL_IDLE_TIMEOUT (seconds; 0 disables).

CLI Reference

Kernel

uv run scripts/cli.py warehouse list      # List warehouses
uv run scripts/cli.py start [-w name]     # Start kernel (with optional warehouse)
uv run scripts/cli.py exec "..."          # Execute Python code
uv run scripts/cli.py status              # Kernel status
uv run scripts/cli.py restart             # Restart kernel
uv run scripts/cli.py stop                # Stop kernel
uv run scripts/cli.py install <pkg>       # Install package

Concept Cache

uv run scripts/cli.py concept lookup <name>                     # Look up
uv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn
uv run scripts/cli.py concept list                               # List all
uv run scripts/cli.py concept import -p /path/to/warehouse.md   # Bulk import

Pattern Cache

uv run scripts/cli.py pattern lookup "question"                                      # Look up
uv run scripts/cli.py pattern learn <name> -q "..." -s "..." -t "TABLE" -g "gotcha"  # Learn
uv run scripts/cli.py pattern record <name> --success                                # Record outcome
uv run scripts/cli.py pattern list                                                   # List all
uv run scripts/cli.py pattern delete <name>                                          # Delete

Table Schema Cache

uv run scripts/cli.py table lookup <TABLE>            # Look up schema
uv run scripts/cli.py table cache <TABLE> -c '[...]'  # Cache schema
uv run scripts/cli.py table list                       # List cached
uv run scripts/cli.py table delete <TABLE>             # Delete

Cache Management

uv run scripts/cli.py cache status                # Stats
uv run scripts/cli.py cache clear [--stale-only]  # Clear

References

astronomerのその他のスキル

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
カスタムOpenLineage抽出器。未対応のAirflowオペレーターや複雑な系列シナリオ向け。2つのアプローチ:所有するオペレーターに直接OpenLineageメソッドを追加する方法(推奨)、または変更できないサードパーティ製オペレーター用にカスタム抽出器を作成する方法。抽出器は3つの時点でオペレーターの実行をインターセプトします:静的な系列のための実行前、実行時に決定される出力のための成功後、およびオプションで部分的な系列のための失敗後。抽出器はairflow.cfgまたは環境変数経由で登録...
debugging-dags
astronomer
失敗したAirflow DAGに対する体系的な根本原因分析と修正、構造化された調査ワークフローを提供。4段階の診断プロセス(障害の特定、エラー詳細の抽出、コンテキスト情報の収集、実行可能な修正手順の提示)をガイド。障害を4つのタイプ(データ、コード、インフラストラクチャ、依存関係)に分類し、調査を集中させ適切な修正を提案。ログ取得、実行比較、タスククリア、DAG...のための即時使用可能なCLIコマンドを提供。
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
AirflowのDAGやプロジェクトをデプロイします。ユーザーがコードのデプロイ、DAGのプッシュ、CI/CDの設定、本番環境へのデプロイ、またはデプロイ戦略について質問した場合に使用します…
deploying-go-sdk-bundles
astronomer
コンパイルされたAirflow Go SDKバンドルをビルド、パッケージ化、デプロイし、ExecutableCoordinatorが実行できるようにします。ユーザーがGoタスクバンドルをコンパイルしたい場合、または…と尋ねた場合に使用します。
testing-dags
astronomer
Airflow DAGに対する反復的なテスト・デバッグ・修正サイクルと、包括的な障害診断を提供します。af runs trigger-wait <dag_id> でDAGを実行し完了を待機します。事前チェックは不要です。失敗時はaf runs diagnoseで包括的な障害サマリーを取得し、af tasks logsで特定タスクのエラー詳細を確認できます。カスタム設定、タイムアウト、リトライ試行に対応。成功、失敗、タイムアウトの各シナリオを明確な応答解釈で処理します。迅速な検証が可能です...
tracing-downstream-lineage
astronomer
テーブルやDAGを変更する前に、下流のデータ系列を追跡して変更影響を評価します。ソースコード検索、ビュー依存関係、BIツール接続を通じて、対象テーブルまたはDAGの直接的な消費者を特定します。テーブルからダッシュボード、MLモデルに至るまで、すべての下流影響をマッピングする完全な依存関係ツリーを構築します。依存関係を重要度(クリティカル、高、中、低)で分類し、ステークホルダーへの連絡とテストの優先順位付けを行います。リスク評価と影響を受けるものを含む影響レポートを生成します。