analyzing-data

작성자: astronomer

데이터 웨어하우스에 질의하여 캐시된 패턴과 개념 매핑을 통해 비즈니스 질문에 답변합니다. 반복되는 질문 유형에 대한 패턴 조회 및 캐싱을 지원하며, 결과 기록을 통해 향후 질의를 개선합니다. 개념-테이블 매핑 캐시와 INFORMATION_SCHEMA 또는 코드베이스 grep을 통한 테이블 스키마 탐색을 포함합니다. 분석을 위해 Polars 또는 Pandas DataFrame을 반환하는 run_sql() 및 run_sql_pandas() 커널 함수를 제공합니다. 개념, 패턴 및 테이블 캐시를 관리하기 위한 CLI 명령어와 추가 기능을 포함합니다.

npx skills add https://github.com/astronomer/agents --skill analyzing-data

Data Analysis

Answer business questions by querying the data warehouse. The kernel auto-starts on first exec call.

All CLI commands below are relative to this skill's directory. Before running any scripts/cli.py command, cd to the directory containing this file.

Workflow

  1. Pattern lookup — Check for a cached query strategy:

    uv run scripts/cli.py pattern lookup "<user's question>"
    

    If a pattern exists, follow its strategy. Record the outcome after executing:

    uv run scripts/cli.py pattern record <name> --success  # or --failure
    
  2. Concept lookup — Find known table mappings:

    uv run scripts/cli.py concept lookup <concept>
    
  3. Table discovery — If cache misses, search the codebase (Grep pattern="<concept>" glob="**/*.sql") or query INFORMATION_SCHEMA. See reference/discovery-warehouse.md.

  4. Execute query:

    uv run scripts/cli.py exec "df = run_sql('SELECT ...')"
    uv run scripts/cli.py exec "print(df)"
    
  5. Cache learnings — Always cache before presenting results:

    # Cache concept → table mapping
    uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>
    # Cache query strategy (if discovery was needed)
    uv run scripts/cli.py pattern learn <name> -q "question" -s "step" -t "TABLE" -g "gotcha"
    
  6. Present findings to user.

Kernel Functions

FunctionReturns
run_sql(query, limit=100)Polars DataFrame
run_sql_pandas(query, limit=100)Pandas DataFrame
run_sql_many(queries, limit=100)List of Polars DataFrames (one per query)

pl (Polars) and pd (Pandas) are pre-imported.

Run independent queries together with run_sql_many — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:

uv run scripts/cli.py exec "dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])"

run_sql_many is fail-fast: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate run_sql calls if you need partial results.

Timeouts: exec waits up to 120s by default, then interrupts the query and returns a "client stopped waiting" message (the query may still finish server-side). Raise it for known long-running queries: uv run scripts/cli.py exec "..." -t 600.

Idle kernel: the kernel self-terminates after 2h idle (preserving state until then). Override with ASTRO_KERNEL_IDLE_TIMEOUT (seconds; 0 disables).

CLI Reference

Kernel

uv run scripts/cli.py warehouse list      # List warehouses
uv run scripts/cli.py start [-w name]     # Start kernel (with optional warehouse)
uv run scripts/cli.py exec "..."          # Execute Python code
uv run scripts/cli.py status              # Kernel status
uv run scripts/cli.py restart             # Restart kernel
uv run scripts/cli.py stop                # Stop kernel
uv run scripts/cli.py install <pkg>       # Install package

Concept Cache

uv run scripts/cli.py concept lookup <name>                     # Look up
uv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn
uv run scripts/cli.py concept list                               # List all
uv run scripts/cli.py concept import -p /path/to/warehouse.md   # Bulk import

Pattern Cache

uv run scripts/cli.py pattern lookup "question"                                      # Look up
uv run scripts/cli.py pattern learn <name> -q "..." -s "..." -t "TABLE" -g "gotcha"  # Learn
uv run scripts/cli.py pattern record <name> --success                                # Record outcome
uv run scripts/cli.py pattern list                                                   # List all
uv run scripts/cli.py pattern delete <name>                                          # Delete

Table Schema Cache

uv run scripts/cli.py table lookup <TABLE>            # Look up schema
uv run scripts/cli.py table cache <TABLE> -c '[...]'  # Cache schema
uv run scripts/cli.py table list                       # List cached
uv run scripts/cli.py table delete <TABLE>             # Delete

Cache Management

uv run scripts/cli.py cache status                # Stats
uv run scripts/cli.py cache clear [--stale-only]  # Clear

References

astronomer의 다른 스킬

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
지원되지 않는 Airflow 연산자와 복잡한 계보 시나리오를 위한 맞춤형 OpenLineage 추출기. 두 가지 접근 방식: 소유한 연산자에 직접 OpenLineage 메서드를 추가(권장)하거나, 수정할 수 없는 타사 연산자를 위한 맞춤형 추출기를 생성합니다. 추출기는 세 지점에서 연산자 실행을 가로챕니다: 정적 계보를 위한 실행 전, 런타임에 결정된 출력을 위한 성공 후, 그리고 선택적으로 부분 계보를 위한 실패 후. airflow.cfg 또는 환경을 통해 추출기를 등록합니다...
debugging-dags
astronomer
체계적인 근본 원인 분석 및 구조화된 조사 워크플로를 통한 실패한 Airflow DAG의 문제 해결. 4단계 진단 프로세스를 안내합니다: 실패 식별, 오류 세부 정보 추출, 컨텍스트 정보 수집, 실행 가능한 수정 단계 제공. 실패를 네 가지 유형(데이터, 코드, 인프라, 종속성)으로 분류하여 조사에 집중하고 적절한 수정을 제안합니다. 로그 검색, 실행 비교, 작업 정리, DAG...을 위한 즉시 사용 가능한 CLI 명령을 제공합니다.
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
Airflow DAG 및 프로젝트를 배포합니다. 사용자가 코드를 배포하거나, DAG를 푸시하거나, CI/CD를 설정하거나, 프로덕션에 배포하거나, 배포 전략에 대해 질문할 때 사용하세요.
deploying-go-sdk-bundles
astronomer
컴파일된 Airflow Go SDK 번들을 빌드, 패킹 및 배포하여 ExecutableCoordinator가 실행할 수 있도록 합니다. 사용자가 Go 태스크 번들을 컴파일하려고 하거나 요청할 때 사용합니다.
testing-dags
astronomer
포괄적인 실패 진단 기능을 갖춘 Airflow DAG의 반복적인 테스트-디버그-수정 주기. af runs trigger-wait <dag_id>로 시작하여 DAG를 실행하고 완료를 기다립니다. 사전 점검은 필요하지 않습니다. 실패 시 af runs diagnose를 사용하여 포괄적인 실패 요약을 확인하고, af tasks logs를 사용하여 특정 태스크의 오류 세부 정보를 검사합니다. 사용자 정의 구성, 시간 제한 및 재시도 횟수를 지원하며, 명확한 응답 해석으로 성공, 실패 및 시간 초과 시나리오를 처리합니다. 빠른 검증 가능...
tracing-downstream-lineage
astronomer
테이블이나 DAG를 수정하기 전에 다운스트림 데이터 계보를 추적하여 변경 영향을 평가합니다. 소스 코드 검색, 뷰 종속성, BI 도구 연결을 통해 대상 테이블 또는 DAG의 직접적인 소비자를 식별합니다. 테이블에서 대시보드, ML 모델에 이르기까지 모든 다운스트림 영향을 매핑하는 전체 종속성 트리를 구축합니다. 종속성을 중요도(심각, 높음, 중간, 낮음)별로 분류하여 이해관계자 커뮤니케이션 및 테스트의 우선순위를 지정합니다. 위험 평가, 영향을 받는 항목이 포함된 영향 보고서를 생성합니다...