airflow-java-sdk

Guide for contributing to the Airflow Java SDK (AIP-108). Use this skill whenever a contributor is working in the `java-sdk/` directory or on the Java…

npx skills add https://github.com/astronomer/airflow --skill airflow-java-sdk

Airflow Java SDK contributor guide

The Java SDK lets Airflow tasks execute JVM code (Java, Kotlin, or any JVM language). You are helping a contributor work in one or both of these locations:

  • java-sdk/ — the JVM-side library (Kotlin source, published to Maven)
  • task-sdk/src/airflow/sdk/coordinators/java/ — the Python coordinator that launches the JVM subprocess

Read these two documents early in every session — they contain the authoritative reference material:

  • airflow-core/docs/authoring-and-scheduling/language-sdks/java.rst — user-facing guide: annotation vs. interface API, XCom type mapping, Gradle/Maven steps, coordinator config.
  • java-sdk/README.md — contributor guide: repository layout, detailed execution walkthrough, Gradle + Breeze test commands, coding conventions, common tasks, and PR checklist.

SDK package architecture

The JVM-side library is split into two packages with distinct visibility rules:

  • org.apache.airflow.sdk — public, user-facing API. Classes here (e.g. Client, Bundle, BundleBuilder, Server) are stable contracts that DAG authors and task implementers import directly. Changes to this package are breaking changes.
  • org.apache.airflow.sdk.execution — internal implementation detail. Everything in this package (CoordinatorComm, LogSender, Log, Client in execution/, generated schema models, etc.) is not intended to be imported by users. It may change between releases without notice.

When reviewing or writing code, enforce this boundary: user task code and BundleBuilder subclasses must only import from org.apache.airflow.sdk; any import of org.apache.airflow.sdk.execution.* in user-facing API surface is a red flag.


Bundle composition and coordinator discovery

A bundle is a directory of JAR files (typically build/bundle/) placed on the coordinator's jars_root. The coordinator scans the directory at task-dispatch time to find:

  1. Main-Class (standard JAR manifest attribute) — the fully-qualified class name of the entry point that the coordinator invokes with java -classpath … <Main-Class> --comm … --logs …. This must be a class with a public static void main(String[] args) method; the Gradle plugin org.apache.airflow.sdk writes it automatically from airflowBundle { mainClass = "…" } and validates that the class exists and has the right signature at build time.

  2. Airflow-Supervisor-Schema-Version (Airflow-specific manifest attribute) — the wire protocol version the JVM side expects when talking to the Python supervisor. In fat-JAR mode (the default), the Gradle plugin reads this value from the airflow-sdk JAR in runtimeClasspath and copies it into the shadow JAR manifest. In thin-JAR mode (fatJar = false), the value stays in the airflow-sdk JAR deployed alongside the bundle JAR.

The Python coordinator (JavaCoordinator) scans every JAR under jars_root with _JarInfo.find(), reads META-INF/MANIFEST.MF out of each ZIP, and collects Main-Class and Airflow-Supervisor-Schema-Version from whichever JARs carry them. The resolved schema version is then passed as the schema_version return value from _build_execute_task_command, which the base SubprocessCoordinator uses to negotiate the supervisor wire protocol.

If main_class is set explicitly on the JavaCoordinator instance (via [sdk] coordinators kwargs), the scan uses it as a filter; otherwise the first JAR with a Main-Class attribute wins. Either way, Airflow-Supervisor-Schema-Version must be present in at least one JAR in jars_root or startup fails.


Key files to know

FilePurpose
java-sdk/sdk/.../Client.ktPublic API (Variables, Connections, XCom)
java-sdk/sdk/.../execution/Client.ktSupervisor wire calls
java-sdk/sdk/.../execution/Comm.kt4-byte-prefix MessagePack framing
java-sdk/sdk/.../Server.ktEntry-point; drives the execution loop
java-sdk/processor/.../BuilderProcessor.ktKapt annotation processor
java-sdk/plugin/.../AirflowSdkPlugin.ktGradle bundle plugin
task-sdk/.../coordinators/java/coordinator.pyPython side — spawns the JVM
task-sdk/.../schema/schema.jsonWire protocol definition (both sides)

Running tests

Always use ./gradlew from inside java-sdk/; never run Gradle via apt's gradle. See java-sdk/README.md#testing for the full list of Gradle commands.

For the Python coordinator, use Breeze (never pytest directly on the host):

breeze testing task-sdk-tests -- task_sdk/coordinators/java

End-to-end test suite:

E2E_TEST_MODE=java_sdk uv run --project airflow-e2e-tests pytest \
    tests/airflow_e2e_tests/java_sdk_tests/ -xvs

Updating the Python coordinator

coordinator.py extends SubprocessCoordinator. The only method subclasses must implement is _build_execute_task_command, which returns (argv, schema_version). Look at the existing implementation for how jars_root, java_executable, jvm_args, and main_class are assembled into the command. Do not reach into the JVM process from Python beyond what this method provides.


Upgrading Supervisor Schema client

When upgrading to a newer Supervisor Schema version:

  • Regenerate models with ./gradlew generateJsonSchema2Pojo
  • Modify execution/Client.kt to handle changes

The java-sdk/README.md#contributing section walks through the full "adding a new Client method" sequence step by step.

Thêm skills từ astronomer

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
Các extractor OpenLineage tùy chỉnh cho các toán tử Airflow không được hỗ trợ và các kịch bản lineage phức tạp. Hai cách tiếp cận: thêm các phương thức OpenLineage trực tiếp vào các toán tử bạn sở hữu (khuyến nghị), hoặc tạo các extractor tùy chỉnh cho các toán tử bên thứ ba mà bạn không thể sửa đổi. Extractor can thiệp vào quá trình thực thi toán tử tại ba điểm: trước khi thực thi để lấy lineage tĩnh, sau khi thành công để lấy đầu ra được xác định trong thời gian chạy, và tùy chọn sau khi thất bại để lấy lineage một phần. Đăng ký extractor thông qua airflow.cfg hoặc môi trường...
debugging-dags
astronomer
Phân tích nguyên nhân gốc rễ có hệ thống và khắc phục cho các DAG Airflow bị lỗi với quy trình điều tra có cấu trúc. Hướng dẫn qua quy trình chẩn đoán bốn bước: xác định lỗi, trích xuất chi tiết lỗi, thu thập thông tin ngữ cảnh và đưa ra các bước khắc phục khả thi. Phân loại lỗi thành bốn loại (dữ liệu, mã, cơ sở hạ tầng, phụ thuộc) để tập trung điều tra và đề xuất các bản sửa lỗi phù hợp. Cung cấp các lệnh CLI sẵn sàng sử dụng để truy xuất nhật ký, so sánh lần chạy, xóa tác vụ và DAG...
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
Triển khai Airflow DAGs và các dự án. Sử dụng khi người dùng muốn triển khai mã, đẩy DAGs, thiết lập CI/CD, triển khai lên môi trường sản xuất hoặc hỏi về các chiến lược triển khai…
deploying-go-sdk-bundles
astronomer
Xây dựng, đóng gói và triển khai các gói Airflow Go SDK đã biên dịch để ExecutableCoordinator có thể chạy chúng. Sử dụng khi người dùng muốn biên dịch một gói tác vụ Go, yêu cầu…
testing-dags
astronomer
Các chu trình kiểm tra-gỡ lỗi-sửa lỗi lặp đi lặp lại cho Airflow DAG với chẩn đoán lỗi toàn diện. Bắt đầu bằng af runs trigger-wait <dag_id> để chạy một DAG và chờ hoàn tất; không cần kiểm tra trước khi chạy. Khi gặp lỗi, sử dụng af runs diagnose để có bản tóm tắt lỗi toàn diện và af tasks logs để kiểm tra chi tiết lỗi từ các tác vụ cụ thể. Hỗ trợ cấu hình tùy chỉnh, thời gian chờ và số lần thử lại; xử lý các tình huống thành công, thất bại và hết thời gian chờ với diễn giải phản hồi rõ ràng. Có sẵn tính năng xác th
tracing-downstream-lineage
astronomer
Truy xuất dòng dữ liệu xuôi dòng để đánh giá tác động của thay đổi trước khi sửa đổi bảng hoặc DAG. Xác định người tiêu dùng trực tiếp của bảng hoặc DAG mục tiêu thông qua tìm kiếm mã nguồn, phụ thuộc view và kết nối công cụ BI. Xây dựng cây phụ thuộc đầy đủ ánh xạ tất cả tác động xuôi dòng, từ bảng đến dashboard đến mô hình ML. Phân loại phụ thuộc theo mức độ quan trọng (quan trọng, cao, trung bình, thấp) để ưu tiên giao tiếp với các bên liên quan và kiểm thử. Tạo báo cáo tác động với đánh giá rủi ro, các thành phần bị