deploying-java-sdk-bundles

作者: astronomer

构建并部署编译后的Airflow Java SDK包,以便工作节点能够运行它们。当用户想要将JVM任务包打包为JAR,或询问有关…时使用。

npx skills add https://github.com/astronomer/agents --skill deploying-java-sdk-bundles

Deploying Java SDK Bundles

A Java SDK deployment has one artifact: a bundle — your compiled task classes plus the SDK, packaged as a JAR (or a thin JAR alongside its dependency JARs). You build it with Gradle or Maven, then place it in a directory that the JavaCoordinator scans (jars_root) on every worker. This skill is platform-neutral; it shows the build once, then both an open-source and an Astro deployment path.

Experimental. The Java SDK is in preview. Artifact versions below are shown as ${version}; while the SDK is pre-release you may need to build the artifacts into your local Maven repository yourself (see the preview builds section).

Order of operations: build the bundle (this skill) → place it where jars_root points → configure the coordinator (configuring-airflow-language-sdks). The task code itself is authoring-java-sdk-tasks.


Build with Gradle (recommended)

Apply the SDK's Gradle plugin and declare dependencies in build.gradle:

plugins {
    id("org.apache.airflow.sdk") version "${version}"
}

repositories {
    mavenCentral()
}

dependencies {
    annotationProcessor("org.apache.airflow:airflow-sdk-processor:${version}")  // annotation API only
    implementation("org.apache.airflow:airflow-sdk:${version}")
    // Optional logging integration, e.g.:
    // implementation("org.apache.airflow:airflow-sdk-jpl:${version}")
}

airflowBundle {
    mainClass = "com.example.Main"   // your BundleBuilder entry point
    // fatJar = false                // opt out of the single-JAR build (see below)
}

Build it:

./gradlew bundle

The build/bundle/ directory then holds all required JAR(s). Notes:

  • The annotationProcessor line is needed only if you use the annotation-based API. The interface-based API doesn't need it.
  • By default the plugin produces a fat JAR (via the Shadow plugin) — one self-contained file, which avoids cross-project dependency clashes. Set fatJar = false in airflowBundle for thin JARs; you then deploy every dependency JAR too.
  • The Gradle plugin validates that mainClass exists at build time (verifyBundleMainClass).

Build with Maven

Import the BOM so artifact versions and the supervisor schema version are managed in one place:

<dependencyManagement>
  <dependencies>
    <dependency>
      <groupId>org.apache.airflow</groupId>
      <artifactId>airflow-sdk-bom</artifactId>
      <version>${version}</version>
      <type>pom</type>
      <scope>import</scope>
    </dependency>
  </dependencies>
</dependencyManagement>

<dependencies>
  <dependency>
    <groupId>org.apache.airflow</groupId>
    <artifactId>airflow-sdk</artifactId>   <!-- version from the BOM -->
  </dependency>
</dependencies>

Wire the annotation processor through maven-compiler-plugin (annotation API only) so it stays off the runtime classpath. Then pick a packaging option:

  • Fat JAR (recommended): use maven-shade-plugin. In its ManifestResourceTransformer, set <mainClass> to your BundleBuilder and add the manifest entry Airflow-Supervisor-Schema-Version resolved from the BOM property ${airflow.supervisor.schema.version} (don't hard-code it). mvn package writes the JAR to target/.
  • Thin JAR: use maven-jar-plugin to set Main-Class and maven-dependency-plugin (copy-dependencies) to collect runtime JARs into target/bundle/. Here Airflow-Supervisor-Schema-Version is not needed — Airflow reads it from the airflow-sdk JAR on the classpath.

Unlike Gradle, Maven does not validate mainClass at build time; a wrong value only fails at runtime.


Logging integration

For task log records to reach Airflow's log store (and the task log view in the UI), the bundle must include exactly one SDK logging artifact per logging facade you use. Versions are managed by airflow-sdk-bom; Maven users apply the same artifact IDs.

Choosing a facade. For a greenfield project, prefer JPL (System.Logger) — it is built into the JDK, so your tasks need no extra logging API. Pick another facade only when the libraries you integrate with already log through it, so their records reach Airflow too. Preference order: JPL > SLF4J = Log4j 2 > JUL; treat JUL as legacy integration only, not a choice for new code.

FacadeArtifactSetup beyond the dependency
System.Logger (JPL)airflow-sdk-jplNone — the provider is discovered via ServiceLoader.
SLF4J 2.xairflow-sdk-slf4jNone — the binding is discovered automatically (pulls in slf4j-api for you).
Log4j 2airflow-sdk-log4j2log4j-core on the runtime classpath + AirflowAppender declared in log4j2.xml (below).
java.util.logging (JUL)airflow-sdk-julCall AirflowJulHandler.setup() in main() (below), or use a logging.properties file (see configuring-airflow-language-sdks).

Log4j 2 — log4j-core hosts the plugin loader that discovers the appender (log4j-api comes in transitively):

implementation("org.apache.airflow:airflow-sdk-log4j2:${version}")
runtimeOnly("org.apache.logging.log4j:log4j-core:${log4jVersion}")
<Configuration>
  <Appenders>
    <AirflowAppender name="Airflow"/>
  </Appenders>
  <Loggers>
    <Root level="info">
      <AppenderRef ref="Airflow"/>
    </Root>
  </Loggers>
</Configuration>

JUL — call AirflowJulHandler.setup() before any task runs. It clears the root logger's existing handlers (the default ConsoleHandler writes to stderr, which Airflow would otherwise capture as task.stderr at ERROR level, duplicating each record):

public static void main(String[] args) {
    AirflowJulHandler.setup();
    Server.create(args).serve(new MyBundle().build());
}

Don't double up providers. A second System.LoggerFinder implementation alongside airflow-sdk-jpl, or a second SLF4J binding (logback-classic, slf4j-simple) alongside airflow-sdk-slf4j, makes provider selection unpredictable.


Preview builds (before a stable release)

Skip this section if you depend on a stable release. Once you pin a released version (e.g. 1.0.0) published to Maven Central, the mavenCentral() repository in the build snippets above is enough.

While the SDK is pre-release, the documented path is to build the artifacts and the Gradle plugin from the Airflow repo into your local Maven repository:

# in apache/airflow's java-sdk/ directory
./gradlew publishToMavenLocal -PskipSigning=true

Then add mavenLocal() in your project, in both pluginManagement (in settings.gradle) and project repositories (in build.gradle) — this is how the SDK's own example project resolves it.

Once -SNAPSHOT artifacts are published to Apache's snapshot Nexus, that repository can stand in for the local build (same two places):

maven {
    name = "apacheSnapshots"
    url = "https://repository.apache.org/content/repositories/snapshots/"
    mavenContent { snapshotsOnly() }
}

Snapshots move; force a refresh with ./gradlew bundle --refresh-dependencies. (For Maven, add the same repository to <repositories> and <pluginRepositories>.)


Place the bundle where the coordinator scans

The coordinator scans jars_root recursively and builds the classpath automatically, so you copy the whole output directory:

cp build/bundle/* /opt/airflow/jars/    # /opt/airflow/jars == jars_root

The worker also needs a JRE 17+. Wiring the coordinator to this directory is covered in configuring-airflow-language-sdks.


Deployment paths

Astronomer tooling is not required — the SDK runs on any Airflow with the Task SDK. Choose the path that matches the user's setup.

Open-source (Docker / Kubernetes)

Bake the JRE and the bundle into your Airflow image, or mount them:

FROM apache/airflow:3          # pin a specific 3.x in production
USER root
RUN apt-get update \
    && apt-get install -y --no-install-recommends default-jre-headless \
    && apt-get clean && rm -rf /var/lib/apt/lists/*
RUN mkdir -p /opt/airflow/jars
COPY build/bundle/ /opt/airflow/jars/
USER airflow

On Kubernetes (Helm chart), bake the JAR into a custom image as above, or mount it via a shared volume; set the [sdk] config through environment variables on the worker/scheduler. See deploying-airflow for the broader Docker Compose and Helm workflow.

Astro (one option, not required)

If the user is on Astronomer's Astro CLI, the same idea maps onto an Astro project:

  1. Build the bundle, then stage it in the project: mkdir -p include/jars && cp ../java-bundle/build/bundle/*.jar include/jars/.

  2. Edit the project Dockerfile to install a JRE and copy the JARs to the coordinator's directory:

    FROM quay.io/astronomer/astro-runtime:<version>
    USER root
    RUN apt-get update \
        && apt-get install -y --no-install-recommends default-jre-headless \
        && apt-get clean && rm -rf /var/lib/apt/lists/*
    RUN mkdir -p /opt/airflow/jars
    COPY include/jars/ /opt/airflow/jars/
    USER airflow
    
  3. Put the coordinator config in the project's .env (loaded automatically) — see configuring-airflow-language-sdks for the AIRFLOW__SDK__* values.

  4. astro dev start (or astro dev restart after changes) builds the image and starts Airflow locally; deploy with astro deploy as usual.

Don't pin Astro Runtime / Airflow versions from memory — read the generated Dockerfile or check current docs. While the SDK and Airflow 3.3 are in preview, a beta/dev Astro Runtime image may be required.


Deploy checklist

  • Bundle built (./gradlew bundle or mvn package) and mainClass points at your BundleBuilder.
  • annotationProcessor present iff you use the annotation API.
  • JAR(s) copied into the worker's jars_root directory; with thin JARs, dependency JARs too.
  • JRE 17+ available on the worker.
  • Coordinator + queue_to_coordinator configured (configuring-airflow-language-sdks).
  • If multiple executable JARs exist under jars_root, set main_class explicitly.

Related Skills

  • authoring-java-sdk-tasks: Write the Java task code and the matching Python stubs.
  • configuring-airflow-language-sdks: Register the coordinator and route the queue.
  • deploying-airflow: General Airflow deployment (Astro, Docker Compose, Kubernetes).
  • setting-up-astro-project: Initialize and configure an Astro project.

来自 astronomer 的更多技能

airflow-state-store
astronomer
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (`task_state_store`, `asset_state_store`) and the…
creating-openlineage-extractors
astronomer
针对不受支持的Airflow运算符及复杂血缘场景的自定义OpenLineage提取器。提供两种方案:建议在自有运算符中直接添加OpenLineage方法,或为无法修改的第三方运算符创建自定义提取器。提取器在三个执行节点进行拦截:执行前获取静态血缘、成功后获取运行时输出、可选在失败后获取部分血缘。通过airflow.cfg或环境变量注册提取器...
debugging-dags
astronomer
针对失败的Airflow DAG进行系统性根因分析与修复,提供结构化调查工作流。引导完成四步诊断流程:识别故障、提取错误详情、收集上下文信息、提供可操作的修复步骤。将故障分为四类(数据、代码、基础设施、依赖),以聚焦调查并建议适当的修复方案。提供即用型CLI命令,用于日志检索、运行对比、任务清除及DAG...
delegating-to-otto
astronomer
Drives Astronomer's Otto agent (`astro otto`) as a delegated sub-agent for Airflow, dbt, and data-engineering work. Use when the user explicitly asks to "use…
deploying-airflow
astronomer
部署Airflow DAG和项目。当用户想要部署代码、推送DAG、设置CI/CD、部署到生产环境,或询问部署策略时使用…
deploying-go-sdk-bundles
astronomer
编译、打包并部署已编译的Airflow Go SDK包,以便ExecutableCoordinator能够运行它们。当用户想要编译Go任务包时使用,询问…
testing-dags
astronomer
针对Airflow DAG的迭代式测试-调试-修复循环,提供全面的故障诊断。首先使用af runs trigger-wait <dag_id>运行DAG并等待完成,无需预检。失败时,使用af runs diagnose获取全面的故障摘要,并通过af tasks logs查看特定任务的错误详情。支持自定义配置、超时和重试次数;处理成功、失败和超时场景,并给出清晰的响应解读。提供快速验证功能...
tracing-downstream-lineage
astronomer
追踪下游数据血缘,在修改表或DAG前评估变更影响。通过源代码搜索、视图依赖和BI工具连接识别目标表或DAG的直接消费者,构建完整的依赖树,映射从表到仪表盘再到机器学习模型的所有下游影响。按关键性(关键、高、中、低)对依赖进行分类,以优先安排利益相关者沟通和测试。生成包含风险评估、受影响...的影响报告。