pulumi-debug-failed-operation

作者: pulumi

調試失敗的 Pulumi 更新或預覽:讀取 Pulumi 已記錄的失敗資訊,找出原因並修復。當使用者要求…時載入此技能。

npx skills add https://github.com/pulumi/agent-skills --skill pulumi-debug-failed-operation

Debug a failed Pulumi operation

A Pulumi operation has failed. Find what caused it and fix it. The user usually points you at it, so start by working out which operation to debug, and confirm it with the user before doing anything else. Pulumi recorded the error when the operation failed, so once you know which operation it is you can read the error from that record without running anything again.

The commands below reach Pulumi Cloud with pulumi api, a subcommand of the Pulumi CLI that you run in your shell. Each one targets a stack by the explicit {orgName}/{projectName}/{stackName} path you pass it, so you do not need that stack selected locally to read its record. Selecting the stack matters later, when you go to apply a fix.

Start from the operation the user gave you

The user usually supplies the operation as a set of fields: the org, project, stack, and update version (or preview id) — most often stated in prose, for example "debug update 161 of vvm-dev". You need these to address the API: {orgName}, {projectName}, {stackName}, and the version or preview id.

Fill any missing field from context. Take the org, project, or stack from the currently selected stack (pulumi stack --show-name, pulumi stack ls) or Pulumi.yaml. A missing version means the most recent update on that stack.

Briefly confirm which operation you landed on, its version or preview id and the stack, before reading further. Keep it lightweight; they already told you.

Read what failed

A failed update and a failed preview both record engine events, and the error is in the diagnostic messages inside those events. Using the fields you settled on above, fetch the events and pull the messages out.

For a failed update, use the update path with the version number:

pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/update/<version>/events \
  | jq -r '.events[].diagnosticEvent | select(. != null) | "[\(.severity)] \(.message)"' \
  | sed 's/<{%reset%}>//g'

For a failed preview, use the preview path with the preview id:

pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/preview/<preview-id>/events \
  | jq -r '.events[].diagnosticEvent | select(. != null) | "[\(.severity)] \(.message)"' \
  | sed 's/<{%reset%}>//g'

Read every message, not only the ones tagged severity == "error". A provider error carries that error tag, but a program error, which is the common case when a preview fails, arrives as a stderr diagnostic tagged info#err. The trailing sed strips terminal color codes that Pulumi embeds in the text, which otherwise show up as <{%reset%}>.

Find the cause and where the fix belongs

An operation can fail with errors from more than one resource, so read all of the diagnostics first, then work through each error. Trace every error back to the resource that raised it (its URN and type), to where that resource is declared in the program, and to the inputs that feed it.

The error text tells you what kind of problem it is, and that points to where the fix belongs. A Pulumi fix lands in one of three places, and naming the right one keeps you from editing code that was never the problem.

  • The program. The code is wrong: a bad reference, a wrong type, an input the provider rejected, or a value used before it had resolved. This is what a failed preview usually reports, because the plan could not be built. Fix it by editing the code.
  • The state. The code is correct, but the stored state and the real cloud resources disagree. Reconcile drift with pulumi refresh, and bring a resource that already exists outside the state under management with pulumi import rather than recreating it. Note that an operation which failed partway through applying may have already changed some resources, so check the current state before you decide.
  • The environment. The problem is outside Pulumi: credentials, permissions, OIDC, or a quota. Fix the role, the ESC environment, or the capacity that the provider rejected, rather than the resource code.

When a diagnostic is empty or too thin to act on, the real error usually isn't in the record — it's in the log of whatever the resource shelled out to. Read it there. If reaching it needs access you don't have (a token, a run, a not-found), stop and tell the user what you're blocked on and the one thing you need from them.

Fix the cause

Make the smallest change that addresses the root cause. How you confirm the fix, and how you deliver it, whether as a local edit or as a pull request, follow your mode's workflow, not this skill.

If the user didn't say which operation

When the user gives you nothing to go on, debug their most recent operation on the stack. The update list does not record who ran each update, so find it through the API:

  1. Run pulumi whoami to get the current user's login.
  2. Read the latest update and who requested it with pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/updates/latest, and compare its requestedBy.githubLogin to the login from step 1.
  3. If they match, that update is the one to debug. If they do not, walk back one version at a time with pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/updates/<n> until requestedBy.githubLogin matches the user.

Tell the user which operation you landed on, its version, kind, and result, and confirm it is the one they mean before going further.

來自 pulumi 的更多技能

package-usage
pulumi
追蹤 Pulumi 組織中各堆疊使用特定套件及其版本的情況。用於跨堆疊審計,識別過時或未維護的…
official
pulumi-automation-api
pulumi
跨多個堆疊與應用程式的 Pulumi 基礎設施操作之程式化編排。支援本地來源(現有 Pulumi 專案)與內嵌來源(嵌入式程式)架構,實現從簡單到複雜多堆疊場景的靈活部署模式。處理具相依性排序的多堆疊編排、平行獨立部署,以及跨堆疊輸出傳遞,以達成協調的基礎設施佈建。提供程式化...
official
pulumi-best-practices
pulumi
撰寫可靠、可維護的 Pulumi 基礎設施程式碼的全面最佳實踐。避免在 apply() 回呼中建立資源;直接將 Output 物件作為輸入傳遞,以保留依賴追蹤與預覽可見性。使用 ComponentResource 類別將相關資源分組為可重複使用的邏輯單元,並透過 parent: this 建立正確的父子層級。從一開始就使用 --secret 標誌或 config.requireSecret() 加密機密資訊,以防止憑證在狀態檔案中洩漏...
official
pulumi-component
pulumi
可重複使用的基礎架構元件,支援多語言、合理的預設值與組合模式。需具備四個核心要素:繼承 ComponentResource、接受標準參數、為所有子資源設定 parent: this,並在建構函式結尾呼叫 registerOutputs()。Args 介面必須使用 Input<T> 包裝器,避免聯合型別與函式,並保持結構扁平以支援多語言 SDK 生成。僅將必要的輸出暴露為公開屬性;隱藏...
official
pulumi-esc
pulumi
集中式機密、配置及動態憑證管理,適用於Pulumi基礎設施與應用程式。支援透過匯入與分層進行環境組合,並保留 environmentVariables 、 pulumiConfig 及 files 的保留鍵。透過OIDC為AWS、Azure及GCP產生短期憑證;可整合AWS Secrets Manager、Azure Key Vault、HashiCorp Vault及1Password。核心CLI指令包含 pulumi env init 、 pulumi env edit 、 pulumi env open (顯示...
official
pulumi-neo-handoff
pulumi
將當前線程以單向傳輸方式移交給新的 Pulumi Neo 任務。當用戶明確要求移交、發送、轉移或繼續當前…時使用。
official
pulumi-overview
pulumi
使用此技能處理任何建立、修改、檢查或銷毀雲端基礎設施或SaaS配置的任務,範圍從一次性CLI操作到完整…
official
pulumi-terraform-to-pulumi
pulumi
將 Terraform/OpenTofu 專案遷移至 Pulumi,包括轉譯 HCL 原始碼及/或將 Terraform 狀態匯入 Pulumi 堆疊。當使用者…
official