pulumi-debug-failed-operation

作者: pulumi

调试失败的Pulumi更新或预览:读取Pulumi已记录的错误,找出原因并修复。当用户要求…时加载此技能。

npx skills add https://github.com/pulumi/agent-skills --skill pulumi-debug-failed-operation

Debug a failed Pulumi operation

A Pulumi operation has failed. Find what caused it and fix it. The user usually points you at it, so start by working out which operation to debug, and confirm it with the user before doing anything else. Pulumi recorded the error when the operation failed, so once you know which operation it is you can read the error from that record without running anything again.

The commands below reach Pulumi Cloud with pulumi api, a subcommand of the Pulumi CLI that you run in your shell. Each one targets a stack by the explicit {orgName}/{projectName}/{stackName} path you pass it, so you do not need that stack selected locally to read its record. Selecting the stack matters later, when you go to apply a fix.

Start from the operation the user gave you

The user usually supplies the operation as a set of fields: the org, project, stack, and update version (or preview id) — most often stated in prose, for example "debug update 161 of vvm-dev". You need these to address the API: {orgName}, {projectName}, {stackName}, and the version or preview id.

Fill any missing field from context. Take the org, project, or stack from the currently selected stack (pulumi stack --show-name, pulumi stack ls) or Pulumi.yaml. A missing version means the most recent update on that stack.

Briefly confirm which operation you landed on, its version or preview id and the stack, before reading further. Keep it lightweight; they already told you.

Read what failed

A failed update and a failed preview both record engine events, and the error is in the diagnostic messages inside those events. Using the fields you settled on above, fetch the events and pull the messages out.

For a failed update, use the update path with the version number:

pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/update/<version>/events \
  | jq -r '.events[].diagnosticEvent | select(. != null) | "[\(.severity)] \(.message)"' \
  | sed 's/<{%reset%}>//g'

For a failed preview, use the preview path with the preview id:

pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/preview/<preview-id>/events \
  | jq -r '.events[].diagnosticEvent | select(. != null) | "[\(.severity)] \(.message)"' \
  | sed 's/<{%reset%}>//g'

Read every message, not only the ones tagged severity == "error". A provider error carries that error tag, but a program error, which is the common case when a preview fails, arrives as a stderr diagnostic tagged info#err. The trailing sed strips terminal color codes that Pulumi embeds in the text, which otherwise show up as <{%reset%}>.

Find the cause and where the fix belongs

An operation can fail with errors from more than one resource, so read all of the diagnostics first, then work through each error. Trace every error back to the resource that raised it (its URN and type), to where that resource is declared in the program, and to the inputs that feed it.

The error text tells you what kind of problem it is, and that points to where the fix belongs. A Pulumi fix lands in one of three places, and naming the right one keeps you from editing code that was never the problem.

  • The program. The code is wrong: a bad reference, a wrong type, an input the provider rejected, or a value used before it had resolved. This is what a failed preview usually reports, because the plan could not be built. Fix it by editing the code.
  • The state. The code is correct, but the stored state and the real cloud resources disagree. Reconcile drift with pulumi refresh, and bring a resource that already exists outside the state under management with pulumi import rather than recreating it. Note that an operation which failed partway through applying may have already changed some resources, so check the current state before you decide.
  • The environment. The problem is outside Pulumi: credentials, permissions, OIDC, or a quota. Fix the role, the ESC environment, or the capacity that the provider rejected, rather than the resource code.

When a diagnostic is empty or too thin to act on, the real error usually isn't in the record — it's in the log of whatever the resource shelled out to. Read it there. If reaching it needs access you don't have (a token, a run, a not-found), stop and tell the user what you're blocked on and the one thing you need from them.

Fix the cause

Make the smallest change that addresses the root cause. How you confirm the fix, and how you deliver it, whether as a local edit or as a pull request, follow your mode's workflow, not this skill.

If the user didn't say which operation

When the user gives you nothing to go on, debug their most recent operation on the stack. The update list does not record who ran each update, so find it through the API:

  1. Run pulumi whoami to get the current user's login.
  2. Read the latest update and who requested it with pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/updates/latest, and compare its requestedBy.githubLogin to the login from step 1.
  3. If they match, that update is the one to debug. If they do not, walk back one version at a time with pulumi api /api/stacks/{orgName}/{projectName}/{stackName}/updates/<n> until requestedBy.githubLogin matches the user.

Tell the user which operation you landed on, its version, kind, and result, and confirm it is the one they mean before going further.

来自 pulumi 的更多技能

package-usage
pulumi
追踪Pulumi组织中哪些堆栈使用了特定包及其版本。用于跨堆栈审计,识别过时或未维护的…
official
pulumi-automation-api
pulumi
跨多个堆栈和应用程序对Pulumi基础设施操作进行编程化编排。支持本地源(现有Pulumi项目)和内联源(嵌入式程序)架构,实现从简单到复杂多堆栈场景的灵活部署模式。处理具有依赖顺序的多堆栈编排、并行独立部署以及跨堆栈输出传递,以实现协调的基础设施配置。提供编程化...
official
pulumi-best-practices
pulumi
编写可靠、可维护的Pulumi基础设施代码的全面最佳实践。避免在apply()回调中创建资源;直接将Output对象作为输入传递,以保留依赖跟踪和预览可见性。使用ComponentResource类将相关资源分组为可复用的逻辑单元,并通过parent: this建立正确的父子层级。从一开始就使用--secret标志或config.requireSecret()加密机密,防止状态文件中泄露凭据...
official
pulumi-component
pulumi
可复用的基础设施组件,支持多语言、提供合理默认值并采用组合模式。需满足四个核心要素:继承ComponentResource、接收标准参数、为所有子资源设置parent: this、在构造函数末尾调用registerOutputs()。Args接口必须使用Input<T>包装器,避免联合类型和函数,保持结构扁平以支持多语言SDK生成。仅将必要输出暴露为公共属性;隐藏...
official
pulumi-esc
pulumi
集中式机密、配置和动态凭据管理,适用于Pulumi基础设施和应用程序。支持通过导入和分层进行环境组合,包含环境变量、pulumiConfig和文件的保留键。通过OIDC为AWS、Azure和GCP生成短期凭据;与AWS Secrets Manager、Azure Key Vault、HashiCorp Vault和1Password集成。核心CLI命令包括pulumi env init、pulumi env edit、pulumi env open(显示...
official
pulumi-neo-handoff
pulumi
将当前线程单向移交给新的 Pulumi Neo 任务。当用户明确要求移交、发送、转移或继续当前…时使用。
official
pulumi-overview
pulumi
将此技能用于任何创建、修改、检查或销毁云基础设施或SaaS配置的任务,从一次性CLI操作到完整的…
official
pulumi-terraform-to-pulumi
pulumi
将Terraform/OpenTofu项目迁移到Pulumi,包括将HCL源代码转换和/或将Terraform状态导入到Pulumi堆栈中。当用户…
official