cluster-update-advisor

작성자: openshift

OpenShift 클러스터 업데이트(업그레이드) 준비 상태와 위험을 평가합니다. 클러스터 업데이트가 안전한지 평가할 때, 업데이트가 가능할 때, 또는...

npx skills add https://github.com/openshift/agentic-skills --skill cluster-update-advisor

Cluster Update Advisor

Purpose

Assess cluster update readiness and produce a structured risk report with actionable prerequisites, blockers, and recommendations.

The proposal request includes pre-collected cluster readiness data (JSON) gathered by the Cluster Version Operator. Analyze this data, classify findings, and produce a decision with evidence. Do not re-collect cluster data — it is already in the request.

Inputs

The proposal request contains:

  • Current and target version metadata
  • Channel and update path information
  • Cluster readiness JSON — cluster health checks with context relevant to preparing for the update

The readiness JSON is embedded in the request between ```json markers under the "Cluster Readiness Data" heading. Parse it to begin analysis.

Readiness JSON structure:

{
  "current_version": "4.21.5",
  "target_version": "4.21.8",
  "checks": {
    "cluster_conditions":    { "_status": "ok", "summary": {...}, ... },
    "operator_health":       { "_status": "ok", "summary": {...}, ... },
    "api_deprecations":      { "_status": "ok", "summary": {...}, ... },
    "node_capacity":         { "_status": "ok", "summary": {...}, ... },
    "pdb_drain":             { "_status": "ok", "summary": {...}, ... },
    "etcd_health":           { "_status": "ok", "summary": {...}, ... },
    "network":               { "_status": "ok", "summary": {...}, ... },
    "crd_compat":            { "_status": "ok", "summary": {...}, ... },
    "olm_operator_lifecycle": { "_status": "ok", "summary": {...}, ... }
  }
}

Each check contains _status (ok or error) and check-specific data with a summary section for quick parsing.

Evaluation

Parse readiness data

Extract the JSON from the proposal request. Count checks with _status ok vs error for completeness.

Verify data completeness

Any check with _status error represents a gap in visibility. Note incomplete areas — they reduce confidence.

Evaluate findings in detail

If the system prompt includes organization-specific policy (thresholds, scheduling preferences, risk tolerance), apply those constraints. Otherwise use sensible defaults. Walk through each check's summary and detail data:

  • Compare numeric thresholds (node headroom, etcd backup age)
  • Evaluate conditional update risks against cluster state
  • Identify compounding risks (e.g., paused MCP + cert expiry)
  • Estimate update duration (~10 min/node)

Classify findings

Assign each finding a severity per the classification table.

CheckBlocker if...Warning if...
Cluster conditionsUpgradeable=False (non-z-stream)Update already in progress
API deprecationsWorkloads use APIs removed in targetWorkloads use deprecated APIs
Operator healthAny operator has Upgradeable=FalseAny operator is Degraded=True
MachineConfigPoolAny MCP paused or degradedMCP updating or not all machines ready
Node capacityHeadroom < 20%Headroom < 40%
PDB configPDB blocks ALL replicas from drainingPDB has maxUnavailable: 0
etcd healthAny member unhealthyNo recent backup (within 24h)
Network pluginSDN in use and target requires OVN (4.17+)Using deprecated SDN (< 4.17)
CRD compatibilityStored version not served; operator maxOpenShiftVersion < targetDeprecated versions still served
OLM operator lifecycleInstalled operator incompatible with target OCP; operator product EOLOperator has pending update; operator product in Maintenance Support

For other checks, treat an issue as a blocker if would cause data loss, a performance regression, or a failed update. Treat the issue as a warning if would cause temporary disruption or slow updates.

Investigate with other skills

If additional information or context is needed to classify a finding, these skills may be useful:

  • openshift-docs — Read official OpenShift update docs for version-specific procedures and breaking changes.

  • prometheus — Query cluster metrics for trend analysis (etcd latency, CPU headroom, firing alerts).

  • jira — Search Red Hat Jira for bugs and known issues affecting the target version.

  • product-lifecycle — Query Red Hat Product Life Cycle API to check support status and OCP compatibility for installed operators. Use the operator's package name from OLM readiness data to look up entries via the package field (exact match). Flag operators whose product version is End of life or whose openshift_compatibility does not include the target OCP version.

Classify overall recommendation

Aggregate finding classification, and and make a decision on the overall assessment:

  • escalate — insufficient data for confident assessment.
  • block — findings must be resolved before update.
  • warn — findings exist but manageable with prerequisites.
  • recommend — all checks pass within acceptable thresholds.
BlockersWarningsDecision
Unable to assessanyescalate
1+anyblock
01+warn
00recommend

Produce a structured risk report

The output schema is enforced by the OlsAgent CR's outputSchema field — the operator handles structured output compliance via the LLM API.

Failure Modes — What NOT to Do

  1. Never recommend updating without analyzing the readiness data. The JSON in the request is the source of truth.

  2. Never dismiss conditional update risks. If the update path is conditional, evaluate each risk against the cluster.

  3. Never skip the API deprecation check. Workloads using removed APIs will break after the update.

  4. Never assume etcd is healthy. Always check member health in the readiness data.

  5. Never fabricate Jira issue keys, KB article IDs, or CVE numbers. Use the redhat-support skill to get real data.

  6. Never recommend skipping an update version unless the readiness data shows that path exists.

  7. Never recommend force-updating. If the standard path is blocked, report it.

openshift의 다른 스킬

openshift-docs
openshift
OpenShift Container Platform 문서를 마크다운 형식으로 검색하고 읽습니다. 사용자가 OpenShift 기능, 구성, 설치 등에 대해 질문할 때 사용합니다.
triage-leaked-infra
openshift
AWS VPC 또는 HyperShift CI의 인프라 세트가 삭제해도 안전한지 평가합니다. 사용자가 cleanleaked 출력을 붙여넣고 '이거 삭제해도 되나요?', '이거...'라고 물을 때 사용합니다.
openshift-expert
openshift
OpenShift 플랫폼 및 Kubernetes 전문가로, 클러스터 아키텍처, 오퍼레이터, 네트워킹, 스토리지, 문제 해결 및 CI/CD 파이프라인에 대한 깊은 지식을 보유하고 있습니다. 사용…
Konflux Archived PipelineRuns
openshift
KubeArchive를 통해 보관된 Konflux PipelineRun, TaskRun 및 파드 로그에 접근합니다. Konflux PipelineRun 결과를 확인하거나 조사할 때 자동으로 적용됩니다.
backport
openshift
메인 브랜치에서 릴리스 브랜치로 커밋이나 PR을 백포트합니다. 사용자가 브랜치 간 변경 사항을 백포트, 체리픽, 포팅하거나 해결을 요청할 때 사용합니다.
rebase
openshift
현재 브랜치를 기본 브랜치 위로 리베이스하고, 모든 충돌을 해결한 뒤 린트, i18n, 빌드가 통과하는지 확인합니다. 사용자가 리베이스, 업데이트, 또는 동기화를 요청할 때 사용합니다…
Build CPO Image
openshift
컨트롤 플레인 오퍼레이터 컨테이너 이미지를 빌드하고 푸시합니다. 라이브 클러스터에 배포가 필요한 CPO 변경 사항을 테스트할 때 자동으로 적용됩니다.
find-complexity
openshift
순환 복잡도가 높거나, 길이가 지나치게 길거나, 매개변수가 너무 많은 함수와 메서드를 찾습니다. 사용자가 복잡한 코드나 복잡도를 찾아 달라고 요청할 때 사용하세요.