production-incident-responder

โดย kotlin

แนะนำการตอบสนองต่อเหตุการณ์ในการผลิตสำหรับบริการ Kotlin และ Spring ตั้งแต่การแจ้งเตือนครั้งแรกจนถึงการบรรเทา การวินิจฉัย และการติดตามผล ใช้เมื่ออัตราข้อผิดพลาดพุ่งสูงขึ้น…

npx skills add https://github.com/kotlin/kotlin-backend-agent-skills --skill production-incident-responder

Production Incident Responder

Source mapping: Tier 2 high-value skill derived from Kotlin_Spring_Developer_Pipeline.md (SK-24).

Mission

Restore service safely before chasing perfect explanations. Keep mitigation, diagnosis, communication, and evidence preservation disciplined and explicit.

First Principles

  • Mitigate first.
  • Prefer reversible actions over heroic code changes.
  • Preserve evidence while the system is still exhibiting the problem.
  • Separate confirmed facts from working hypotheses.

Inputs To Gather

  • Current alert state, user impact, and blast radius.
  • Recent deploys, config changes, feature-flag changes, and dependency incidents.
  • Key dashboards: latency, error rate, saturation, dependency health, queue depth, pool usage.
  • Correlated traces and logs for the failing path.
  • Known runbooks, rollback mechanisms, and feature flags.

Response Sequence

  1. State impact and likely severity.
  2. Stop unsafe changes and identify the fastest reversible mitigation:
    • rollback
    • disable feature
    • reduce concurrency
    • shed load
    • rate-limit callers
    • isolate or degrade a dependency
  3. Preserve high-signal evidence while the symptom still exists.
  4. Compare timeline of incident onset with recent changes.
  5. Localize the failing layer: application, database, downstream dependency, queue, infrastructure, or configuration.
  6. Propose long-term corrective actions only after the service is stable.

Advanced Incident Heuristics

  • Restarting everything can destroy the best evidence and amplify a connection storm. Use restarts deliberately, not reflexively.
  • Scaling the app tier does not help when the database or a downstream service is the bottleneck.
  • Rate-limiting or queue pausing may protect core flows better than full rollback when only one feature path is toxic.
  • A config-only incident can look like a code regression; compare effective runtime values before patching application code.
  • A healthy dependency at low volume can still fail under retry storms from your own fleet.
  • If the system uses caches, verify whether bad cache fill, stampede, or stale data amplified the incident.
  • If the incident is intermittent, preserve timing and hypothesis logs. Races and saturation patterns are easy to lose after mitigation.
  • Post-incident work must include detection and prevention, not only the code fix.

Incident Command Nuances

  • One person should own technical command during a serious incident. Parallel debugging without a decision owner often slows mitigation.
  • Communication cadence matters. Operators and stakeholders need regular updates even when the technical picture is incomplete.
  • Rollback is not always safe if data shape or side effects have already changed. Assess rollback safety before pressing the button.
  • Canary comparison, feature-flag cohort analysis, and effective-config diffing often localize incidents faster than code inspection.
  • Preserve version, commit, config, and infrastructure fingerprints in the incident notes while they are still recoverable.

Expert Heuristics

  • Choose the first mitigation that reduces blast radius and buys time, not the one that feels most technically satisfying.
  • Prefer mitigations that also test a hypothesis when that can be done safely.
  • If the incident spans several layers, identify the current bottlenecked layer first. Solving secondary symptoms wastes the window of action.
  • A good postmortem action item changes detection, defaults, rollout strategy, or operational safety nets, not just one line of code.

Output Contract

Return these sections:

  • Impact: who or what is affected and how badly.
  • Immediate mitigation: the safest reversible action to reduce pain now.
  • Evidence: the strongest signals collected so far.
  • Working hypothesis: the leading explanation plus uncertainty.
  • Next diagnostic step: the most informative next action once stable.
  • Follow-up: long-term fix, monitoring change, and postmortem actions.

Guardrails

  • Do not recommend code changes as the very first incident action when a reversible mitigation exists.
  • Do not claim root cause certainty without evidence.
  • Do not optimize for elegance over containment during an outage.
  • Do not forget operator communication and blast-radius tracking while debugging.

Quality Bar

A good run of this skill reduces user pain quickly and leaves the team with a cleaner path to root cause. A bad run jumps to speculative code fixes while the service remains unstable and evidence disappears.

Skills เพิ่มเติมจาก kotlin

kotlin-backend-jpa-entity-mapping
kotlin
คลาสข้อมูลของ Kotlin เหมาะกับ DTO แต่เป็นอันตรายสำหรับเอนทิตี้ JPA Hibernate อาศัยความหมายของเอกลักษณ์ที่คลาสข้อมูลทำลาย: equals / hashCode ที่ครอบคลุมทุกฟิลด์ทำให้สมาชิกภาพของ Set / Map เสียหายหลังการเปลี่ยนแปลงสถานะ และ copy() ที่สร้างอัตโนมัติจะสร้างสำเนาที่แยกออกจากเอนทิตี้ที่ถูกจัดการ
kotlin-tooling-agp9-migration
kotlin
Android Gradle Plugin 9.0 ทำให้ปลั๊กอินแอปพลิเคชันและไลบรารีของ Android ไม่เข้ากันกับปลั๊กอิน Kotlin Multiplatform ในโมดูลเดียวกัน ทักษะนี้จะแนะนำคุณตลอดการย้ายข้อมูล
kotlin-tooling-cocoapods-spm-migration
kotlin
โยกย้ายโปรเจกต์ KMP จาก CocoaPods (kotlin("native.cocoapods")) ไปยัง Swift Package Manager (swiftPMDependencies DSL) — แทนที่ pod() ด้วย swiftPackage(),…
kotlin-tooling-immutable-collections-0-5-x-migration
kotlin
ย้ายโค้ด Kotlin (และ Java) จาก kotlinx.collections.immutable 0.3.x / 0.4.x ไปยังเวอร์ชันล่าสุด 0.5.x เส้นทาง 0.5.x เปลี่ยนชื่อทุกเมธอดที่คืนค่าสำเนาบน...
kotlin-tooling-java-to-kotlin
kotlin
แปลงไฟล์ต้นฉบับ Java เป็น Kotlin ที่เป็นธรรมชาติ โดยใช้วิธีการแปลงแบบ 4 ขั้นตอนที่มีระเบียบวินัย พร้อมตรวจสอบ 5 เงื่อนไขคงที่ในแต่ละขั้นตอน รองรับการแปลงที่คำนึงถึงเฟรมเวิร์ก ซึ่งจัดการเป้าหมายตำแหน่งของแอนโนเทชัน สำนวนของไลบรารี และการคงไว้ซึ่ง API
kotlin-tooling-native-build-performance
kotlin
วินิจฉัยและแก้ไขการคอมไพล์และการลิงก์ Kotlin/Native ที่ช้าในโปรเจกต์ Kotlin Multiplatform ที่กำหนดเป้าหมายเป็น iOS ใช้เมื่อผู้ใช้รายงานว่า iOS หรือ... ช้า
kotlin-spring-proxy-compatibility
kotlin
Diagnose and prevent Kotlin plus Spring proxy failures around `@Transactional`, `@Cacheable`, `@Async`, method security, retry, configuration proxies, and JPA…
ci-cd-containerization-advisor
kotlin
ออกแบบไปป์ไลน์การสร้าง อิมเมจ และการปรับใช้ที่ทำซ้ำได้สำหรับแอปพลิเคชัน Kotlin และ Spring รวมถึงการตรวจสอบ CI คอนเทนเนอร์แบบเลเยอร์ ความปลอดภัยในการเปิดตัว…