observability-integrator

작성자: kotlin

Kotlin과 Spring 서비스에 대해 로그, 메트릭, 트레이싱, 헬스 엔드포인트 전반에 걸쳐 실행 가능한 관찰성을 설계합니다. 서비스를 계측할 때 사용합니다.

npx skills add https://github.com/kotlin/kotlin-backend-agent-skills --skill observability-integrator

Observability Integrator

Source mapping: Tier 2 high-value skill derived from Kotlin_Spring_Developer_Pipeline.md (SK-17).

Mission

Make the service explain its own behavior under normal load and under failure. Instrument for operational questions and decisions, not for vanity dashboards.

Read First

  • Business-critical user journeys and SLO or SLA targets.
  • Existing metrics, logs, traces, and actuator exposure.
  • Platform stack: Prometheus, Grafana, OpenTelemetry, ELK, Loki, vendor APM, or mixed.
  • Service topology, downstream dependencies, async boundaries, and coroutine usage.
  • Existing runbooks or alerting gaps.

Design Sequence

  1. Define the most important operational questions:
    • is the service healthy
    • which journey is slow
    • which dependency is failing
    • where saturation is growing
  2. Instrument the critical path before the nice-to-have path.
  3. Add metrics, traces, and logs that answer those questions together.
  4. Add health indicators and alerts with clear actionability.
  5. Re-check cardinality, PII, and exposure risk.

Metrics Rules

  • Prefer metrics tied to user journeys, dependencies, pool saturation, queue depth, and retry behavior.
  • Choose labels with a cardinality budget in mind.
  • Favor histogram or timer metrics where latency distributions matter.
  • Separate success, client error, server error, and dependency failure semantics clearly.
  • Include outcome metrics for retries, circuit breakers, cache behavior, and scheduler work when they affect operations.

Logging Rules

  • Use structured logs with stable field names.
  • Include correlation identifiers such as trace id or request id when the platform supports them consistently.
  • Log business-relevant events at service boundaries and failure points, not every line of code.
  • Redact or avoid PII, secrets, tokens, and credentials by policy, not by luck.
  • Prefer stable event names or codes over prose-only log statements when operations depend on searchability.

Tracing And Health Rules

  • Propagate trace context across HTTP, messaging, async execution, and coroutine boundaries.
  • Sample traces deliberately. Full sampling is not always affordable or necessary.
  • Distinguish liveness, readiness, and startup health semantics.
  • Keep actuator exposure minimal and authenticated where needed.
  • Include dependency health only when the signal is actionable and does not create cascading false alarms.

Advanced Observability Traps

  • High-cardinality labels can make a metric unusable and expensive at the same time.
  • Trace propagation may silently fail across executors, coroutines, or listeners even when HTTP tracing looks fine.
  • Logs without stable correlation are often worse than fewer logs with consistent context.
  • A metric that never drives an alert, dashboard, or investigation path is probably noise.
  • Readiness that depends on every optional downstream can create self-inflicted outages.
  • Health endpoints that expose secrets or internal topology are security risks, not observability wins.
  • Poor sampling decisions can hide the exact slow or failing traces operators care about.

SLO And Cost Nuances

  • RED and USE perspectives complement each other. User-facing latency and error metrics do not replace resource saturation visibility, and vice versa.
  • Burn-rate alerting is often more actionable than static threshold alerting for SLO-backed services.
  • Histogram bucket choice affects both storage cost and usefulness. Buckets should reflect user-facing latency objectives, not library defaults.
  • Exemplars or trace links can shorten incident diagnosis dramatically when supported by the platform.
  • Metric names and labels become quasi-APIs for operators. Renaming them casually creates observability drift across dashboards and alerts.
  • Observability cost is part of the design. Sampling, retention, and cardinality are architectural choices, not cleanup work for later.

Expert Heuristics

  • Instrument the path that paged someone last time before instrumenting the path that is merely interesting.
  • Prefer a smaller set of trusted dashboards and alerts over a broad telemetry surface nobody uses.
  • If correlation breaks across async boundaries, fix that before adding more log lines.
  • Good observability makes rollback, mitigation, and capacity decisions faster. Favor signals that support those decisions directly.

Output Contract

Return these sections:

  • Operational questions: what the instrumentation must answer.
  • Metrics plan: the key metrics and label strategy.
  • Logging plan: structure, correlation, and redaction rules.
  • Tracing plan: propagation points and sampling guidance.
  • Health and alerting: readiness, liveness, startup, and actionable alerts.
  • Minimal implementation plan: the smallest set of instrumentation changes that materially improves operability.

Guardrails

  • Do not instrument everything.
  • Do not expose actuator or debug endpoints casually.
  • Do not emit sensitive data in logs or traces.
  • Do not add cardinality-heavy labels such as raw user ids, full URLs, or free-form exception messages.
  • Do not create alerts with no obvious operator action.

Quality Bar

A good run of this skill gives operators clear signals, low-noise alerts, and fast incident localization. A bad run produces a large telemetry bill, noisy dashboards, and no practical improvement in diagnosis.

kotlin의 다른 스킬

kotlin-backend-jpa-entity-mapping
kotlin
Kotlin의 data class는 DTO에 자연스럽지만 JPA 엔티티에는 위험합니다. Hibernate는 data class가 깨뜨리는 identity 의미론에 의존합니다. 모든 필드에 대한 equals/hashCode는 상태 변경 후 Set/Map 멤버십을 손상시키고, 자동 생성된 copy()는 관리되는 엔티티의 분리된 복제본을 만듭니다.
kotlin-tooling-agp9-migration
kotlin
Android Gradle Plugin 9.0은 동일한 모듈에서 Android 애플리케이션 및 라이브러리 플러그인을 Kotlin Multiplatform 플러그인과 호환되지 않게 만듭니다. 이 스킬은 마이그레이션 과정을 안내합니다.
kotlin-tooling-cocoapods-spm-migration
kotlin
KMP 프로젝트를 CocoaPods(kotlin("native.cocoapods"))에서 Swift Package Manager(swiftPMDependencies DSL)로 마이그레이션 — pod()를 swiftPackage()로 대체,…
kotlin-tooling-immutable-collections-0-5-x-migration
kotlin
Kotlin(및 Java) 코드를 kotlinx.collections.immutable 0.3.x / 0.4.x에서 최신 0.5.x로 마이그레이션합니다. 0.5.x 라인은 모든 복사본을 반환하는 메서드의 이름을 변경합니다…
kotlin-tooling-java-to-kotlin
kotlin
Java 소스 파일을 체계적인 4단계 변환 방법론을 사용하여 관용적인 Kotlin으로 변환하며, 각 단계에서 5가지 불변 조건을 확인합니다. 애노테이션 사이트 대상, 라이브러리 관용구, API 보존을 처리하는 프레임워크 인식 변환을 지원합니다.
kotlin-tooling-native-build-performance
kotlin
Kotlin Multiplatform 프로젝트에서 iOS를 대상으로 할 때 느린 Kotlin/Native 컴파일 및 링크를 진단하고 수정합니다. 사용자가 느린 iOS 또는…을 보고할 때 사용하세요.
kotlin-spring-proxy-compatibility
kotlin
Diagnose and prevent Kotlin plus Spring proxy failures around `@Transactional`, `@Cacheable`, `@Async`, method security, retry, configuration proxies, and JPA…
ci-cd-containerization-advisor
kotlin
재현 가능한 빌드, 이미지 및 배포 파이프라인을 설계합니다. Kotlin 및 Spring 애플리케이션을 대상으로 하며, CI 검증, 계층형 컨테이너, 롤아웃 안전성 등을 포함합니다.