dd-monitors

작성자: datadog-labs

모니터 관리 - 목록, 검색, 파일 기반 생성 및 알림 모범 사례.

npx skills add https://github.com/datadog-labs/agent-skills --skill dd-monitors

Datadog Monitors

Create, manage, and maintain monitors for alerting.

Prerequisites

This requires pup in your path. See Setup Pup.

Command Execution Order (Token-Efficient)

For scoped commands, use this order:

  1. Check context first (prior outputs, conversation, saved values).
  2. If a required value is missing, run a discovery command first.
  3. If still ambiguous, ask the user to confirm.
  4. Then run the target command.
  5. Avoid speculative commands likely to fail.

Quick Start

pup auth login

Common Operations

List Monitors

pup monitors list
pup monitors list --tags "team:platform"

Get Monitor

pup monitors get <id>

Create Monitor

pup monitors create --file monitor.json

Silence Alerts (Downtime)

# No pup monitors mute/unmute commands.
# Use downtime payloads to silence monitor notifications.
pup downtime create --file downtime.json
pup downtime cancel <downtime_id>

Monitor Creation Best Practices

1. Avoid Alert Fatigue

RuleWhy
No flapping alertsUse last_Xm not last_1m
Meaningful thresholdsBased on SLOs, not guesses
Actionable alertsIf no action needed, don't alert
Include runbook@runbook-url in message
# WRONG - will flap constantly
query = "avg(last_1m):avg:system.cpu.user{*} > 50"  # ❌ Too sensitive

# CORRECT - stable alerting
query = "avg(last_5m):avg:system.cpu.user{env:prod} by {host} > 80"  # ✅ Reasonable window

2. Use Proper Scoping

# WRONG - alerts on everything
query = "avg(last_5m):avg:system.cpu.user{*} > 80"  # ❌ No scope

# CORRECT - scoped to what matters
query = "avg(last_5m):avg:system.cpu.user{env:prod,service:api} by {host} > 80"  # ✅

3. Set Recovery Thresholds

monitor = {
    "query": "avg(last_5m):avg:system.cpu.user{env:prod} > 80",
    "options": {
        "thresholds": {
            "critical": 80,
            "critical_recovery": 70,  # ✅ Prevents flapping
            "warning": 60,
            "warning_recovery": 50
        }
    }
}

4. Include Context in Messages

message = """
## High CPU Alert

Host: {{host.name}}
Current Value: {{value}}
Threshold: {{threshold}}

### Runbook
1. Check top processes: `ssh {{host.name}} 'top -bn1 | head -20'`
2. Check recent deploys
3. Scale if needed

@slack-ops @pagerduty-oncall
"""

NEVER Delete Monitors Directly

Use safe deletion workflow (same as dashboards):

def safe_mark_monitor_for_deletion(monitor_id: str, client) -> bool:
    """Mark monitor instead of deleting."""
    monitor = client.get_monitor(monitor_id)
    name = monitor.get("name", "")
    
    if "[MARKED FOR DELETION]" in name:
        print(f"Already marked: {name}")
        return False
    
    new_name = f"[MARKED FOR DELETION] {name}"
    client.update_monitor(monitor_id, {"name": new_name})
    print(f"✓ Marked: {new_name}")
    return True

Monitor Types

TypeUse Case
metric alertCPU, memory, custom metrics
query alertComplex metric queries
service checkAgent check status
event alertEvent stream patterns
log alertLog pattern matching
compositeCombine multiple monitors
apmAPM metrics

Audit Monitors

# Find monitors without owners
pup monitors list | jq '.[] | select(.tags | contains(["team:"]) | not) | {id, name}'

# Find noisy monitors (high alert count)
pup monitors list | jq 'sort_by(.overall_state_modified) | .[:10] | .[] | {id, name, status: .overall_state}'

Downtime vs Muting

UseWhen
DowntimeAny planned silence window
Monitor editQuery/threshold behavior changes
# Downtime (preferred)
pup downtime create --file downtime.json

Failure Handling

ProblemFix
Alert not firingCheck query returns data, thresholds
Too many alertsIncrease window, add recovery threshold
No data alertsCheck agent connectivity, metric exists
Auth errorpup auth refresh

References

datadog-labs의 다른 스킬

dd-audit-compliance-report
datadog-labs
Datadog Audit Trail에서 SOC 2 및 PCI DSS에 대한 감사자 준비 완료 규정 준수 증거를 생성합니다. 프레임워크 컨트롤을 특정 쿼리 패턴에 매핑하고 다음을 생성합니다…
experiment-analyzer
datadog-labs
LLM 실험 결과를 분석합니다. 단일 또는 비교 실험, 탐색적 또는 Q&A 모드를 처리합니다. 사용자가 "실험 분석", "비교…"라고 말할 때 사용하세요.
dd-logs
datadog-labs
로그 관리 - 검색, 파이프라인, 아카이브 및 비용 제어.
dd-monitors
datadog-labs
모니터 관리 - 생성, 업데이트, 음소거 및 알림 모범 사례.
dd-audit-cost-spike-investigation
datadog-labs
Datadog 제품 사용량 또는 비용 급증을 조사하기 위해 사용량 측정 데이터(언제/무엇이 급증했는지)를 감사 추적 구성 변경 사항(누가 무엇을 변경했는지)과 연관시킵니다.
dd-audit-key-compromise
datadog-labs
잠재적으로 유출된 Datadog API 키를 조사합니다 — 작업 타임라인, 지리/IP 분석, 호출된 엔드포인트, 이상 징후 플래그 및 복구 단계.
eval-trace-rca
datadog-labs
프로덕션 LLM 트레이스에서 평가 판정 결과를 신호로 사용하여 근본 원인을 분석합니다. 사용자 애플리케이션이 실패하는 이유를 진단합니다. 사용자가 "eval…"이라고 말할 때 사용하세요.
dd-account-setup
datadog-labs
사용자가 Datadog 설정 또는 계측을 시작하기 전에 올바른 리전에서 유효한 DD_API_KEY로 인증된 Datadog 계정을 보유하고 있는지 확인합니다. 기존…