cron-watchdog-debug

por vercel

Depuración de cron y watchdog para vercel-openclaw: autenticación de cron de Vercel, trabajos persistentes de OpenClaw, claves de activación de cron, actualización de tokens, oráculo de restauración, respaldo en caliente y…

npx skills add https://github.com/vercel-labs/vercel-openclaw --skill cron-watchdog-debug

Cron Watchdog Debug

Use this skill for cron wake, watchdog, scheduled OpenClaw job, and restore-oracle incidents.

Non-Negotiables

Before proposing a fix, produce:

  1. Deployment-state proof.
  2. Watchdog path diagram with every edge marked unknown, verified-good, or verified-bad.
  3. Hypothesis table with fastest falsifier and status.
  4. Cron Watchdog Handoff.

Do not treat "cron route ran" as proof that an OpenClaw scheduled job ran. Keep these states separate: Vercel Cron invoked, watchdog authorized, persisted wake due, sandbox woke, AI Gateway token refreshed, OpenClaw cron scheduler loaded jobs, user-visible delivery happened.

Evidence First

For live deployments, create one artifact root:

RUN_TS="$(date -u +%Y%m%dT%H%M%SZ)"
ART=".agent-runs/cron-debug/$RUN_TS"
mkdir -p "$ART"/{admin,vercel,sandbox,workflow}

Collect before editing code:

  • .vercel/project.json versus the intended project/team/deployment.
  • git rev-parse HEAD and git ls-remote origin main.
  • Live GET /api/status.
  • Live GET /api/admin/sandbox-diag.
  • Live GET /api/admin/logs filtered for watchdog., sandbox., gateway., cron, and restore.
  • Live GET /api/admin/launch-verify if the incident involves launch readiness.
  • Live GET or POST /api/cron/watchdog only with valid cron authorization.

If a release env file is supplied, source it only in the shell running live commands. Never print or save raw secrets.

Watchdog Path Diagram

Use this shape and mark each edge:

Vercel Cron tick -> /api/cron/watchdog auth -> runSandboxWatchdog
  -> deployment contract -> metadata read -> busy/stale/running/stopped branch
  -> cron wake key read -> ensureSandboxReady? -> token refresh
  -> OpenClaw cron jobs present in sandbox -> scheduled job executes -> delivery visible

For a running sandbox due soon, include the parallel branch:

running sandbox -> gateway probe -> cronNextWakeMs <= now + 60s
  -> force AI Gateway token refresh -> restore oracle cycle -> no duplicate wake

Runtime Probes

Read admin surfaces first. Then, only if needed, inspect the sandbox with npx sandbox or the app admin SSH/exec fallback.

Read-only sandbox checks:

set -eu
echo "== process =="
ps -eo pid,ppid,comm,args | grep -E "[o]penclaw|[n]ode" || true
echo "== ports =="
(ss -ltnp || netstat -ltnp) 2>/dev/null | grep -E ":(3000|8787)\\b" || true
echo "== cron files =="
ls -la /home/vercel-sandbox/.openclaw/cron 2>/dev/null || true
echo "== cron jobs shape =="
node - <<'NODE'
const fs = require('fs');
for (const p of [
  '/home/vercel-sandbox/.openclaw/cron/jobs.json',
  '/home/vercel-sandbox/.openclaw/cron/jobs-state.json',
]) {
  try {
    const j = JSON.parse(fs.readFileSync(p, 'utf8'));
    console.log(p, JSON.stringify({
      version: j.version ?? null,
      jobCount: Array.isArray(j.jobs) ? j.jobs.length : Object.keys(j.jobs || {}).length,
      jobIds: Array.isArray(j.jobs) ? j.jobs.map((x) => x.id) : Object.keys(j.jobs || {}),
    }));
  } catch (err) {
    console.log(p, 'unreadable', err.message);
  }
}
NODE

Do not print job payload text if it may contain user data. Prefer shape, counts, ids, and next-run timestamps.

Store Evidence

Cron wake depends on store keys written through the lifecycle layer. Do not hardcode Redis prefixes in app code. For debugging, use app/admin surfaces when available; if direct store inspection is unavoidable, record key names and redacted shapes only.

High-signal facts:

  • cronNextWakeMs exists and is due, future, missing, or malformed.
  • cronJobsJson structured record exists, has a SHA-256, job count, and source.
  • lastRestoreMetrics.cronRestoreOutcome is present or missing after wake.
  • watchdog report contains cron.wake, token.refresh, probe, and restore.prepare checks.

Hypotheses To Compare

Keep all rows visible as they are ruled out:

HypothesisEvidence forEvidence againstFastest falsifierStatus
Vercel Cron did not invoke the routeVercel logs for /api/cron/watchdog
Cron auth rejected the requestroute response/log status 401 vs report
Wake key was never persistedstore/admin evidence for cronNextWakeMs
Wake key is future or stalecompare cronNextWakeMs to current UTC
Watchdog intentionally skipped idle sandboxcron.wake check message and metadata status
Sandbox woke but token refresh failedtoken.refresh check and AI Gateway token metadata
Jobs file missing or malformed in sandboxsanitized cron file shape from sandbox
OpenClaw scheduler loaded jobs but delivery failedOpenClaw logs plus channel lastForward/user-visible evidence
Restore oracle or hot spare side effect changed cron staterestore.prepare check and restore metrics

Safe Fix Boundaries

  • src/app/api/cron/watchdog/route.ts owns cron-route auth and response shape.
  • src/server/watchdog/run.ts owns watchdog decision order and report checks.
  • src/shared/watchdog.ts owns report types.
  • src/server/sandbox/lifecycle.ts owns cron persistence keys, stop/touch persistence, and wake behavior.
  • src/server/sandbox/cron-persistence.test.ts and src/server/watchdog/run.test.ts own regression coverage.
  • lat.md/sandbox-lifecycle.md owns the architecture narrative for cron wake.

Do not edit channel routes while debugging cron unless evidence proves the scheduled job reached channel delivery and failed there. Switch to the relevant channel skill for that phase.

Verification

Pick the smallest verification that covers the changed edge:

node scripts/verify.mjs --steps=test,typecheck
node --test src/server/watchdog/run.test.ts
node --test src/server/sandbox/cron-persistence.test.ts
lat check

For live fixes, include before/after evidence from /api/cron/watchdog, /api/admin/logs, and sandbox cron file shape.

Handoff

Use references/handoff-template.md for incident reports.

Más skills de vercel

benchmark-sandbox
vercel
Ejecuta escenarios de evaluación de vercel-plugin en Vercel Sandboxes en lugar de paneles locales de WezTerm. Aprovisiona microVMs efímeras con Claude Code y el plugin preinstalado,…
official
emil-design-eng
vercel
Esta habilidad codifica la filosofía de Emil Kowalski sobre el pulido de la interfaz de usuario, el diseño de componentes, las decisiones de animación y los detalles invisibles que hacen que el software se sienta genial.
official
vercel-react-best-practices
vercel
Directrices de optimización de rendimiento para React y Next.js de Vercel Engineering. Esta habilidad debe usarse al escribir, revisar o refactorizar React/Next.js…
official
vercel-react-best-practices
vercel
Directrices de optimización de rendimiento para React y Next.js de Vercel Engineering. Esta habilidad debe usarse al escribir, revisar o refactorizar React/Next.js…
official
write-guide
vercel
Producir una guía técnica que enseñe un caso de uso del mundo real mediante ejemplos progresivos. Los conceptos se introducen solo cuando el lector los necesita.
official
release
vercel
Liberar vercel-plugin — ejecutar gates, aumentar versión, generar artefactos, commit y push. Usar cuando se pida "release", "ship", "bump and push" o "cut a release".
official
deepsec
vercel
Ejecutar DeepSec contra un checkout de proyecto de Vercel desde dev3000. Usar para configuración de DeepSec con un clic, arranque de contexto de proyecto, procesamiento limitado de primera pasada y…
official
backport-pr
vercel
Hacer backport de un pull request fusionado de Next.js desde canary a una rama de versión anterior, como next-16-2. Usar cuando el usuario solicite hacer backport, cherry-pick o abrir un…
official