seed-test-data

bởi langfuse

Khởi tạo dữ liệu kiểm thử Langfuse cục bộ bằng một lệnh duy nhất: cây quan sát lớn/phân nhánh (sự kiện v3 và v4), phiên dài, dấu vết hàng loạt để kiểm tra hiệu suất danh sách. Dùng…

npx skills add https://github.com/langfuse/langfuse --skill seed-test-data

Seed Test Data

One-shot deterministic test data for local Langfuse. The CLI handles env loading, preflight checks, ClickHouse/Postgres writes, readback verification, and prints UI deep links plus a machine-readable JSON summary (last stdout line).

If anything fails, run doctor first

pnpm run seed -- doctor

Prints PASS/WARN/FAIL per dependency (Postgres, migrations, project, ClickHouse, v4 dev tables, Redis, MinIO, web app) with the exact fix command for every failure. Do not debug Docker/ClickHouse manually before running this.

Need → command

I need...Command
A very complex observation tree (v3)pnpm run seed -- trace-tree --observations 5000 --depth 12 --breadth 500
The same tree readable in the v4 events UIadd --v4 (writes events_full; events_core fills via MV)
Async parents whose subtree outlives their own span (subtree wall-clock duration badge)add --async-parents to trace-tree (root + hub end immediately while children keep running)
A realistic agent flow over a timeline (graph view + scrubbable timeline)pnpm run seed -- agent-timeline --turns 6 --v4 (LangGraph refine loop planner→retriever→generator→critic→loop, staggered in time; add --timing-only for the pure timing fallback)
A trace that is large as a GRAPH (many distinct node names + connections, the trace-graph layout stress)pnpm run seed -- agent-graph --v4 (~1,350 distinct connections from 350 observations; --nodes 120 --steps 100 --parallel 8 crosses the layout ceiling, --nodes 80 --steps 30 --parallel 4 is small-but-dense)
A dozen SMALL traces, each a different timeline shape (the everyday case, not a stress test)pnpm run seed -- timeline-shapes --v4 (12 hand-timed traces of 4-25 observations: rag answer, streamed chat with a TTFT split, 8-way fan-out, retry backoff with widening gaps, a 13-minute wait on a human, one slow tool dwarfing everything, an error cascade with failover, in-flight spans with no end time, zero-duration checkpoints, a ten-level ladder, 24 flat siblings, a three-turn agent loop with think time; --shape <slug> for one)
A row carrying every annotation at once, to judge whether the timeline is overloadedpnpm run seed -- timeline-annotated --v4 (ONE trace, 12 observations: one/two/four scores — the last collapsing into +1 — one and twelve comments, a streaming first-token mark, costs and durations spread so the heat map paints some rows, plain rows beside them for contrast, and all three label placements on screen together)
A multi-HOUR agent run (wall clock dominated by waiting, as production ones are)pnpm run seed -- agent-timeline --turns 120 --turn-gap-ms 60000 --v4 (~2.5h; the gap is jittered up to 2x, so the work lands at a few percent of the trace instead of packing every span into a few seconds)
A demo-grade, real-looking agent trace (videos, screenshots, docs)pnpm run seed -- support-agent --v4 --id-prefix <hex> (one fixed, fully handcrafted support-copilot refund run: guardrails, parallel context fan-out, 3-turn ReAct loop with real payloads/costs; deterministic — reseed with a FRESH prefix for a clean take; the prefix is the trace id, so a hex prefix reads like production)
A demo-grade, real-looking multi-user SESSION (session timeline videos, screenshots, docs)pnpm run seed -- incident-session --id-prefix inc-4471-checkout-latency (one fixed, handcrafted incident channel: 7 turns, 4 named users, accumulating chat history with reasoning and markdown, tool calls paired to TOOL rows plus two untraced ones, a WARNING and an ERROR tool, a nested sub-agent, a refusal, typed session scores and comments; v4 on by default; the prefix is the session id)
A plain trace with no agentic types (collapsed-by-default graph panel)add --plain to trace-tree (SPAN/GENERATION/EVENT only)
A trace with MORE observations than the detail view loads (v4 caps the tree at 10k, startTime ASC — the chronological tail is missing)pnpm run seed -- trace-tree --observations 12000 --stride-ms 10 --v4 (--stride-ms starts each observation index × N ms in, so start times are unique and strictly increasing: observation index < 10000 loads, -obs-10000 and up fall past the cap. Without it thousands of rows share one millisecond and the boundary is arbitrary. Re-runs that CHANGE timing flags need a fresh --id-prefix: start_time is an events ORDER BY key, so both versions persist in the same trace)
An extremely DEEP single-chain trace (tree depth = observation count; layout stress)pnpm run seed -- deep-chain --v4 (1401 sequential generations, each the sole child of the previous — the mis-parented-instrumentation shape from LFE-10959 that collapses tree/timeline layouts; --observations N to change depth)
A super tough session (v3 legacy session view)pnpm run seed -- long-session --traces 300 --observations-per-trace 8
Diverse v4 session shapes (chat / coding-agent / mixed / media) for the session-detail viewpnpm run seed -- session-shapes --shape all (the agent shape has I/O on AGENT/TOOL with no GENERATION — pre-LFE-10520 the "first generation" default rendered empty cards for it; the current "All observations with I/O" default renders it correctly; v4 on by default)
Many sessions for the sessions TABLE and its filters / search bar (not the detail view)pnpm run seed -- session-variety --sessions 120 --days 14 (searchable topic ids, 1-3 userIds and 1-3 tags per session, four environments, session metadata tier/region/channel/topic, numeric + categorical + boolean session scores, comments on a quarter; v4 on by default. Widen the time-range picker past the default 1 day.)
Session messages carrying Langfuse media references (inline image, several refs in one message, bare refs in a content array, link-only payload)pnpm run seed -- session-shapes --shape media — uploads the image/audio/pdf fixtures to MinIO and links them to the observation, so the inline chip and the "Media" strip both resolve (LFE-14815, LFE-9577)
Many traces for list/filter performancepnpm run seed -- many-traces --count 100000 --days 14
Long-window v4 traffic with cost/latency/token OUTLIERS (outlier chart strip, LFE-14451)pnpm run seed -- outlier-traffic --days 90 (diurnal base load + deterministic spikes + hour-long ×8-latency incidents; root AGENT + GENERATION carrying cost + TOOL per trace; v4 on by default)
Scores with spaces in the name (filter/grammar testing)pnpm run seed -- scored-traces --traces 24 --v4
Custom model definitions reachable from a trace (price editor entry points)pnpm run seed -- custom-models --v4 (a tiered model with a condition-gated second tier and one usage type priced at 0, a single-tier model, and a generation whose model matches no definition so its badge opens the create dialog)
Many project-owned evaluators for gallery infinite-scroll testingNEXTAUTH_URL=https://pr-<N>.preview.langfuse.com pnpm run seed -- evaluator-gallery --count 200 (uses the default synthetic seed API key and reconciles deterministic evaluator names; also works against local seeded environments when NEXTAUTH_URL is omitted)
Lots of scores on every node (dense score badges, tree-row overflow testing)add --scores-per-node 12 to trace-tree (N distinct scores per observation; try --depth 2 --breadth 44 for many tall sibling rows)
Extra trace tags, incl. mixed case/accents (tag filter ordering)add --tags "Zebra,apple,Ärger" to trace-tree (comma-separated, appended to the scenario's own tags)
Varied human-annotation queues (annotate UI / keyboard testing)pnpm run seed -- annotation-queue --core-items 12 (creates a "core types" queue covering every score-field render path + an "edge cases" queue with archived/stale/partial scores and observation/session/deleted/completed items)
Huge/malformed/unicode payloadspnpm run seed -- trace-tree --payload-bytes 1000000 --payload-style malformed (styles: json, text, malformed, unicode, bignum, base64)
Big integers beyond 2^53-1 (number-precision testing)pnpm run seed -- trace-tree --observations 1 --payload-style bignum
Huge base64 data-URI in ChatML IO (multimodal crash shape, LFE-10152)pnpm run seed -- trace-tree --observations 30 --payload-bytes 20000000 --payload-style base64 --v4 (one unbroken multi-MB base64 token in trace + root-observation IO; max 50 MB)
See all scenarios and flagspnpm run seed -- list --json
Predict without writingadd --dry-run

For a v4 experiment with chat messages and nested JSON input/output, run pnpm run seed -- experiment-io. It creates one dataset, one experiment, and three items, then prints the experiment results link. Set NEXTAUTH_URL to your local app URL when using a port other than 3000.

Contract

  • Last stdout line is a JSON summary: traceIds, sessionIds, counts, verified (ClickHouse readback), links (UI deep links). Use --json to suppress progress logs. Non-zero exit = data did not land; the error includes a fix: line.
  • Deterministic: same --seed (default 42) and flags → same ids (ids never contain dates), with timestamps anchored to the current UTC day. Re-running within the same day overwrites in place; a later-day re-run updates the same ids with re-anchored timestamps (the previous day's rows persist under their old dates until then). Independent copies come only from --id-prefix.
  • Default project is the seeded 7a88fb47-b4e2-43b8-a06c-a5ce950dc53a (login demo@langfuse.com / password); override with --project.
  • Open the printed links in the browser to verify visually. The v4 events-backed UI is the per-user "Fast (Preview)" sidebar toggle, or LANGFUSE_MIGRATION_V4_WRITE_MODE=events_only server-side.

Extending

Add a scenario in packages/shared/scripts/seeder/scenarios/: a plain function using the deterministic Rng (never Math.random), register it in scenarios/index.ts, and update the table in packages/shared/scripts/seeder/AGENTS.md and this skill. Scenario names, flags, and JSON keys are additive-only contracts. Design rationale: packages/shared/scripts/seeder/README.md.

Thêm skills từ langfuse

frontend-browser-review
langfuse
Sử dụng kỹ năng này khi một thay đổi ảnh hưởng đến những gì người dùng nhìn thấy hoặc thực hiện trong trình duyệt.
frontend-large-feature-architecture
langfuse
Sử dụng khi xây dựng, thay đổi hoặc tái cấu trúc các tính năng frontend lớn của Langfuse, danh sách ảo hóa, bảng lớn, thành phần điều khiển, tính năng cục bộ…
skill-developer
langfuse
Tạo và quản lý các kỹ năng Claude Code theo các phương pháp tốt nhất của Anthropic. Sử dụng khi tạo kỹ năng mới, sửa đổi skill-rules.json, hiểu trình kích hoạt…
langfuse-prompt-migration
langfuse
Di chuyển các prompt được viết cứng sang Langfuse để quản lý phiên bản và lặp lại mà không cần triển khai. Sử dụng khi người dùng muốn ngoại hóa prompt, chuyển prompt sang Langfuse,…
incident-alert-tickets
langfuse
Read and, after human approval, update the Linear `incident-alert` knowledge base. Use before and after investigating a named Datadog monitor,…
refactor-react-effects
langfuse
Tái cấu trúc việc sử dụng React useEffect không cần thiết trong mã frontend Langfuse. Sử dụng khi thêm, xem xét hoặc xóa các effect; khởi tạo biểu mẫu hoặc trạng thái UI cục bộ từ...
sentry-instrumentation
langfuse
Decide whether and how errors report to Sentry. Use when touching capture or error-handling paths in `web/**`, triaging Sentry noise, or changing Sentry…
posthog-instrumentation
langfuse
Product analytics with posthog. Use when adding a meaningful user action or feature in `web/**`, touching PostHog capture code, or answering product-usage…