exploring-the-wizard
Run, drive, and explore the PostHog wizard headlessly against an app — boot it on the app and decide each screen yourself over the wizard-ci MCP tools…
npx skills add https://github.com/posthog/wizard --skill exploring-the-wizardExploring the wizard as an agent
Drive a real wizard run yourself: boot it on an app, read each screen, decide, act, snapshot.
The MCP server
Everything goes through the wizard-ci MCP server, registered in this
repo's .mcp.json (npx tsx scripts/wizard-ci-mcp.no-jest.ts) — it runs the
wizard from this checkout's source, so whatever branch you're on is what
you're testing. If the tools (open_app, read_state, …) aren't available,
the server isn't approved yet — ask the user to approve wizard-ci, then
retry. For how the harness works underneath, read
e2e-harness/ARCHITECTURE.md.
What it can do
- Boot the real TUI on any app directory and walk every pre-auth screen: framework detection, gatherContext, feature discovery, warehouse scan, intro, setup questions, health check, auth.
- Run the real integration (
run_agent): OAuth-equivalent bootstrap, the full agent run, outro, MCP install, skills — creating real PostHog resources (dashboard + insights) in the target project. - Show you every screen (
render_screen) exactly as a user sees it. - What it can't do: drive two apps at once (
open_appreplaces the active wizard), or advanceauth/runwithout credentials.
Credentials
Two modes — pick by how far the run must go:
- Detection-only (no credentials).
open_appneeds just{ appDir, projectId }. Everything up to and including theauthscreen runs credential-free — enough to regression-test detection, setup questions, and screen flow. End the run atauth. - Full run (credentials required) — anything past
auth, i.e.run_agent. Prompt the user for three things before starting:- Key — ask "What's the path to your phx key file?" and pass it as
keyFile(preferred: keeps the key out of logs). If they hold the key in an environment variable instead, have them write it to a file first (printenv THEIR_VAR > /tmp/phx-key— you never echo it) or passapiKeyinline as a last resort. Never print or commit the key. - Project id — the PostHog project the key is scoped to (
projectId). - Region —
us(default) oreu(region).
- Key — ask "What's the path to your phx key file?" and pass it as
Always copy the target app to a throwaway /tmp copy (never a real
fixture) — the run edits files.
How to drive it
open_app({ appDir, projectId, keyFile?, region? })— boots a live wizard on the app and returns the first screen.read_state— current screen, run phase, secret-free session, tasks, and the actions legal right now. Call after every move.perform_action({ action, params? })— commit a decision:confirm_setup,dismiss_outage,choose(a setup question, e.g.{ key, value }),set_mcp_outcome,dismiss_slack,keep_skills.render_screen— render the current TUI to ANSI so you can see it.run_agent— kicks off the real integration in the background and returns immediately; it bootstraps credentials, so it's what advancesauthandrun. Then pollread_state—runPhasegoesrunning → completedand the screen advances tooutro.
A typical full walk:
open_app → intro → perform_action confirm_setup
read_state → health-check → perform_action dismiss_outage
read_state → auth → run_agent (returns at once; integration runs in background)
read_state (poll) → runPhase running → completed, screen → outro
outro → perform_action dismiss_outro → … → keep_skills
A detection-only walk (no key):
open_app → read_state (poll until detectionComplete) →
check integration / detectedFrameworkLabel → confirm_setup → auth → done, next app
Snapshot with render_screen at each key moment and save each frame to a
numbered file — /tmp/wz-explore-snaps/NN-<screen>.txt, incrementing NN in
visit order — so the run leaves a readable, ordered record you and the user
can review afterward (the same shape the CI route's .txt frames take).
Capture the run screen as it progresses, not just on screen changes.
Sweeping the workbench
The fixture library lives at
wizard-workbench/apps/basic-integration/<framework>/<app> (sibling repo).
To regression-test detection across every framework, loop the detection-only
walk over each app. Learned the hard way:
- Copy with
rsync -a --exclude node_modules --exclude .git(plus vendor/venv/Pods/build/dist). A plaincp -Rof the workbench fills the disk, and a copy that dies mid-write leaves a truncated app that detects asnull— a false regression. If detection returnsnullunexpectedly, check the fixture (package.jsonpresent?) before blaming the code. open_appreturns the first paint, sometimes before detection lands (detectionComplete: false,integration: null). Pollread_state— slower detectors (laravel, rails) need a beat.- Logs: every run appends to
/tmp/posthog-wizard.log. Recordwc -c < /tmp/posthog-wizard.logbefore the sweep and slice withtail -c +OFFSETafter — that's the run's own log, greppable for detection lines and[bounded-fs]cap warnings. - Router modes and other
gatherContextresults don't appear inread_stateor the log. An emptysetupQuestionsimplies the mode resolved (ambiguity would raise a question), but for positive proof call the util directly withnpx tsxagainst the same fixture (e.g.getNextJsRouter,getTanStackRouterMode). - Delete fixtures as you go — a full sweep is multiple GB.
Key facts
- State → screen. You never navigate; you commit a decision (an action) and the router re-derives the active screen. Name actions, not keys.
authandrunadvance only viarun_agent. They expose no action and don't self-advance.run_agentreturns immediately and runs the integration in the background — pollread_stateforrunPhase(running → completed). Everything else is an instant commit.run_agentcreates real PostHog resources (a dashboard + insights) in the project; each run duplicates them.- A green run ≠ a valid integration.
runPhase=completedmeans the flow finished, not that the wizard understood the framework (e.g. it'll treat a Wasp app as react-router). Read what it actually changed. - None of this ships. The harness lives in
e2e-harness/, out ofsrc/.