VibeView

Pilotez des appareils iOS, Android, Apple TV et Android TV en direct dans le cloud pour vérifier les modifications d'applications : lisez l'arborescence de l'interface, touchez, saisissez, appuyez sur la télécommande, vérifiez et lisez les journaux d'application depuis n'importe quel client MCP.

Documentation

Let your AI coding agent verify a change by driving a real, running copy of your app on a VibeView cloud device — installing the build, tapping through the UI, reading back what’s on screen, and reporting whether it actually worked. This is the same live-development loop described in Live Development, extended so an agent (not just you) can drive the session from the command line.

Install and authenticate

npm install -g vibeview
vibeview login

vibeview login opens your browser to complete authentication and saves a token locally. In CI or any non-interactive environment, set VIBEVIEW_API_TOKEN instead — every command accepts it as an alternative to being logged in.

Making it available to your agent

Installing the CLI doesn’t by itself tell your coding agent that VibeView exists. One command sets that up:

vibeview agent-setup

It does three things:

  1. Installs the agent skill into ~/.claude/skills/vibeview-agent/, where Claude Code discovers it automatically in every project. The skill describes the whole workflow — build requirements, the verify loop, cleanup — so the agent knows how to use VibeView, not just that it exists. Pass --project to install into the current project’s .claude/skills/ instead (useful when teammates should get it too, since that directory can be committed).
  2. Offers to register the MCP server with Claude Code (claude mcp add vibeview -- vibeview mcp), so the tools are visible in every session. Pass --mcp to register without asking, or --no-mcp to skip. Other MCP clients are configured per client — see Using it with an MCP client below.
  3. Suggests a line for your project’s CLAUDE.md (or AGENTS.md if you use other agents) telling the agent to verify UI changes on a live device. This is the strongest per-project nudge: the skill teaches the agent how, the instruction line tells it when.

Do you need both the skill and MCP? Usually not. If your agent can run shell commands (Claude Code in a terminal, for example), the skill alone is enough — it drives the CLI directly, and skipping MCP keeps tool definitions out of your agent’s context. Register the MCP server when the agent can’t shell out: Claude Desktop, Copilot agent mode, restricted environments where named tools are allowed but arbitrary commands aren’t. For those clients MCP isn’t an extra — it’s the only way in.

Agents other than Claude Code can be pointed at the skill file directly — it’s plain instructions any coding agent can follow — or given the MCP server via their own client configuration.

Other ways to install the skill

The same skill is published at github.com/vibeview/skills, so agents that read skills from GitHub can install it without the CLI’s setup command:

  • Any agent that uses the skills CLI (Claude Code, Cursor, Codex, Gemini CLI, GitHub Copilot and others):
    npx skills add vibeview/skills
    
  • Claude Code, as a plugin — installs the skill and registers the MCP server in one step:
    /plugin marketplace add vibeview/skills
    /plugin install vibeview@vibeview
    

Whichever route you use, the agent still needs the vibeview CLI installed and logged in (vibeview login), because the skill drives the CLI and the plugin’s MCP server runs vibeview mcp.

One honest note: even fully set up, an agent asked only to “implement X” won’t always verify on a device unprompted. The setup above raises the odds considerably, but saying “implement X and verify it on the device” is what makes it reliable.

The dev loop

  1. Upload a debug build of your app (only needed again after a native change — see Preparing Your Build):
    vibeview upload-app ./path/to/app-debug.apk
    
  2. Start Metro in your React Native project (yarn start or equivalent).
  3. Start a session in the background:
    vibeview dev --detach --json
    
    This prints a session_ready event with a session_id your agent uses to target every command that follows (--session <id>). session_ready fires once the session actually accepts commands — the device is up and the app is installed — so the first ui-tree can follow it immediately. The event also includes the URL of a page where you can watch and interact with the session — detached sessions never open a browser themselves, so open that URL yourself if you want to watch along. Its build field says which build was installed and why, and lists the app’s other debug builds (see below). The page shows the live device only to the account that started the session; a teammate who opens it is told whose session it is. To let someone else watch or drive it, turn on collaboration and send them its link.
  4. Drive the session with the command verbs below — always start with ui-tree to see what’s on screen and get element references (@e5, etc.) to act on. If the app is still on its splash or loading screen, the tree can come back empty or near-empty — that’s the app booting, not a broken session. Wait a few seconds and fetch it again.
  5. When you’re done, stop the session:
    vibeview dev-stop
    

Every command also works against a session you already started from the VibeView dashboard — pass its session ID with --session <id> instead of starting a new one with dev --detach. Stop such a session with vibeview stop <session-id>.

To run a specific upload rather than the app’s newest debug build, list the app’s builds and pin one — it must be a debug build (KIND debug):

vibeview list-builds <app-id>
vibeview dev --detach --json --app <app-id> --build <build-id>

With no --build <id>, vibeview dev runs the app’s newest debug build, unless the project’s own app version can be read (from the Expo config, including a dynamic app.config.ts when Expo is installed in the project, or from the native project) and an older debug build has that version while the newest does not; it then runs that one and says why. If the version can’t be read, it says so. When the session starts, it prints the build it installed (id, version, upload time and note), the newest debug build of each other version, and the project’s React Native version. If the app shows a red error screen right after it starts, the build may have been made from other native code than your project: run a build of your project’s version with --build <id>.

Watching along — and taking over

Every session has a live, fully interactive page (the url in session_ready, or page_url from the MCP dev_start tool). The bundled skill instructs agents to share it with you the moment a session starts, so you can watch the agent work in real time — and drive the device yourself whenever you want. The agent’s commands and your input coexist; keep it to a single browser tab per session.

That takeover is also how agents get unstuck. The skill tells them to stop and ask you whenever a screen needs something only a human should provide — signing in with real credentials, a 2FA code, a CAPTCHA, or a risky confirmation — rather than guessing. Do your part on the session page, tell the agent you’re done, and it re-reads the screen (verifying the blocker is actually gone) before continuing.

A detached session records its state in a .vibeview/ directory in your project root (that’s how later commands find it without --session). It’s local machine state, not something to commit: when your project has a .gitignore, the CLI adds .vibeview/* to it, plus !.vibeview/test-spec.json so a test spec in that folder is still committed. Don’t ignore the whole folder with .vibeview/: Git can’t re-include a file inside an ignored folder.

Commands

Every command below accepts --session <id> (targets a specific session; falls back to the current project’s dev --detach session if omitted) and --json (prints the raw result as one JSON line instead of human-readable text).

CommandPurpose
ui-treeFetch the current screen’s UI elements, each with a reference (@e5) to act on.
logsRead recent app logs from the device — JS console output, native errors, and crash messages.
screenshotCapture a screenshot of the current screen and save it to disk. Its pixels are the coordinates tap and drag take (see Coordinates).
tap <target>Tap an element by reference (e.g. @e5) or by coordinates (e.g. 100,200), in the space described under Coordinates.
long-press <ref>Long-press (touch and hold) an element.
swipe <direction>Swipe up, down, left, or right.
scroll <direction>Scroll the current view up, down, left, or right.
scroll-to <text>Scroll until an element matching that text is visible, stopping at the end of the list.
drag <from> <to>Drag between two points, each an element reference or raw coordinates.
alert <get|accept|dismiss>Inspect or respond to a system alert on iOS or Apple TV.
type <text>Type text into the currently focused input field.
clear-textClear the currently focused input field.
press <button>Press a device button or system gesture. iPhone/iPad: home, back, lock, siri, enter. Android phone/tablet: home, back, lock, enter. Apple TV: dpad_up, dpad_down, dpad_left, dpad_right, dpad_center, back, home, menu. Android TV: dpad_up, dpad_down, dpad_left, dpad_right, dpad_center, back, home, enter. Roku: up, down, left, right, select, back, enter, backspace, play, rewind, forward, replay, info (the dpad_* names also work). vibeview press --help lists them, and a refused press prints the list.
open-url <url>Open a deep link inside the app. On iOS and Apple TV only the app’s own URL scheme is accepted (see below).
relaunch-appClose the app under test and start it again. It is a cold start that keeps the app’s data, so it is the way to check that something survives a restart; the app’s log keeps streaming across it. Not available on Roku.
set-posture <closed|partial|open> or set-posture --angle <0-180>Fold or unfold a foldable device (form factor foldable in list-devices --models — the iPhone Duo, Android foldables such as the Pixel 9 Pro Fold) — to a preset, or to an exact hinge angle in degrees (pass one or the other, not both). Errors on a device without a hinge.
rotate [--degrees <90|180|270>]Rotate the device. A phone or tablet, and an Android foldable, switches between portrait and landscape (only 90 is accepted); the iPhone Duo turns a quarter turn clockwise, 270 turns it back a quarter and 180 turns it upside down. Not available on TV devices.
set-location --lat <n> --lon <n>Set the device’s simulated GPS position, in decimal degrees (west and south are negative), for testing location-based flows. The app reads it like a GPS fix. Phones, tablets and foldables; not available on TV devices.
waitWait a moment before the next action.
find <text>Find an element by its text, optionally constrained to be near/above/below another element.
tap-focused <ref>TV only: move focus to an element and select it in one step.
focus <ref>TV only: move focus to an element without selecting it.

After every action, the response tells you what changed on screen — a full element list, the specific differences, or a note that nothing changed — so an agent always knows the current state before deciding what to do next. That means a normal loop is ui-tree once at the start, then act, act, act: you don’t need to re-fetch the tree between steps.

A few details of that output:

  • The list of differences starts with changed field values (what you just typed or set), and ends with a count of any differences it left out.
  • While the iOS keyboard is up, it appears as one line at the end of the tree (keyboard: visible, with a reference to its return key) instead of one line per key; the app’s own elements stay at the top.
  • On Apple TV, Android TV and Roku, focus, tap-focused and press start with the element that has focus afterwards, for example focused: @e112 button "Camping".
  • When several elements match, find prefers one whose visible label contains the text, and says so when it matched an element’s id instead.
  • If the device returns no element tree at all, even when read again a moment later, ui-tree prints error and exits with code 1 rather than printing an empty tree; run it again.

A command that ran but did not do what it was asked prints error instead of ok and exits with code 1, like a refused one. That includes a scroll-to that never found its text, and a type or clear-text whose field does not hold the expected text afterwards; the reason is printed below the status line. A password field cannot be read back, so typing into one is not checked.

Coordinates

tap and drag take coordinates in the same space as the element rectangles in ui-tree and the results of find. On a phone, tablet or Android TV, screenshot saves its image at that size, so a point you read off a screenshot can be tapped as it is. Don’t scale coordinates yourself. On Android this space can be smaller than the device’s screen resolution (a Pixel 7 Pro’s screenshots are 1152x2496, not 1440x3120), and it always covers the whole screen, status and navigation bars included, even when the app’s own window doesn’t. Coordinates outside it are refused with an error that gives the valid range.

On Roku, the rectangles in ui-tree are in the channel’s own UI coordinates (1920x1080 for a full-HD channel), while a screenshot is the picture the Roku puts out, which can be smaller (1280x720 at 720p). A Roku is driven with the remote, not by coordinates, so compare the two by proportion: on a full-HD channel in a 1280x720 screenshot, divide a rectangle by 1.5. Rectangles can also reach past the screen’s edge, for the parts of a row that are scrolled out of view.

Buttons, Enter and deep links

press takes the buttons of the device you are on. A name the device doesn’t have is refused, and the error lists the valid ones for that device.

press enter presses Enter (Return on iOS), which submits the focused text field, such as a search box or the last field of a form. It works on phones, tablets, Android TV and Roku; on Apple TV, use dpad_center. On phones and tablets a line break in type does the same, so vibeview type $'hello\n' types hello and then submits.

open-url hands the link to the app under test only. On Android, a link the app has no screen for (a web page, or another app’s link) fails with an error that says so; it never opens a browser or another app. On iOS and Apple TV the system decides which app opens a link, so only links in the app’s own URL scheme (for example myapp://settings) are accepted: a web link (https://...) or a system link such as tel: is refused with an error instead of opening Safari or another app.

A press back that leaves the screen exactly as it was still prints ok, with a note that nothing changed: the screen may not respond to the system back gesture (a custom header, for example), so tap the app’s own back button instead.

Foldable devices

Any device with a hinge is a foldable — the iPhone Duo (iOS) and Android foldables such as the Pixel 9 Pro Fold. To run on one, start the session with its model name. A session that asks only for a phone or tablet (the default) never gets a foldable, even when a foldable is the only device free — it waits for a phone like any other busy start. vibeview list-devices --models lists the models you can start, and marks each foldable as foldable:

vibeview list-devices --models
vibeview dev --detach --json --model "iPhone Duo"
vibeview dev --platform android --detach --json --model "Pixel 9 Pro Fold"

They all work the same way: the session starts closed, on the smaller cover screen, and you fold it with a preset or an exact angle, and rotate it in any posture:

vibeview set-posture partial --session <id>
vibeview set-posture --angle 75 --session <id>
vibeview rotate --session <id>

Each response says where the device ended up. After set-posture that is the posture it settled in, the hinge angle, and the size of the screen now showing, as it is shown (width and height swap while the picture is turned sideways, so the open iPhone Duo is 2853 x 2007), and the lit screen, cover or inner. The device decides which screen to light, the way a real one does, and the response waits until it has settled.

The presets move the hinge to the same angles on every foldable — closed 0, partial 120, open 180 degrees — but each device names an exact angle by ranges the device itself defines, so the same angle can read differently. Trust the posture in the response over the angle you asked for. The iPhone Duo and, as an Android example, the Pixel 9 Pro Fold:

Hinge angleiPhone DuoPixel 9 Pro Fold
0closedclosed
1-30usually closed; can stay partialclosed
31-84usually closed; can stay partialpartial
85-149usually partial; can stay closedpartial
150-169usually partial; can stay closedopen
170-180openopen

On the iPhone Duo, which screen is lit between 1 and 169 degrees depends on how the hinge got there. Moved in one go from shut or from fully open, it switches at about 85 degrees in both directions — the “usually” column. Moved in small steps, it keeps the screen it already had: opening a little at a time can keep the cover screen on up to about 160 degrees, and closing a little at a time can keep the inner screen on until the device is shut. Only 0 (the cover) and 170 and up (the inner screen) are the same whatever came before; the device’s posture_ranges mark the other angles with the posture they can also read as. An Android foldable names an angle the same way whichever direction it came from, and uses its inner screen whenever it is not closed — half-open included. After rotate it is the new orientation (portrait, landscape_left, landscape_right or portrait_upside_down) and, on the Duo, how the screen’s contents are turned: an app that supports the new orientation turns its layout, while the Home screen and portrait-only apps turn with the device.

A Duo’s rotate response carries two orientation fields, and they answer different questions:

  • device_orientation is how the device itself is held. It moves a quarter turn with every rotate and starts at portrait.
  • screen_orientation is how the picture is turned on the screen that is lit, measured against that screen’s own upright edge. It is not a second reading of how the device is held, so the two fields can name different orientations even when the app turned with the device.

To check whether your app laid itself out for the new orientation, compare the width and height of the new UI tree’s root rather than either field. An Android foldable’s rotate response carries both fields too, with the posture it was rotated in. Phones and tablets report device_orientation and, as screen_orientation, how the screen actually came out: an app that supports only portrait stays portrait when the device turns, and the response says so (device held landscape_left (screen stayed portrait)) instead of suggesting the screen turned.

Folding switches between two screens of different sizes and rotating turns the picture, so every element reference from before is stale — the response carries the new screen, and ui-tree fetches it again. A device that is still folding refuses a rotate; try it again once the fold has finished.

Use the tree and the screenshot for different things

The UI tree describes structure: which elements exist, what text they carry, and where they sit. It carries no colour, contrast, or stacking information, so a screen can pass every check the tree can express and still be visibly broken — text rendered in a colour that vanishes into the background, or an element drawn underneath another one.

Use the tree to find elements and confirm structure. Take a screenshot whenever a change affects how something looks, and inspect the image itself.

The tree describes your app, not the device’s status and navigation bars. On Android, when the notification shade is pulled down over your app, the tree describes the shade instead, so its buttons can be found and tapped.

Two gaps to know about on iOS. The tree doesn’t show whether a checkbox is checked: the change does appear in the diff after you tap it (for example value "checkbox, unchecked" → "checkbox, checked"), and a screenshot shows it. And a scrollable view’s scroll bars are listed as slider elements ("Vertical scroll bar, 1 page"); they are not controls to act on, so scroll with swipe instead.

Over MCP, screenshot writes the file and returns its path; pass inline: true to receive the image itself. An agent that can’t open files on the machine running the server needs inline: true to see anything at all — it costs context, so it’s worth reserving for the checks where appearance actually matters. The inline image is kept under 1 MB so model clients accept it: a large screenshot comes back as a JPEG, at the same size as the saved PNG, so a point read off it is still the coordinate tap and drag take. With --json, the CLI prints where it wrote the file and its size, not the image data.

A worked example

Adding a heading to a screen and confirming it renders, with hot reload already running:

vibeview ui-tree                      # what's on screen now?
# ...edit your component in your editor; Fast Refresh pushes it in ~1s...
vibeview ui-tree                      # the new heading appears in the tree
vibeview screenshot --out ./check.png # ...and actually looks right
vibeview tap @e7                      # keep going: tap through the flow

No rebuild and no re-upload — a JavaScript or asset change reaches the device through Metro. You only rebuild when native dependencies change.

Element references go stale as soon as the screen changes. Using a stale one is refused and the action does not run; because nothing happened, the response reports the screen as unchanged rather than returning a fresh element list, so run ui-tree again to pick up live references before retrying.

A few commands are worth knowing about specifically:

  • scroll-to is the right way to reach something off-screen — it keeps scrolling until an element whose text or accessibility label matches is visible, and gives up when the content stops moving, instead of you guessing how many scroll calls are needed. It works out which way to scroll on its own for ordinary vertical lists; pass --direction for horizontal rows of items. On an iPhone Duo that is open or partly open, scroll-to is not available on the inner screen and returns an error saying so; use scroll or swipe and check the tree again.
  • drag is the only command for a precise two-point gesture — dragging a slider to a value, reordering a list by dragging an item, or panning a map. swipe and scroll only take a direction and can’t express those. Each end can be an element reference or coordinates, so you can mix them. Options let you slow the drag down for precision, press and hold before the drag starts (for reorder gestures), and hold at the destination before releasing.
  • alert handles system alerts such as permission prompts, which sit above your app and block it. Use alert get to read the message and the exact button labels, then alert accept or alert dismiss — optionally with --button to choose a specific label. On Apple TV the same commands cover tvOS prompts, including the “Open in …?” confirmation a deep link raises; the chosen button is answered with the remote, so nothing in your app is touched.
  • logs is how an agent finds out why something broke. It returns the app’s own output — JS console lines, native errors, crash messages — scoped to the app under test, not the whole device. When the app crashes or freezes, the screen alone can’t explain it; the log usually can. Each response ends with a cursor value: pass it back as --since to receive only lines that arrived after your previous read, so acting and then checking logs --since <cursor> shows exactly what that action logged. --tail caps how many lines come back (default 100, maximum 500, newest kept). On Roku the log is the channel’s console: its print output, the reason it exited, and the error plus backtrace when it stops on a runtime error.

For Apple TV and Android TV apps, start the session with --platform tvos or --platform androidtv and navigate with the d-pad via press (dpad_up, dpad_down, dpad_left, dpad_right, dpad_center). press home leaves the app on both; the next ui-tree describes what is on screen (the TV home screen, or a prompt sitting over it), not an empty tree.

Roku sessions start from vibeview dev --platform roku (with --detach if the agent runs the loop itself) and take the same verbs, with a few Roku-specific details. The remote’s own names are press up, down, left, right, select and back (the dpad_* names work as aliases), plus enter, backspace and the media keys play, rewind, forward, replay and info, and press home is refused, because the session is confined to your channel. Any other name is refused too, with the list above. type sends the text straight into a focused on-screen keyboard, and the text it typed is read back from the field so the result tells you whether it landed. open-url is not available. A channel moves by focus, not touch, so tap, long-press, swipe, scroll, drag and scroll-to are refused on Roku with a message naming what to use instead: press to move, focus to walk to an element, type for text. logs returns the channel’s console output — everything it prints, why it exited, and the error and backtrace after a runtime error — so a channel that fails to start is diagnosed from logs alongside screenshot and ui-tree.

Using it with an MCP client

The same commands are also available as MCP tools, so any MCP-compatible agent can call them directly instead of shelling out to the CLI. MCP is an open standard, so this works with any model and any client that speaks it — not just one vendor.

The server runs over stdio: the command is vibeview and the argument is mcp. Every client expresses that slightly differently.

Claude Code (command line):

claude mcp add vibeview -- vibeview mcp

Cursor — .cursor/mcp.json in your project (or the global one in ~/.cursor/):

{
  "mcpServers": {
    "vibeview": {
      "command": "vibeview",
      "args": ["mcp"]
    }
  }
}

Claude Desktop — claude_desktop_config.json, same mcpServers shape as above.

VS Code (GitHub Copilot agent mode) — .vscode/mcp.json:

{
  "servers": {
    "vibeview": {
      "command": "vibeview",
      "args": ["mcp"]
    }
  }
}

Anything else — point the client at command vibeview with args ["mcp"]. If the client can’t find the binary, use the absolute path from which vibeview (a global npm install usually resolves without it).

Authentication comes from the same place as the CLI: run vibeview login once, or set VIBEVIEW_API_TOKEN in the environment your client launches the server with. If your client supports per-server environment variables, add it there:

{
  "mcpServers": {
    "vibeview": {
      "command": "vibeview",
      "args": ["mcp"],
      "env": { "VIBEVIEW_API_TOKEN": "your-token" }
    }
  }
}

Tool names are the command names above with hyphens replaced by underscores (ui_tree, tap, scroll_to, long_press, and so on), and each takes its arguments by name plus an optional session_id.

Eight tools exist only over MCP, covering the steps a CLI user would do with ordinary shell commands, so an agent with no shell access can still run the whole loop:

ToolPurpose
upload_appUpload a build by path (a zipped iOS or tvOS simulator .app, an Android or Android TV .apk, or a Roku channel .zip) and get back the app ID to start a session with — the equivalent of vibeview upload-app.
list_appsList your organization’s apps with their IDs and platforms, optionally for one platform — the equivalent of vibeview list-apps.
dev_startStart a session that stays open for the lifetime of the connection — the equivalent of vibeview dev --detach. It returns once the app is up on the device, and says which device it got: model, OS version, category and, for a foldable, the posture and lit screen. It also says which build it installed (id, version, upload time and note), lists the app’s other debug builds, and names the React Native version the local project uses, so a red error screen right after start (such as “React Native version mismatch”) can be traced to a build made from other native code. With no build_id, it chooses the build the same way vibeview dev does: the newest debug build, unless the project’s own app version can be read and an older debug build has that version while the newest does not; it then runs that one and says why, and it says when the version can’t be read. If every device of the model is busy, it waits in the queue and reports its place and progress while it waits (to clients that ask for progress), and how long it waited. Pass model (e.g. "iPhone Duo") to run on an exact device model, and build_id to run a specific build instead. Pass standalone: true to run a build as it is, with no Metro: a release build, or a CI or cloud build you just want to open. A debug build doesn’t need it, even one with its JavaScript bundled in. A standalone session with no build_id runs the app’s newest release build. An app for another platform is refused before any device is used.
list_buildsList an app’s uploaded builds with their kind (debug, release or unknown) — the equivalent of vibeview list-builds.
list_device_modelsList the device models you can start, with how many of each are free and busy right now, and a foldable marker for every device with a hinge (the iPhone Duo, Android foldables such as the Pixel 9 Pro Fold) — the equivalent of vibeview list-devices --models. A dev_start on a model with none free waits in the queue for one.
dev_stopStop the session dev_start started. With none running, it says so.
stop_sessionStop a session by its ID — the equivalent of vibeview stop <session-id>. A session that had already ended is reported as such. A teammate’s session is refused unless you are an organization admin.
dev_reloadRoku only: re-package the channel with the build command in vibeview.json, upload it, and restart it in the session dev_start holds (or, with no held session, this project’s vibeview dev --detach session) — the equivalent of pressing r in vibeview dev. A session_id naming any other session is refused. Other platforms hot-reload through Metro on their own.

Note that dev_start is stateful: once it succeeds, every later tool call that doesn’t name a session_id targets that session. Only one can run at a time, and it keeps billing streaming minutes until dev_stop is called or the connection ends. When the client quits, closes the connection, or stops the server, the server stops the session dev_start started — including one still waiting in the queue — before it exits. A session it only targeted by session_id, or a vibeview dev --detach session, keeps running. If the server disappears without that (it is killed or crashes, or the computer goes to sleep), its session ends within about 90 seconds, unless a browser tab is still watching it. The idle timeout applies to a dev_start session too, because holding the connection open is not use. Device actions (tap, type, swipe, press and the other action tools, and dev_reload on Roku) keep the session up. Reading it with ui_tree, screenshot, logs or find, or waiting with wait, does not. When the session is about a minute from ending for inactivity, the next tool result ends with a line such as Session … ends in 40 s for inactivity, so the agent can act or stop. If the session ends some other way (stopped from the dashboard, the idle timeout, or a failure), the next tool call says the session has ended and why, and dev_start starts a new one. If another VibeView dev connection takes the session over, the server lets it go without stopping it, and the next tool call says so. After the computer wakes from sleep, that message says the session ended because the computer stopped checking in.

When a client connects, the server also sends a short primer (the MCP instructions field) describing this loop, which clients such as Claude Desktop show to the model.

Roku channels run the same loop over MCP: dev_start with platform roku starts the session, and dev_reload after each edit puts the new package on the device. From a shell, vibeview dev --platform roku --detach and vibeview dev-reload are the same two steps.

Driving an embed session

A session someone started from an embedded demo can be driven with the same commands. You target the session that visitor is already using: pass its id with --session <id> (or session_id over MCP) and every command lands on the device in front of them. Agent Control never starts a second session for this, and no additional device is used.

Getting the session id

The embedded page announces it to your page. Listen for the session:started event — it carries the id — and hand that to your agent:

window.addEventListener('message', (event) => {
  if (event.data?.source !== 'vibeview-embed') return;
  if (event.data.type === 'session:started') {
    const sessionId = event.data.sessionId; // pass this to Agent Control
  }
});

If the page identifies its visitors — see Identifying your visitors — the id you get is that visitor’s own session, and — with Max sessions per visitor at its default of 1 — it stays the same across their reloads: they come back to the session already running rather than starting a new one, so an agent you pointed at it earlier is still pointed at the device in front of them. Raise that limit and a reload starts a session with a NEW id until the visitor is at their limit, so re-read the id on every session:started rather than assuming it is the one you already have. Each load announces the session to your page again, so you’ll see session:started carrying the same id as before — after a short wait, if that session is still being set up, or straight away if it’s ready. On an anonymous embed, every load is a different session with a different id.

If the visitor lands in a queue you’ll get session:queued first — that one carries only their place in line, not an id, because no device has been handed out yet. Wait for session:started. See Page events for the full list.

Keep your credential on your own server. The embed page never receives one, and your page shouldn’t hold one either — it only needs the session id.

Two things have to be true first:

  1. Agent Control is allowed on that embed key. It’s a per-key setting under Settings → Embeds, off by default, and only an organization admin can turn it on — see Allow agent control.
  2. You’re calling as a Developer or Admin in the organization that owns the key. Viewer credentials are refused, and a token never carries more authority than the member it belongs to.

People viewing your embed in a browser never gain any of this. There’s no command surface on the embedded page — driving the session is something your organization does with its own credentials, from your own tooling — and turning the setting on changes nothing about what a visitor can do.

If either condition isn’t met the command is refused outright and nothing runs on the device. The same applies the moment the key stops being active — turning the setting off, disabling the key, or revoking it all end agent access at the very next command, even mid-session.

What’s available on an embed session

An embed deliberately keeps visitors inside your app, and Agent Control holds that same line. Everything you need to read and drive a screen is there; the things that would take a visitor out of your app, or restart it under them, are not.

CommandOn an embed session
ui-tree, screenshot, logsRead the screen and the app’s own output, exactly as on any other session.
tap, long-press, drag, swipe, scroll, scroll-toAvailable, by element reference or by coordinates.
type, clear-textAvailable.
alertAvailable — respond to a system alert sitting over your app.
find, waitAvailable.
set-posture, rotateAvailable — folding or turning the device changes how the visitor holds it, not the app they’re in.
set-locationAvailable — it changes where the app thinks the device is, not the app the visitor is in.
tap-focused, focusAvailable on TV sessions.
press <button>Limited to back and the d-pad (dpad_up, dpad_down, dpad_left, dpad_right, dpad_center). home, lock, and siri are refused: they would drop the visitor out of your app, blank their screen, or hand them to a system assistant.
open-urlNot available. Opening a deep link or URL is the one command whose whole purpose is leaving the app.
relaunch-appNot available. The embed keeps your app running for the visitor and restarts it itself when it has to.

A refused command says what isn’t allowed and what is, so an agent can correct itself instead of retrying blindly — and the refusal happens before anything reaches the device.

The visitor is using the screen too

This is the one place where you aren’t alone on the device. A visitor can tap between two of your commands, so the screen you read a moment ago may not be the screen you’re acting on, and an element reference you were holding can be stale by the time you use it.

When that happens the command comes back unsuccessful without acting, and the response carries a fresh reading of the current screen — including whatever the visitor just changed. Treat that as ordinary rather than as a failure to retry: re-read the screen and pick the element again from what’s there now. Repeating the same reference will keep failing, because it no longer points at anything.

Sharing the screen is the point of this, not a flaw to work around. The visitor keeps tapping while you work, and they watch your actions happen live on their own device.

The skill

The npm package also ships a ready-made skill at skills/vibeview-agent/SKILL.md describing this entire workflow (build requirements, the loop, verifying, cleanup, and the full command reference) in a form coding agents with filesystem access can read and follow directly. It covers the same ground as this page.

To see what an agent does with it, read An AI agent shipped an Expo app to TestFlight in 43 minutes: a coding agent built the Kitlist sample app, verified it on cloud iOS and Android devices with these commands, and submitted it to both stores, with the run log quoted throughout.

Billing

A session started this way is an ordinary VibeView session — it bills streaming minutes the same as any session started from the dashboard. Remember to run vibeview dev-stop (or stop the session your agent started) when you’re done so it doesn’t keep running unattended.