dsh-cua — Windows Computer Use

Windows 컴퓨터 사용 MCP 서버: 접근성 우선 요소 작업, 스카이샷 텍스트 트리, 보호된 원시 입력, 그리고 인간에게 양보하는 중재자.

문서

dsh-cua

ci PyPI Hutusion/dsh-cua MCP server LINUX DO

English · 中文

An MCP server + agent skill for computer use on Windows: accessibility element actions come first and screenshots are only the fallback. It ships a cross-session arbiter — when several agents share one machine it serializes them, and it yields while you are actually using the computer yourself.

This repository contains only the MCP server and the skill. It stays neutral toward any stdio MCP client (dsh / Claude Code / Codex / Cursor / Cline / ZCode …) — nothing here requires dsh.

Platform semantics (0.3.1 and later): the tools only work on Windows — they drive user32/kernel32 and UI Automation. But the package also imports elsewhere and the server starts there, answering tools/list as usual, so any client or directory crawler can enumerate all 19 tools with their full schemas; actually calling a tool returns a clear "requires Windows" error rather than the process failing to start at all. In 0.3.0 the import itself raised, which made such crawlers unable to see the server at all (tests/linux-handshake.py is the regression test for this property, and CI runs it on ubuntu-latest).

What it is

A stdio MCP server exposing 19 tools. Every tool name is prefixed tool_, exactly as tools/list returns it.

  • Observe (read-only, callable at any time): tool_skyshot (reads a window as a compact, diffable text tree — three orders of magnitude smaller than a screenshot), tool_element_at_point, tool_read_element, tool_find_elements, tool_capture_window (DPI-aware, cropped to the client area), tool_list_windows / tool_find_window / tool_get_window_rect, tool_list_displays, tool_cursor_position, tool_clipboard_read, tool_coexistence_status
  • Element actions (soft gate: serialized across agents, no physical input injected): tool_element_action / tool_element_action_at — press / set_value / select / toggle / expand / collapse / scroll_into_view / focus, delivered straight to the UIA element, so they never steal focus and never care about z-order
  • Other mutating calls (soft gate too: the mutex, but no human-contention yield): tool_type_text (targeted PostMessage with effect_verified), tool_clipboard_write, tool_open_application. They synthesize no physical input, so they do not wait for you to stop working — and a clipboard write still destroys whatever you last copied, so announce it when you do.
  • Physical input (hard gate: serialized across agents and yields to the human): exactly two things share your one cursor and one keyboard — tool_click_at's raw_event path and tool_send_keys' global hotkeys. The gate waits for the machine to go input-quiet, then refuses with user-active rather than fight you for the cursor. tool_click_at tries its element path first (ax_press), which injects no physical input and therefore takes the mutex only; the receipt's method field says which path actually ran.

"Read-only" here means it takes no mutating action and synthesizes no input, so it is safe to call while someone is using the machine. Two of them have a side effect worth knowing: tool_capture_window writes the screenshot to disk (save_path; a temp file when omitted), and tool_skyshot updates the server-side diff baseline it diffs the next shot against.

A small tree is not proof of an empty window. tool_skyshot and tool_find_elements report minimized, because a minimized window may be showing "nothing is displayed" rather than "nothing is there" — and how much it hides depends on the application, which is why the receipt reports the state instead of guessing at the cause. Measured here: an Edge window minimized before anything read it exposes its browser chrome and no page at all (and include_offscreen does not recover it), Explorer falls from 40 elements to 8 whenever it is minimized, and Notepad is unaffected. Restore the window once and read it again if the result matters.

Every action returns a receipt rather than a self-reported success: action_sent / effect_verified / foreground_changed / user-active / arbiter-busy. "The call was accepted" and "the effect happened" are two different things, and the tool separates them for the agent.

How it differs

There are already several mature open-source Windows implementations. dsh-cua's differences are concentrated on one thing: sharing a machine with a human.

dsh-cuacua-driverahk-mcplean-computer-use-mcp
Element actions delivered as UIA patterns (no focus steal, z-order irrelevant)✅✅ (ax mode)❌ reads via UIA, acts by coordinate clickvia cua-driver
Recent human input → refuse✅ user-active❌❌❌
Cross-agent serialization (multi-process)✅ named mutex❌❌❌
Per-action effect assertion✅ three-state effect_verifiedreports a delivery tier❌❌ state_changed heuristic only
Foreground-steal side effect measured✅ foreground_changed❌❌❌
Tool count1959156

The key distinction is two things that are routinely conflated:

  • "No focus steal" is a mechanism guarantee — either a UIA pattern or a targeted PostMessage, so the cursor and keyboard focus are physically never touched. cua-driver has it (ax mode). ahk-mcp does not, and the distinction is narrower than "no UIA": it reads through UIA (ahk_uia_tree / ahk_uia_find / ahk_uia_url), but it has no UIA pattern action — per its README it acts with coordinate clicks or synthetic keys, so an action does move the real cursor.
  • "Yield the moment you move" is a timing guarantee — it reads the age of the human's last input via GetLastInputInfo, waits when it sees you using the machine, and on timeout refuses (user-active) instead of barging in. As of 2026-09-25 a pattern search across the other three codebases in that table found no equivalent — that is search evidence, not proof, and it covers input-age detection only: cua-driver does have human-facing guards of a different kind (a consent requirement, and foreground-steal detection with restore).

effect_verified is likewise something the alternatives lack: it splits "the call was accepted" from "the effect happened" and gives three states (true changed as expected / false accepted but unchanged, downgraded to a failure / null no comparable state, i.e. unconfirmed). The usual alternative is to re-observe once after the action and leave the judgement to the model.

What dsh-cua does not do (stated up front to avoid misunderstanding): no grounding of its own — the server does not analyse pixels, so a text-only model cannot drive interfaces that a tree cannot express (canvas, games, remote desktop). With a vision-capable model the pixel path is supported end to end: capture_window returns the image together with a verified image→screen mapping (bounds, scale, dpi_verified), and the model supplies the grounding. Also not provided: record-and-replay, and an isolation sandbox. There are better-suited tools for those.

FAQ

Why not run the agent on a second desktop or a virtual display, so it never touches mine?

Because Windows has nothing to build that on, and the mobile design that does work rests on exactly the missing piece. On Android an app can create a VirtualDisplay and address input at it — an input event carries a display id, so the agent's taps are routed to its own screen and the human's touchscreen never notices. Windows routes input per desktop, not per display: a desktop has one input queue and one cursor position.

That makes the obvious analogues dead ends:

IdeaWhy it does not isolate
Add a virtual monitor (an indirect display driver)Another canvas, not another cursor — the pointer still has a single position across all monitors
A Windows virtual desktop (Win+Ctrl+D)A view switch inside the same session: same input queue, same cursor
A second session (RDP, or another user)Isolation is real, but client Windows allows one interactive session per user at a time — connecting remotely locks the console, so the human loses their screen, which was the whole point
A hidden Win32 desktop (CreateDesktop)The agent would get its own input queue and cursor, but its windows are invisible, so it can only drive instances it launched — not the program you are looking at. (Inferred from the window-station/desktop model; not measured here.)

That table is about input, and it is worth being exact about what a second display does buy, because the answer is not "nothing". Measured on Edge with an isolated profile (tests/verify-visible-vs-foreground.py): Chromium builds a page's accessibility tree for a window that is merely shown. A window minimized before any query reported no page; the same window reported a named page as soon as it was shown — while not the active window (GetGUIThreadInfo(hwndActive) false) and fully covered (0% of its area on top). Every state was read back from the OS rather than assumed, and a ForegroundWatch around each walk (1.4-2.7M samples, foreground unchanged) shows the read takes no focus at all. A browser window parked on a second display is therefore readable with skyshot while your focus never leaves your screen.

Input does not survive the move. A WM_CHAR posted to that same window is accepted by PostMessage and changes nothing — covered or uncovered — while the identical post to the same window when it is active inserts the character. So a second display buys the read and not the act, and the surfaces that need acting on (Chromium internals, canvas, games) are exactly the ones needing the activation it cannot provide. The constraint is input, and for input there is still one cursor and one foreground window per desktop.

One qualification, because "my browser is on a second display and works fine" is a true report of a different path. All of the above is about OS-level injection, and it does not generalise to the browser's own protocol. Measured on the same machine, with a headed Edge window driven over CDP (Input.dispatchMouseEvent, which is what Playwright's page.mouse.click sends): the click arrived as a trusted event, the system cursor did not move by a single pixel (GetCursorPos identical before and after), the foreground window did not change, and a further click still landed after the window had been minimized (IsIconic true, foreground on another window). CDP injects into the browser's own input pipeline, above the OS input queue, so the one-cursor/one-foreground rule does not reach it at all — and neither does it reach anything else addressed through an application's own API instead of through the screen. (The page's document.visibilityState still read visible while minimized; that part is an artefact of the flags Playwright launches Chromium with, not general Chromium behaviour. Delivering the click is a protocol property, independent of it.)

What decides the outcome is which layer the input enters at:

RouteCursor / foreground neededCrosses a display or occlusion boundary
An application's own protocol (CDP, Playwright, a DOM/JS event, an app API)no — nothing OS-level is injectedyes, and the window need not even be visible
element_action (a UIA pattern delivered to an element)no — but the window must exist and expose an elementyes
PostMessage to a window (type_text)measured on Chromium: yes — the target must be active for the text to appearno
Physical injection (SendInput, the raw-event path of click_at, send_keys)yes — one cursor and one foreground window for the whole desktopno

So for a page, "put the browser on the virtual display" is a good answer, and this project does not compete with it: a page is already its own addressable, scriptable surface. dsh-cua is for the targets that have no such surface — Explorer, native dialogs, Office, canvas and games, applications with no scripting API — where OS input is the only route there is, and that route is exactly the one a second display does not help. Both statements are true.

So the conflict is not a gap in this implementation, it is an OS constraint: Windows has no second cursor. Given that, yielding is the only correct response — and most calls never get near the problem, because they inject no input at all:

  • element_action / element_action_at — a UIA pattern is delivered to the element: no physical input, no cursor movement, no focus change. These are the paths that "can run while the user types".
  • type_text — a window-targeted PostMessage, not global keystrokes. It resolves the text control inside the window first (a top-level window does not forward WM_CHAR to its child edit) and reads that control back, so the receipt carries effect_verified rather than only reporting that something was posted.
  • Every read-only tool (skyshot, element_at_point, capture_window, …) — touches nothing.

Exactly two paths inject physical input, and they exist because canvas-, game- and Chromium-internal surfaces expose no element to address: the raw-event path of tool_click_at, and the global hotkeys of tool_send_keys. Those two are what the arbiter guards.

Install

You need Windows x64 + an interactive desktop session + Python ≥3.10 to actually drive a desktop. (The package installs and starts on Linux/macOS too, tools/list answers normally, and a tool call then reports "requires Windows" — see "platform semantics" above.)

# Option 1: uvx, zero install (recommended)
uvx dsh-cua                      # runs the stdio MCP server directly

# Option 2: pip
pip install dsh-cua

# Option 3: from source
pip install git+https://github.com/Hutusion/dsh-cua.git

However you install it, start the server with python -m dsh_cua:

python -m dsh_cua                # depends on no executable being on PATH

Why the README does not say dsh-cua-server: pip installs console scripts into the interpreter's Scripts directory, and that directory is not necessarily on PATH — measured on a stock python.org 3.12 install, neither the User nor the Machine PATH contained it, so pip install dsh-cua succeeded while dsh-cua-server reported command not found. python -m needs no PATH entry at all. The console script is still shipped and works when PATH does contain it.

Options 1 and 2 both work today: the package is published on PyPI (https://pypi.org/project/dsh-cua/). If uvx/pip ever 404s, use option 3 — it always works.

Wiring it up

Any MCP client; name the server win32 (the skill's tool-name convention is mcp__win32__*).

python -m (no PATH dependency, recommended):

{ "mcpServers": { "win32": { "command": "python", "args": ["-m", "dsh_cua"] } } }

uvx:

{ "mcpServers": { "win32": { "command": "uvx", "args": ["dsh-cua"] } } }

More shapes are in examples/: Claude Code / generic clients / a dsh cordis.patch.yml fragment / the route modality declaration you need if you want the model to read screenshots (tr-route-settings.yml).

Skill (optional but strongly recommended)

skill/computer-use/SKILL.md is the companion doctrine for using these tools: the observe → locate → act → verify loop, receipt semantics, retry safety, and the discipline of coexisting with a human. The model can use the tools without it, but with it the model picks the right path by itself — the measured difference is large.

The skill lives in this repository, not in the package — uvx and pip do not put a skill/ directory on your disk, so the cp below only works from a checkout. Without one, fetch the file:

mkdir -p ~/.dsh/skills/computer-use
curl -fsSL https://raw.githubusercontent.com/Hutusion/dsh-cua/main/skill/computer-use/SKILL.md \
  -o ~/.dsh/skills/computer-use/SKILL.md      # use ~/.agents/skills/ for Claude Code

From a clone, copy the directory instead:

# Claude Code / generic agents
cp -r skill/computer-use ~/.agents/skills/
# dsh
cp -r skill/computer-use ~/.dsh/skills/

Security model

TierOperationsGate
Read-onlythe 12 observe toolsno gate, callable at any time
Softelement actions, PostMessage typing, clipboard write, launching applicationscross-agent mutex (named mutex, multi-process, automatic serialization)
Hardraw clicks, global hotkeysmutex + GetLastInputInfo yielding: if the user typed recently it waits, and on timeout refuses with user-active instead of stealing the cursor

Honest boundaries: yielding is a cooperation protocol, not a hard guarantee (the tight check 150 ms before injection narrows the window as much as possible); a few applications self-activate even on set_value (the receipt reports foreground_changed truthfully); and two operators on the same window has no technical solution — do not drive the same window the agent is driving.

Tests

python tests/verify-coexistence.py    # 25 checks: zero-input proof / cross-process mutex / synthetic human contention / kill switch
python tests/verify-p0-fixes.py       # the three P0s fixed in 0.2.0: each fails before the fix

The tests need no human cooperation — "user input" is synthesized with one real 1-pixel cursor move, and the cursor is restored afterwards.

The tests need a real interactive desktop session (some checks create windows and address them through UIA), so they cannot run on a GitHub-hosted runner. What CI does cover is the part that needs no desktop: packaging and installation, module import, regressions for the diff index and tree-line escaping, and the arbiter's decision logic — see .github/workflows/ci.yml.

python tests/ci-desktop-free.py       # the local equivalent of the above, no desktop needed

License

MIT