npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@harusame64/desktop-touch-mcp

v2.0.0

Published

Let Claude, Cursor, or any MCP client see and operate your Windows 10/11 desktop. 32 tools for screenshots, UI Automation, Chrome CDP, keyboard/mouse, terminal, with semantic discover-then-act targeting and per-action perception guards that avoid wrong-wi

Readme

desktop-touch-mcp

desktop-touch-mcp MCP server

日本語

Computer-use MCP server for Windows. Lets Claude, Cursor, or any MCP client see and operate your Windows 10/11 desktop — screenshots, UI Automation, Chrome CDP, keyboard / mouse, terminal — with semantic discover-then-act targeting that avoids pixel-coordinate guessing, and per-action perception guards that catch wrong-window typing before it happens.

npx -y @harusame64/desktop-touch-mcp

32 tools, native Rust engine (UIA in 2 ms), zero-config PowerShell fallback, full CJK support, MIT licensed. Add the snippet above to your Claude / Cursor / VS Code Copilot config and Claude can drive Notepad, Excel, Chrome, Windows Terminal, and any other app on your machine.

Why this over pixel-clicking? Two ideas run through every tool: discover-then-act — desktop_discover returns interactive entities with short-lived leases instead of raw coordinates, so desktop_act operates on what you mean, not where it was — and per-action perception guards that verify the target window's identity and bounds before input lands, catching wrong-window typing and stale-coordinate clicks before they happen.

Under the hood: an 82× average speedup from the Rust native engine (UIA focus queries in 2 ms, SSE2-accelerated image diffing at 13–15×), with a transparent PowerShell fallback when the engine is absent. The npm launcher fetches only the GitHub Release tag matching the installed version and verifies the Windows runtime zip before extraction.


Features

  • ⚡ High-performance Rust Native Core — The UIA bridge and image-diff engine are written in Rust (napi-rs + windows-rs) and loaded as a native .node addon. Direct COM calls from a dedicated MTA thread eliminate PowerShell process spawning — getFocusedElement completes in 2 ms (160× faster), and getUiElements returns full trees in ~100 ms with a batch BFS algorithm that minimizes cross-process RPC. Image-diff operations use SSE2 SIMD for 13–15× throughput. When the native engine is unavailable, every function transparently falls back to PowerShell — zero config required.
  • 🎯 Set-of-Marks (SoM) visual fallback — Games, RDP sessions, and non-accessible Electron apps return clickable elements even when UIA is completely blind. screenshot(detail="text") automatically detects UIA sparsity and activates a Hybrid Non-CDP pipeline: Rust-powered grayscale + bilinear upscale → Windows OCR → clustering → red bounding-box annotation with numbered badges ([1], [2]…). Two parallel representations returned: a visual PNG for spatial orientation and a semantic elements[] list with clickAt coords — no CDP required.
  • 🔁 One-call confirmation on visual-only targets — On UIA-blind targets (Electron, PWAs, games, custom canvases, RDP windows), desktop_act can fold the post-action confirmation into its own response: an optional roiCapture carrying a PNG crop of just the region that changed plus a lease-less preview of the controls now visible there. The agent confirms what its click did and finds the next target without a separate desktop_state + screenshot. On visual-only targets it is on by default for a visible change (returnCapture:"on-change"); pass returnCapture:"never" to suppress it, or "always" to force it. Never attached on structured targets (browser/CDP, UIA-rich native), where desktop_state is cheaper and exact — so those responses are unchanged.
  • 🔐 Key Locker — the terminal autofills your SSH / sudo passwords — Save a credential once into the locker's own secure dialog (stored encrypted on your machine with Windows DPAPI; never shown to the assistant), then run ssh / sudo in a console opened by key_locker(action='launch_console') — the password is filled in automatically when the hidden prompt appears, with a per-fill confirmation prompt by default. See Key Locker.
  • LLM-native design — Built around how LLMs think, not how humans click. run_macro batches multiple operations into a single API call; diffMode sends only the windows that changed since the last frame. Minimal tokens, minimal round-trips.
  • Reactive Perception Graph — Register a lensId for a window or browser tab, pass it to action tools, and get guard-checked post.perception feedback after each action. It reduces repeated screenshot / desktop_state calls and prevents wrong-window typing or stale-coordinate clicks.
  • Full CJK support — Uses Win32 GetWindowTextW for window titles, avoiding nut-js garbling. IME bypass input supported for Japanese/Chinese/Korean environments.
  • 3-tier token reduction — detail="image" (~443 tok) / detail="text" (~100–300 tok) / diffMode=true (~160 tok). Send pixels only when you actually need to see them.
  • 1:1 coordinate mode — dotByDot=true captures at native resolution (WebP). Image pixel = screen coordinate — no scale math needed. With origin+scale passed to mouse_click, the server converts coords for you — eliminating off-by-one / scale bugs.
  • Browser capture data reduction — grayscale=true (~50% size), dotByDotMaxDimension=1280 (auto-scaled with coord preservation), and windowTitle + region sub-crops help exclude browser chrome and other irrelevant pixels. Typical reduction for heavy captures: 50–70%.
  • Chromium smart fallback — detail="text" on Chrome/Edge/Brave auto-skips UIA (prohibitively slow there) and runs Windows OCR. hints.chromiumGuard + hints.ocrFallbackFired flag the path taken.
  • UIA element extraction — detail="text" returns button names and clickAt coords as JSON. Claude can click the right element without ever looking at a screenshot.
  • Auto-dock CLI — window_dock(action='dock') snaps any window to a screen corner with always-on-top. Set DESKTOP_TOUCH_DOCK_TITLE='@parent' to auto-dock the terminal hosting Claude on MCP startup — the process-tree walker finds the right window regardless of title.
  • Emergency stop (Failsafe) — Park the mouse in the top-left corner of the primary monitor (within 10px of 0,0) for 500ms to trigger the emergency stop.

Requirements

| | | |---|---| | OS | Windows 10 / 11 (64-bit) | | Node.js | v20+ recommended (tested on v22+) — to develop or run the test suite, ^22.12 || ^24 || >=26 — the test runner's own range since #658, which excludes odd majors such as 23 and 25 | | PowerShell | 5.1+ (bundled with Windows) — used only as fallback when the Rust native engine is unavailable | | Claude CLI | claude command must be available |

Note: nut-js native bindings require the Visual C++ Redistributable. Download from Microsoft if not already installed.

Note (Key Locker): The credential helper Key Locker uses is an unsigned executable, so on some machines Windows SmartScreen or antivirus may show an "unknown publisher" warning the first time it runs. This is expected — the helper ships with desktop-touch-mcp and runs locally on your machine; you can allow it to proceed. Code signing is planned for a future release.


Installation

npx -y @harusame64/desktop-touch-mcp

The npm launcher resolves runtime strictly by npm package version. For package X.Y.Z, it fetches only GitHub Release tag vX.Y.Z, downloads desktop-touch-mcp-windows.zip, verifies its SHA256 digest, and only then expands it under %USERPROFILE%\.desktop-touch-mcp. Verified cached releases are reused on later runs.

Set DESKTOP_TOUCH_MCP_HOME to override the cache root directory.

On a shared or CI network? The first run reads the GitHub Releases API to locate the runtime zip. The anonymous limit is 60 requests/hour per IP, which a shared public address (CI runners, office NAT) can exhaust before your download even starts. Set GITHUB_TOKEN (or GH_TOKEN) in the environment and the launcher authenticates the request, raising the limit to 5,000 requests/hour. No token is needed on an ordinary home connection.

Running the launcher from a source checkout? A source build's bin/launcher.js carries a placeholder integrity hash (sha256: "PENDING") instead of a finalized one. Rather than download and run an unverified runtime, the launcher fails closed — this guard stops an accidentally published or unfinalized launcher from silently starting unverified code. Published npm releases always ship a real SHA256, so end users never see this. If you are intentionally running the launcher from source, set DESKTOP_TOUCH_MCP_ALLOW_UNVERIFIED=1 to skip integrity verification (development only).

Does your host give up before the launcher finishes? Some desktop hosts allow a plugin a fixed budget — 60 seconds is common — to become ready, and a launcher waiting on an unreachable GitHub can spend all of it. Two environment variables cover that case.

DESKTOP_TOUCH_MCP_FETCH_TIMEOUT_MS (default 15000) bounds how long the launcher waits without hearing from GitHub. It applies to the release lookup and to the download; for the download it counts silence rather than total time, so a large runtime still installs over a slow connection. A value that is not a positive number of milliseconds is ignored with a warning.

DESKTOP_TOUCH_MCP_OFFLINE_FALLBACK=1 lets the launcher start a release that is already installed when GitHub cannot be reached at all. It is off by default. GitHub is always contacted first, so a reachable network still re-downloads and repairs a damaged install; only a network failure reaches the fallback, which starts the copy of your version on disk — without re-verification — or, when that version was never installed, the newest older release that completed a verified install. Answers that are not network failures (a 404, the API rate limit, a mismatched integrity hash) still stop startup loudly. Leave it off unless a host timeout forces your hand: while it is set, a corrupted install of your current version is reused instead of being repaired.

The two work together: with the fallback on, startup still waits out the timeout before falling back, so lower DESKTOP_TOUCH_MCP_FETCH_TIMEOUT_MS if your host's budget is tight. Note also that a download which is still arriving, however slowly, is never interrupted — the fallback answers when the network has gone silent, not when it is merely slow.

Register with Claude CLI

Add to ~/.claude.json under mcpServers:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"]
    }
  }
}

No system prompt needed. The command reference is automatically injected into Claude via the MCP initialize response's instructions field.

Register with other clients (HTTP mode)

Clients that require an HTTP endpoint (GPT Desktop, VS Code Copilot, Cursor, etc.) can use the built-in Streamable HTTP transport:

npx -y @harusame64/desktop-touch-mcp --http
# or with a custom port:
npx -y @harusame64/desktop-touch-mcp --http --port 8080

The server starts at http://127.0.0.1:23847/mcp (localhost only). Register the URL in your MCP client settings. A health check is available at http://127.0.0.1:<port>/health.

In HTTP mode the system tray icon shows the active URL and provides quick-copy and open-in-browser shortcuts.

Development install

git clone https://github.com/Harusame64/desktop-touch-mcp.git
cd desktop-touch-mcp
npm install

Build after install:

npm run build

For a local checkout, register the built server directly:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "node",
      "args": ["D:/path/to/desktop-touch-mcp/dist/index.js"]
    }
  }
}

Note: Replace D:/path/to/desktop-touch-mcp with the actual path where you cloned this repository.


Tools (32 Optimized Tools)

📖 Full Reference: docs/system-overview.md — Exhaustive guide on parameters, return schemas, and coordinate math.

🌐 World-Graph V2 (Primary Path)

| Tool | Description | |---|---| | desktop_discover | Observe the desktop. Returns interactive entities with leases (UIA, CDP, Terminal, Visual SoM). | | desktop_act | Perform actions (click, type, drag) on entities via lease validation. Returns semantic diffs — plus an optional roiCapture (changed-region PNG + next-target preview) on visual-only targets. |

👁️ Observation & State

| Tool | Description | |---|---| | desktop_state | Lightweight check of focus, active window, cursor, and Auto-Perception attention signal. | | screenshot | Multi-mode capture: detail='text' (UIA/OCR), diffMode (P-frame), dotByDot (1:1), and background. Returns a cheap screenshot://by-ref/{id} link to the saved image instead of inlining pixels every time. | | screenshot_query / screenshot_gc | Inspect and prune the on-disk screenshot cache behind the by-ref links: screenshot_query lists saved captures without re-reading pixels; screenshot_gc reclaims space by retention policy (dry-run by default). | | workspace_snapshot | Instant session orientation: all window thumbnails + UI summaries in one call. | | server_status | Diagnostic check for native engine health and feature activation. |

⌨️ Input & Control

| Tool | Description | |---|---| | keyboard | Send keyboard input. Supports background input (WM_CHAR) and IME-safe clipboard bypass. | | mouse_click / mouse_drag | Precision coordinate-based interaction with homing and force-focus protection. | | scroll | Multi-strategy: raw (notches), to_element, smart (virtual lists), and capture (stitch). | | click_element | Legacy UIA-based click by name/ID (fallback when entities are unavailable). |

🌐 Browser CDP (Chrome/Edge/Brave)

| Tool | Description | |---|---| | browser_open / browser_navigate | Idempotent debug-mode launch and reliable navigation. | | browser_click / browser_fill / browser_form | High-level DOM interaction stable across repaints and framework re-renders. | | browser_eval | Deep inspection via js (scripting), dom (HTML), and appState (SPA data extraction). | | browser_overview / browser_search / browser_locate | Semantic discovery, grep-like DOM search, and pixel-accurate coordinate lookup. |

🛠️ Utilities & Workflow

| Tool | Description | |---|---| | terminal | Unified command execution: run (send + wait + read), read (OCR/UIA), and send. run completion modes: quiet, pattern, and exit (waits for the command to finish + returns its exit code — see Terminal command completion). | | wait_until | Efficient server-side polling for window, focus, text, or URL state changes. | | window_dock / focus_window | Window management: pin (always-on-top), unpin, dock (corner snap), and focus. | | workspace_launch | Launch apps and auto-detect new HWNDs (supports localized titles). | | run_macro | Batch up to 50 operations into a single round-trip for maximum efficiency. | | clipboard / notification_show | System-level text exchange and user alerts. | | key_locker | Manage credentials the terminal autofills for you (SSH key passphrases, sudo / login passwords). Secrets are entered once into the locker's own secure dialog and stored encrypted on this machine (Windows DPAPI); they are never shown to the assistant. action='launch_console' opens an autofill-capable console (returns a paneId to drive ssh/sudo into via terminal); save / list / forget / set_policy / status manage bindings. Autofill only fires in a console opened by launch_console. Disable with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1. |

📊 Office (Excel)

| Tool | Description | |---|---| | excel | Author and run Excel VBA macros via COM. action='run_vba' writes a macro into a managed Trusted Location and runs it; action='check_access_vbom' is a read-only preflight. Runs VBA where formula-only tools cannot. One-time setup: node scripts/enable-access-vbom.mjs. |


Standard workflow (v1.0.0)

The v2 World-Graph surface (desktop_discover / desktop_act) is the recommended dispatch path. The four-call shape works for native apps, browsers, and terminals identically.

desktop_state          → orient: focused window/element, modal, attention signal
desktop_discover       → find actionable entities (returns lease + windows[])
desktop_act(lease, …)  → act on entity (returns attention + post.perception)
desktop_state          → confirm the world changed as expected

Clicking — priority order:

browser_click(selector)               → Chrome / Edge (CDP, stable across repaints)
desktop_act(lease, action='click')    → native / dialog / visual (entity-based; use after desktop_discover)
click_element(name | automationId)    → native UIA fallback if desktop_act returns ok:false
mouse_click(x, y, origin?, scale?)    → pixel last resort; origin+scale from dotByDot screenshots only

Recovery hints — read response.attention after every observation and response.warnings[] on desktop_discover / desktop_act. Common reasons:

  • lease_expired / lease_generation_mismatch / lease_digest_mismatch / entity_not_found → re-call desktop_discover
  • modal_blocking → response.blockingElement (when present) names the blocking modal. role: "dialog" means a separate dialog window has disabled the target's window: blockingElement.hwnd is that dialog — re-call desktop_discover with target.hwnd = blockingElement.hwnd, answer it there, then retry (name is its title, which may be empty or shared). Any other role: a window the desktop_discover snapshot holds, where the OS could not say whether it blocks this entity — with blockingElement.hwnd, re-call desktop_discover with target.hwnd = blockingElement.hwnd and answer it there; without it, dismiss via click_element(name=blockingElement.name). Then re-call desktop_discover on the original target and act on the new lease — this refusal came from that snapshot, so the same lease is refused again
  • entity_outside_viewport → the element moved off screen: scroll(action='to_element' | 'raw'), or re-call desktop_discover if its window moved or closed
  • origin_window_not_visible → the element's window is minimised or hidden, so nothing is drawn where it was found: focus_window(windowTitle) to restore it, then re-call desktop_discover
  • coordinate_outside_reachable_bounds → the coordinate is not on any connected monitor. Coordinate-based mouse input (mouse_click / mouse_drag / scroll / browser_click, and the mouse route inside desktop_act) now works on every monitor, including monitors placed left of or above the primary one, so this error normally means the coordinates are stale — the window moved or closed after they were read. Re-run desktop_discover and act on the new coordinates. If the server is running without its built-in Windows input module, mouse input falls back to the primary monitor only; the error message says so, and moving the window onto the primary monitor (or reinstalling the server) is the fix
  • cursor_placement_blocked → the coordinate is on a monitor, but the pointer could not be placed there, so nothing was clicked. This happens while another app confines the cursor to its own window (common in full-screen games), while a remote-desktop session is disconnected or locked, while another program keeps repositioning the pointer, or right after a monitor is added or removed. Leave the app holding the cursor, reconnect the session, or — after a monitor change — re-run desktop_discover, then retry. click_element acts through the accessibility API without moving the cursor and works meanwhile
  • keyboard_target_unsafe → a type was refused, because the characters would not have reached the field you named: the keyboard focus is on a different control or in a different window, the control that would receive them — or the field you named — is read-only, or the field you named — or its window — is disabled. Nothing was typed, and if_unexpected.detail says which. For a disabled field, answer or wait out whatever disabled it, then re-run desktop_discover and type again; clicking it does not help, and desktop_discover does not list a disabled field, so while it is missing there it is still disabled. A field you named that is read-only does not take text; typing again will not change that. For another control or window, put the focus on the field you named, then type again; if_unexpected.detail names the way back for the road the act took. On a window named by title, desktop_act with action='click' on the same entity does it. On a window named by handle nothing here moves the focus to a text field yet, so re-run desktop_discover by the window's title and click the field there — a common dialog's title resolves to a handle as well, so that road does not open there. For another window, bring the field's window forward first (focus_window): it comes forward with the focus it last had, and the window holding the focus is usually drawn over the field. Do not retry with a foreground keyboard type: whatever holds the focus would take the characters
  • executor_failed → fall back to click_element / mouse_click / browser_click

A successful type can carry landing: { confirmed: false, why }. The write took the background route, but the server could not confirm that it reached the field you named — for example, in a WPF window, whose fields have no window of their own. This is a report, not a state that can be resolved here: nothing in the response establishes whether the characters arrived, reading the field back does not settle it (desktop_state answers about the foreground, and may come back with no value at all — hints.focusedElementValueAbsent names the road that dropped it, view_road_has_no_value or masked_on_this_road, and no hint is not evidence a value was there — or name a field in another window with the same title), diff.value_changed is not delivery either, its baseline being your desktop_discover snapshot rather than the write, and retrying a nonempty write is not a repeat — a background write lands at the caret and replaces the selection, exactly as typing does.

Lease lifecycle:

  • Each desktop_discover response carries softExpiresAtMs (≈ 60 % of the TTL window). Past that timestamp the LLM should consider re-calling desktop_discover even though the lease is still technically valid — lease.expiresAtMs is the only correctness wall.
  • TTL adapts to view mode (action/explore/debug), entity count, and response payload size. Cap is 60 s.
  • Set DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2=1 to fall back to the v1 tool surface (get_windows / get_ui_elements / set_element_value) for troubleshooting only — V2 is the recommended default.

Terminal command completion (until)

terminal(action='run') sends a command, waits for it to complete, and reads the output in one call. How it decides "complete" is controlled by until:

| Mode | Waits for | Best for | |---|---|---| | quiet (default) | output to fall silent for quietMs | short interactive commands | | pattern | a string/regex you expect in the output | long commands with a known final marker | | exit | the command to actually finish | when you need completion or the exit code |

Anchoring caveat (#384): a command whose final line has no trailing newline glues the marker to the next prompt with no line boundary (printf X → Xuser@host:~$), so an end-anchored pattern (X\s*\n / X$) can never bind. For completion use mode:'exit'; for content matching use a bare marker (no \n/$). mode:'pattern' also accepts an optional quietMs settle fallback: until:{mode:'pattern', pattern, quietMs:1000} completes with reason:'quiet' (no matchedPattern) once output is stable for that long without a match — instead of hanging until timeoutMs. It is opt-in (omit quietMs to keep waiting for the pattern; long commands with mid-run silent gaps are unaffected).

until:{mode:'exit'} — real completion + exit code

The heuristic modes can misfire on the common "append a sentinel" idiom (some-task; echo DONE matched by DONE): the sentinel also shows up in the echoed command line, and for multi-line commands there is no reliable way to tell that echo apart from real output. mode:'exit' removes the guesswork — the server appends its own completion marker whose printed form differs from its typed form, so it never matches the echoed command (even for multi-line input), and it returns the real process exit code:

terminal({
  action: 'run',
  windowTitle: 'pwsh',
  input: 'npm run build',
  until: { mode: 'exit', shell: 'powershell' },
})
// → completion: { reason: 'exited', exitCode: 0, elapsedMs: … }
//   output: just the command's real output (the injected marker is stripped)
  • Pass shell explicitly ('bash' or 'powershell'). shell:'auto' detects the shell from the terminal window, but it cannot see a shell running inside SSH or WSL — the window still looks like its local host — so for remote/nested sessions pass the remote side's shell (auto otherwise warns and may pick the outer shell). A window whose process is genuinely unidentifiable (e.g. Windows Terminal) returns ExitModeShellAmbiguous.
  • First-class shells: bash and powershell. cmd.exe is not supported yet (ExitModeShellUnsupported).
  • Unsafe input is rejected up front (ExitModeUnsafeInput) rather than hanging: a command ending mid-construct (unterminated quote, here-doc, $(…), a trailing \ or PowerShell backtick).
  • Exit mode controls its own delivery, so delivery-shaping sendOptions (method / preferClipboard / pressEnter / chunkSize / pasteKey) are rejected with InvalidArgs; focus options remain accepted.

Key Locker (terminal credential autofill)

Running ssh user@host or sudo … normally stops at a hidden password prompt an assistant can't safely type into. Key Locker stores your SSH key passphrases and sudo / login passwords encrypted on your machine (Windows DPAPI, current user) and fills them in automatically when a bound command reaches its prompt. The secret is typed once into the locker's own secure dialog — it is never shown to the assistant and never travels through the MCP channel.

// 1. Save the credential once — opens a secure dialog on your desktop
key_locker({ action:'save', uri:'ssh://user@host:22' })

// 2. Open an autofill-capable console (returns its paneId)
key_locker({ action:'launch_console' })   // → { paneId:'12345678', windowTitle:'…' }

// 3. Run the command through that pane — the password is filled at the prompt
terminal({ action:'send', paneId:'12345678', input:'ssh user@host' })
  • Autofill only fires in a console opened by launch_console — a pre-existing terminal is never autofilled. The console is a classic visible Windows console, so you can watch it and take over at any prompt yourself.
  • Every autofill asks you to confirm by default; opt a binding out with set_policy. list / status / forget manage saved credentials.
  • terminal read / send accept paneId as an alternative to windowTitle — it targets that exact window even after an ssh login renames its title.
  • Supported binding URIs: ssh://user@host:22, sudo://host/user, https-cred://host, and SSH key passphrases (sshkey:SHA256:…). An ssh save needs the host key already in known_hosts (connect to the host once first).
  • Windows only. Disable the whole feature with DESKTOP_TOUCH_DISABLE_KEY_LOCKER=1. The secure dialog is an unsigned helper executable — Windows SmartScreen may show an "unknown publisher" warning on first run (see the note under Requirements).

Browser CDP automation

For web automation, connect Chrome or Edge with the remote debugging port enabled — no Selenium or Playwright needed.

# Launch Chrome in CDP mode
chrome.exe --remote-debugging-port=9222 --user-data-dir=C:\tmp\cdp
browser_open({launch:{}})                          → spawn-if-needed Chrome in debug mode + list tabs (idempotent)
browser_open()                                     → connect-only (fail if no CDP endpoint live)
browser_locate({selector:"#submit"})               → CSS selector → physical screen coords
browser_click({selector:"#submit"})                → find + click in one step (auto-focuses browser)
browser_eval({action:"js", expression:"document.title"})  → evaluate JS, returns result
browser_eval({action:"dom", selector:"#main", maxLength:5000})  → outerHTML, truncated to maxLength chars
browser_eval({action:"appState"})                  → one-shot SPA state (Next/Nuxt/Remix/Apollo/GitHub react-app/Redux SSR)
browser_fill({selector:"#email", value:"[email protected]"})  → fill React/Vue/Svelte controlled input (state-safe)
browser_overview()                                 → links/buttons/inputs + ARIA toggles + viewportPosition per element
browser_search({by:"text", pattern:"..."})         → grep DOM with confidence ranking
browser_navigate({url:"https://example.com"})      → navigate via CDP (no address bar interaction)

For chained calls in the same tab, pass includeContext:false to omit the activeTab/readyState annotation (~150 tok/call saved). Boolean / object params accept the LLM-friendly string spellings ("true", "{}").

Coordinates returned by browser_locate account for the browser chrome (tab strip + address bar height) and devicePixelRatio, so they can be passed directly to mouse_click without any scaling.

Recommended web workflow:

browser_open({launch:{}}) → browser_eval({action:"dom"}) → browser_locate(selector) → browser_click(selector)

Auto-dock CLI on startup

Keep Claude CLI visible while operating other apps full-screen. Set env vars in your MCP config and the docked window auto-snaps into place every MCP startup.

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_DOCK_TITLE": "@parent",
        "DESKTOP_TOUCH_DOCK_CORNER": "bottom-right",
        "DESKTOP_TOUCH_DOCK_WIDTH": "480",
        "DESKTOP_TOUCH_DOCK_HEIGHT": "360",
        "DESKTOP_TOUCH_DOCK_PIN": "true"
      }
    }
  }
}

| Env var | Default | Notes | |---|---|---| | DESKTOP_TOUCH_DOCK_TITLE | (unset = off) | @parent walks the MCP process tree to find the hosting terminal — immune to title / branch / project changes. Or use a literal substring. | | DESKTOP_TOUCH_DOCK_CORNER | bottom-right | top-left / top-right / bottom-left / bottom-right | | DESKTOP_TOUCH_DOCK_WIDTH / HEIGHT | 480 / 360 | px ("480") or ratio of work area ("25%") — 4K/8K auto-adapts | | DESKTOP_TOUCH_DOCK_PIN | true | Always-on-top toggle | | DESKTOP_TOUCH_DOCK_MONITOR | primary | Monitor id from desktop_state({includeScreen:true}) | | DESKTOP_TOUCH_DOCK_SCALE_DPI | false | If true, multiply px values by dpi / 96 (opt-in per-monitor scaling) | | DESKTOP_TOUCH_DOCK_MARGIN | 8 | Screen-edge padding (px) | | DESKTOP_TOUCH_DOCK_TIMEOUT_MS | 5000 | Max wait for the target window to appear |

Input routing gotcha: when a pinned window is active (e.g. Claude CLI), keyboard(action='type') / keyboard(action='press') send keys to it, not the app you wanted to type into. Always call focus_window(title=...) before keyboard operations, then verify isActive=true via screenshot(detail='meta').

Screenshot cache (by-ref storage)

screenshot and the other visual results return a cheap screenshot://by-ref/{id} link to an image saved on disk instead of inlining the pixels every time, so routine look-act-confirm loops cost far fewer tokens. The cache bounds itself automatically and screenshot_query / screenshot_gc let you inspect and prune it. Tune the storage with:

| Env var | Default | Notes | |---|---|---| | DESKTOP_TOUCH_SCREENSHOTS_DIR | (per-user cache dir) | Pin the cache to a specific folder. If the default folder can't be created or written (e.g. corporate policy blocking new folders under your profile), the server auto-probes this → the runtime dir → an OS temp folder and uses the first writable one instead of giving up on the cache. | | DESKTOP_TOUCH_SCREENSHOT_MAX_COUNT | 200 | Keep at most this many captures in the cache. | | DESKTOP_TOUCH_SCREENSHOT_MAX_BYTES | 256 MiB | Cap the total cache size on disk. | | DESKTOP_TOUCH_SCREENSHOT_MAX_AGE_MS | (off) | Drop captures older than this many milliseconds (opt-in). | | DESKTOP_TOUCH_SCREENSHOT_AUTOPRUNE | on | Auto-trim the cache as new captures are saved. Set 0 to disable. | | DESKTOP_TOUCH_SCREENSHOT_MIN_EVICT_AGE_MS | 60000 | Never auto-evict a capture younger than this (ms), so a by-ref link you were just handed survives long enough to open even when another AI/process on the same PC is also capturing. 0 disables. |

Multi-monitor screenshots

screenshot(displayId=…) and screenshot(region=…) capture any monitor, including one placed left of or above the primary — those have negative desktop coordinates, and you pass them exactly as screenshot(detail='meta') reports them. screenshot() with no region is the primary monitor, as it has always been.

A region that cannot be captured comes back as RegionOutsideCapturableBounds rather than a raw Windows error, and the message says which of three things happened. The region may be on no monitor at all, which usually means the coordinates went stale because the window moved or closed — take a fresh screenshot and use the new numbers. It may overlap a monitor but stretch past the edge of the screen area, in which case the coordinates are fine and the region is simply too big: ask for a smaller one, or capture the window itself with screenshot(windowTitle=…). Or this server may be limited to the primary monitor, which the message says outright — along with why, because that decides the fix: if an env override pinned it, screenshot(windowTitle=…) normally still works on every monitor, whereas if the built-in capture module is missing then window capture usually needs that same module and fails too, so move the window onto the primary monitor or reinstall the server. Whole-screen capture and single-window capture are separate parts of that module, though, and a server can end up with one but not the other — so rather than working it out from the cause, read the message: it says plainly whether screenshot(windowTitle=…) is available on this server.

If Windows returns no pixels at all — a locked screen, a UAC prompt, a disconnected remote-desktop session — you get CaptureBackendFailed; capturing the window itself with screenshot(windowTitle=…) usually still works, because it reads through a different Windows API.

| Env var | Default | Notes | |---|---|---| | DESKTOP_TOUCH_CAPTURE_BACKEND | (unset = automatic) | Diagnostic override for the screen-capture path. Set to nutjs to force the older capture backend, which can only read the primary monitor — useful for isolating a capture problem. The server picks the backend once at startup, so change this in your MCP client config and restart. Any other value is ignored. |

Auto Perception (always-on)

Phase 4 privatizes the explicit perception_* tool family — the v0.12 Auto Perception layer attaches an attention signal to every desktop_state and desktop_act response automatically. Action tools also auto-guard when given a windowTitle. There is no longer a need to register / read / forget lenses manually.

# desktop_state always returns the attention signal
desktop_state() → {focusedWindow, focusedElement, modal, attention:"ok", ...}

# Action tools auto-guard when windowTitle is given:
keyboard({action:"type", text:"hello", windowTitle:"Notepad"})
→ post.perception:{status:"ok"}  // unsafe input blocked if guards fail

# When attention is dirty / stale / settling, refresh with desktop_state:
desktop_state()  // re-evaluates attention via Auto Perception

For advanced pinned-target workflows, the lensId parameter remains on action tools (keyboard, mouse_click, mouse_drag, click_element, browser_click, browser_navigate, browser_eval, desktop_act). Omit lensId for the normal Auto Perception path. The underlying registry, hot target cache, and sensor loop are unchanged; only the explicit perception_register / perception_read / perception_forget / perception_list tools were retired.


Mouse homing correction

When Claude calls screenshot(detail='text') to read coordinates and then mouse_click seconds later, the target window may have moved. The homing system corrects this automatically.

| Tier | How to enable | Latency | What it does | |------|--------------|---------|--------------| | 1 | Always-on (if cache exists) | <1ms | Applies (dx, dy) offset when window moved | | 2 | Pass windowTitle hint | ~100ms | Auto-focuses window if it went behind another | | 3 | Pass elementName/elementId + windowTitle | 1–3s | UIA re-query for fresh coords on resize |

# Tier 1 only (automatic)
mouse_click(x=500, y=300)

# Tier 1 + 2: also bring window to front if hidden
mouse_click(x=500, y=300, windowTitle="Notepad")

# Tier 1 + 2 + 3: also re-query UIA if window resized
mouse_click(x=500, y=300, windowTitle="Notepad", elementName="Save")

# Traction control OFF — no correction
mouse_click(x=500, y=300, homing=false)

The homing parameter is available on mouse_click, mouse_drag, and scroll. The cache is updated automatically on every screenshot(), desktop_discover(), focus_window(), and workspace_snapshot() call.

mouse_click image-local coords (origin + scale)

When you take a dotByDot screenshot with dotByDotMaxDimension, the response prints the origin and scale values. Instead of computing screen coords manually, copy them into mouse_click:

# Screenshot response:
#   origin: (0, 120) | scale: 0.6667
#   To click image pixel (ix, iy): mouse_click(x=ix, y=iy, origin={x:0, y:120}, scale=0.6667)

mouse_click(x=640, y=300, origin={x:0, y:120}, scale=0.6667, windowTitle="Chrome")
# Server converts: screen = (0 + 640/0.6667, 120 + 300/0.6667) = (960, 570)

This eliminates a whole class of off-by-one and scale bugs. Without origin/scale, x/y remain absolute screen pixels (unchanged behavior).


screenshot key parameters

detail="image"          — PNG/WebP pixels (default)
detail="text"           — UIA element JSON + clickAt coords (no image, ~100–300 tok)
detail="meta"           — Title + region only (cheapest, ~20 tok/window)
dotByDot=true           — 1:1 WebP; image_px + origin = screen_px
dotByDotMaxDimension=N  — cap longest edge (response includes scale for coord math)
grayscale=true          — ~50% smaller for text-heavy captures (code/AWS console)
region={x,y,w,h}        — with windowTitle: window-local coords (exclude browser chrome)
                          without: virtual screen coords
diffMode=true           — I-frame first call, P-frame (changed windows only) after (~160 tok)
ocrFallback="auto"      — detail='text' auto-fires Windows OCR on uiaSparse or empty

Recommended Chrome combo (50–70% data reduction):

screenshot(windowTitle="Chrome",
           dotByDot=true, dotByDotMaxDimension=1280, grayscale=true,
           region={x:0, y:120, width:1920, height:900})  # skip browser chrome

Recommended workflow:

workspace_snapshot()                     → full orientation (resets diff buffer)
screenshot(detail="text", windowTitle=X) → get actionable[].clickAt coords
mouse_click(x, y)                        → click directly, no math needed
screenshot(diffMode=true)                → check only what changed (~160 tok)

Security

Emergency stop (Failsafe)

Park the mouse in the top-left corner of the primary monitor (within 10px of 0,0) for 500ms continuously to trigger the emergency stop.

  • The trigger corner is on the primary monitor only. Areas that used to trigger the stop in older versions (monitors left of or above the primary) no longer do; if the cursor dwells there, a one-time balloon notification points you to the right corner.
  • While a tool call is running: the server exits (exit code 1) — the runaway-automation brake. A balloon notification and a diagnostic log entry (with cursor coordinates) record why it stopped. Only a call that is actually mid-flight triggers the exit; in the rare case where that call finishes during the ~1 second the notification takes, the server stays up instead and a follow-up balloon corrects the first one.
  • While idle: the server stays up and refuses new tool calls until the cursor leaves the corner. Background credential autofill (key_locker) is cancelled before any of its dialogs open — while you hold the corner, no credential prompt dialog appears and no credential is typed. It does not pick up again by itself: move the cursor away from the corner and run the command again.
  • Per-tool check: runs before every tool handler. Background monitor: 500ms polling as a backup for long-running operations. Trigger radius: 10px.
  • DESKTOP_TOUCH_FAILSAFE_HOLD_MS — dwell time in ms before the stop fires (default 500; 0 = fire immediately on corner entry).

Blocked operations

workspace_launch blocklist: cmd.exe, powershell.exe, pwsh.exe, wscript.exe, cscript.exe, mshta.exe, regsvr32.exe, rundll32.exe, msiexec.exe, bash.exe, wsl.exe are blocked. Script extensions (.bat, .ps1, .vbs, etc.) are rejected. Arguments containing ;, &, |, `, $(, ${ are also rejected.

keyboard(action='press') blocklist: Win+R (Run dialog), Win+X (admin menu), Win+S (search), Win+L (lock screen) are blocked.

PowerShell injection protection

All -like patterns in the UIA bridge PowerShell fallback path are sanitized with escapeLike(), which escapes wildcard characters (*, ?, [, ]) before they reach PowerShell. When the Rust native engine is active, PowerShell is not invoked for UIA operations.

Allowlist for workspace_launch

Shell interpreters are blocked by default. To allow specific executables, create an allowlist file:

File locations (searched in order):

  1. Path in DESKTOP_TOUCH_ALLOWLIST environment variable
  2. ~/.claude/desktop-touch-allowlist.json
  3. desktop-touch-allowlist.json in the server's working directory

Format:

{
  "allowedExecutables": [
    "pwsh.exe",
    "C:\\Tools\\myapp.exe"
  ]
}

Changes take effect immediately — no restart needed.


Mouse movement speed

All mouse tools (mouse_click, mouse_drag, scroll) accept an optional speed parameter:

| Value | Behavior | |---|---| | Omitted | Uses the configured default (see below) | | 0 | Instant teleport — setPosition(), no animation | | 1–N | Animated movement at N px/sec |

Default speed is 1500 px/sec. Change it permanently via the DESKTOP_TOUCH_MOUSE_SPEED environment variable:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_MOUSE_SPEED": "3000"
      }
    }
  }
}

Common values: 0 = teleport, 1500 = default gentle, 3000 = fast, 5000 = very fast.


Force-Focus (AttachThreadInput)

Windows foreground-stealing protection can prevent SetForegroundWindow from succeeding when another window (such as a pinned Claude CLI) is in the foreground. This causes subsequent keystrokes or clicks to land in the wrong window — a silent failure.

mouse_click, keyboard(action='type'), keyboard(action='press'), and terminal(action='send') all accept a forceFocus parameter that bypasses this protection using AttachThreadInput:

{
  "name": "mouse_click",
  "arguments": {
    "x": 500,
    "y": 300,
    "windowTitle": "Google Chrome",
    "forceFocus": true
  }
}

If the force attempt is refused despite AttachThreadInput, the response is ok:false with code: "ForegroundRestricted" (issue #202 unification — same shape as focus_window, keyboard, terminal_send, mouse_click). The action itself is suppressed so the keystrokes / click never land on the wrong window. Recover via focus_window's auto-escalate ladder before retrying. The legacy hints.warnings: ["ForceFocusRefused"] shape is no longer emitted.

Global default via environment variable:

{
  "mcpServers": {
    "desktop-touch": {
      "env": {
        "DESKTOP_TOUCH_FORCE_FOCUS": "1"
      }
    }
  }
}

Setting DESKTOP_TOUCH_FORCE_FOCUS=1 makes forceFocus: true the default for all four tools without changing each call.

Known tradeoffs:

  • During the ~10ms AttachThreadInput window, key state and mouse capture are shared between the two threads. In rapid macro sequences this can cause a race condition (rare in practice).
  • Disable forceFocus (or unset the env var) when the user is manually operating another app to avoid unexpected focus shifts.

Auto Guard

Action tools (mouse_click, mouse_drag, keyboard(action='type'/'press'), click_element, desktop_act, browser_click, browser_navigate) automatically guard each action when you pass windowTitle / tabId:

  • Verifies target window identity (process restart / HWND replacement detected)
  • Confirms click coordinates are inside the target window rect
  • Returns post.perception.status on every response — including failures — so the LLM can recover without a screenshot

Keyboard writes must name a destination. keyboard(action='type'/'press'/'sequence') requires either windowTitle or hwnd. Without one there is no target to guard, and the keys would land on whatever window is foreground at that instant — including one you just clicked into yourself. Such a call is refused with code:"DestinationRequired" before any key is sent, and a windowTitle that is empty or only spaces counts as no target at all. A window that has no title can be addressed by hwnd, but only while it is already the foreground window — keyboard focus and guarding cannot target a titleless window yet, so bring it forward with focus_window first if it is not in front. Such a call also comes back with a warning saying the input was delivered unguarded.

| Variable | Default | Meaning | |---|---|---| | DESKTOP_TOUCH_REQUIRE_DESTINATION | (unset = required) | Set to 0 to type into the current foreground window on purpose. The refusal becomes a warning on the response instead of an error — never a silent pass. | | DESKTOP_TOUCH_AUTO_GUARD | (unset = on) | Set to 0 to turn the whole guard layer off, the destination check included. |

Disabling auto guard — set DESKTOP_TOUCH_AUTO_GUARD=0 to restore v0.11.12 behavior (no auto guard):

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_AUTO_GUARD": "0"
      }
    }
  }
}

When auto guard is enabled (default), post.perception.status will be one of:

| Status | Meaning | |---|---| | ok | Guard passed — target verified | | unguarded | windowTitle not provided; action ran without guard | | ambiguous_target | Multiple windows matched; pass hwnd to name one exactly, or use a more specific title | | target_not_found | No window matched the given title | | identity_changed | Window was replaced (process restart / HWND change) | | blocked_by_modal | A modal dialog is in the way — dismiss it, then retry | | unsafe_coordinates | Click coordinates are outside the target window rect | | browser_not_ready | The browser tab is still loading — wait, then retry | | needs_escalation | Use browser_click or specify windowTitle | | destination_required | A keyboard write named no target. Refused before the guard runs, so it arrives as code:"DestinationRequired" with this status under context.guard rather than in post.perception — pass windowTitle or hwnd |

When unsafe_coordinates or identity_changed is returned, the response may include a suggestedFix.fixId. Pass that fixId to the relevant tool call to approve the recovery:

{ "name": "mouse_click",           "arguments": { "fixId": "fix-..." } }
{ "name": "keyboard(action='type')",         "arguments": { "fixId": "fix-...", "text": "hello" } }
{ "name": "click_element",         "arguments": { "fixId": "fix-..." } }
{ "name": "browser_click", "arguments": { "fixId": "fix-..." } }

The fix is one-shot and expires in 15 seconds. The server revalidates the target process identity before executing.


Diagnostic log

The server keeps an append-only log of events that never reach a tool response, at %USERPROFILE%\.desktop-touch-mcp\logs\diagnostic.log (one JSON object per line). It records crashes and slow calls, and — since the diagnostic log became the place to look when input lands in the wrong place — how each windowTitle was resolved and where each write went:

  • a resolve record per title lookup: how many windows matched, which one was picked, the ones that lost, and a flag when the terminal process-name fallback fired because nothing matched by title;
  • a dispatch_sink record per input dispatch — keyboard, terminal, scroll and desktop_act's background writes: which channel was used, which window it was addressed to, and which window was in the foreground at that moment;
  • a correlation id shared by all records from one tool call, so a resolution can be matched to the write it produced even when calls overlap.

If an input call ever seems to type into the wrong window, this is the file that says which window it picked and why. A record is written immediately before the write leaves the process, so a dispatch that is refused or fails first is not on record as having happened.

The log rolls over: once diagnostic.log passes 64 MiB it becomes diagnostic.log.1, and at most two rolled generations are kept. With one server running, the newest records are always in diagnostic.log, but when you are searching for something that happened a while ago, search diagnostic.log* rather than the one file. That glob also catches diagnostic.log.<pid>.rotating, which is where a server parks the live file for the moment it is being rolled. One of these left behind means a roll did not finish: the server was killed partway through, or the roll failed after the file was parked and the server could not put it back — it will not if a fresh diagnostic.log has been started in the meantime, and the move back can fail for the same reason the roll did. Nothing in it is lost. The pid in the name says which server it belonged to, and it is filed back into the numbered generations by the next roll — by that server if it is still running, and otherwise by any other server once the original has exited or, because process ids are reused, once the file has been parked for an hour — so a crashed server's log is not left sitting on disk forever.

Every record is measured against the limit before it is written, so a server left running for days rolls the file as it goes — there is no scheduled job, nothing to restart, and nothing to clean up by hand. A server sitting idle never rolls anything, because the check only runs when there is something to write.

The ceiling is a size, not an age. Three generations hold 192 MiB of records, and how far back that reaches depends entirely on how busy the machine is: on the install that prompted this limit, averaging roughly 170 MB a day, it is a little over one day. If you want to keep a particular incident, copy the file out rather than expecting to find it next week; if you would rather trade disk space for reach, raise DESKTOP_TOUCH_DIAGNOSTIC_LOG_MAX_BYTES.

Two situations go past that figure, and both are worth knowing about:

  • Several servers sharing one log. Every MCP client starts its own server, and by default they all write to the same file. Each tracks the bytes it has written itself and only re-measures the real file every few MB, so the live file can overshoot before one of them rolls it. The overshoot grows with the number of servers running, not without limit. Two servers can also roll at the same moment and step on each other's rename: that costs a generation, and can leave one server's newest record in diagnostic.log.1 instead of the live file. Grepping diagnostic.log* rather than the one file covers both.
  • A live file that cannot be renamed — held open by another program, or permission denied. Rotation then cannot happen and the log keeps growing at full speed; a parked .rotating file that is held open stops a roll the same way, and is checked before any numbered generation is touched. Nothing is lost — a roll that fails leaves the live file where it was, or at worst parked under the .rotating name above for a later roll to file — but this is the one case the limit does not cover, so it is not silent: a log_rotation_failed record is written into the log itself, once per stretch of failed rolls rather than once per line. It is the first thing to grep for if you find an oversized diagnostic.log after updating.

One record is never allowed to be larger than the file it lives in, so an event carrying an unusually large payload is written as a shortened stand-in: same kind, plus record_truncated, the original size, and a head field holding the first few KB of what it would have been.

| Variable | Default | Meaning | |---|---|---| | DESKTOP_TOUCH_RESOLVE_LOG_RAW | (unset = off) | Window titles and the titles you search for are recorded as a short hash plus their length, because a title can contain a file name, a mail subject, or a browser page title. Set to 1 to also record the text in clear (the hash stays, so a log with both is still readable end to end). | | DESKTOP_TOUCH_DIAGNOSTIC_LOG_DISABLE | (unset = on) | Set to 1 to stop writing the log entirely. | | DESKTOP_TOUCH_DIAGNOSTIC_LOG_PATH | (per-user log dir) | Write the log somewhere else. A symbolic link works: the roll follows it, so the link keeps pointing at the live log and the rolled generations appear beside the real file rather than beside the link. | | DESKTOP_TOUCH_DIAGNOSTIC_LOG_MAX_BYTES | 67108864 (64 MiB) | Roll the live log to diagnostic.log.1 once it passes this size. Two rolled generations are kept, so the log directory ordinarily holds about three times this value — see above for the two situations that go past it. A value below 1 MiB is raised to 1 MiB and one above 1 GiB is lowered to 1 GiB, and anything that is not a positive whole number falls back to the default — a typo here cannot switch rotation off in either direction, whether you mean bytes and write MiB or the other way round. To stop logging entirely, use DESKTOP_TOUCH_DIAGNOSTIC_LOG_DISABLE. |


Advanced response options

browser_eval Structured Mode

Pass withPerception: true to receive a structured JSON response with post.perception instead of raw text:

{ "name": "browser_eval", "arguments": { "expression": "document.title", "withPerception": true } }

Returns { ok: true, result: "...", post: { perception: { status: "ok", ... } } }.

mouse_drag Cross-Window Guard

mouse_drag now guards both start and end coordinates. Drags that cross window boundaries (or reach the desktop wallpaper) are blocked by default. To allow intentional cross-window or range-selection drags:

{ "name": "mouse_drag", "arguments": { "startX": 100, "startY": 100, "endX": 900, "endY": 900, "allowCrossWindowDrag": true } }

Performance (v0.15 — Rust Native Engine)

The Rust native engine (@harusame64/desktop-touch-engine) replaces PowerShell process spawning with direct COM calls over a persistent MTA thread. It loads automatically as a .node addon — no configuration needed.

UIA Benchmark (vs PowerShell baseline)

| Function | Rust Native | PowerShell | Speedup | |---|---|---|---| | getFocusedElement | 2.2 ms | 366 ms | 163.9× | | getUiElements (Explorer, ~60 elements) | 106.5 ms | 346 ms | 3.3× | | Weighted average | | | ~82× |

Image Diff Benchmark (SSE2 SIMD)

| Function | Rust (SSE2) | TypeScript | Speedup | |---|---|---|---| | computeChangeFraction (1920×1080) | 0.26 ms | 3.8 ms | ~15× | | dHash (perceptual hash) | 0.09 ms | 1.2 ms | ~13× |

Architecture

Claude CLI / MCP Client
    │  stdio or HTTP (MCP protocol)
    ▼
desktop-touch-mcp (TypeScript)
    │
    ├── Rust Native Engine (.node addon)          ← NEW in v0.15
    │   ├── UIA: 13 functions via napi-rs + windows-rs 0.62
    │   │   └── Dedicated COM thread (MTA) + batch BFS algorithm
    │   └── Image: SSE2 SIMD pixel diff + perceptual hashing
    │
    └── PowerShell Fallback (automatic)
        └── Activates transparently if .node is unavailable

Why getUiElements is 3.3× (not 160×)

The 160× speedup on getFocusedElement comes from eliminating PowerShell process startup (~200 ms) and .NET assembly loading. For getUiElements, the bottleneck shifts to the UIA provider inside the target application (e.g., Explorer) — it must enumerate its UI tree regardless of who asks. The Rust engine uses a batch BFS algorithm (FindAllBuildCache + TreeScope_Children) that minimizes cross-process RPC calls and supports maxElements early exit, making it dramatically faster on large trees (VS Code, browsers with 1000+ elements).


UI Operating Layer (V2)

Status: Default ON since v0.17. desktop_discover and desktop_act are available out of the box.

V2 introduces two new tools that replace coordinate-based clicking with entity-based interaction:

| Tool | Description | |---|---| | desktop_discover | Observe a window or browser tab. Returns interactive entities with leases — no raw screen coordinates. Supports UIA (native), CDP (browser), terminal, and visual GPU lanes. | | desktop_act | Interact with an entity returned by desktop_discover. Validates the lease before executing. Returns a semantic diff (entity_disappeared, modal_appeared, focus_shifted, …). When diffUnchecked is present, the diff did not look for the kinds it lists, so their absence from the diff does not mean they did not happen. On visual-only targets a successful act can bundle a roiCapture (a PNG crop of the changed region + a lease-less next-target preview) so you confirm the result and find the next target in one call — controlled by returnCapture (on-change, the default on a visible change; never to suppress; always to force). |

Clicking — priority order

When multiple tools could perform the same click, prefer them in this order:

  1. browser_click(selector) — Chrome / Edge over CDP (stable across repaints)
  2. desktop_act(lease) — native windows, dialogs, visual-only targets (entity-based; use after desktop_discover)
  3. click_element(name | automationId) — native UIA fallback when desktop_act returns ok:false
  4. mouse_click(x, y) — pixel-level last resort (origin + scale from dotByDot screenshots only)

Disabling V2 (kill switch)

To hide desktop_discover / desktop_act from the tool catalog, add the disable flag and restart:

{
  "mcpServers": {
    "desktop-touch": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@harusame64/desktop-touch-mcp"],
      "env": {
        "DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2": "1"
      }
    }
  }
}

All V1 tools continue to work without interruption — no reinstall required. Remove the env entry and restart to re-enable.

Flag semantics (exact-match: only the literal string "1" counts):

| DISABLE_FUKUWARAI_V2 | V2 state | |---|---| | unset / not "1" | ON (default) | | "1" | OFF (kill switch) |

Removed: DESKTOP_TOUCH_ENABLE_FUKUWARAI_V2

This was the opt-in switch in v0.16.x. V2 is on by default since v0.17, so the flag no longer has any effect and is safe to delete from your config. To turn V2 off, set DESKTOP_TOUCH_DISABLE_FUKUWARAI_V2=1.

Recovery when V2 fails

If desktop_act returns ok: false, read reason and follow the built-in recovery hints in the tool description. Common paths:

  • lease_expired / *_mismatch / entity_not_found → re-call desktop_discover
  • modal_blocking → response.blockingElement (when present) carries { name, role, automationId?, hwnd? }. With role: "dialog" the blocker is a separate dialog window that has disabled the target's window, and hwnd is it — desktop_discover with target.hwnd = blockingElement.hwnd, answer it, then retry. Any other role: a window the discover snapshot holds, which the OS could not confirm — desktop_discover with target.hwnd = blockingElement.hwnd when hwnd is present, otherwise click_element(name=blockingElement.name); then re-discover the original target and act on the new lease (the same lease is refused again)
  • entity_outside_viewport → the element moved off screen: scroll / scroll(action='to_element') when it scrolled out of its own window, or re-call desktop_discover when the window itself moved or closed
  • origin_window_not_visible → focus_window(windowTitle) to restore the minimised / hidden window, then re-call desktop_discover
  • coordinate_outside_reachable_bounds → the target is on no connected monitor — usually stale coordinates: re-run desktop_discover. (Without the built-in Windows input module, only the primary monitor is reachable; the message says so.)
  • cursor_placement_blocked → the pointer could not be placed there (an app is holding the cursor, or the session is not interactive), so nothing was clicked: free the cursor or reconnect the session, or use click_element (UIA invoke, cursor-free)
  • keyboard_target_unsafe → nothing was typed: the characters would have gone to a different control or window, or to a read-only control, or the field you named (or its window) is disabled (if_unexpected.detail says which; for disabled, wait out whatever disabled it and re-discover). Put the focus on the field you named, then type again — not through a foreground keyboard type. By title: desktop_act with action='click'. By handle: re-discover by title first — except for a common dialog, which a title resolves to by handle as well, where nothing here can focus its text field yet. For another window: focus_window first
  • executor_failed → fall back to click_element / mouse_click / browser_click

For desktop_discover warnings (visual_provider_unavailable, visual_provider_warming, cdp_provider_failed, …), the coordinate-based tools (screenshot(detail='text'), click_element, mouse_click, terminal, …) remain available as an escape hatch.


Known limitations

| Limitation | Detail | Workaround | |---|---|---| | Games / video players may return black or hang in PrintWindow capture | DirectX fullscreen apps may not redraw under PW_RENDERFULLCONTENT. Window-targeted screenshot(detail='image') already falls back to BitBlt automatically when PrintWindow returns no data or an all-black + zero-variance frame, but DirectX surfaces that hang the call don't surface as fallback. | Retry with screenshot({mode:'background', fullContent:false}) to switch to the legacy PrintWindow flag; if still black, the BitBlt fallback path (default mode='normal') will at least return the on-screen rect — hints.captureFallbackReason will say printwindow-all-black | | UIA call overhead | ~2 ms (focus) / ~100 ms (tree) via Rust native engine; ~300 ms via PowerShell fallback | Rust engine loads automatically; workspace_snapshot uses a 2 s timeout internally | | Chrome / WinUI3 UIA elements are empty | Chromium exposes only limited UIA | screenshot(detail='text') auto-detects Chromium and falls back to Windows OCR (hints.chromiumGuard=true). For richer DOM access use browser_open + browser_locate | | Chromium title-regex misses when sites rewrite document.title | Guard relies on the - Google Chrome suffix being present; some sites push it off the end of a long title | Title is treated as plain Chrome (UIA runs). OCR path is still reachable via ocrFallback='always' or when UIA returns <5 elements (uiaSparse) | | browser_* CDP tools need Chrome launched with --remote-debugging-port | If Chrome is already running on the default profile without the flag, browser_open fails. The CDP