SLOPSHOPPER

mobile-mcp

Mobile automation for Claude Code: the mobile-mcp server, and /mobile-mirror, a live, clickable device screen in a pane.

newpanecommandtoast
v1.0.7Apache-2.0updated 2026-10-08Yomiamy/jev-mobile-mcp/plugin
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · mobile-mcp
│ ┃ mobile-mirror ✕ › fix the failing auth test and add an audit log call │ ┃ Waiting for the first frame… │ ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /mobile-mirror │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · mobile-mirror
Waiting for the first frame…
README

jev-mobile-mcp

English | 繁體中文

Mobile agent testing by jev decision and ocr detection

An MCP server built on mobile-next/mobile-mcp that lets AI agents drive Android / iOS devices, emulators and simulators. Every upstream feature is kept; this repo adds an OCR layer between the accessibility tree and screenshots, so agents need screenshots far less often to find what to tap.

Why OCR

Upstream locates elements in only two ways:

  1. mobile_list_elements_on_screen: reads the accessibility tree and returns refs, coordinates and labels. Fast, cheap, accurate.
  2. mobile_take_screenshot: when the tree lacks the target, the model eyeballs a screenshot and converts positions by the scale ratio. Slow, token-heavy, and only as accurate as the model's vision.

Many screens never put their text into the accessibility tree: a Flutter Drawer without semantics, canvas-drawn UI, text baked into images. Those used to fall straight through to screenshots.

This repo adds OCR on two paths, so a screenshot becomes the last resort:

mobile_tap (with a TypeSafe key): OCR first. OCR costs about 1–1.5 s, while reading the tree of a Flutter debug build costs 6–10 s, so the tree is read only when OCR is not enough.

mobile_tap(target)
        │
        ▼
OCR (screenshot + Vision)          ← read first; Jev picks the text
        │ no confident match
        ▼
tree + OCR merged                  ← tree read only now; Jev asked again
        │ still no confident match
        ▼
nothing tapped, candidates + element list returned → agent picks one and clicks it

mobile_list_elements_on_screen: tree first, OCR on request. list is the most frequent call, so OCR is never turned on by the server.

list_elements_on_screen            ← accessibility tree (default)
        │ target text not found
        ▼
list_elements_on_screen(ocr: true) ← tree + OCR supplement (added here)
        │ target is an icon, OCR cannot read it
        ▼
take_screenshot                    ← model eyeballs it (last resort)

Changes

FileChange
src/ocr.tsThe OCR itself: screenshot → macOS Vision → screen coordinates → dedupe against the tree
src/server.tsNew ocr parameter on mobile_list_elements_on_screen; an empty tree hints to retry with ocr: true
src/format-elements.tsElements without a ref (OCR results, legacy mode) get a center point tap=x,y
skills/mobile-automation/SKILL.mdTells the agent when to use OCR
src/compact-elements.tsCompacts the element list for every user, with or without a TypeSafe key: inside the current viewport only (see Screen bounds and rotation), no empty containers, a repeated multi-line (merged) label once; about 60% shorter
src/jev.tsJev element choice behind mobile_tap (optional, see below)

The ocr parameter

// mobile_list_elements_on_screen
{ "device": "Pixel_6", "ocr": true }   // defaults to false

OCR results are appended after the tree elements as OcrText:

@e65 Button at=11,139 size=126x126
@e66 Header text="尋找餐廳" label="尋找餐廳" at=189,162 size=252x79
OcrText text="關鍵字過濾" at=150,623 size=216x48 tap=258,647
OcrText text="我的位置" at=118,1063 size=202x50 tap=219,1088
OcrText text="設定" at=146,1360 size=91x46 tap=192,1383
  • tap=x,y is the precomputed center; pass it straight to mobile_click_on_screen_at_coordinates. at= is the top-left corner, so the model no longer has to compute the center itself.
  • Coordinates are already screen coordinates: Vision returns boxes normalized to the screenshot, and the screenshot covers the whole screen, so multiplying by the current viewport is enough. Screenshot scaling and iOS points vs. pixels do not matter, and a rotated screen maps correctly (see Screen bounds and rotation).
  • Dedupe: an OCR box is dropped when its center lies inside a tree element with the same text (letters and digits only, ignoring ·/•, spaces and the like), so only what the tree lacks remains.
  • OCR elements have no ref and can only be tapped by coordinates.

When it is used

The agent decides from the tool description; the server never turns OCR on by itself for list (it is the most frequent call, and running OCR every time would slow everything down). mobile_tap is different: it reads OCR first on every tap (see below). For list:

  • The text to tap is missing from the previous listing → list again with ocr: true.
  • When the accessibility tree is completely empty, the result includes a Retry with ocr: true hint.

Design, trade-offs and full test records: spec · plan.

Screen bounds and rotation

The element list keeps only what lies inside the current viewport, and OCR maps its boxes onto the same viewport. The viewport is decided in this order:

  1. The window root in the dump: an element at 0,0 whose size equals the screen size or its swap, such as android:id/content or Flutter's root box.
  2. The orientation the robot reports, used to turn the reported size the right way.
  3. The reported size as is.

The robot's orientation is only a fallback because it is not reliable: on a Pixel 6 emulator with auto-rotate on, Chrome in landscape (ROTATION_270, 2400x1080) is still reported as portrait 1080x2400 by mobilecli. Trusting it dropped all 36 elements past x=1080; with the window root all of them are kept.

Tap by description with Jev (optional)

With a TypeSafe API key, the server also registers mobile_tap. The agent describes the target instead of reading the whole element list; Jev, TypeSafe's System One model, picks the element (the same approach as jev-ultrafast).

// mobile_tap
{ "device": "Pixel_6", "target": "menu button at the top left" }
// → Tapped @e65 Button "" at 74,202 (confidence 0.62, from OCR + accessibility tree)
{ "device": "Pixel_6", "target": "關鍵字過濾" }
// → Tapped OcrText "關鍵字過濾" at 257,663 (confidence 0.81, from OCR)
  1. Take a screenshot, read its text with OCR and ask Jev which text matches target (one Choice question: one option per element, plus NONE). The orientation comes from the screenshot itself. OCR is cheap (about 1–1.5 s), while reading the accessibility tree takes 6–10 s on a Flutter debug build (see below).
  2. No match or low confidence → read the accessibility tree and compact it (drop off-screen elements and empty containers, keep a multi-line (merged) label repeated by child nodes only once), merge in the OCR elements from step 1 and ask once more. OCR does not run a second time.
  3. When the server does not run on macOS (no OCR), or the screenshot or OCR fails, go straight to step 2 with the accessibility tree only.
  4. Still no confident match → nothing is tapped; the closest candidates are returned together with the element list Jev chose from (the same format as mobile_list_elements_on_screen, OCR text included), so the agent can pick a ref or tap= coordinates for mobile_click_on_screen_at_coordinates without reading the screen again.
  5. If the screen changed between reading and tapping (stale ref), read it again, including the viewport, and retry once.

An element with a ref is tapped by ref; one without (an OCR element) is tapped at the center of its visible part, kept inside the viewport. Jev can only pick an observed element, so the model never makes up coordinates. The agent sends one short phrase instead of reading 3,000–7,000 characters of element list per step.

Setup: without TYPESAFE_API_KEY, mobile_tap is not registered and nothing is sent to TypeSafe. Put the server name before -e, otherwise -e swallows the name as another variable:

claude mcp add jev-mobile-mcp -e TYPESAFE_API_KEY=<your key> -- npx -y github:Yomiamy/jev-mobile-mcp#main

For a shared .mcp.json, reference the variable instead of committing the key: "env": { "TYPESAFE_API_KEY": "${TYPESAFE_API_KEY}" }. TYPESAFE_MODEL overrides the model (default jev-latest).

Privacy: every mobile_tap sends the on-screen text read by OCR to TypeSafe, and the text of the accessibility tree elements too when the tree is read. Screens can contain personal data such as account emails; enable it only where that is acceptable.

Limits:

  • Icons missing from the accessibility tree cannot be found (OCR reads text only).
  • Icon targets and native screens (the launcher's tree reads in 0.7 s) pay about 1–1.5 s more per tap: OCR runs first, is unsure, and then the tree is read.
  • A text target that OCR alone matches confidently is tapped by coordinates, even when the tree has a ref for it (e.g. a dialog's "取消"), so it loses the stale-ref protection below.
  • Two identical targets on screen (e.g. the same app icon on the home screen and in the dock) split the probability and are refused; describe the target more precisely, e.g. by position.
  • Buttons without a label are picked by position only, with lower confidence (0.60–0.66 in the field test). A tooltip / Semantics(label:) in the app fixes that.
  • The confidence threshold (0.5) is hand-picked and should be tuned on recorded runs.
  • Only taps by ref can be rejected when the screen changed (mobilecli reports a ref that is not on the current screen; whether it catches every change depends on how mobilecli numbers refs, which is not confirmed). OCR elements and legacy robots are tapped by coordinates, which are not re-checked.

Design, trade-offs and full test records: spec · plan.

Field test

A 10-step flow on the Flutter app "FindRestaurant" on a Pixel 9a emulator (Android 17): terminate all apps → tap the app on the launcher → wait for load → scroll 100 px → open side menu → "關鍵字過濾" → "取消" → open side menu → "我的位置" → wait for reload. Budget 300 s per run, at most 2 retries per step. Five runs per server and per model, each run in a fresh Claude Code subagent (Opus 5.5, or Haiku 5.5 for the Haiku column) so context does not accumulate across runs.

upstream mobile-mcp 1.0.8 ³jev-mobile-mcp (Opus 5.5)jev-mobile-mcp (Haiku 5.5) ⁴
Passed5 / 55 / 53 / 5
Avg time137.2 s70.0 s (−49%)180.0 s (+31%)
Avg model requests per run (Claude API) ²23.616 (−32%)24.6 (+4%)
Avg input-equivalent tokens per run ¹295.3k166.3k (−44%)270.2k (−8%)
Steady state, runs 3–5 ¹202k–336k141k–152k244k–372k
Avg output tokens per run1,575517 (−67%)596 (−62%)
Retries, all runs2 (launcher tap, dialog "取消")08

¹ From the usage of every model request in the subagent transcript, priced relative to plain input: cache read × 0.1 + cache write × 1.25 + input. Raw totals are much larger (1.2–2.6 M tokens per run) because every request resends the whole context, about 63k of which is fixed overhead (system prompt, tool definitions, project rules) before the first step. Runs that write that context to the cache (run 1 of each server, and upstream run 4) are higher.

² One request to the Claude Messages API, i.e. one model turn; counted as distinct assistant messages in the transcript. Not the number of MCP tool calls or device actions: one request can issue several tool calls, and one mobile_batch_commands can run many device steps.

³ Upstream 1.0.8, installed as the Claude Code plugin (/plugin install mobile-mcp@mobile-mcp), run on 2026-10-03. The jev-mobile-mcp (Opus 5.5) column was measured earlier and not rerun.

⁴ jev-mobile-mcp with Haiku 5.5 as the subagent model, run on 2026-10-08 with the same prompt and method. The five runs were executed one after another on Pixel_9a. Runs 2 and 4 failed partway (step 5 and step 4), so their times are partial and their figures are included only so the average covers all five runs. Run 4 was marked FAIL for the step 4 scroll distance, but runs 1, 3 and 5 had the same kind of error (a 100 px swipe moved the list about 300–375 px) and were marked PASS, so the pass count depends on that judgment.

  • Where the saving comes from: the four text targets ("FindRestaurant", "關鍵字過濾", "取消", "我的位置") were each tapped by one mobile_tap (OCR, confidence 0.91–0.99), with no screenshot to locate them first. Fewer model requests means fewer resends of the context, which dominates the cost.
  • Why upstream took longer: on this Flutter screen mobile_list_elements_on_screen usually did not show the open side menu, so the agent took screenshots and tapped coordinates read off them; screenshots often still showed the previous screen and had to be taken again. Every run also listed the 22–25 installed apps and terminated each one, since there is no tool for listing running apps.
  • Launcher tap: with upstream, the agent tapped the icon by ref; the first tap was ignored in 1 of 5 runs and worked when retried at the icon's coordinates. mobile_tap hit the label below the icon (923,1524) and launched the app on the first tap every time. Observed, root cause not verified.
  • Wrong element by ref: in upstream run 3, tapping the dialog's "取消" by ref (@e77) opened the Flutter Inspector instead of closing the dialog; the retry at screenshot coordinates worked. Root cause not verified.
  • The hamburger button has no label, so both servers tapped it by coordinates or ref.
  • Not a strictly equal comparison: the jev-mobile-mcp prompt gave the hamburger button's coordinates, the upstream prompt did not, and the two columns were measured at different times; part of the difference may come from that.
Per-run data

upstream mobile-mcp 1.0.8 ³:

RunTimeModel requestsCache readCache writeOutputInput-equivalent ¹Retries
197 s252,157,10586,1161,378323.4k0
2219 s242,380,58653,5232,656305.0k0
3189 s282,543,59843,9251,298309.3k2 (launcher tap; "取消" by ref opened the Flutter Inspector)
493 s221,977,021110,9591,229336.4k0
588 s191,592,62534,2861,312202.2k0

jev-mobile-mcp:

RunTimeModel requestsCache readCache writeOutputInput-equivalent ¹Retries
170 s161,179,74585,262475224.6k0
278 s171,331,51423,022619162.0k0
377 s161,242,28821,766485151.5k0
461 s151,155,36820,749511141.5k0
564 s161,243,85621,952494151.9k0

Plain input was 30–56 tokens per run and is left out.

Haiku 5.5 ⁴:

RunResultTimeModel requestsCache readCache writeOutputInput-equivalent ¹Retries
1PASS168 s211,690,97796,790331290.1k2 (step 5 by ref after a 0.44-confidence tap; step 7 stale ref, re-listed)
2FAIL at step 5142 s151,177,54128,342678153.2k2 (step 5: tap and ref click did not open the drawer; coordinate retry also failed)
3PASS197 s242,027,31632,539382243.5k0 (step 5 drawer not in element list, confirmed by screenshot)
4FAIL at step 4225 s353,163,45044,589623372.2k3 (step 4: 100 px swipe moved ~342 px, then two corrective swipes; step 5: stale ref, coordinate tap worked)
5PASS168 s282,439,72538,203966291.8k1 (step 2: grid icon tap missed, hotseat icon worked)

Haiku notes:

  • Step 4 swipe: the swipe tool did not honor the 100 px distance in any Haiku run (about 300–375 px). The single Opus 5.5 spot check above showed the same, so the "scroll 100 px" step is approximate for both.
  • Step 2 launch: the launcher grid icon tap failed in run 5 and the hotseat icon worked.
  • Stale refs: refs from an earlier listing went stale after the screen changed (runs 1 and 4) and were not reused.
Re-run on 2026-10-08 (jev-mobile-mcp, single run)

All 10 steps passed on the same Pixel 9a emulator, run directly in the main session rather than a subagent, so no usage data was captured. Token and request figures are therefore not comparable and are left out; the upstream column was not rerun.

  • Time: not timed precisely. The device clock read 9:08 when the app's first screen appeared and 9:10 at the end, so the run fit well within the 300 s budget.
  • Retries (3, all within the 2-per-step limit):
  • Launcher tap (mobile_tap, step 2): TypeSafe returned HTTP 529, then a timeout, both with nothing tapped; the third attempt matched the icon label by OCR (confidence 0.60).
  • Hamburger button (step 5): the first tap by coordinates did not open the drawer; the tap by ref (@e72) did.
  • Load wait (step 10): after "我的位置" the list showed skeleton placeholders for about 30 s while a network request completed.

Why reading the screen is slow on Flutter debug builds

For a debuggable Flutter app, mobilecli (1.0.16) does not use the Android accessibility dump. It walks the whole render tree over the Dart VM service, with 10–25 calls per render object, including rows rendered off screen and routes behind the current one. There is no option to turn this off.

Casedump ui time
Native screen (launcher)0.7 s
Flutter debug build (VM service walk)6.3–10.2 s

A profile or release build makes mobilecli fall back to the accessibility dump, which should be much faster and also exposes button tooltips as labels; this is not measured yet.

Limitations

  • macOS only: uses the built-in Vision framework (called through osascript JXA, so nothing to compile and no new npm dependency). On other platforms ocr: true returns an error; everything else is unaffected.
  • Text only, no icons: icon-only buttons (heart, hamburger menu) still need a screenshot or known coordinates. The real fix is a tooltip / Semantics(label:) in the app.
  • Noise: icons, star ratings and low-contrast text may be misread (e.g. $$ as $s). Results are not filtered by confidence, because Vision's confidence cannot tell noise from valid targets; the model picks by meaning.
  • Slower: each ocr: true adds roughly 1–1.5 seconds (including the screenshot).
  • Rotation: landscape is tested on an Android emulator only; iOS simulators are not tested yet. Without a full-screen window root in the dump (an app that is not edge-to-edge, split screen, legacy WDA which filters root types out), the viewport falls back to the reported orientation, which mobilecli and the legacy Android robot (user_rotation) can get wrong.

Installation

From GitHub (for teams)

Requires read access to this repo. prepare builds on install.

claude mcp add jev-mobile-mcp -- npx -y github:Yomiamy/jev-mobile-mcp#main

Or commit a .mcp.json at your project root so teammates are prompted to enable it when they open the project:

{
  "mcpServers": {
    "jev-mobile-mcp": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "github:Yomiamy/jev-mobile-mcp#main"]
    }
  }
}

Updating: npx reuses its cached install and does not pick up new commits on the branch. After pushing a new version, delete the cache and reconnect:

grep -l 'jev-mobile-mcp.git' ~/.npm/_npx/*/package-lock.json   # find the cache dir
rm -rf ~/.npm/_npx/<that-dir>
# then in Claude Code: /mcp → jev-mobile-mcp → Reconnect

Local development

npm ci && npm run build
claude mcp add jev-mobile-mcp -- node /path/to/jev-mobile-mcp/lib/index.js

After changing src/, run npm run build and reconnect in /mcp.

If the upstream mobile-mcp is installed too, both expose the same tool names. Disable one while testing so the agent does not call the wrong version.

Syncing with upstream

Upstream is tracked through an upstream remote and merged in, keeping its history. See .claude/skills/gen-sync-mobile-mcp/SKILL.md for the procedure.

License

Upstream mobile-mcp is licensed under Apache-2.0; its original license is kept in LICENSE-mobile-mcp.

Source 5 files
hooks/register.tsx 243 lines
1import { atom, read, update } from "claude-code";
2import type { EngineInterface, Register } from "claude-code";
3
4import type { Shot, Target } from "../types";
5import { keysToActions } from "./keys";
6import { fitInside, pngSize } from "./png";
7
8// the key of the MCP server in this plugin's manifest
9const MCP_SERVER = "mobile-mcp";
10const COMMAND = "mobile-mirror";
11const PANE = "mobile-mirror";
12const VIEW_KEY = "view";
13const TAP_KEY = "tap";
14// frames rotate through a few files, so the terminal never reads one being rewritten
15const FRAME_FILES = 3;
16// large enough that the terminal never upscales it on a Retina pane; smaller is faster
17const SHOT_MAX_SIZE = 1600;
18const RETRY_AFTER_ERROR_MS = 1000;
19// the toolbar row, plus the URL field's row while it shows
20const TOOLBAR_ROWS = 1;
21const URL_FIELD_ROWS = 1;
22
23// kept in session state, not module variables, so they survive a reload of the mod
24const target = atom({ plugin: "mobile-mcp", key: "target" } as const, null as Target | null);
25// bumped to stop the running frame loop: a loop runs only while it holds the current value
26const streamId = atom({ plugin: "mobile-mcp", key: "streamId" } as const, 0);
27const isAskingUrl = atom({ plugin: "mobile-mcp", key: "isAskingUrl" } as const, false);
28const shot = atom({ plugin: "mobile-mcp", key: "shot" } as const, { file: "", generation: 0, width: 1, height: 1 } as Shot);
29
30type Device = { id: string; platform: string; state: string };
31type TapMessage = { x: number; y: number };
32type KeysMessage = { keys: string[] };
33// the runtime has Uint8Array.fromBase64; the TS lib does not declare it yet
34type Base64Decoder = { fromBase64(base64: string): Uint8Array };
35type McpResult = Awaited<ReturnType<EngineInterface["mcp"]["call"]>>;
36
37function firstText(content: unknown): string {
38	const block = (content as { type: string; text?: string }[]).find(b => b.type === "text");
39
40	return block?.text ?? "";
41}
42
43function frameFile(generation: number): string {
44	return `/tmp/mobile-mirror-${generation % FRAME_FILES}.png`;
45}
46
47// calls a mobile-mcp tool by the name this session runs the server under, connecting it on first use
48async function callMobileMcp($: EngineInterface, tool: string, args: Record<string, unknown> = {}): Promise<McpResult> {
49	const connection = await $.mcp.connect(MCP_SERVER);
50	if (!connection.isConnected) {
51		throw new Error(`mobile-mcp is not connected: ${connection.message}`);
52	}
53
54	return $.mcp.call(connection.server, tool, args);
55}
56
57// runs a toolbar action, and toasts only when it fails
58async function runAction($: EngineInterface, what: string, tool: string, args: Record<string, unknown>): Promise<void> {
59	try {
60		const result = await callMobileMcp($, tool, args);
61		if (result.isError) {
62			$.ui.toast(`${what} failed: ${firstText(result.content)}`);
63		}
64
65	} catch (error) {
66		$.ui.toast(`${what} failed: ${String(error)}`);
67	}
68}
69
70function pickDevice(devices: Device[], wanted: string | undefined): Device | undefined {
71	const online = devices.filter(d => d.state === "online");
72	if (wanted) {
73		return online.find(d => d.id === wanted);
74	}
75
76	return online.find(d => d.platform === "android") ?? online[0];
77}
78
79export const register: Register = on => {
80	// key batches are sent one after another, so typed text keeps its order
81	let typing: Promise<void> = Promise.resolve();
82
83	on("session.start", async ($, e, next) => {
84		await $.command.register({ name: COMMAND, description: "Mirror a device screen: live picture, taps, keys and buttons", argumentHint: "[device-id]" });
85
86		return next(e);
87	});
88
89	on("command.run", { command: COMMAND }, async ($, e) => {
90		const wanted = e.args.trim() || undefined;
91		const listed = await callMobileMcp($, "mobile_list_available_devices");
92		const { devices } = JSON.parse(firstText(listed.content)) as { devices: Device[] };
93		const device = pickDevice(devices, wanted);
94		if (!device) {
95			const online = devices.filter(d => d.state === "online").map(d => d.id);
96			return { text: `No online device${wanted ? ` "${wanted}"` : ""}. Online: ${online.join(", ") || "none"}` };
97		}
98
99		const sized = await callMobileMcp($, "mobile_get_screen_size", { device: device.id });
100		const match = /(\d+)x(\d+)/.exec(firstText(sized.content));
101		if (!match) {
102			return { text: `Could not read the screen size of ${device.id}.` };
103		}
104
105		await update($, target, () => ({
106			id: device.id,
107			platform: device.platform === "android" ? "android" : "ios",
108			screenWidth: Number(match[1]),
109			screenHeight: Number(match[2]),
110		}));
111		await update($, isAskingUrl, () => false);
112		const id = await update($, streamId, n => n + 1);
113
114		const captureFrame = async (generation: number): Promise<boolean> => {
115			const file = frameFile(generation);
116			const saved = await callMobileMcp($, "mobile_save_screenshot", { device: device.id, saveTo: file, maxSize: SHOT_MAX_SIZE });
117			if (saved.isError) {
118				return false;
119			}
120
121			const { base64 } = await $.fs.read(file, { as: "bytes" });
122			const size = pngSize((Uint8Array as unknown as Base64Decoder).fromBase64(base64));
123			const current = await read($, shot);
124			// a new shape (first frame, rotation) needs a full redraw; otherwise swap the pixels in place
125			if (current.generation === 0 || current.width !== size.width || current.height !== size.height) {
126				await update($, shot, () => ({ file, generation, ...size }));
127			} else {
128				await $.ui.blit({ requestId: PANE, key: VIEW_KEY, source: { file, format: "png", generation } }).catch(() => {});
129			}
130
131			return true;
132		};
133
134		const stream = async () => {
135			// the next frame is asked for as soon as the previous one is shown
136			for (let generation = 1; id === (await read($, streamId)); generation++) {
137				const isShown = await captureFrame(generation).catch(() => false);
138				if (!isShown) {
139					await $.clock.sleep(RETRY_AFTER_ERROR_MS);
140				}
141
142			}
143		};
144
145		await update($, shot, () => ({ file: "", generation: 0, width: 1, height: 1 }));
146		void stream();
147		await $.ui.open({ id: PANE, title: `mobile-mirror · ${device.id}` });
148
149		return { text: `Mirroring ${device.id}. Click the picture to tap, then type to send keys.` };
150	});
151
152	on("ui.message", async ($, e) => {
153		const device = await read($, target);
154		if (e.element !== TAP_KEY || !device) {
155			return {};
156		}
157
158		if ("keys" in (e.data as object)) {
159			const actions = keysToActions((e.data as KeysMessage).keys, device.platform);
160			typing = typing.then(async () => {
161				for (const action of actions) {
162					if (action.type === "text") {
163						await runAction($, "Typing", "mobile_type_keys", { device: device.id, text: action.text, submit: false });
164					} else {
165						await runAction($, `Press ${action.button}`, "mobile_press_button", { device: device.id, button: action.button });
166					}
167
168				}
169			});
170
171			return {};
172		}
173
174		const tap = e.data as TapMessage;
175		const x = Math.round(tap.x * device.screenWidth);
176		const y = Math.round(tap.y * device.screenHeight);
177		await runAction($, "Tap", "mobile_click_on_screen_at_coordinates", { device: device.id, x, y });
178
179		return {};
180	});
181
182	on("ui.close", async ($, e, next) => {
183		if (e.id === PANE) {
184			await update($, streamId, n => n + 1);
185		}
186
187		return next(e);
188	});
189
190	on("ui.render", { component: "Pane", requestId: PANE }, async ($, e) => {
191		const ui = $.ui.resolve(e);
192		const { Box, Text } = ui;
193		if (e.surface !== "terminal" || !("Image" in ui)) {
194			return <Text>mobile-mirror needs the terminal, in kitty or Ghostty.</Text>;
195		}
196
197		const { file, generation, width, height } = await read($, shot);
198		if (generation === 0) {
199			return <Text dimColor>Waiting for the first frame…</Text>;
200		}
201
202		const device = await read($, target);
203		if (!device) {
204			return <Text dimColor>Run /mobile-mirror to pick a device.</Text>;
205		}
206
207		const isAndroid = device.platform === "android";
208		const isAsking = await read($, isAskingUrl);
209		const roomForImage = e.props.scroll.bodyRows - TOOLBAR_ROWS - (isAsking ? URL_FIELD_ROWS : 0);
210		const size = fitInside(e.props.bodyColumns, roomForImage, width, height);
211		const { Button, Input } = ui;
212
213		const press = (button: string) => () => runAction($, `Press ${button}`, "mobile_press_button", { device: device.id, button });
214
215		const openUrl = async (url: string) => {
216			await update($, isAskingUrl, () => false);
217			if (!url.trim()) {
218				return;
219			}
220
221			await runAction($, `Open ${url.trim()}`, "mobile_open_url", { device: device.id, url: url.trim() });
222		};
223
224		return (
225			<Box flexDirection='column'>
226				<Box flexDirection='row' gap={1}>
227					<Button key='home' label='Home' onPress={press("HOME")} />
228					{isAndroid && <Button key='back' label='Back' onPress={press("BACK")} />}
229					{isAndroid && <Button key='app-switch' label='App Switch' onPress={press("APP_SWITCH")} />}
230					<Button key='url' label='URL' onPress={() => update($, isAskingUrl, asking => !asking)} />
231				</Box>
232				{isAsking && <Input key='url-field' label='URL ' placeholder='https://example.com' submitLabel='open' autoFocus onSubmit={openUrl} />}
233				<Box>
234					<ui.Image key={VIEW_KEY} source={{ file, format: "png", generation }} columns={size.columns} rows={size.rows} alt='device screen' />
235					<Box position='absolute' top={0} left={0}>
236						<ui.Client key={TAP_KEY} module='./tap.tsx' width={size.columns} height={size.rows} />
237					</Box>
238				</Box>
239			</Box>
240		);
241	});
242};
243
hooks/keys.ts 70 lines
1import type { Platform } from "../types";
2
3export type KeyAction = { type: "text"; text: string } | { type: "button"; button: string };
4
5// special key names the terminal reports, mapped to mobile_press_button buttons
6const ANDROID_BUTTONS: Record<string, string> = {
7	return: "ENTER",
8	backspace: "BACKSPACE",
9	delete: "BACKSPACE",
10	up: "DPAD_UP",
11	down: "DPAD_DOWN",
12	left: "DPAD_LEFT",
13	right: "DPAD_RIGHT",
14};
15
16const IOS_BUTTONS: Record<string, string> = {
17	return: "ENTER",
18};
19
20// ponytail: iOS has no BACKSPACE button in mobilecli, so it is typed as \b; unverified on a device
21const IOS_TEXT: Record<string, string> = {
22	backspace: "\b",
23	delete: "\b",
24};
25
26function actionFor(key: string, platform: Platform): KeyAction | undefined {
27	const buttons = platform === "android" ? ANDROID_BUTTONS : IOS_BUTTONS;
28	const button = buttons[key];
29	if (button) {
30		return { type: "button", button };
31	}
32
33	const text = platform === "ios" ? IOS_TEXT[key] : undefined;
34	if (text) {
35		return { type: "text", text };
36	}
37
38	if (key === "space") {
39		return { type: "text", text: " " };
40	}
41
42	// a single character is typed as is; any other named key is dropped
43	if ([...key].length === 1) {
44		return { type: "text", text: key };
45	}
46
47	return undefined;
48}
49
50// turns pressed keys into device actions, joining consecutive text into one type call
51export function keysToActions(keys: string[], platform: Platform): KeyAction[] {
52	const actions: KeyAction[] = [];
53	for (const key of keys) {
54		const action = actionFor(key, platform);
55		if (!action) {
56			continue;
57		}
58
59		const last = actions[actions.length - 1];
60		if (action.type === "text" && last?.type === "text") {
61			actions[actions.length - 1] = { type: "text", text: last.text + action.text };
62		} else {
63			actions.push(action);
64		}
65
66	}
67
68	return actions;
69}
70
hooks/png.ts 34 lines
1// width and height from a PNG's IHDR chunk, big-endian at bytes 16..23
2export function pngSize(bytes: Uint8Array): { width: number; height: number } {
3	const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
4
5	return { width: view.getUint32(16), height: view.getUint32(20) };
6}
7
8const MAX_CELLS = 255;
9// cell height / cell width of the terminal font, measured 17.3px / 8.1px; no API reports it, so tune per font
10export const CELL_ASPECT = 2.14;
11
12function clampCells(n: number): number {
13	return Math.max(1, Math.min(MAX_CELLS, Math.floor(n)));
14}
15
16// the largest box of cells that keeps the picture's aspect and fits inside maxColumns x maxRows
17export function fitInside(columnsLimit: number, rowsLimit: number, width: number, height: number): { columns: number; rows: number } {
18	const maxColumns = clampCells(columnsLimit);
19	const maxRows = clampCells(rowsLimit);
20	// a broken PNG header can report zero; without an aspect to keep, fill the box instead of dividing by zero
21	if (width <= 0 || height <= 0) {
22		return { columns: maxColumns, rows: maxRows };
23	}
24
25	const rowsAtFullWidth = (maxColumns * height) / width / CELL_ASPECT;
26	if (rowsAtFullWidth <= maxRows) {
27		return { columns: clampCells(maxColumns), rows: clampCells(rowsAtFullWidth) };
28	}
29
30	const columnsAtFullHeight = (maxRows * width * CELL_ASPECT) / height;
31
32	return { columns: clampCells(columnsAtFullHeight), rows: clampCells(maxRows) };
33}
34
hooks/tap.tsx 43 lines
1import type { ClientKeyEvent, ClientPointerEvent, ClientSurface } from "claude-code";
2
3type Pending = { keys: string[] };
4
5// a post per frame at most, and a later one replaces an undelivered one, so keys are sent in batches
6const FLUSH_MS = 50;
7
8// Laid over the screenshot: posts each left click as a fraction of the picture (0..1 on both axes),
9// and, once clicked, the keys typed while it holds the focus
10export default function TapInput(_props: unknown, surface: ClientSurface<Pending>) {
11	if (surface.state === undefined) {
12		const pending: Pending = { keys: [] };
13
14		surface.onPointer((e: ClientPointerEvent) => {
15			if (e.type !== "down" || e.button !== "left" || surface.columns === 0 || surface.rows === 0) {
16				return;
17			}
18
19			const x = e.fine?.x ?? e.x + 0.5;
20			const y = e.fine?.y ?? e.y + 0.5;
21			surface.post({ x: x / surface.columns, y: y / surface.rows });
22		});
23
24		surface.onKey((e: ClientKeyEvent) => {
25			pending.keys.push(e.key);
26		});
27
28		surface.every(FLUSH_MS, () => {
29			if (pending.keys.length === 0) {
30				return;
31			}
32
33			surface.post({ keys: pending.keys.splice(0) });
34		});
35
36		surface.setState(pending);
37	}
38
39	const { Box } = surface.elements;
40
41	return <Box width='100%' height='100%' />;
42}
43
types/index.d.ts 10 lines
1export type Shot = { file: string; generation: number; width: number; height: number };
2export type Platform = "android" | "ios";
3export type Target = { id: string; platform: Platform; screenWidth: number; screenHeight: number };
4
5declare module "claude-code" {
6	interface PluginState {
7		"mobile-mcp": { shot: Shot; target: Target | null; streamId: number; isAskingUrl: boolean }
8	}
9}
10