SLOPSHOPPER

agent-mobile-observer

手機聊天室的觀測 mod:回合開始/結束與手機直送抵達,POST 給 :8899/api/observe;純觀測不改行為

newnetwork
A shopper browsing a rack in a slop shop
README

agent-mobile-control

A self-hosted, mobile-friendly web UI for driving your local coding agents like a chat app. Each chat room is one agent session, and opening a room resumes that conversation. Typical setup: run it on your desktop/workstation, reach it from your phone over Tailscale or another private VPN.

Supported engines:

EngineTypeStatus
Claude Codelocal agent (reads/writes files, runs commands)tested
Codexlocal agent (reads/writes files, runs commands)tested
Grok / Gemini / ChatGPT / DeepSeek / OpenRouterchat only, via OpenAI-compatible APIwired up but never exercised — see Limitations

Rooms for Claude Code and Codex read the transcripts those tools already write on disk, so your phone and your desktop see the same conversation.

Security warning — read this before you run it

This tool can hand an AI agent full control of your computer. The defaults are deliberately locked down; you decide how far to open them.

Defaults out of the box: binds to 127.0.0.1 only (nothing else on your network can reach it), the agent runs in plan mode (read-only — it can look and talk, never change anything), and GET /api/file can only read the folder this app uploads into. In that state it is roughly as dangerous as a local text editor.

Everything below is about what happens when you loosen those.

  1. Private network only, never the public internet. Setting bind_tailscale puts the server on your tailnet. That is fine for a VPN only you are on. Do not port-forward it, do not give it a public IP, do not run it on a cloud box with an open port, and do not put it behind Tailscale Funnel (Funnel publishes to the whole internet; plain tailscale serve stays inside your tailnet). Treat "someone can reach this port" as equivalent to "someone is sitting at your keyboard."
  2. Authentication is opt-in, and you should turn it on the moment you leave loopback. Set auth_token in config.json (or the CLAUDE_CHAT_TOKEN environment variable). Requests from 127.0.0.1 skip the check; everything else must send the token as an X-Auth-Token header or a ?token= query parameter. With no token set and bind_tailscale on, anyone who can reach the port has full access — they can read your history, send messages as you, and get the agent to run commands. Also note: behind a reverse proxy every request looks like it came from 127.0.0.1, so the loopback exemption would let everyone through. Don't put this behind one.
  3. auto mode lets the agent act without asking. In that mode Claude Code runs with --dangerously-skip-permissions and Codex with --dangerously-bypass-approvals-and-sandbox, so the agent reads files, writes files, and executes shell commands with no prompt. That is what makes "control your computer from your phone" actually work, and it is also how a malicious or careless message does real damage. The alternatives, selectable per room in the in-app Settings sheet: edits maps to Claude Code's acceptEdits, which auto-approves file edits and — per Anthropic's documentation — some filesystem commands too, so it is not a guarantee that nothing executes; ask shows every command / edit / fetch on the phone as an allow-or-deny card before it runs (see "Permission cards"); plan is genuinely read-only, and is the default. Unknown mode values are rejected outright rather than falling back to something permissive.
  4. /api/file and /api/upload are only as safe as the folders you expose. By default /api/file serves nothing but this app's own uploads/ folder. Turning on allow_home_reads exposes your entire home directory to anyone who can reach the server — including ~/.claude/.credentials.json, SSH keys, and every private document under it. That is a convenience trade-off, not a safe default; enable it only on a network where you are the only participant. HTML, SVG, XML and JavaScript are always sent as downloads rather than rendered, so a file the agent generated cannot execute as a page inside this app's origin.

If any of this is unacceptable for your situation, don't run this tool — it was built for a single trusted user on a single trusted private network.

Screenshot

(Add screenshots here, e.g. docs/screenshot-list.png for the chat list and docs/screenshot-chat.png for a conversation.)

Requirements

  • Windows. The server currently relies on a few Windows-specific APIs (ctypes.windll for process-liveness checks, taskkill to stop a run, CREATE_NO_WINDOW, and an auto-detected Tailscale executable path). See Limitations for what porting to macOS/Linux would involve.
  • Python 3.10+ (tested on 3.11).
  • Node.js + the Claude Code CLI, installed and already logged in — i.e. running claude from a terminal works.
  • Tailscale (optional but recommended) if you want to reach the server from your phone. Without it, the server only binds to 127.0.0.1 (local machine only).

Installation

git clone <this-repo-url> agent-mobile-control
cd agent-mobile-control
python -m venv .venv
.venv\Scripts\pip install -r requirements.txt

Regenerating the app icons is optional — the repo already ships with static/icon-*.png — but if you want to change the look (Pillow comes from requirements-dev.txt, together with the lint and test tools):

.venv\Scripts\pip install -r requirements-dev.txt
.venv\Scripts\python make_icon.py

Usage

Start the server:

.venv\Scripts\pythonw.exe server.py

or just double-click start-server.cmd (starts it hidden, in the background).

Open it:

  • Locally: http://127.0.0.1:8899
  • From your phone, over Tailscale: http://<your-tailscale-ip>:8899 (find your Tailscale IPv4 address with tailscale ip -4 on the machine running the server)

On iPhone Safari: Share → Add to Home Screen, and it behaves like a standalone app (icon, no browser chrome), backed by the included manifest.webmanifest.

Keep it running:

  • The included start-server.cmd is the simplest way to start it manually.
  • For "always on," register it as a Windows scheduled task (or any process supervisor you prefer) that runs pythonw.exe server.py at logon.
  • To restart: double-click restart-server.cmd. It kills only the pythonw.exe whose command line points at this folder (not "whatever is listening on 8899" — that once took out an unrelated service), starts it again, and prints the new process's start time. Look at that time: a health-check 200 alone proves nothing, because a stale process answers 200 too. Double-clicked, it waits for a key so you can read the result; run from a script or a background shell, it exits on its own instead of hanging. If you registered a scheduled task, re-trigger that instead.
  • Logs are written to logs\server.log.

Chat rooms:

The chat list is simply every ~/.claude/projects/*/*.jsonl file Claude Code has already written to disk, sorted by last activity. Opening a room resumes that session (claude -p --resume <session-id>); tapping "+" starts a brand-new session in a project folder you pick.

Archive & delete:

Swipe left on a room (or long-press / right-click) to reveal Archive and Delete.

  • Archive just hides it from the list — the transcript file is untouched. The mark is stored in web-archive.json (created automatically, gitignored). "Show all" in Settings brings archived rooms back into view.
  • Delete moves the transcript into a local trash\ folder (created automatically) instead of destroying it. To restore, move the file back into ~\.claude\projects\<project-slug>\. A room that's currently running a message can't be deleted.

File preview:

Files the agent hands over with the desktop app's SendUserFile tool are shown too (the tool call is turned into a text line with the caption and the full paths). Windows-style file paths that show up in a reply (e.g. C:\Users\you\Desktop\screenshot.png) are automatically turned into inline previews — images render inline, video and audio get a player. Anything that a browser could execute as a page (HTML, SVG, XML, JS) is always sent as a download instead, so a file the agent generated can't run as script inside this app's origin.

Tapping a file never navigates away from the app. That matters once you add it to the Home Screen: a standalone web app has no browser chrome, so a plain link to /api/file would replace the whole app with the file and leave you no way back. Instead:

  • Images open in a full-screen viewer: close with ✕, by tapping the dark background, or by swiping down; tap the image to toggle actual size (the page disables pinch-zoom, so this is how you read a screenshot).
  • Other files open in a viewer sheet with a title bar and ✕: .md is rendered with the chat's own Markdown renderer, .txt/.csv/.json as monospaced text (JSON pretty-printed, capped at the first million characters), .pdf in an iframe, video/audio in a player. .html is shown in an <iframe sandbox> with no allow-scripts and no allow-same-origin, so you see the layout but nothing in it runs and it cannot reach the app.
  • Opening a viewer pushes a history entry, so any back action (the ‹ button, the system edge-swipe, Android's back button) closes the viewer first and only a second back leaves the conversation.

Swipe back: in a conversation, swipe right anywhere to return to the list. The screen follows your finger; past a third of the width or a quick flick it leaves, otherwise it springs back. It stays out of the way inside horizontally scrollable blocks (wide code, tables), text fields, media players, and while a viewer is open. The browser's own overscroll-to-go-back is disabled (overscroll-behavior-x: none) so it can't fire while you scroll a code block sideways.

This is served by GET /api/file?path=..., which only reads from an allow-list:

  • By default the allow-list contains exactly one folder: this app's own uploads/. Nothing else on your disk is reachable through this endpoint.
  • To allow more folders, add absolute paths to extra_file_roots in config.json.
  • Setting allow_home_reads: true adds your entire home directory. That is convenient — the agent can show you anything it produced anywhere — but it also exposes ~/.claude/.credentials.json, SSH keys, and every private document under your home folder to whoever can reach the server. Only do this on a network where you are the only participant.

Background tasks: commands the agent runs in the background (run_in_background, or a command that hit its time limit and was moved there), background subagents and workflows used to be invisible from the phone — the room just looked stuck. While a Claude room is open the phone now polls GET /api/bg/{slug}/{sid} every 5 seconds and, whenever there is something to show, puts a status bar under the context bar: "⏳ 2 background tasks running". If the quietest of them has produced no output for more than two minutes the bar turns orange and says for how long — usually the answer to "why is it stuck". Tap it for the list: kind (command / subagent / workflow), label, running time, time since it last showed signs of life, the last 8 lines of a command's output (a subagent shows its last 5 actions), and the completion notice for finished ones, which stay listed for 6 hours. Detection reads the transcript incrementally and is deliberately strict about two things that produced false positives: a start message only counts at the very beginning of a tool result (the same text merely quoted inside some other command's output does not), and a task is only reported as running if it was started after the process that is currently alive for that session — a background task dies with the process that started it and leaves no notice, so anything older is shown as unknown. Output files named in the transcript are read only from Claude Code's own temp task directory (workflow transcripts only from under ~/.claude/projects), so a tampered tool result cannot make the server hand your phone an arbitrary file.

Sending images and documents to Claude:

Tap "+" next to the composer to attach a photo (camera or library) or a document, multiple at once. It uploads to a local uploads\ folder (created automatically) and the message text gets a [phone photo, please use Read to view: <path>] (or [phone file, please use Read to read: <path>]) note appended so Claude knows to look at it. Accepts png/jpg/gif/webp and pdf/txt/md/csv/json (magic bytes are checked), 25 MB max per file.

Rename & full-text search:

  • Long-press a room → Rename. Stored in web-titles.json (gitignored); clear the name to fall back to the original title.
  • Type two or more characters in the search box and, besides matching titles, the server searches the content of every visible conversation (GET /api/search?q=, last 1.5 MB of each transcript). Matching rooms show the hit as their preview line, prefixed with 🔎.

Appearance & per-room settings:

  • Theme: Settings → Appearance (Auto / Light / Dark). All colors are CSS custom properties in static/style.css (:root for dark, html[data-theme="light"] for light).
  • Each room has its own model/effort override (top-right pill button in a conversation), stored in localStorage; "Default" falls back to the global choice in Settings.
  • The model picker offers every model the current Claude Code CLI accepts as --model (Fable 5.1/5, Opus 5.5/5/4.8/4.7/4.6, Sonnet 5/4.6, Haiku 4.5). The list lives in exactly one place — MODEL_LIST at the top of static/js/state.js (labels plus the one-line plain-language description shown in Settings) — and the short key → full model ID map is MODELS in claude_chat/config.py; change both when a new model ships. The descriptions are in Traditional Chinese; edit them to taste.
  • Voice input: the microphone button uses the Web Speech API; unsupported browsers get a hint to use the keyboard's built-in dictation key instead.
  • Usage panel (Settings → "plan limits & usage", collapsed by default): the top half shows your Anthropic plan limits (5-hour / weekly, same source as the desktop app's OAuth usage API, using the local ~/.claude/.credentials.json token, cached 60s); the bottom half is real token usage computed by scanning your local .jsonl transcripts (today / last 7 days, deduplicated by request ID, includes scheduled/subagent runs).
  • Context bar at the top of a conversation shows current context usage for that session (turns orange above 70%, red above 85%), estimated from the model name and the last reported token usage.

Notes:

  • Don't type into the same session from your desktop Claude Code and this phone UI at the same time — the list shows a "desktop open" tag as a reminder.
  • Locking your phone doesn't interrupt anything — the run keeps going on the server; reopen the room later to see the result.
  • One room can only run one message at a time; the whole server allows at most 4 concurrent runs (MAX_CONCURRENT_RUNS in claude_chat/config.py).
  • For an https:// URL instead of http://, enable Tailscale Serve on your tailnet and run tailscale serve --bg 8899 — optional, and it does not change who can reach the server (Serve stays inside your tailnet). Do not use Tailscale Funnel, which would publish it to the open internet.

Two-way sync with the Claude Code desktop app

If the Claude desktop app is installed on the server machine, the room list lines up with it in both directions. None of this is documented by Anthropic — it was worked out from the app bundle (version 1.40609.1) and verified by hand, so a future desktop release may break it; everything degrades to "no desktop app" behaviour when it does.

Desktop → phone (automatic, live). The desktop app keeps one JSON file per session under %APPDATA%\Claude\claude-code-sessions\<account>\<org>\local_<id>.json; its cliSessionId is the .jsonl file name in ~/.claude/projects. The server reads those files (5-second cache) for titles, archive state and project paths, so archiving or renaming a session on the desktop shows up on the phone a few seconds later. The IndexedDB copy is never touched.

Phone → desktop (claude:// deep link). Sessions started from the phone don't exist in the desktop app's list. The desktop app registers the claude:// URL scheme, and claude://resume?session=<cli-session-id> imports an on-disk transcript into its list — same file, nothing copied, and both sides keep appending to it afterwards. (This is what the CLI's /desktop command uses too.)

  • Settings → Sync to desktop app: auto (a new phone-started room is imported as soon as its first turn finishes; the desktop app switches to that session once), manual, or off. Stored as desktop_sync in config.json.
  • Long-press a room → Open in desktop app: imports it if needed, otherwise just switches the desktop app to it via claude://code/continue?session=local_<id>. Re-importing the same session does not create a duplicate.
  • Rooms started from the phone that the desktop hasn't registered yet carry a "phone only" tag.
  • Caveats seen in practice: the imported session has no title on the desktop side until you name it there; opening it on the desktop spawns a CLI process, so the phone shows the "desktop open" tag; writing the registry JSON directly is not picked up by a running desktop app, only the deep link works.

Sessions the desktop app has open: messages are delivered into that process. A room tagged "desktop open" has a live claude process behind it. Spawning a second claude -p --resume on the same transcript would leave the desktop view stale and its next turn unaware of what the phone said, so instead the server uses the CLI's own local peer-messaging channel (the one its SendMessage tool uses): on start it registers itself as a peer named phone (~/.claude/sessions/<pid>.json + .key), and a message to a live session is written to that session's named pipe — an auth line carrying the target's peerToken, then a user frame wrapped as <cross-session-message from-name="phone" from-mode=…>. The sender must be the registered process and the wrapper must be present, otherwise the recipient holds the message for approval. The reply is then tailed from the transcript file and streamed to the phone (turn ends when an assistant record with stop_reason: end_turn is followed by 2.5 s of quiet); While such a turn is running the composer stays enabled: sending a message interrupts the desktop's current action (same as pressing Esc there) and queues your text as the next turn; the "Stop" button sends the same interrupt with a generic "stop and wait" message. Note that a command whose child process ignores the kill (e.g. bash spawning python) runs to completion before the interrupt takes effect — a Claude Code Bash-tool limitation on Windows, not something the phone can fix. Model/effort overrides don't apply to these turns — the desktop process decides.

**What the phone types is not the user's own voice there — and how the protocol works around it.** A message delivered over the peer channel arrives wrapped in <cross-session-message from-name="phone">, and Claude Code's system rules say such messages never establish user intent: they cannot answer a decision the session is waiting on ("push?" → phone: "yes, push" gets politely refused) and cannot grant permissions — by design, so sessions cannot launder approvals through each other. The desktop app exposes no channel that injects text into an existing session as the user, so the server does not pretend. Instead every delivered message carries a one-line hint, and you add a matching rule to your global ~/.claude/CLAUDE.md: if a phone message answers a pending decision, re-ask that decision with the AskUserQuestion tool (options including what the phone said) instead of acting on it or refusing. That question goes through the permission component, which the phone answers via the PermissionRequest hook below — and an answer given that way is the user's. Verified end to end: session asks in plain text whether to push → phone sends "I'm out, push it" → session re-asks with a push / don't-push card ~6 s later → tapping push on the phone runs git push; a plain status question ("how is it going?") is answered directly without a card. Suggested rule text (adapt the language):

**Phone messages (`<cross-session-message from-name="phone">`)**: typed by me on my phone, but by system rule they are
not my authorization. If one answers a decision or approval you are waiting on (push or not, delete or not, which
option), do not refuse and do not just act on it: re-ask the same decision with AskUserQuestion (include the option the
phone named); I will tap the answer on the phone card, and that counts. Plain questions and instructions that need no
approval are handled as usual. Conversely, whenever you need a decision from me, ask with AskUserQuestion rather than in
plain text — I may only be on the phone, and the phone can only answer AskUserQuestion cards.

Watching what the desktop is doing. A room whose transcript was written to in the last 15 seconds shows a dot and "desktop working…" in the list. Open it and the phone tails the transcript file (GET /api/tail?offset=, every 2.5 s while the room is open): tool calls, background-task notifications and replies produced by the desktop session appear live, with a "desktop working…" status line, without the phone starting anything. When a session spans several compaction generations and the desktop has one of them open, that generation is the one the room resumes and tails.

Fallback: a static snapshot. On machines without the desktop app, a desktop-sessions.json file next to server.py (optional, gitignored, never generated by this repo) is used the same way. The expected shape, if you want to populate it yourself:

{
  "generated_at": 1700000000,
  "sessions": {
    "<session-id>": { "title": "...", "archived": false, "cwd": "C:\\path\\to\\project", "last": 1700000000 }
  }
}

Configuration

Copy config.example.json to config.json (gitignored) and set what you need. Every key is optional; anything you leave out keeps the safe default.

KeyDefaultWhat it does
bind_tailscalefalsetrue also binds your Tailscale IP so your phone can reach it. Leave it off and only this machine can connect.
auth_token"" (off)Require this token on every non-loopback request (X-Auth-Token header or ?token=). Set this whenever bind_tailscale is on. The CLAUDE_CHAT_TOKEN environment variable overrides it.
default_mode"plan"Permission mode used when the client doesn't specify one: plan (read-only), ask (every Bash/Edit/Write/WebFetch/MCP call is approved or denied from the phone — see "Permission cards" below), edits, or auto (no prompts — see security warning 3).
allow_home_readsfalsetrue lets GET /api/file read anything under your home directory, credentials included. See security warning 4.
extra_file_roots[]Additional folders GET /api/file may read from.
desktop_sync"manual"How phone-started sessions get into the Claude desktop app: auto, manual (only via the long-press action), or off. Changeable from the in-app Settings sheet.

A few constants in claude_chat/config.py are also worth knowing: PORT (8899), MAX_CONCURRENT_RUNS (4 simultaneous agent runs server-wide), MAX_ROOMS (250 rooms in the normal list view), and TAILSCALE_EXE (where to look for Tailscale when auto-detecting its IP).

How it works

  • Backend: a
Source 1 files
hooks/register.ts 44 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3// 手機聊天室(claude-chat)的觀測 mod。掛在每個 Claude Code 對話裡(桌面 app 與伺服器起的 -p 都會載),
4// 只做三件回報、不改任何行為:
5//   turn.start / turn.complete → 伺服器知道這個對話「確認執行中」(比看 jsonl 有沒有在動準)
6//   session.receive(手機直送的訊息抵達)→ 伺服器把等了 6 秒的回條標成「已送達」
7// 伺服器沒開、埠不對、任何錯誤:安靜略過,對話照常。
8// ponytail: 埠寫死 8899(與 claude_chat/config.py 的 PORT 一致);要改成可設定再讀 options.userConfig。
9const OBSERVE_URL = 'http://127.0.0.1:8899/api/observe'
10
11async function report($: EngineInterface, body: Record<string, unknown>): Promise<void> {
12  try {
13    await $.http.fetch(OBSERVE_URL, {
14      method: 'POST',
15      headers: { 'Content-Type': 'application/json' },
16      // sid 每次現讀(async):/clear 之後行程換了 session id,但不會再 fire session.start
17      body: JSON.stringify({ sid: await $.session.id(), ts: Date.now() / 1000, ...body }),
18    })
19  } catch {
20    // 伺服器不在:觀測失敗不能影響對話
21  }
22}
23
24export const register: Register = (on) => {
25  on('turn.start', async ($, e, next) => {
26    await report($, { kind: 'turn.start', turnId: e.turnId })
27    return next(e)
28  })
29  on('turn.complete', async ($, e, next) => {
30    await report($, { kind: 'turn.complete', turnId: e.turnId, reason: e.reason })
31    return next(e)
32  })
33  on('session.receive', async ($, e, next) => {
34    // 只認手機聊天室直送的(伺服器登記的 peer 名字叫 phone);別的 session 的訊息不關我們的事。
35    // 型別裡 peer 類有兩個 kind('peer'=另一個對話的模型、'peer-send-message'=走 prompt 那條),
36    // 手機那封實際落在哪個沒有文件說,兩個都收,真正的判別是 from-name。
37    const k = e.origin.kind
38    if ((k === 'peer' || k === 'peer-send-message') && /from-name="phone"/.test(e.text)) {
39      await report($, { kind: 'receive' })
40    }
41    return next(e)
42  })
43}
44