Speak Persian, Claude Code gets English. Hold Space, talk in Persian, English or both, and a clean English prompt appears in the prompt box

Hold <kbd>Space</kbd> → say it in Persian, English, or both → release.<br> A clean English prompt is waiting in your prompt box.
╭──────────────────────────────────────────────────────────────╮
│ ● REC 0:07 ▁▂▅▇▅▃▂▁▂▃▅▆ release space to finish │
│ تستها رو اجرا کن و اونایی که خراب شدن رو درست کن │
│ → Run the tests and fix the ones that fail │
╰──────────────────────────────────────────────────────────────╯
| You say | Claude Code receives |
|---|---|
| تستها رو اجرا کن و اونایی که خراب شدن رو درست کن | Run the tests and fix the ones that fail. |
این parseOrder توی auth.ts وقتی order خالیه کرش میکنه، درستش کن | Fix the crash in parseOrder in auth.ts when the order is empty. |
| اممم… همون باگی که قبلاً گفتم رو درست کن | Fix the null check in parseOrder in auth.ts. (in chat mode, it reads the conversation) |
auth.ts and parseOrder are spelled correctly.It is free. You need no account and no API key. Speech recognition runs on your Mac with Whisper. You need Claude Code, a Mac, Python 3.10+ and Homebrew. On Apple Silicon the install below uses mlx-whisper. On an Intel Mac it uses faster-whisper on the CPU.
Windows or a gaming laptop? The same requirements-local.txt installs faster-whisper, which uses an NVIDIA GPU when CUDA 12 and cuDNN 9 are installed (pip install nvidia-cublas-cu12 nvidia-cudnn-cu12 is the easy way), and the CPU if not. The speech engine is ready, but the rest of the mod is not: it reads the microphone with macOS AVFoundation and uses macOS paths. Windows is not supported yet and is untested. See Speech engines.
# 1. Get the mod and the free local speech engine
git clone https://github.com/Ariiima/claude-code-persian-to-english-voice.git ~/.claude/persian-voice
cd ~/.claude/persian-voice
python3 -m venv .venv && .venv/bin/pip install -r requirements-local.txt
brew install ffmpeg
# 2. Start Claude Code with the mod
claude --plugin-dir ~/.claude/persian-voice
In Claude Code, run /voice off once (the built-in dictation also uses <kbd>Space</kbd>), then run /fa engine local. Hold <kbd>Space</kbd> and talk. The first time, macOS asks for microphone access for your terminal. Allow it.
The first recording downloads a 1.6 GB model, so it takes a while. After that it works offline, and a short recording is read in about 1 second.
Want live text while you speak, or a faster first start? See the other speech engines.
To load the mod in every session, add it to ~/.claude/settings.json (use your own home path):
{ "env": { "CLAUDE_CODE_PLUGIN_DIRS": "/Users/YOU/.claude/persian-voice" } }
| To do this | Do this |
|---|---|
| Record | Hold <kbd>Space</kbd>, speak, release |
| Record without holding | Tap <kbd>Space</kbd> 3 times, speak, tap once |
| Use your own shortcut | /fa key ctrl+x v, then press it to start and again to stop |
| Cancel a recording | Type any letter, or click ✕ Cancel |
| Get back what you said | /fa last |
| Change the cleanup mode | /fa mode, or click ✨ mode |
| Send as soon as you finish | /fa send, or click ⏎ send |
| All settings | /fa, or click ⚙ Settings |
A single tap of <kbd>Space</kbd> still types a space.
exact) to a full task spec (spec) or a git commit message (commit). See Modes.Change the engine with /fa engine local, /fa engine google or /fa engine soniox, or in /fa settings.
| Engine | Cost | Key | Live text | Internet | Status |
|---|---|---|---|---|---|
| Local Whisper (recommended) | Free | None | No, text appears after you release | Only for the first model download | Built in. Tested on Persian, English and mixed speech on Apple Silicon (mlx-whisper). The faster-whisper version for Intel, Windows and NVIDIA GPUs runs on the CPU in our test; the GPU is untested. |
| Google Gemini | Free key, no card: aistudio.google.com/apikey | Yes | No, text appears after you release | Yes | Built in. Tested. Save the key in ~/.config/gemini/key. |
| Soniox | Paid | Yes | Yes, with live English translation | Yes | Built in. Save the key in ~/.config/soniox/key. |
| Groq Whisper | Free key, no card, with daily limits | Yes | No | Yes | Not built in yet. Groq did not answer from Iran when we tried, so we could not test it. |
Local Whisper and Gemini only transcribe. Claude makes the English draft from the transcript. In our test both read Persian with English code words well, and both can swap look-alike words (for example parseOrder). Add such words to your word list.
/fa engine google) only while you record. With /fa engine local, audio never leaves your Mac. Recording stops by itself after 5 minutes.~/.config/… or your environment. They are never written to the mod folder.Mod, not plugin. This is a Claude Code mod: code that runs inside Claude Code and hooks into its events and interface. A plain plugin cannot catch the <kbd>Space</kbd> key or draw the live view.
| What | Why |
|---|---|
| macOS | The microphone is read through AVFoundation. |
| Claude Code with plugin hooks (tested on v2.1.287) | The mod runs inside Claude Code. |
| Python 3.10 or later | Runs the audio stream (stt/stream.py). |
| ffmpeg | Reads the microphone. |
| Apple Silicon Mac | Best for the local Whisper engine. An Intel Mac works too, on the CPU, and is slower. |
| A speech engine | Local Whisper needs no key. Google needs a free Gemini key. Soniox needs a paid Soniox key. See Speech engines. |
The Google engine is gemini-3.5-transcribe-live. Soniox translates by itself.
Hold to talk
Tap mode
Tapped keys do not repeat, so the cursor does not move while you record.
Cancel
Nothing goes into the prompt box, and nothing is sent. /fa last still shows what you said. <kbd>Esc</kbd> cannot cancel a recording: Claude Code does not pass the <kbd>Esc</kbd> key to mods. You can also run /fa rec to start, and run /fa rec again, or press ⏹ Stop, to finish.
You can use hold-Space, your own shortcut, or both.
/fa key ctrl+x v
The shortcut must have a modifier (ctrl+r, meta+k) or be a chord (ctrl+x v). A plain key would type a character. Choose a combination that Claude Code does not already use. The keybindings docs list the defaults.
meta+k, the recording stops when you release it. /fa space off
To remove the shortcut, run /fa key off. This also turns hold-Space on again. The shortcut is written to ~/.claude/keybindings.json. Other bindings in the file are kept.
Type /fa (or click ⚙ Settings above the prompt) to open the settings dialog:
How to start a recording
Hold Space [ ● On ] hold Space to talk, release to finish
Shortcut type one, e.g. ctrl+x v, then Enter
What happens to your words
Cleanup mode Prompt ▾
Claude rewrites what you said into a clear prompt…
Auto-send [ ○ Off ] goes into the prompt box; you press Enter
Microphone and words
Microphone System default ▾
Word list 12 words, 1 translation, 4 learned · ~/.config/persian-voice/terms.txt
Use <kbd>Tab</kbd> to move, <kbd>Enter</kbd> to change, and <kbd>Esc</kbd> to close. Changes are saved at once, for all projects.
| Command | What it does | |
|---|---|---|
/fa | Open the settings dialog. | |
/fa rec | Start or stop a recording without holding a key. | |
/fa last | Show your last recording (Persian and English) and put it in the prompt box again. | |
/fa key [shortcut] | Show or set your own shortcut, for example /fa key ctrl+x v. /fa key off removes it. | |
| `/fa space [on\ | off]` | Turn hold-Space to talk on or off. |
/fa mode [name] | Show or set the cleanup mode. Without a name, it moves to the next mode. | |
/fa polish | Turn the cleanup on (prompt) or off (exact). | |
/fa send | Turn auto-send on or off. | |
/fa mic | List the microphones. /fa mic 1 selects microphone 1. | |
/fa terms | Show your word list and the words the mod learned. | |
/fa help | Show all commands. |
With auto-send on, Claude Code shows the prompt as "The persian-voice plugin sent a message". Claude Code adds this label to every prompt that a plugin sends, and a mod cannot remove it. Claude still treats the text as your request.
The mode sets what happens to your words after you stop speaking.
| Mode | Shown as | Result | Uses a model |
|---|---|---|---|
auto | Auto | Like prompt, but JEV can also select a spec or a commit message. Without a JEV key, it works like prompt. | As the selected kind |
prompt (default) | Prompt | A clear prompt in the form of the task: bug fix, feature, question, and others. | Only when needed (see below) |
chat | Prompt (reads this chat) | Like prompt, but Claude also reads this conversation to replace "that bug" with the real name. Slower. | Yes |
spec | Spec | A task spec with Goal, Context, Requirements and Done when. Good for thinking aloud. | Yes |
commit | Commit msg | A git commit message. | Yes |
exact | Exact | The translation only, with fillers removed. | No |
In prompt mode, Claude rewrites only when JEV finds a problem (and always for a story analysis). Without JEV, it rewrites only long or self-corrected requests. The rewrite uses Claude Sonnet 5.5 at low effort, through your Claude Code login. If it takes more than 8 seconds or fails, you get the plain translation. In chat mode, after 8 seconds it falls back to the normal prompt mode.
Each type of task needs different information. For example, a bug fix needs the symptom, the location and what "fixed" looks like. So the rewrite uses a different prompt for each type of task. With a JEV key, JEV selects the prompt for each recording. Without a key, the rewrite uses the general prompt.
| Kind | What the rewrite gives |
|---|---|
| General | The goal first, then the context, constraints and expected result. |
| Bug fix | The goal ("Fix …"), the symptom with the exact error message, when it occurs, where to look, and what "fixed" looks like. |
| Feature | The goal, where it goes and which pattern to follow, the requirements, the constraints, and how to check it. |
| Refactor | The goal, the scope, what must stay the same, and the target structure. |
| Tests | The code under test, the cases and edge cases, the constraints, and how to run the tests. |
| Review | What to review, what to look for, and the form of the report. |
| Question | Your question, which stays a question, with the exact code and sources it is about. |
| Story editor | The role "Act as an experienced fiction editor.", then the text, the aspects to analyze, and what to return. |
| Spec | A task spec with Goal, Context, Requirements, Out of scope and Done when. |
| Commit msg | A git commit message: an imperative subject of at most 72 characters. |
Each prompt includes only the parts that you said. The only addition is the role sentence of the story editor. The prompts follow Anthropic's prompting guidance and the Claude Code best practices. They are in hooks/prompts.ts.
JEV is a fast decision model from TypeSafe. It does not write text. It answers yes/no and choice questions with probabilities, in about 1 second. When you set a JEV key, the mod asks JEV these questions about each recording:
| Check | What the mod does |
|---|---|
| Does the text have fillers, a vague reference, rambling, or a translation error? | It runs the Claude rewrite only when the answer is yes. A clear request goes into the prompt box at once. |
| Does the recent chat make a vague reference clear ("fix that bug")? | It rewrites in the chat mode, so Claude replaces the reference with the real name. |
| Is the text a request for Claude? | When JEV is sure that it is not (for example, you talk to another person), the mod puts the text in the prompt box, but does not rewrite it or send it. |
| Does the request do something that you cannot undo (delete, force-push, deploy)? | With auto-send on, the mod does not send the prompt. It puts the text in the prompt box. Press <kbd>Enter</kbd> to send it. |
| Which kind of task is it? | It uses the rewrite prompt for that kind. |
| After the rewrite: did Claude add, change or remove something that you said? | It asks Claude for one more rewrite and tells it the problem. If the second rewrite also has a problem, it uses your own words. |
To turn on the checks, save your TypeSafe key:
mkdir -p ~/.config/typesafe
printf '%s' 'YOUR_JEV_API_KEY' > ~/.config/typesafe/key
chmod 600 ~/.config/typesafe/key
You can also set the JEV_API_KEY environment variable. Without a key, or when JEV does not answer in 2.5 seconds, the mod uses its old rules. You do not lose a recording.
/fa last to see it again and to put it back in the prompt box.flowchart LR
A["Hold Space<br/>or your shortcut"] --> B["stream.py<br/>ffmpeg reads the mic"]
B -- "audio" --> C["Soniox<br/>speech to text + translation"]
C -- "words + English, live" --> D["REC panel<br/>above the prompt"]
D -- "release Space" --> E["Rule cleanup<br/>remove fillers"]
E --> F{"JEV (or rules):<br/>needs a rewrite?"}
F -- "no" --> G["Prompt box"]
F -- "yes" --> H["Claude rewrite<br/>(the selected mode)"]
H --> J{"JEV: same<br/>meaning?"}
J -- "yes" --> G
J -- "no: your words" --> G
G -- "Enter, or auto-send" --> I["Claude Code"]
hooks/register.tsx) watches the prompt box. A held key sends the same character many times. When spaces repeat fast, the mod starts a recording. A shortcut is a Claude Code keybinding that presses the mod's Talk button.stt/stream.py reads the microphone with ffmpeg and sends the audio to Soniox. It prints the live text and the microphone level 6–7 times each second.stream.py to stop. Soniox then confirms the last words and their translation.Your word list
Add words that the recognizer gets wrong to ~/.config/persian-voice/terms.txt. Write one word or name on each line. To set a translation, write persian = english:
# Names and terms
Kubernetes
useEffect
# Translations
دیپلوی = deploy
For words that belong to one project, create .persian-voice-terms.txt in the project folder. It has the same format. The mod reads it together with the global list, and only in that project. Commit the file to share it with your team.
The mod also adds the project name, the git branch and the project's file names by itself.
Where settings are saved
Your mode, auto-send, microphone, hold-Space and shortcut choices are saved in Claude Code's plugin store. They apply to all projects and sessions.
Environment variables
stt/stream.py reads these variables. The mod sets most of them for you.
| Variable | Meaning |
|---|---|
SONIOX_API_KEY | The Soniox key. If it is not set, the key is read from ~/.config/soniox/key. |
FA_ENGINE | soniox (default), google or local. |
FA_LOCAL_MODEL | The Whisper model for local. Default on Apple Silicon: mlx-community/whisper-large-v3-turbo. Default elsewhere (faster-whisper): large-v3-turbo. Smaller models (mlx-community/whisper-small-mlx, small) are faster but make many mistakes in Persian. |
FA_LOCAL_LANG | A language code for local, for example fa. Default: detect it. Forcing fa turned English-only speech into garbage in tests. |
GEMINI_API_KEY | The Gemini key for google. If it is not set, the key is read from ~/.config/gemini/key. |
FA_GEMINI_MODEL | The Gemini model. Default gemini-3.5-transcribe-live. |
FA_GEMINI_LANGS | BCP-47 codes to favour, comma separated. Default: fa-IR. It still reads English words and English-only speech. Empty means the model detects the language, which split parseOrder into "parse order" in tests. |
FA_MIC | The microphone: an AVFoundation index, or default. |
FA_CONTEXT | Project words for Soniox, as JSON. |
FA_STOP | The file that tells the stream to finish. |
FA_INPUT | An audio file to use instead of the microphone (for tests). |
| Problem | Solution |
|---|---|
A warning says the built-in /voice is on | Run /voice off once. Each /voice command toggles it. |
| "Voice: nothing heard" | Check the microphone with /fa mic. Check that your terminal has microphone access in System Settings → Privacy & Security → Microphone. |
| "No key: set SONIOX_API_KEY …" | Save your key (step 2 of Quick start). |
| Recording does not start | Start the recording on an empty prompt, or press <kbd>Space</kbd> twice quickly in text. Make sure .venv exists in the mod folder. |
| The cursor moves back and forth at the start of a hold | Claude Code draws each key before a mod can remove it. When a hold starts, the mod adds the chord space space to ~/.claude/keybindings.json. Then Claude Code takes the held <kbd>Space</kbd> itself, and the cursor stops. Claude Code reads the changed file after about 2 seconds, so the first 2 seconds of a hold still flicker. The mod removes the chord when you release <kbd>Space</kbd>. To avoid the flicker completely, use a shortcut and run /fa space off. |
| A key that you type right after a recording does not appear | After a hold, Claude Code can still wait for the second key of the space space chord for up to about 3 seconds, and it drops the next key. Wait until the text is in the prompt box, then type. |
| <kbd>Space</kbd> does not type in the prompt box | The space space chord stayed in ~/.claude/keybindings.json (for example, after a crash). The mod removes it when a session starts. To remove it now, run /reload-plugins, or delete the "space space" line from the file. |
| The shortcut does nothing | Run /fa key to see it. Check that no other binding in ~/.claude/keybindings.json uses the same keys. |
| "This does not look like a request for Claude" | JEV decided that you did not talk to Claude. Your text is in the prompt box. Press <kbd>Enter</kbd> to send it, or delete it. |
| "The translation did not finish" | The recording was long, and Soniox did not finish in time. Claude translated your words instead. Run /fa last to see the Persian text. |
| You cannot find what you said | Run /fa last. |
| "Not sent: this
hooks/register.tsx 1142 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Engine, Live, Mode, Prefs, Undo } from '../types'
5import { JEV_QUESTIONS } from './jev'
6import { PROMPTS } from './prompts'
7
8const STOP = '/tmp/persian-voice.stop' // stream.py finishes cleanly when this file appears
9// ponytail: macOS paths for Homebrew's ffmpeg; the module has no Node/os access to look it up
10const PATH = '/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin'
11
12type Paths = { home: string; python: string; script: string; terms: string }
13let cachedPaths: Paths | null = null
14
15// The plugin's own files, and the person's word list (user data, kept outside the plugin).
16async function paths($: EngineInterface): Promise<Paths> {
17 if (cachedPaths) return cachedPaths
18 const home = (await $.env.get('HOME').catch(() => undefined)) ?? ''
19 const root = $.plugin.root
20 cachedPaths = {
21 home,
22 python: `${root}/.venv/bin/python`,
23 script: `${root}/stt/stream.py`,
24 terms: `${home}/.config/persian-voice/terms.txt`,
25 }
26 return cachedPaths
27}
28
29// ponytail: a terminal sends no key-release, so "hold" = the key's auto-repeat presses; it ends when they stop
30const HOLD_GAP_MS = 1000 // a 2nd press within this of the 1st is a repeat, not a tap (covers the OS repeat delay)
31const RELEASE_MS = 700 // no repeat for this long = released
32const FINISH_MS = 5000 // after Stop, wait this long at most for Soniox to finalize the translation
33const LOCAL_FINISH_MS = 120000 // local Whisper reads the whole recording after Stop (and downloads its model on the first run)
34const MAX_MS = 5 * 60_000 // a forgotten tap-mode session stops by itself
35const UNDO_MS = 20_000 // how long "Use plain translation" stays
36const LAST = 'lastDictation' // store key: the last recording's words, for /fa last
37const POLISH_MODEL = 'claude-sonnet-5-5'
38// Picked with tools/rewrite_eval.py, 3 runs each: low and medium both passed 33/33, with API-time medians
39// of 2.1–2.7 s and 2.1–3.5 s. Same quality and speed, so low: it skips thinking on most short rewrites.
40const POLISH_EFFORT = 'low'
41const POLISH_MS = 8000 // slower than this: use the plain translation
42
43const live = atom({ plugin: 'persian-voice', key: 'live' } as const, null as Live | null)
44const DEFAULT_PREFS: Prefs = { engine: 'soniox', mode: 'prompt', autoSend: false, mic: 'default', holdSpace: true, shortcut: null }
45const prefs = atom({ plugin: 'persian-voice', key: 'prefs' } as const, DEFAULT_PREFS)
46const undo = atom({ plugin: 'persian-voice', key: 'undo' } as const, null as Undo | null)
47
48// ---------- Polish: rules first (instant), a model only when the text needs judgement ----------
49
50// label: the button and the dialog; about: one line for the person.
51const MODES: Record<Mode, { label: string; about: string }> = {
52 auto: {
53 label: 'Auto',
54 about: 'Like Prompt, but JEV can also pick Spec or Commit msg from what you said. Without a JEV key it works like Prompt.',
55 },
56 prompt: {
57 label: 'Prompt',
58 about: 'Claude rewrites what you said into a clear prompt, shaped for the task JEV detects (bug fix, feature, question, story …). Short, clear requests stay as they are.',
59 },
60 chat: {
61 label: 'Prompt (reads this chat)',
62 about: 'Like Prompt, but Claude always reads this conversation, so "fix that bug" becomes "fix the null check in parseOrder". Slower.',
63 },
64 spec: { label: 'Spec', about: 'For thinking aloud: Claude turns it into a task spec with Goal, Context, Requirements and Done when.' },
65 commit: { label: 'Commit msg', about: 'Claude turns what you said into a git commit message.' },
66 exact: { label: 'Exact', about: 'Only the translation, with fillers like "um" removed. No AI rewrite.' },
67}
68const MODE_ORDER: Mode[] = ['auto', 'prompt', 'chat', 'spec', 'commit', 'exact']
69
70// Who recognizes the speech. Soniox also translates; Gemini only transcribes, so Claude makes the English.
71const ENGINES: Record<Engine, string> = { soniox: 'Soniox', google: 'Google (Gemini 3.5 Transcribe)', local: 'Local Whisper (free, offline)' }
72
73// The rewrite prompts per task kind live in prompts.ts (tools/rewrite_eval.py runs that same object).
74// JEV's `kind` question (jev.ts) picks one; its options are these names.
75const KINDS = {
76 general: 'Prompt',
77 bug: 'Bug fix',
78 feature: 'Feature',
79 refactor: 'Refactor',
80 test: 'Tests',
81 review: 'Review',
82 question: 'Question',
83 story: 'Story editor',
84 spec: 'Spec',
85 commit: 'Commit msg',
86} as const
87export type Kind = keyof typeof KINDS
88export const KIND_NAMES = Object.keys(KINDS) as Kind[]
89// These reshape the text into another form, so they run even on a clean request.
90const ALWAYS_REWRITE: Kind[] = ['spec', 'commit', 'story']
91
92const FILLERS = /\b(?:u+m+|u+h+m*|e+r+m+|h+m+|a+h+)\b[,.]?\s*/gi
93const REPEATS = /\b(\w+)(?:\s+\1\b)+/gi
94const CORRECTIONS = /\b(?:no wait|sorry|i mean|actually|scratch that|rather|no no)\b/i
95
96/** Removes fillers and stutters ("um", "the the") without a model. */
97export function preclean(text: string) {
98 const t = text
99 .replace(FILLERS, '')
100 .replace(REPEATS, '$1')
101 .replace(/\s+([,.?!])/g, '$1')
102 .replace(/\s{2,}/g, ' ')
103 .trim()
104 return t.charAt(0).toUpperCase() + t.slice(1)
105}
106
107/** Prompt mode skips the model for a short request with no self-correction in it. */
108export function needsModel(mode: Mode, raw: string) {
109 if (mode === 'exact') return false
110 if (mode !== 'prompt') return true
111 return raw.split(/\s+/).length > 8 || CORRECTIONS.test(raw)
112}
113
114// ---------- JEV (TypeSafe): fast typed judgments before and after the Claude rewrite ----------
115// The questions live in jev.ts; tools/jev_eval.py scores that same object on labelled cases.
116// One proposition per question, and code combines the answers (TypeSafe's guidance).
117
118const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
119const JEV_MS = 2500 // slower than this: carry on without JEV (most calls take ~1 s)
120const JEV_THRESHOLD = 0.5 // ponytail: JEV's yes/no midpoint; retune with tools/jev_eval.py
121const JEV_ASIDE_BELOW = 0.2 // is_for_agent: below this the text stays in the box, not sent and not rewritten
122const JEV_HOLD_AT = 0.3 // is_irreversible: a missed one costs more than one extra Enter
123const JEV_CHAT = 4 // recent messages sent with the prompt (JEV's state limit is ~32k tokens)
124
125type JevSet = keyof typeof JEV_QUESTIONS
126
127/** JEV's answer per question: a noul's probability or a choice's option. */
128export type Answers = Record<string, number | string>
129
130/** Reads JEV's response; null when any asked question has no usable answer. */
131export function parseJev(stdout: string, names: string[]): Answers | null {
132 try {
133 const a = JSON.parse(stdout)?.answers ?? {}
134 const out: Answers = {}
135 for (const k of names) {
136 const isChoice = a[k]?.type === 'choice'
137 const v = isChoice ? a[k].choice : a[k]?.noul
138 if (typeof v !== (isChoice ? 'string' : 'number')) return null
139 out[k] = v
140 }
141 return out
142 } catch {
143 return null
144 }
145}
146
147/**
148 * What happens to the dictated text. Nothing here deletes it: `aside` (JEV: not for the assistant)
149 * and `hold` (cannot be undone) only keep it in the box unsent. `kind`: the rewrite prompt, or null for none.
150 */
151export type Plan = { aside: boolean; hold: boolean; kind: Kind | null; useChat: boolean }
152
153/**
154 * Turns the "before" answers into a plan; without them (no key, slow, error), the old rules.
155 * `isComplete` is false when the stream stopped before Soniox finished: the draft may miss the end.
156 */
157export function plan(mode: Mode, raw: string, a: Answers | null, isComplete = true): Plan {
158 const aside = !!a && (a.is_for_agent as number) < JEV_ASIDE_BELOW
159 const hold = !!a && (a.is_irreversible as number) >= JEV_HOLD_AT
160 const picked: Kind = KIND_NAMES.includes(a?.kind as Kind) ? (a?.kind as Kind) : 'general'
161 // Spec and Commit msg set by hand win; Auto takes any kind; Prompt and Chat leave spec and commit to the person.
162 const kind: Kind = mode === 'spec' || mode === 'commit' ? mode : mode === 'auto' || (picked !== 'spec' && picked !== 'commit') ? picked : 'general'
163 const yes = (k: string) => !!a && (a[k] as number) >= JEV_THRESHOLD
164 const vague = yes('has_vague_reference') || yes('chat_resolves_reference')
165 const flaw = a ? vague || yes('has_noise') || yes('is_rambling') || yes('mistranslated') : needsModel(mode === 'auto' ? 'prompt' : mode, raw)
166 const rewrite = mode !== 'exact' && !aside && (!isComplete || flaw || ALWAYS_REWRITE.includes(kind))
167 // the chat fork reads the whole transcript, so it can name what "that bug" means
168 return { aside, hold, kind: rewrite ? kind : null, useChat: rewrite && (mode === 'chat' || vague) }
169}
170
171type RetryProblem = 'adds_request' | 'changes_fact' | 'drops_fact'
172
173/** The "after" answers: what the rewrite added, changed or dropped; none when JEV did not answer. */
174export function rewriteProblems(a: Answers | null): RetryProblem[] {
175 const keys: RetryProblem[] = ['adds_request', 'changes_fact', 'drops_fact']
176 return a ? keys.filter(k => (a[k] as number) >= JEV_THRESHOLD) : []
177}
178
179/** Asks one question set of jev.ts; null when JEV cannot answer in time. */
180async function askJev($: EngineInterface, set: JevSet, state: Record<string, string>): Promise<Answers | null> {
181 const { home } = await paths($)
182 const key = ((await $.env.get('JEV_API_KEY').catch(() => undefined)) ?? (await $.fs.read(`${home}/.config/typesafe/key`).catch(() => ''))).trim()
183 if (!key) return null
184 const questions = JEV_QUESTIONS[set]
185 // State holds evidence only, as JSON; the instructions refer to its fields by backticked name.
186 const recent_chat = (await $.session.messages().catch(() => []))
187 .slice(-JEV_CHAT)
188 .filter(m => m.text.trim())
189 .map(m => ({ role: m.role, text: m.text.slice(0, 1500) }))
190 const body = { model: 'jev-latest', state: { recent_chat, ...state }, questions }
191 const r = await $.process
192 .run(
193 ['sh', '-c', `curl -sS -m ${JEV_MS / 1000} ${JEV_URL} -H "Authorization: Bearer $JEV_API_KEY" -H "Content-Type: application/json" --data @-`],
194 { env: { PATH, HOME: home, JEV_API_KEY: key }, stdin: JSON.stringify(body), timeoutMs: JEV_MS + 1000 },
195 )
196 .catch(() => null)
197 return r && r.exitCode === 0 ? parseJev(r.stdout, Object.keys(questions)) : null
198}
199
200/** The rewrite's system prompt for one kind; tools/rewrite_eval.py builds the same string. */
201export function systemPrompt(kind: Kind) {
202 return `${PROMPTS.rules.join('\n')}\n\n<task>\n${PROMPTS.kinds[kind]}\n</task>`
203}
204
205/** A first rewrite that JEV found faults in, and those faults: the second try fixes them. */
206export type Retry = { previous: string; problems: RetryProblem[] }
207
208/** The rewrite's user message; tools/rewrite_eval.py builds the same string. */
209export function userPrompt(fa: string, en: string, retry?: Retry) {
210 const input = `<spoken>\n${fa}\n</spoken>\n<draft>\n${en}\n</draft>`
211 if (!retry) return `${input}\nRewrite the draft now.`
212 const r = PROMPTS.retry
213 const list = retry.problems.map(p => `- ${r[p]}`).join('\n')
214 return `${input}\n<previous_rewrite>\n${retry.previous}\n</previous_rewrite>\n${r.ask}\n${list}\n${r.end}`
215}
216
217// Returns the polished text; the caller falls back to the plain text on a throw.
218async function polish($: EngineInterface, kind: Kind, useChat: boolean, fa: string, en: string, retry?: Retry) {
219 const system = systemPrompt(kind)
220 const prompt = userPrompt(fa, en, retry)
221 if (useChat) {
222 // The main thread's own transcript (prompt-cached), so "that bug" can be resolved.
223 const r = await Promise.race([
224 $.model.fork({ prompt: `${system}\n\n${PROMPTS.chat}\n\n${prompt}` }),
225 $.clock.sleep(POLISH_MS).then(() => null),
226 ])
227 if (r?.isAnswered && r.text.trim()) return r.text.trim()
228 }
229 const r = await $.model.complete({
230 model: POLISH_MODEL,
231 effort: POLISH_EFFORT,
232 system,
233 prompt,
234 maxTokens: 2000, // room for a long spec, and for thinking, which counts toward this limit
235 timeoutMs: POLISH_MS,
236 })
237 return r.isAnswered && r.text.trim() ? r.text.trim() : en
238}
239
240// ---------- Settings (global: $.store is one file under the Claude config dir) ----------
241
242async function loadPrefs($: EngineInterface): Promise<Prefs> {
243 const p = await $.store.get('prefs').catch(() => undefined)
244 if (p && typeof p === 'object') return { ...DEFAULT_PREFS, ...(p as Partial<Prefs>) }
245 const old = await $.store.get('isPolishOn').catch(() => undefined) // v0.2 setting
246 return { ...DEFAULT_PREFS, mode: old === false ? 'exact' : 'prompt' }
247}
248
249// A copy of prefs.holdSpace the prompt.edit hook reads without waiting (see there).
250let isHoldSpace = DEFAULT_PREFS.holdSpace
251
252async function applyPrefs($: EngineInterface, p: Prefs) {
253 isHoldSpace = p.holdSpace
254 await update($, prefs, () => p)
255}
256
257async function savePrefs($: EngineInterface, change: Partial<Prefs>) {
258 const p = { ...(await loadPrefs($)), ...change }
259 await $.store.set('prefs', p)
260 await applyPrefs($, p)
261 return p
262}
263
264// ---------- Custom shortcut: a keybinding to an engine action the talk Button names ----------
265
266// ponytail: a plugin cannot define its own keybinding action, so it borrows one with no default key
267// whose handler only lives in Claude Code's old diff panel (off unless cc-plugin-diff is disabled)
268const ACTION = 'app:toggleDiffPreSession'
269// Claude Code fires a Button's action only for a chord or a key with a modifier; a bare key would type.
270const SHORTCUT = /^(?:ctrl|meta|alt|shift|cmd)\+\S+(?: \S+)?$/
271
272type Keybindings = { bindings?: { context: string; bindings: Record<string, string | null> }[]; [k: string]: unknown }
273
274// Binds `chord` (or nothing, for null) to ACTION in ~/.claude/keybindings.json, keeping the rest.
275async function bindShortcut($: EngineInterface, chord: string | null) {
276 const file = `${(await paths($)).home}/.claude/keybindings.json`
277 const text = await $.fs.read(file).catch(() => '')
278 const kb: Keybindings = typeof text === 'string' && text.trim() ? JSON.parse(text) : {}
279 kb.bindings ??= []
280 for (const block of kb.bindings) {
281 for (const [key, action] of Object.entries(block.bindings)) if (action === ACTION) delete block.bindings[key]
282 }
283 if (chord) {
284 let global = kb.bindings.find(b => b.context === 'Global')
285 if (!global) kb.bindings.push((global = { context: 'Global', bindings: {} }))
286 global.bindings[chord] = ACTION
287 }
288 await $.fs.write(file, `${JSON.stringify(kb, null, 2)}\n`)
289}
290
291// While a hold-Space recording runs, the chord `space space` presses the plugin's button (ACTION).
292// Claude Code then takes the held Space before the prompt draws it, so the cursor stays still, and
293// every 2nd repeat reaches press(), which keeps the recording alive. Tested live (Claude Code 2.1.287):
294// the binding applies ~2 s after the write; a lone Space left open at release is dropped, not typed,
295// and a key typed within ~3 s after it can be dropped too. Off at every exit, and at session start
296// in case a crash left it on (it would delay every Space in the prompt).
297const SPACE_CHORD = 'space space'
298let isChordArmed = false // this recording turned the chord on
299
300async function setSpaceChord($: EngineInterface, on: boolean) {
301 const file = `${(await paths($)).home}/.claude/keybindings.json`
302 const text = await $.fs.read(file).catch(() => '')
303 const kb: Keybindings = typeof text === 'string' && text.trim() ? JSON.parse(text) : {}
304 kb.bindings ??= []
305 let chat = kb.bindings.find(b => b.context === 'Chat')
306 if ((chat?.bindings[SPACE_CHORD] === ACTION) === on) return // already so: no write, no reload
307 if (on) {
308 if (!chat) kb.bindings.push((chat = { context: 'Chat', bindings: {} }))
309 chat.bindings[SPACE_CHORD] = ACTION
310 } else if (chat) delete chat.bindings[SPACE_CHORD]
311 await $.fs.write(file, `${JSON.stringify(kb, null, 2)}\n`)
312}
313
314// One write at a time, in order: an off that overtook a pending on would leave the chord on.
315let chordWrites: Promise<unknown> = Promise.resolve()
316const queueChord = ($: EngineInterface, on: boolean) => (chordWrites = chordWrites.then(() => setSpaceChord($, on)).catch(() => {}))
317
318function armChord($: EngineInterface) {
319 if (isChordArmed) return
320 isChordArmed = true
321 void queueChord($, true)
322}
323
324function disarmChord($: EngineInterface) {
325 if (!isChordArmed) return
326 isChordArmed = false
327 void queueChord($, false)
328}
329
330async function keyCommand($: EngineInterface, arg: string) {
331 if (!arg) {
332 const p = await loadPrefs($)
333 return p.shortcut
334 ? `⌨ Shortcut: ${p.shortcut} (press to start, press again to stop). Remove it: /fa key off`
335 : '⌨ No shortcut. Set one, for example: /fa key ctrl+x v'
336 }
337 if (arg === 'off') {
338 await bindShortcut($, null)
339 const p = await savePrefs($, { shortcut: null, holdSpace: true })
340 return `⌨ Shortcut removed. Hold Space is ${p.holdSpace ? 'on' : 'off'}.`
341 }
342 const chord = arg.toLowerCase()
343 if (!SHORTCUT.test(chord)) {
344 return `"${arg}" cannot be a shortcut. Use a key with ctrl, meta, alt or shift (ctrl+r), or a chord (ctrl+x v).`
345 }
346 await bindShortcut($, chord)
347 await savePrefs($, { shortcut: chord })
348 return `⌨ Shortcut: ${chord}. Press it to start, and again to stop.\nIt is saved in ~/.claude/keybindings.json. Choose a combination Claude Code does not already use.\nTo stop holding Space: /fa space off`
349}
350
351async function spaceCommand($: EngineInterface, arg: string) {
352 const cur = await loadPrefs($)
353 const isOn = arg === 'on' ? true : arg === 'off' ? false : !cur.holdSpace
354 const p = await savePrefs($, { holdSpace: isOn })
355 const other = p.shortcut ? `Use ${p.shortcut} or /fa rec.` : 'Set a shortcut with /fa key ctrl+x v, or use /fa rec.'
356 return p.holdSpace ? '␣ Hold Space to talk: on' : `␣ Hold Space to talk: off. Space only types now. ${other}`
357}
358
359async function setMode($: EngineInterface, mode?: Mode) {
360 const cur = (await loadPrefs($)).mode
361 const next = mode ?? MODE_ORDER[(MODE_ORDER.indexOf(cur) + 1) % MODE_ORDER.length]
362 return savePrefs($, { mode: next })
363}
364
365// ---------- Recognition context: project words + the person's words + learned words ----------
366
367let context: { cwd: string; json: string } | null = null
368
369const SKIP_FILES = /\.(png|jpe?g|gif|svg|ico|webp|lock|map|woff2?|ttf|otf|mp[34]|wav|zip|gz|pdf)$/i
370
371// The global list plus `.persian-voice-terms.txt` in the project folder (same format; commit it to share it).
372async function readTerms($: EngineInterface) {
373 const cwd = await $.session.cwd().catch(() => '')
374 const read = (file: string) => $.fs.read(file).catch(() => '')
375 const texts = await Promise.all([read((await paths($)).terms), cwd ? read(`${cwd}/.persian-voice-terms.txt`) : ''])
376 const lines = texts.flatMap(t => (typeof t === 'string' ? t : '').split('\n')).map(l => l.trim()).filter(l => l && !l.startsWith('#'))
377 const pairs = lines.filter(l => l.includes('=')).map(l => l.split('=').map(s => s.trim()) as [string, string])
378 return { words: lines.filter(l => !l.includes('=')), pairs: pairs.filter(([s, t]) => s && t) }
379}
380
381async function learnedTerms($: EngineInterface) {
382 const l = await $.store.get('learned').catch(() => [])
383 return Array.isArray(l) ? l.filter((w): w is string => typeof w === 'string') : []
384}
385
386// Builds Soniox's `context` for the session's project; cached per working directory.
387async function refreshContext($: EngineInterface) {
388 try {
389 const cwd = await $.session.cwd()
390 const { home } = await paths($)
391 const run = (argv: string[]) =>
392 $.process.run(argv, { cwd, env: { PATH, HOME: home } }).then(r => (r.exitCode === 0 ? r.stdout : '')).catch(() => '')
393 const [branch, files, mine, learned] = await Promise.all([
394 run(['git', 'rev-parse', '--abbrev-ref', 'HEAD']),
395 run(['git', 'ls-files']),
396 readTerms($),
397 learnedTerms($),
398 ])
399 const project = cwd.split('/').pop() ?? ''
400 const fileNames = files.split('\n').map(f => f.split('/').pop() ?? '').filter(f => /^\w[\w.-]{2,40}$/.test(f) && !SKIP_FILES.test(f))
401 const own = [...new Set([project, branch.trim(), ...mine.words, ...learned, ...mine.pairs.map(p => p[1])])].filter(t => t && t !== 'HEAD')
402 const json = JSON.stringify({
403 general: [
404 { key: 'domain', value: 'Software development' },
405 { key: 'topic', value: 'Spoken instructions for an AI coding agent (Claude Code)' },
406 { key: 'project', value: project },
407 { key: 'style', value: 'Code identifiers, file names and technical terms stay in English, Latin script' },
408 ],
409 terms: [...new Set([...own, ...fileNames])].slice(0, 150),
410 translation_terms: [
411 ...mine.pairs.map(([source, target]) => ({ source, target })),
412 ...own.slice(0, 60).map(t => ({ source: t, target: t })), // names stay as they are
413 ],
414 })
415 context = { cwd, json }
416 } catch {
417 context = null // dictation still works, only without the project words
418 }
419}
420
421/** Technical-looking words the person added while editing the dictated text before sending. */
422export function newTerms(filled: string, sent: string) {
423 const split = (s: string) => s.split(/[\s,;:!?()"'`]+/).map(w => w.replace(/\.+$/, '')).filter(Boolean)
424 const a = split(filled)
425 const b = split(sent)
426 const said = new Set(a)
427 if (b.filter(w => said.has(w)).length < said.size / 2) return [] // not an edit of the dictated text
428 // Only one corrected span counts: the words before and after it are unchanged, and it replaced some dictated words.
429 let head = 0
430 while (head < a.length && head < b.length && a[head] === b[head]) head++
431 let tail = 0
432 while (tail < a.length - head && tail < b.length - head && a[a.length - 1 - tail] === b[b.length - 1 - tail]) tail++
433 const removed = a.length - head - tail
434 const added = b.slice(head, b.length - tail)
435 if (removed < 1 || added.length > 4) return [] // a pure addition is new text, not a correction
436 // camelCase, snake_case, a path or a digit; a plain capital word like "STE" is an acronym, not a misheard term
437 const isTechnical = (w: string) => /^[A-Za-z_$][\w$.\-/]*$/.test(w) && /[a-z][A-Z]|[_$./\d]/.test(w.slice(1))
438 return [...new Set(added.filter(w => !said.has(w) && isTechnical(w)))].slice(0, 10)
439}
440
441async function learn($: EngineInterface, filled: string, sent: string) {
442 const add = newTerms(filled, sent)
443 if (!add.length) return
444 const merged = [...new Set([...add, ...(await learnedTerms($))])].slice(0, 200)
445 await $.store.set('learned', merged)
446 $.ui.toast(`📚 Learned: ${add.join(', ')} (wrong? /fa forget <word>)`)
447 void refreshContext($)
448}
449
450// ---------- Recording ----------
451
452let isActive = false
453let isTentative = false // started on the 1st space of an empty prompt; cancelled when no repeat follows
454let isCancelled = false
455let isStopping = false
456let sendNow = false // the Send button: send this recording even with auto-send off
457let isFinishing = false
458let isPolishing = false
459let polishLabel = '' // what the spinner says: the rewrite kind, or "Translating"
460let isHeld = false
461let lastPress = 0
462let lastSpace: number | null = null // time of the last space typed with nothing else between
463let startedAt = 0
464let frame = 0 // animation step
465let levels: number[] = [] // recent mic levels for the meter
466let latest: Live = { fa: '', en: '' } // stream.py's last line, also while tentative (not drawn yet)
467let lastFill: string | null = null // dictated text in the box, to learn from the person's edits
468
469let isSpaceHold = false // this recording was started by holding Space (not the shortcut, button or /fa)
470
471// Synchronous on purpose: the prompt.edit hook calls it before it answers the key (see there).
472// Each recording's number. abort() frees the plugin at once, so the next recording can start while the
473// aborted one still winds down; that one then checks isMine() before it touches the shared state.
474let gen = 0
475let abortedGen = -1
476
477// Discards the recording that runs now, in any phase (recording, finishing, rewriting). Its words stay in /fa last.
478function abort($: EngineInterface) {
479 if (!isActive) return
480 abortedGen = gen
481 isActive = isTentative = isPolishing = isFinishing = false
482 disarmChord($)
483 // The stop file now, before any new recording can start (a new stream.py deletes an old one when it
484 // starts); start()'s poll loop also kills stream.py. A stop file written later could stop the next stream.
485 background(() => $.fs.write(STOP, ''))
486 background(() => update($, live, () => null))
487 $.ui.toast('🎙 Cancelled: nothing was put in or sent. Start again when you are ready. /fa last shows what you said.')
488}
489
490// `taps`: the spaces this start already counts as taps (see spaceWhileRecording).
491function begin($: EngineInterface, held: boolean, tentative = false, tapsSoFar = 0) {
492 gen++
493 isActive = true
494 isSpaceHold = held
495 isTentative = tentative
496 taps = tapsSoFar
497 isTapMode = false
498 isCancelled = isStopping = sendNow = false
499 isHeld = held
500 lastPress = 0 // start() sets the clock times
501 levels = []
502 latest = { fa: '', en: '' }
503 // A throw before start()'s own cleanup would leave isActive set for good (Space swallowed, no new recording).
504 background(() =>
505 start($).catch(async e => {
506 isActive = isTentative = isPolishing = isFinishing = false
507 $.ui.toast(`Voice error: ${String(e).slice(0, 120)}`)
508 await update($, live, () => null)
509 }),
510 )
511}
512
513// The talk / Stop button, its shortcut, and `/fa`: press to start, press again to stop.
514// A shortcut held down repeats; presses closer than HOLD_GAP_MS are that hold, and it ends on release.
515async function press($: EngineInterface) {
516 const now = await $.clock.now()
517 if (!isActive) {
518 begin($, false)
519 lastPress = now
520 } else if (isHeld || now - lastPress < HOLD_GAP_MS) {
521 isHeld = true
522 lastPress = now
523 } else isStopping = true
524}
525
526// Tap Space 3 times to record without holding; tap it once more to finish. A space this long after
527// the one before is a tap; a shorter gap is the key repeat of a hold.
528// ponytail: fixed gap; a slow macOS key-repeat setting (180 ms or more) reads as taps, then raise it
529const TAP_MIN_MS = 140
530const TAPS_TO_START = 3
531let taps = 0
532let isTapMode = false // started by taps: no release to wait for, the next tap finishes
533let lastSpaceAt = 0 // every space's time is read the same way, so the gaps compare
534
535// A space while a Space recording runs: a key repeat (hold), a tap toward tap mode, or the finishing tap.
536// `isFirst`: the space that started the recording; only its time is kept.
537async function spaceWhileRecording($: EngineInterface, isBurst: boolean, isFirst = false) {
538 if (isTapMode) {
539 isStopping = true
540 return
541 }
542 const now = await $.clock.now()
543 const gap = now - lastSpaceAt
544 lastSpaceAt = now
545 if (isFirst) return
546 if (isBurst || gap < TAP_MIN_MS) return repeat($)
547 lastPress = now // a tap keeps a tentative start alive too
548 if (++taps < TAPS_TO_START) return
549 isTapMode = true
550 isHeld = false // nothing to release: the poll loop waits for isStopping
551 if (isTentative) {
552 isTentative = false
553 await update($, live, () => ({ ...latest, ms: now - startedAt }))
554 } else await redraw($)
555}
556
557// A held key's repeat: keep the recording alive (and show it, if it was tentative).
558async function repeat($: EngineInterface) {
559 const now = await $.clock.now()
560 lastPress = Math.max(lastPress, now)
561 armChord($) // a repeat means a real hold: from now on Claude Code takes the held Space (see SPACE_CHORD)
562 if (isTentative) {
563 isTentative = false
564 await update($, live, () => ({ ...latest, ms: now - startedAt }))
565 }
566}
567
568// Work left running after a hook returns; a hot reload aborts it, which is fine.
569const background = (work: () => Promise<unknown>) => void work().catch(() => {})
570
571const redraw = ($: EngineInterface) => update($, live, l => l && { ...l })
572
573// Runs a model step with the spinner and its label in the REC panel.
574async function spinning<T>($: EngineInterface, label: string, work: () => Promise<T>): Promise<T> {
575 isPolishing = true
576 polishLabel = label
577 background(async () => {
578 while (isPolishing) {
579 await $.clock.sleep(120)
580 frame++
581 await redraw($)
582 }
583 })
584 try {
585 return await work()
586 } finally {
587 isPolishing = false
588 }
589}
590
591// Stream mic -> Soniox until Stop, clean up, then put the text in the prompt box (or send it).
592async function start($: EngineInterface) {
593 const myGen = gen // read before the first await: begin() calls this right after gen++
594 const isMine = () => gen === myGen && abortedGen !== myGen
595 isFinishing = isPolishing = false
596 startedAt = await $.clock.now()
597 lastPress = Math.max(lastPress, startedAt)
598 if (!isTentative) await update($, live, () => ({ fa: '', en: '', ms: 0 }))
599 const p = await paths($)
600 const pref = await read($, prefs)
601 const finishMs = pref.engine === 'local' ? LOCAL_FINISH_MS : FINISH_MS
602 const env: Record<string, string> = { PATH, HOME: p.home, FA_STOP: STOP, FA_MIC: pref.mic, FA_ENGINE: pref.engine }
603 if (context) env.FA_CONTEXT = context.json
604 if (pref.engine === 'google') {
605 const g = await $.env.get('GEMINI_API_KEY').catch(() => undefined)
606 if (g) env.GEMINI_API_KEY = g // else stream.py reads ~/.config/gemini/key
607 }
608 const proc = $.process.spawn({ argv: [p.python, p.script], env })
609 let last: Live = { fa: '', en: '' }
610 let err = ''
611 let isRunning = true
612 // Polls for cancel / Stop / release on its own, so it works even while stream.py prints nothing.
613 background(async () => {
614 let stopAt: number | null = null
615 try {
616 while (isRunning) {
617 await $.clock.sleep(100)
618 if (!isMine()) {
619 // Aborted: kill stream.py now. No stop file here: written this late, it could stop the next stream.
620 if (isRunning) await proc.return({ code: null, signal: null })
621 return
622 }
623 const now = await $.clock.now()
624 if (isCancelled || (isTentative && now - lastPress > HOLD_GAP_MS)) {
625 isCancelled = true // start() then drops the text
626 await $.fs.write(STOP, '') // stream.py exits by itself...
627 await $.clock.sleep(finishMs)
628 if (isRunning) await proc.return({ code: null, signal: null }) // ...or is killed
629 return
630 }
631 if (stopAt === null && now - startedAt > MAX_MS) {
632 isStopping = true
633 $.ui.toast('🎙 Stopped after 5 minutes')
634 }
635 // A tentative start is not released but cancelled (above): no repeat came, so nothing was held.
636 if (stopAt === null && (isStopping || (isHeld && !isTentative && now - lastPress > RELEASE_MS))) {
637 stopAt = now
638 disarmChord($) // released: Space types again (after Claude Code reloads the file)
639 isFinishing = true
640 await redraw($)
641 await $.fs.write(STOP, '') // stream.py closes the mic, waits for the last words, then exits
642 }
643 if (stopAt !== null && now - stopAt > finishMs) {
644 await proc.return({ code: null, signal: null }) // kills stream.py
645 return
646 }
647 }
648 } catch {
649 // If this poller dies, nothing else would ever stop stream.py: kill it so start() can finish.
650 if (isRunning) await proc.return({ code: null, signal: null }).catch(() => {})
651 }
652 })
653 try {
654 let buf = ''
655 for await (const { stream, text } of proc) {
656 if (stream === 'stderr') err += text
657 else {
658 buf += text
659 const lines = buf.split('\n')
660 buf = lines.pop() ?? ''
661 for (const line of lines) {
662 try { last = JSON.parse(line) } catch { continue }
663 if (!isMine()) continue // aborted: the shared view belongs to the next recording
664 latest = last
665 frame++
666 levels = [...levels.slice(-15), last.lvl ?? 0]
667 const ms = (await $.clock.now()) - startedAt
668 if (!isTentative) await update($, live, () => ({ ...last, ms }))
669 }
670 }
671 }
672 } catch (e) {
673 err += String(e)
674 } finally {
675 isRunning = false
676 if (isMine()) isFinishing = false
677 }
678 try {
679 const fa = last.fa.trim()
680 let raw = last.en.trim()
681 if (!isMine()) {
682 // Aborted by the person: nothing goes in, but the words stay in /fa last.
683 if (fa || raw) await $.store.set(LAST, { fa, en: raw }).catch(() => {})
684 return
685 }
686 // Only a tentative start is cancelled (one tap of Space, then typing): nothing was dictated, and saving
687 // its noise would overwrite the real last dictation.
688 if (isCancelled) return
689 // The safety copy comes before every other step, so none of them can lose what was said: /fa last shows it.
690 if (fa || raw) await $.store.set(LAST, { fa, en: raw }).catch(() => {})
691 if (!fa && !raw) {
692 $.ui.toast(`Voice: ${err.trim().split('\n').pop() || 'nothing heard'}`)
693 return
694 }
695 const p = await loadPrefs($)
696 // Gemini only transcribes: Claude makes the English draft, then the normal JEV / rewrite steps run on it.
697 if (!raw && p.engine !== 'soniox') raw = await spinning($, 'Translating', () => polish($, 'general', false, fa, '')).catch(() => '')
698 if (!raw) {
699 // A long recording can end before Soniox translates it: Claude translates the Persian, else the Persian goes in.
700 $.ui.toast('🎙 The translation did not finish, so Claude translates your words. /fa last shows them.')
701 const t = await spinning($, 'Translating', () => polish($, 'general', false, fa, '')).catch(() => '')
702 if (isMine()) await put($, t || fa, false) // not when aborted while Claude translated
703 return
704 }
705 const plain = preclean(raw)
706 const isComplete = last.done === true
707 const j = plan(p.mode, raw, await askJev($, 'before', { spoken: fa, prompt: plain }), isComplete)
708 if (j.aside) $.ui.toast('🎙 This does not look like a request for Claude, so it was not sent or rewritten.')
709 else if (!isComplete && !j.kind) $.ui.toast('🎙 The translation may miss your last words. /fa last shows what you said.')
710 let text = plain
711 if (j.kind) {
712 const kind = j.kind
713 const check = async (t: string) => rewriteProblems(await askJev($, 'after', { spoken: fa, draft: plain, rewritten: t }))
714 text = await spinning($, KINDS[kind], async () => {
715 const t = await polish($, kind, j.useChat, fa, plain).catch(() => plain)
716 const problems = t === plain ? [] : await check(t)
717 if (!problems.length) return t
718 // One more try, told what JEV found (costs time only when the first rewrite failed).
719 if (!isMine()) return plain
720 const t2 = await polish($, kind, j.useChat, fa, plain, { previous: t, problems }).catch(() => plain)
721 if (t2 !== plain && !(await check(t2)).length) return t2
722 $.ui.toast('The rewrite changed what you said, so your own words are used')
723 return plain
724 })
725 }
726 if (!isMine()) return // aborted while JEV or Claude worked: nothing goes in, nothing is sent
727 const wantsSend = p.autoSend || sendNow
728 const send = wantsSend && !j.hold && !j.aside
729 if (wantsSend && j.hold && !j.aside) $.ui.toast('⚠ Not sent: this asks for something that cannot be undone. Check it, then press Enter.')
730 await put($, text, send)
731 if (text !== plain && !send) {
732 await update($, undo, () => ({ raw: plain, polished: text }))
733 background(async () => {
734 await $.clock.sleep(UNDO_MS)
735 await update($, undo, u => (u?.polished === text ? null : u))
736 })
737 }
738 } finally {
739 if (gen === myGen) {
740 // still the latest recording (aborted or not): reset; a newer one owns the shared state otherwise
741 disarmChord($)
742 isActive = isTentative = isPolishing = false
743 await update($, live, () => null)
744 }
745 }
746}
747
748// Inserts at the cursor with a separating space; with auto-send, sends the whole draft.
749async function put($: EngineInterface, text: string, autoSend: boolean) {
750 const box = await $.prompt.read()
751 const before = box.text.slice(0, box.cursor)
752 const piece = (before && !/\s$/.test(before) ? ' ' : '') + text
753 if (autoSend) {
754 const full = before + piece + box.text.slice(box.cursor)
755 await $.prompt.fill({ text: '', mode: 'replace' })
756 lastFill = null
757 // asUser: the model reads the dictation bare, without the "plugin sent a message" frame
758 background(async () => {
759 const r = await $.prompt.submit({ text: full, asUser: true }).catch(() => null)
760 if (r && r.drop === undefined) return
761 // Not sent (a hook dropped it, or the submit failed): the text goes back in the box, not lost.
762 await $.prompt.fill({ text: full, mode: 'insert' })
763 $.ui.toast('The prompt was not sent, so it is back in the prompt box')
764 })
765 return
766 }
767 // fill, not submit: Enter sends it as the person's own prompt
768 await $.prompt.fill({ text: piece, mode: 'insert' })
769 lastFill = text
770}
771
772async function usePlain($: EngineInterface) {
773 const u = await read($, undo)
774 await update($, undo, () => null)
775 if (!u) return
776 const box = await $.prompt.read()
777 if (!box.text.includes(u.polished)) {
778 $.ui.toast('The text was edited, so it was not swapped')
779 return
780 }
781 await $.prompt.fill({ text: box.text.replace(u.polished, u.raw), mode: 'replace' })
782 lastFill = u.raw
783}
784
785// ---------- Commands ----------
786
787const HELP = `Persian voice
788 hold space talk, release to finish
789 tap space 3x talk without holding; tap space once more to finish
790 a letter while a Space recording runs (or is translated): cancel it, nothing goes in
791 /fa settings: see and change everything below
792 /fa rec start / stop a recording without holding a key
793 /fa last show your last recording again and put it in the prompt box
794 /fa key [k] your own shortcut, e.g. /fa key ctrl+x v (press to start, again to stop); /fa key off
795 /fa space hold Space to talk on / off (off: Space only types)
796 /fa mode [m] cleanup: ${MODE_ORDER.join(' · ')}
797 /fa polish cleanup on / off (prompt <-> exact)
798 /fa send auto-send on / off
799 /fa engine [e] speech engine: ${Object.keys(ENGINES).join(' · ')} (google = Gemini 3.5 Transcribe, key in ~/.config/gemini/key; local = Whisper on this Mac, no key)
800 /fa mic [n] list / choose the microphone
801 /fa terms your word list for recognition (one per line, or "persian = english")
802 /fa forget w remove a word that the plugin learned by mistake
803Modes:
804${MODE_ORDER.map(m => ` ${m.padEnd(8)} ${MODES[m].about}`).join('\n')}
805Tip: speak in short, complete sentences. Agents follow spoken-formal input better than casual speech.`
806
807// The microphones ffmpeg sees, as [index, name].
808async function listMics($: EngineInterface) {
809 const r = await $.process
810 .run(['ffmpeg', '-hide_banner', '-f', 'avfoundation', '-list_devices', 'true', '-i', ''], {
811 env: { PATH, HOME: (await paths($)).home },
812 })
813 .catch(() => ({ stderr: '' }))
814 const out = r.stderr.split('audio devices:')[1] ?? ''
815 return [...out.matchAll(/\[(\d+)\] (.+)/g)].map(m => [m[1] ?? '', (m[2] ?? '').trim()] as const)
816}
817
818async function micCommand($: EngineInterface, arg?: string) {
819 if (arg) return `🎙 Microphone: ${(await savePrefs($, { mic: arg })).mic}`
820 const mics = await listMics($)
821 const cur = (await loadPrefs($)).mic
822 return mics.length
823 ? `${mics.map(([n, name]) => `${n === cur ? '▶' : ' '} ${n} ${name}`).join('\n')}\nChoose one: /fa mic <number>`
824 : 'No microphones found (is ffmpeg installed?)'
825}
826
827async function termsCommand($: EngineInterface) {
828 const [mine, learned] = await Promise.all([readTerms($), learnedTerms($)])
829 return `Word list: ${(await paths($)).terms}\n ${mine.words.length} words, ${mine.pairs.length} translations; ${learned.length} learned from your edits${
830 learned.length ? `: ${learned.slice(0, 12).join(', ')}` : ''
831 }\nProject words: .persian-voice-terms.txt in the project folder (same format).\nThe project's file names, name and branch are added by themselves.`
832}
833
834// ---------- Settings dialog (/fa, or ⚙ in the band) ----------
835
836const SETTINGS = 'fa-settings'
837// Read once when the dialog opens: ffmpeg and the word file are too slow to read on every redraw.
838let micOptions: { value: string; label: string }[] = []
839let wordsInfo = ''
840
841async function openSettings($: EngineInterface) {
842 const [mics, mine, learned, p] = await Promise.all([listMics($), readTerms($), learnedTerms($), paths($)])
843 micOptions = [{ value: 'default', label: 'System default' }, ...mics.map(([n, name]) => ({ value: n, label: name }))]
844 wordsInfo = `${mine.words.length} words, ${mine.pairs.length} translations, ${learned.length} learned · ${p.terms.replace(p.home, '~')}`
845 await applyPrefs($, await loadPrefs($))
846 return $.ui.open({ id: SETTINGS, title: 'Persian Voice · settings', focus: true, closeOnEscape: true, rows: 24 })
847}
848
849async function shortcutFromDialog($: EngineInterface, value: string) {
850 const reply = await keyCommand($, value.trim())
851 $.ui.toast(reply.split('\n')[0] ?? reply)
852}
853
854// ---------- Drawing ----------
855
856const SPIN = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏']
857const BARS = ' ▁▂▃▄▅▆▇█'
858const RED = '#e5534b'
859const AMBER = '#d4a72c'
860const ORANGE = '#d97757'
861const GREEN = '#57ab5a'
862
863const clockText = (ms = 0) => `${Math.floor(ms / 60_000)}:${String(Math.floor(ms / 1000) % 60).padStart(2, '0')}`
864
865export const register: Register = on => {
866 on('session.start', async ($, e, next) => {
867 await $.command.register({
868 name: 'fa',
869 description: 'Persian voice: settings · /fa rec records without holding · /fa help lists all commands',
870 })
871 await applyPrefs($, await loadPrefs($))
872 await queueChord($, false) // a crash or a reload mid-recording can leave it on
873 const { terms } = await paths($)
874 if (!(await $.fs.exists(terms).catch(() => true))) {
875 await $.fs.write(terms,'# Persian voice: words for speech recognition, one per line.\n# A line "persian = english" sets a translation, for example:\n# دیپلوی = deploy\n')
876 }
877 void refreshContext($)
878 return next(e)
879 })
880
881 // Hold space: on an empty prompt it starts at once (tentative, so the first words are kept) and
882 // is cancelled when no key-repeat follows; in text, a 2nd space right after the 1st is the repeat.
883 on('prompt.edit', async ($, e, next) => {
884 const isBurst = !e.key && /^ {2,}$/.test(e.inputText) // repeats folded into one edit
885 const isSpace = isBurst || (e.key ? e.key.key === ' ' || e.key.key === 'space' : e.inputText === ' ')
886 if (isSpace && !isHoldSpace) return next(e) // `/fa space off`: Space only types
887 if (!isSpace) {
888 lastSpace = null
889 if (isActive && isTentative) isCancelled = true // it was a leading space, then typing
890 else if (isActive && isSpaceHold && (e.inputText !== '' || e.end > e.start)) {
891 // A letter (or a deletion) during a Space recording cancels it, and does not type. Escape cannot:
892 // Claude Code gives a plugin no event for it. A Backspace in an empty box gives none either.
893 abort($)
894 return { text: e.text, cursor: e.cursor }
895 }
896 return next(e)
897 }
898 // The editor draws each key before this hook answers, so a swallowed space shows for the
899 // time the answer takes: the cursor jumps right, then back. Held space repeats ~30 times a
900 // second, so these paths answer at once and keep the clock work for afterwards.
901 if (isActive) {
902 if (!isSpaceHold) return next(e) // started by the shortcut, button or /fa: typing stays normal
903 background(() => spaceWhileRecording($, isBurst))
904 return next({ ...e, inputText: '' })
905 }
906 if (e.text.trim() === '') {
907 begin($, true, !isBurst, isBurst ? 0 : 1)
908 background(() => spaceWhileRecording($, isBurst, true))
909 return next({ ...e, inputText: '' }) // a leading space is useless anyway
910 }
911 // In text a single space types as usual, so the time check costs no jump here.
912 const now = await $.clock.now()
913 if (isBurst || (lastSpace !== null && now - lastSpace < HOLD_GAP_MS)) {
914 lastSpace = null
915 begin($, true, false, isBurst ? 0 : 2) // a hold, or the 2nd of 3 taps: the 3rd space tells
916 background(() => spaceWhileRecording($, isBurst, true))
917 // Delete the first space, which typed before the 2nd showed it was a hold or taps.
918 const c = e.cursor
919 if (c > 0 && e.text[c - 1] === ' ' && e.start === c && e.end === c) {
920 return next({ ...e, text: e.text.slice(0, c - 1) + e.text.slice(c), cursor: c - 1, start: c - 1, end: c - 1, inputText: '' })
921 }
922 return next({ ...e, inputText: '' })
923 }
924 lastSpace = now
925 return next(e)
926 })
927
928 on('session.end', async ($, e, next) => {
929 isChordArmed = false
930 await queueChord($, false) // quit mid-recording: Space must type in the next session
931 return next(e)
932 })
933
934 on('prompt.submit', async ($, e, next) => {
935 const r = await next(e)
936 if (r.drop !== undefined) return r
937 if (lastFill && e.origin?.kind === 'composer') {
938 const filled = lastFill
939 background(() => learn($, filled, e.text))
940 lastFill = null
941 }
942 return r
943 })
944
945 on('command.run', { command: 'fa' }, async ($, e) => {
946 const [sub = '', arg] = e.args.trim().split(/\s+/)
947 const rest = e.args.trim().slice(sub.length).trim() // a chord has a space: "ctrl+x v"
948 if (sub === '' || sub === 'settings') {
949 const r = await openSettings($)
950 return { text: r.isPlaced ? '⚙ Persian Voice settings (Esc closes)' : HELP }
951 }
952 if (sub === 'rec') {
953 const wasListening = isActive
954 if (wasListening) isStopping = true
955 else await press($)
956 return { text: wasListening ? 'Stopping…' : '🎙 Listening. Run /fa rec again to stop.' }
957 }
958 if (sub === 'mode') {
959 if (arg && !(arg in MODES)) return { text: `Unknown mode "${arg}". Modes: ${MODE_ORDER.join(', ')}` }
960 const p = await setMode($, arg as Mode | undefined)
961 return { text: `✨ Mode: ${MODES[p.mode].label}` }
962 }
963 if (sub === 'polish') {
964 const p = await savePrefs($, { mode: (await loadPrefs($)).mode === 'exact' ? 'prompt' : 'exact' })
965 return { text: `✨ Mode: ${MODES[p.mode].label}` }
966 }
967 if (sub === 'send') {
968 const p = await savePrefs($, { autoSend: !(await loadPrefs($)).autoSend })
969 return { text: p.autoSend ? '⏎ Auto-send on: the prompt is sent when you finish' : '⏎ Auto-send off: press Enter to send' }
970 }
971 if (sub === 'last') {
972 const d = (await $.store.get(LAST).catch(() => undefined)) as { fa?: string; en?: string } | undefined
973 if (!d?.fa && !d?.en) return { text: 'No recording saved yet.' }
974 // The command's own reply always shows the words; the box gets the English (or the Persian) to send.
975 background(() => $.prompt.fill({ text: d.en || d.fa || '', mode: 'insert' }))
976 return { text: `↩ Your last recording (also put in the prompt box):\n\n${d.fa ?? ''}\n\n→ ${d.en || '(no translation)'}` }
977 }
978 if (sub === 'engine') {
979 if (arg && !(arg in ENGINES)) return { text: `Unknown engine "${arg}". Engines: ${Object.keys(ENGINES).join(', ')}` }
980 const cur = (await loadPrefs($)).engine
981 const p = await savePrefs($, { engine: (arg as Engine | undefined) ?? (cur === 'soniox' ? 'google' : 'soniox') })
982 return { text: `🎙 Speech engine: ${ENGINES[p.engine]}` }
983 }
984 if (sub === 'key') return { text: await keyCommand($, rest) }
985 if (sub === 'space') return { text: await spaceCommand($, rest) }
986 if (sub === 'mic') return { text: await micCommand($, arg) }
987 if (sub === 'terms') return { text: await termsCommand($) }
988 if (sub === 'forget') {
989 const learned = await learnedTerms($)
990 if (!arg || !learned.includes(arg)) return { text: `Not a learned word: "${arg ?? ''}". Learned: ${learned.join(', ') || '(none)'}` }
991 await $.store.set('learned', learned.filter(w => w !== arg))
992 void refreshContext($)
993 return { text: `Forgot "${arg}".` }
994 }
995 return { text: sub === 'help' ? HELP : `Unknown: /fa ${sub}\n\n${HELP}` }
996 })
997
998 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
999 if (e.props.hasSurvey) return next(e)
1000 const [l, p, u, settings] = await Promise.all([read($, live), read($, prefs), read($, undo), $.settings.read()])
1001 const { Box, Button, Text } = $.ui.resolve(e)
1002
1003 if (!l || isTentative) {
1004 const isVoiceOn = (settings.voice as { enabled?: boolean } | undefined)?.enabled === true
1005 return (
1006 <Box flexDirection="column">
1007 {isVoiceOn && p.holdSpace && (
1008 <Text color={AMBER}>⚠ Built-in /voice is on and takes Space: run /voice off (once) to use Persian hold-space.</Text>
1009 )}
1010 <Box gap={1}>
1011 {p.holdSpace && <Text color={RED}>🎙</Text>}
1012 {p.holdSpace && <Text dimColor>Hold space to talk</Text>}
1013 {/* With a shortcut, its chord presses this Button from the prompt (the Button must be drawn). */}
1014 <Button key="talk" label={p.shortcut ? `🎙 Talk (${p.shortcut})` : '🎙 Talk'} action={ACTION} onPress={() => press($)} />
1015 <Button key="mode" label={`✨ ${MODES[p.mode].label}`} onPress={() => setMode($)} />
1016 <Button key="send" label={p.autoSend ? '⏎ Auto-send' : '⏎ Manual send'} onPress={() => savePrefs($, { autoSend: !p.autoSend })} />
1017 <Button key="settings" label="⚙ Settings" onPress={() => openSettings($)} />
1018 {u && <Button key="undo" label="↩ Use plain translation" onPress={() => usePlain($)} />}
1019 </Box>
1020 </Box>
1021 )
1022 }
1023
1024 const phase = isPolishing ? 'polish' : isFinishing ? 'finish' : 'rec'
1025 const color = phase === 'rec' ? RED : phase === 'finish' ? AMBER : ORANGE
1026 const spin = SPIN[frame % SPIN.length]
1027 const head =
1028 phase === 'rec'
1029 ? `${frame % 6 < 3 ? '●' : '○'} REC ${clockText(l.ms)}`
1030 : phase === 'finish'
1031 ? `${spin} Finishing translation…`
1032 : `${spin} Polishing · ${polishLabel}…`
1033 const meter = levels.map(v => BARS[Math.round(Math.min(1, v) * 8)]).join('').padStart(16, ' ')
1034 const byKey = !isSpaceHold // started by the shortcut, the button or /fa
1035 const hint =
1036 phase !== 'rec'
1037 ? byKey ? '' : 'type a letter to cancel'
1038 : !byKey
1039 ? isTapMode
1040 ? 'tap space to finish · a letter cancels'
1041 : 'release space to finish · a letter cancels'
1042 : isHeld
1043 ? `release ${p.shortcut ?? 'the key'} to finish`
1044 : `${p.shortcut ? `${p.shortcut} or ` : ''}/fa rec to finish`
1045 return (
1046 <Box flexDirection="column" borderStyle="round" borderColor={color} paddingX={1}>
1047 <Box gap={2}>
1048 <Text color={color} bold>{head}</Text>
1049 {phase === 'rec' && <Text color={GREEN}>{meter}</Text>}
1050 <Text dimColor>{hint}</Text>
1051 </Box>
1052 <Text bold wrap="truncate-start">{l.fa || 'Listening…'}</Text>
1053 <Text dimColor wrap="truncate-start">→ {l.en || '…'}</Text>
1054 {/* Drawn while a key is held too: the shortcut's repeats, or held Space's `space space` chord,
1055 reach press() through this Button. */}
1056 <Box gap={2}>
1057 {/* Not a visible choice: the shortcut and held Space press this Button through ACTION. */}
1058 {phase === 'rec' && (!byKey || p.shortcut) && <Button key="hold" label="🎙" plain action={ACTION} onPress={() => press($)} />}
1059 {/* Finish and send in one click, whatever the auto-send setting says (JEV's hold-back still applies). */}
1060 {phase === 'rec' && <Button key="sendnow" label="⏎ Send" onPress={() => { sendNow = true; isStopping = true }} />}
1061 {/* In every phase: until the text is in the box (or sent), a click discards it. */}
1062 <Button key="cancel" label="✕ Cancel" onPress={() => abort($)} />
1063 </Box>
1064 </Box>
1065 )
1066 })
1067
1068 // The settings dialog: every setting, its current value, and a control to change it.
1069 on('ui.render', { component: 'Pane', requestId: SETTINGS }, async ($, e) => {
1070 const p = await read($, prefs)
1071 const els = $.ui.resolve(e)
1072 if (!('Select' in els) || !('Input' in els)) {
1073 const { Text } = els // a surface without pickers: show the commands instead
1074 return <Text>{HELP}</Text>
1075 }
1076 const { Box, Button, Input, Select, Text } = els
1077 const LABEL = 18
1078 const onOff = (isOn: boolean) => (isOn ? '● On ' : '○ Off')
1079 return (
1080 <Box flexDirection="column" paddingX={1}>
1081 <Text bold color={ORANGE}>How to start a recording</Text>
1082 <Box gap={1}>
1083 <Box width={LABEL}><Text>Hold Space</Text></Box>
1084 <Button key="space" label={onOff(p.holdSpace)} onPress={() => savePrefs($, { holdSpace: !p.holdSpace })} />
1085 <Text dimColor>{p.holdSpace ? 'hold Space to talk, release to finish' : 'Space only types'}</Text>
1086 </Box>
1087 <Box gap={1}>
1088 <Box width={LABEL}><Text>Shortcut</Text></Box>
1089 {p.shortcut && <Text bold>{p.shortcut}</Text>}
1090 {p.shortcut && <Button key="unbind" label="Remove" onPress={() => shortcutFromDialog($, 'off')} />}
1091 {!p.shortcut && <Input key="shortcut" placeholder="type one, e.g. ctrl+x v, then Enter" onSubmit={v => shortcutFromDialog($, v)} />}
1092 </Box>
1093 <Box paddingLeft={LABEL + 1}>
1094 <Text dimColor>Press the shortcut to start, and again to stop.</Text>
1095 </Box>
1096
1097 <Box marginTop={1}><Text bold color={ORANGE}>What happens to your words</Text></Box>
1098 <Box gap={1}>
1099 <Box width={LABEL}><Text>Cleanup mode</Text></Box>
1100 <Select
1101 key="mode"
1102 options={MODE_ORDER.map(m => ({ value: m, label: MODES[m].label }))}
1103 value={p.mode}
1104 onSelect={v => setMode($, v as Mode)}
1105 />
1106 </Box>
1107 <Box paddingLeft={LABEL + 1}>
1108 <Text dimColor wrap="wrap">{MODES[p.mode].about}</Text>
1109 </Box>
1110 <Box gap={1}>
1111 <Box width={LABEL}><Text>Auto-send</Text></Box>
1112 <Button key="send" label={onOff(p.autoSend)} onPress={() => savePrefs($, { autoSend: !p.autoSend })} />
1113 <Text dimColor>{p.autoSend ? 'sent as soon as you finish' : 'goes into the prompt box; you press Enter'}</Text>
1114 </Box>
1115
1116 <Box marginTop={1}><Text bold color={ORANGE}>Microphone and words</Text></Box>
1117 <Box gap={1}>
1118 <Box width={LABEL}><Text>Speech engine</Text></Box>
1119 <Select
1120 key="engine"
1121 options={Object.entries(ENGINES).map(([value, label]) => ({ value, label }))}
1122 value={p.engine}
1123 onSelect={v => savePrefs($, { engine: v as Engine })}
1124 />
1125 </Box>
1126 <Box gap={1}>
1127 <Box width={LABEL}><Text>Microphone</Text></Box>
1128 <Select key="mic" options={micOptions} value={p.mic} onSelect={v => savePrefs($, { mic: v })} />
1129 </Box>
1130 <Box gap={1}>
1131 <Box width={LABEL}><Text>Word list</Text></Box>
1132 <Text dimColor wrap="wrap">{wordsInfo}</Text>
1133 </Box>
1134
1135 <Box marginTop={1}>
1136 <Text dimColor>Tab moves · Enter changes · Esc closes · saved for all projects · /fa help lists commands</Text>
1137 </Box>
1138 </Box>
1139 )
1140 })
1141}
1142hooks/jev.ts 109 lines1// The questions the plugin asks JEV (TypeSafe), as the API takes them. One proposition per noul;
2// the code in register.tsx combines the answers. tools/jev_eval.py scores this same object on labelled cases.
3// Question names never reach the model: the instructions and criteria carry the whole meaning.
4export const JEV_QUESTIONS = {
5 // Asked once per recording, before any rewrite. State: { recent_chat, spoken, prompt }.
6 before: {
7 has_noise: {
8 type: 'noul',
9 instructions: 'Does `prompt` contain spoken-language noise: filler words, repeated words, false starts, or a self-correction where the speaker changes what they said?',
10 criteria: {
11 true: 'At least one filler (um, like, you know, whatever), repeated word, abandoned start, or correction (no wait, I mean, sorry, actually).',
12 false: 'Every word carries meaning; the text reads like typed text.',
13 },
14 },
15 has_vague_reference: {
16 type: 'noul',
17 instructions: 'Does `prompt` point at its target only with a vague reference ("that bug", "the thing we did", "it", "the page") that neither `prompt` nor `recent_chat` makes clear?',
18 criteria: {
19 true: 'The main target is a pronoun or vague phrase, and no file, function, feature or error in `prompt` or `recent_chat` tells which one it means.',
20 false: 'The target is named in `prompt`, `recent_chat` makes clear which one it means, or the request needs no target.',
21 },
22 },
23 chat_resolves_reference: {
24 type: 'noul',
25 instructions: 'Does `prompt` point at its target with a vague reference ("that bug", "it", "the file we changed") that `recent_chat` makes clear?',
26 criteria: {
27 true: '`prompt` uses a pronoun or vague phrase for its target, and `recent_chat` names the file, function, feature or error it means.',
28 false: '`prompt` names its target itself, needs no target, or `recent_chat` does not tell which one it means.',
29 },
30 },
31 is_rambling: {
32 type: 'noul',
33 instructions: 'Is `prompt` disorganized: the goal is buried, ideas jump around, or the same point is made more than once?',
34 criteria: {
35 true: 'A reader must reorder or trim the text to find the goal.',
36 false: 'The goal is easy to find and each point appears once, in a sensible order.',
37 },
38 },
39 mistranslated: {
40 type: 'noul',
41 instructions: '`spoken` is what the speaker said (Persian, English or mixed). `prompt` is a machine translation of it into English. Does `prompt` change, add or leave out a fact, name, number, file path, code term or request that `spoken` contains?',
42 criteria: {
43 true: 'At least one fact, name, number, path, code term or request differs between `spoken` and `prompt`, or is missing from `prompt`.',
44 false: '`prompt` keeps every fact, name, number, path, code term and request of `spoken`; only wording, fillers or word order differ.',
45 },
46 },
47 is_for_agent: {
48 type: 'noul',
49 instructions: 'Is `prompt` meant for the AI assistant: a request, an instruction or a question that it could act on or answer, on any topic (code, files, tools, writing, stories, analysis or anything else)?',
50 criteria: {
51 true: 'The speaker asks the assistant to do, analyze, write, explain or answer something, on any topic.',
52 false: 'Talk to another person who is present, background speech, a reaction to something around the speaker, or words with no request or question.',
53 },
54 },
55 is_irreversible: {
56 type: 'noul',
57 instructions: 'Does `prompt` ask the agent to do something that is hard or impossible to undo?',
58 criteria: {
59 true: 'Deletes files or data that git cannot restore, discards uncommitted work, rewrites or force-pushes git history, drops or changes a database, deploys, publishes, sends a message, or spends money.',
60 false: 'Reads, explains, answers, or makes local changes to tracked files that git can undo, such as editing, renaming, committing, or running tests.',
61 },
62 },
63 // The options are the rewrite prompts in prompts.ts (a test checks that the two lists match).
64 kind: {
65 type: 'choice',
66 instructions: "Which kind of task does the speaker give in `prompt`? `spoken` holds the same words in the speaker's language.",
67 criteria: {
68 general: 'Any other request for the coding agent, such as running commands, git operations (commit, push, branch), setup or configuration, or a request that mixes several kinds.',
69 bug: 'Something is broken or wrong and must be fixed or debugged: an error, a crash, a failing test, or wrong behavior.',
70 feature: 'Add or change functionality: a new feature, option, command, endpoint, screen element or behavior.',
71 refactor: 'Restructure or clean up existing code without changing what it does: rename, move, split, merge or simplify.',
72 test: 'Write new tests or extend tests for existing code. A failing test that must be fixed is a bug.',
73 review: 'Review code, a diff, a branch or a pull request, and report problems.',
74 question: 'A question about the code, the project or a tool that asks for an answer or explanation, not for a change.',
75 story: 'Analyze, critique or edit creative writing: a story, chapter, scene, character or dialogue.',
76 spec: 'The speaker thinks aloud about a larger task, with goals, ideas and requirements to organize into a task spec.',
77 commit: 'The speaker wants the text of a git commit message written: they ask for one, or dictate its content. Not a request to run the commit.',
78 },
79 },
80 },
81 // Asked after a rewrite. State: { recent_chat, spoken, draft, rewritten }.
82 after: {
83 adds_request: {
84 type: 'noul',
85 instructions: "`spoken` is the speaker's own words, `draft` a machine translation of them into English, and `rewritten` a cleaned-up version of `draft`. Does `rewritten` ask for a task, step, file, test, option or condition that is in neither `spoken` nor `draft`?",
86 criteria: {
87 true: '`rewritten` contains a request that the speaker never made, in `spoken` or in `draft`.',
88 false: 'Every request in `rewritten` is in `spoken` or in `draft`. Formatting, order, headings, a name taken from `recent_chat` to make a vague reference exact, a goal that the speaker\'s words clearly imply (for example "Fix ..." when the speaker reports a failure), and one expert role sentence at the start (for example "Act as an experienced fiction editor.") do not count.',
89 },
90 },
91 changes_fact: {
92 type: 'noul',
93 instructions: "`spoken` is the speaker's own words, `draft` a machine translation of them into English, and `rewritten` a cleaned-up version of `draft`. Does `rewritten` give a name, number, file path or code term a value that matches neither `spoken` nor `draft`?",
94 criteria: {
95 true: 'A value in `rewritten` matches neither `spoken` nor `draft`, for example a different number or file name.',
96 false: 'Each value in `rewritten` matches `spoken` or `draft`, or comes from `recent_chat` to make a vague reference exact.',
97 },
98 },
99 drops_fact: {
100 type: 'noul',
101 instructions: '`rewritten` is a cleaned-up version of `draft`. Does `rewritten` leave out a fact, name, number, file path, code term or request that `draft` contains?',
102 criteria: {
103 true: 'Something with meaning in `draft` is missing from `rewritten`.',
104 false: "`rewritten` keeps every meaningful item of `draft`. Removed fillers, repetitions, corrected false starts, the speaker's words about the output format (for example \"write a commit message saying\"), and a translation error in `draft` that `rewritten` corrects to match `spoken` (the original speech) do not count.",
105 },
106 },
107 },
108}
109hooks/prompts.ts 53 lines1// The rewrite prompts: shared rules, plus one task per kind of request. JEV's `kind` question (jev.ts)
2// picks the kind; register.tsx builds the system prompt with systemPrompt(). tools/rewrite_eval.py runs
3// this same object through Claude. Follows Anthropic's prompting guidance: a role, the reason behind
4// each rule, XML-tagged input, and positive instructions. tools/prompt_ablation.py measured each part:
5// examples (shared and per kind) changed nothing, so there are none; the rules and the kind tasks did.
6export const PROMPTS = {
7 rules: [
8 "You turn a speaker's dictated words into a prompt for Claude Code, an AI coding agent. You are a rewriter, not an assistant: Claude Code acts on your output later, so you never answer, perform or comment on the request.",
9 '',
10 "The input has two parts. <spoken> holds the speaker's own words (Persian, English or mixed). <draft> holds a machine translation of them into English. The draft can contain translation errors, and it can miss the end of the speech.",
11 '',
12 'Rules:',
13 "- Keep the speaker's intent and every fact, name, file path, number and code token. Claude Code does exactly what your text says, so a requirement, guess or detail that the speaker did not say makes it do work that nobody asked for. Leave out each part of the task structure below that the speaker gave no content for.",
14 '- Use <spoken> as the source of truth. Correct translation errors in <draft>, and add anything from <spoken> that <draft> left out.',
15 '- Write code identifiers, file names, commands and technical terms exactly as the speaker said them, in English (Latin script).',
16 '- Remove fillers, repetitions and false starts. When the speaker corrects themselves, keep only the correction.',
17 '- Keep the speaker\'s level of certainty. Write an idea that the speaker said with "maybe" or "I think" as an option or a guess ("Maybe export as Markdown."), not as a requirement.',
18 '- Write in the speaker\'s own voice, as if they typed the prompt themselves.',
19 '- Write plain English text: short, direct sentences. Use a list only for three or more parallel items. Use headings only when the task below asks for them.',
20 '- Output only the rewritten text, starting with its first word.',
21 ],
22 // Added for the chat fork, which sees the main conversation.
23 chat: 'You also see the conversation so far. Use it only to make vague references exact ("that bug" or "the file we changed" becomes the real name). Add nothing else from the conversation.',
24 kinds: {
25 general:
26 'Write a clear prompt. Start with the goal as one direct sentence. Then give the context, constraints and expected result that the speaker gave. A short, clear request stays short and almost unchanged.',
27 bug: 'The speaker reports a bug. Start with the goal as one sentence ("Fix ..."). Then give, in this order and only from what the speaker said: the symptom (what happens, with any error message copied exactly); when it happens (steps, inputs, conditions); where to look (files, functions, areas); the expected behavior, which tells what "fixed" looks like; and any check the speaker asked for, such as a failing test first. Claude Code fixes a bug best when it knows the symptom, the location and what "fixed" looks like, so keep each of these that the speaker gave.',
28 feature:
29 'The speaker asks for new or changed functionality. Start with the goal as one sentence. Then give, only from what the speaker said: where it goes and any existing file or pattern to follow; the requirements; the constraints (libraries, things not to change); and how to check that it works.',
30 refactor:
31 'The speaker asks to restructure code without changing what it does. Start with the goal as one sentence. Then give, only from what the speaker said: the scope (files, functions, modules); what must stay the same (behavior, public API, outputs, tests); and the target structure.',
32 test: 'The speaker asks for tests. Start with the goal as one sentence that names the code under test. Then give, only from what the speaker said: the cases and edge cases to cover; constraints such as the test framework or no mocks; and how to run the tests.',
33 review:
34 'The speaker asks for a review. Start with the goal as one sentence that names what to review (a diff, file, branch or pull request). Then give, only from what the speaker said: what to look for (for example bugs, security, performance, consistency with the existing code) and the form of the report.',
35 question:
36 'The speaker asks a question. Keep it a question that asks for an answer, not a change to the code. Name the exact code or topic that it is about, and any source that the speaker pointed to (a file, the git history, documentation).',
37 story:
38 'The speaker wants an analysis or edit of creative writing (a story, chapter, scene, character or dialogue). Start with this role sentence: "Act as an experienced fiction editor." A named role focuses the analysis, so this sentence is the one addition that you make. Then give, only from what the speaker said: which text (file, chapter, scene, passage); the aspects to analyze (for example plot logic, stakes, pacing, character motivation, dialogue, point of view, tone); and what to return (a diagnosis, notes with references to the text, or a rewrite). Keep the names of characters and places exactly.',
39 spec: 'The speaker thinks aloud about a larger task. Organize it as a task spec with these headings, each only when the speaker gave content for it: "Goal" (one sentence), "Context", "Requirements" (a list), "Out of scope" (a list), "Done when" (a list of checks).',
40 commit:
41 'Write a git commit message. The first line is the subject: imperative mood, at most 72 characters, no period at the end. Most dictated messages need only the subject. Add a blank line and a short body only when the speaker gave details that the subject does not already say, such as why the change was made. Leave out the speaker\'s words about the message itself, such as "write a commit message saying".',
42 },
43 // A second try when JEV finds that the first rewrite changed what the speaker said (Anthropic's
44 // self-correction pattern: draft, review, refine). One line per JEV `after` question that failed.
45 retry: {
46 adds_request: 'It added a request, step or condition that the speaker did not make. Remove it.',
47 changes_fact: 'It changed a name, number, file path or code term. Use the value from <spoken> or <draft>.',
48 drops_fact: 'It left out a fact, name, number, file path or request that <draft> contains. Put it back.',
49 ask: 'Your previous rewrite is in <previous_rewrite>. A check found these problems in it:',
50 end: 'Rewrite the draft again, and fix these problems. Change nothing else.',
51 },
52}
53types/index.d.ts 30 lines1/** One line of fa.py: what was said, its English version, the mic level; plus time since start. */
2export type Live = { fa: string; en: string; lvl?: number; ms?: number; done?: boolean }
3
4/** How the English is cleaned up before it goes into the prompt box. */
5export type Mode = 'auto' | 'exact' | 'prompt' | 'chat' | 'spec' | 'commit'
6
7/** Who recognizes the speech. Soniox also translates; Google only transcribes, so Claude translates. */
8export type Engine = 'soniox' | 'google' | 'local'
9
10/** The person's settings, global ($.store). */
11export type Prefs = {
12 engine: Engine
13 mode: Mode
14 autoSend: boolean
15 mic: string
16 /** Holding Space records. Off: Space only types. */
17 holdSpace: boolean
18 /** The person's own shortcut (a keybinding), or null for none. */
19 shortcut: string | null
20}
21
22/** The last polished fill, so "Use raw" can swap the plain translation back. */
23export type Undo = { raw: string; polished: string }
24
25declare module 'claude-code' {
26 interface PluginState {
27 'persian-voice': { live: Live | null; prefs: Prefs; undo: Undo | null }
28 }
29}
30