Fetch items over MCP and judge them with Jev without passing the data through context

Configuration files for Bash, Starship, Homebrew, Ghostty, Herdr, Neovim, OpenCode, Claude Code, Codex, and Pi. GNU Stow manages symlinks from stow/* into $HOME. Ghostty, Herdr, Neovim, Pi, OpenCode, Claude Code, and Codex use Tokyo Night with MonoLisaCode 14 pt.
night style with transparent editor, sidebar, float, statusline, and tab-fill surfaces.MonoLisaCode at 14 pt with explicit regular, italic, bold, and bold-italic styles.Symbols Nerd Font fallback. See plans/theme-font-glyph-followups.md.old/ and are not restored or rethemed.| Tool | Minimum Version | Notes |
|---|---|---|
| Neovim | >= 0.11.0 | Required for mason-lspconfig v2 and vim.lsp.config() |
| Git | >= 2.19.0 | Required for lazy.nvim partial clones |
| GNU Stow | >= 2.4.0 | Symlink manager for tracked dotfiles |
| Ghostty | Latest | Uses macos-option-as-alt syntax |
| Herdr | Latest | Agent-aware terminal workspace manager |
| Node.js | LTS | For LSP servers via Mason |
| tree-sitter-cli | >= 0.26.1 | Required for nvim-treesitter main branch parser compilation |
| lazygit | >= 0.40 | Required for snacks.lazygit keymap (<leader>gg) |
| ripgrep | >= 13.0 | Required for nvim-spectre search backend |
| gnu-sed | Latest | Recommended on macOS for nvim-spectre replace engine (brew install gnu-sed) |
| imagemagick | >= 7.0 | Required for snacks.image preview support |
spotify-visualizer is a standalone TypeScript command managed by the bin Stow package. It renders a procedural terminal dot matrix based on the website music visualizer colors, then uses Spotify only for the current track, artist, play state, and track-specific animation seed.
Setup:
# 1. Create or reuse a Spotify developer app.
# 2. Add this redirect URI to that app:
# http://127.0.0.1:8974/callback
# 3. Export the client id before launching the visualizer:
export SPOTIFY_CLIENT_ID=your_spotify_client_id
spotify-visualizer
The command stores OAuth tokens under ~/.cache/dotfiles/spotify-visualizer/. Run it in any Herdr pane or tab when you want a dedicated visualizer screen.
Controls:
| Key | Action |
|---|---|
Space | Toggle Spotify play or pause |
n | Skip to the next track |
p | Skip to the previous track |
s | Toggle shuffle |
r | Cycle repeat off, context, and current track |
q / Ctrl-C | Quit and restore the terminal |
The visualizer shows this key legend in the header. Short notices, such as pressing a playback key before Spotify has an active track, replace the legend for about 3 seconds.
Shuffle and repeat state use compact status tokens in the header. Shuffle uses [S:-] when inactive and yellow [S:*] when active. Repeat uses gray [R:-] when inactive, red [R:all] for repeat context, and red [R:1] for repeat current track.
If Spotify returns 401 after scopes change, remove the cached token and authorize again:
rm ~/.cache/dotfiles/spotify-visualizer/tokens.json
spotify-visualizer
The TypeSafe skill defines when and how to request a Jev judgment. Pi 1.0 uses native Codemode classifiers; it no longer installs a custom TypeSafe SDK adapter. OpenCode retains its typesafe_evaluate adapter under stow/opencode/.config/opencode/plugins/typesafe-ai/. Both send the supplied state and questions to TypeSafe and return typed judgments.
In Pi Codemode:
const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");
if (!jev) throw new Error("Jev classifier is unavailable");
const result = await models.classify(jev, {
state: { message: "A customer asks to cancel today" },
questions: {
urgent: {
type: "bool",
instructions: "Does this need a reply today?",
criteria: { true: "Needs a reply today", false: "Can wait" },
},
},
});
if (result.stopReason !== "stop") throw new Error(result.errorMessage || "Jev classification failed");
return result.answers.urgent;
Native questions use choice, score, or bool; a bool answer supplies a probability, not a boolean decision. Specialized Jev MCP tools remain available for screening, verification, and gates. Do not include credentials, secrets, or unrelated private data.
OpenCode pins the TypeSafe tool in its Code Mode catalog and adds tool invocation guidance to outgoing contexts where the catalog lists it. Shared global rules and the Jev skill own judgment policy; the adapter explains only the Code Mode calling convention and result shape. Use exact search, parsing, arithmetic, and tests for deterministic facts. Inside execute, the tool returns a validated object with answers, model, and usage; no JSON.parse is needed. For example, when the catalog lists tools.typesafe_evaluate:
const result = await tools.typesafe_evaluate({
state: "A customer asks to cancel today",
noul_questions: [{ id: "urgent", instructions: "Does this need a reply today?" }],
});
return result.answers.urgent;
The Noul answer is { type: "noul", noul: probability }, not a boolean. Choice and Score answers include confidence and probability distributions. Missing or malformed provider answers fail explicitly rather than returning an empty success. Tool results retain readable text and model/token metadata for other consumers. Judgments advise; they do not replace permission checks or prove correctness.
Keep TYPESAFE_API_KEY as machine-local state in ~/.config/bash/local.bash:
export TYPESAFE_API_KEY="YOUR_API_KEY"
scripts/bootstrap.sh installs the OpenCode adapters' pinned runtime dependencies. Pi's native classifier needs no separate SDK installation. Apply and reload live configuration only after reviewing and approving the tracked changes.
The opencode Stow package owns the tracked sources under stow/opencode/.config/opencode/.
| Tracked source | Purpose |
|---|---|
opencode.jsonc | Model defaults, permissions, MCP servers, providers, skills, and global instructions |
cli.json | Tokyo Night, TUI layout, permission handling, and keybindings |
commands/ | Slash commands such as /lg |
plugins/typesafe-ai/ | TypeSafe Jev tool |
plugins/tui-conveniences/ | /copy-all, /restart, /update, skill-load confirmations, and the Git status footer |
plugins/request-logger/ | Opt-in private HTTP request capture |
New sessions use openai/gpt-6.1-sol-fast with medium reasoning effort and low response verbosity. Its pinned limits match the running ChatGPT catalog checked on 2026-10-04: 400,000 context, 272,000 input, and 128,000 output tokens. This catalog alias sends gpt-6.1-sol with the priority service tier. The limits preserve the current compaction budget rather than assuming the public API's larger window applies to this connection. The TUI hides the session sidebar and persistent tab strip. The TUI also provides Pi-style navigation shortcuts.
OpenCode can use the current user's filesystem, processes, and network without a permission prompt. Review the tracked configuration before you apply it.
REF_API_KEY, EXA_API_KEY, and TYPESAFE_API_KEY are machine-local state. Do not put real credentials in tracked sources.
Apply only the OpenCode Stow package with:
cd ~/dotfiles
./scripts/stow.sh apply opencode
Bootstrap installs plugin dependencies automatically. Restart OpenCode after an apply.
The last local plugin is request-logger, disabled by default with options.enabled: false. After reviewing the tracked changes and approving live activation, enable it through the plugin's options in opencode.jsonc:
{
"package": "../../code/personal/dotfiles/stow/opencode/.config/opencode/plugins/request-logger",
"options": {
"enabled": true,
"maxFiles": 100
}
}
This is an entry in the existing plugins array, not a replacement configuration. Bootstrap installs its dependencies along with the other local plugins. The logger runs in the background service, so a flag on a new CLI process is not used to enable it. It writes to ~/.local/state/opencode/requests, or an absolute options.directory, with directory mode 0700 and file mode 0600. It refuses a symlink at the log directory and retains the newest 100 logger-owned files by default; maxFiles must be a positive integer. Retention limits file count, not total bytes, and leaves unrelated files alone.
Each JSON file records the timestamp, session, agent, catalog model, request kind, raw body string, and byte counts for instructions, input, and tools when those fields exist. The body shows the actual wire model and provider options; the catalog model can be an alias. Byte counts are not token counts, and the instruction count excludes system messages embedded in the input array. For non-UTF-8 bodies, bodyBase64 preserves the original bytes. The original request remains unchanged and readable; logging failures emit a generic server warning without error details and do not block dispatch.
Logs contain full prompts, tool schemas, and file contents, including any secrets already present in the body. The logger deliberately omits URLs and request headers, but does not redact bodies because that would hide what was sent. Keep logs outside version control, review them before sharing, and disable logging when the investigation ends. Capture covers HTTP requests for primary turns, compaction, titles, transient generation, and retries, not responses or WebSocket frames. Keep the logger after request-mutating plugins; hooks registered later can still change the request after capture. Automatic updates, MCP package updates, and automatic compaction remain unchanged.
OpenCode's codex-chrome MCP server uses codex-control-chrome-mcp to control the existing Chrome profile through the Codex Chrome extension. Unlike the isolated chrome-devtools server, it can use existing signed-in tabs and capture screenshots of localhost apps. It grants access to page contents and browser actions; the community bridge does not enforce per-site permissions.
After approving live configuration changes, install the bridge and register it for Google Chrome only:
npm install -g codex-control-chrome-mcp@1.4.1
codex-control-chrome-mcp install-native-host --browser chrome
codex-control-chrome-mcp status --browser chrome
The tracked MCP entry starts codex-control-chrome-mcp from PATH; on Linux the binary is absent, so only that server fails to start. The installer backs up Chrome's existing native-host manifest and records its original host for proxy mode. Reload the Codex Chrome extension after installation, then reconnect codex-chrome through OpenCode's /mcps menu or restart OpenCode. Automatic registration repair remains enabled: after a Codex update restores its own host registration, starting the bridge re-registers it, and the extension may need another reload.
To restore the previous Chrome native-host registration, disconnect codex-chrome in OpenCode and run:
codex-control-chrome-mcp uninstall-native-host --browser chrome
Also remove the codex-chrome MCP entry if you no longer want OpenCode to start the bridge.
OpenAI GPT-5 models and the verified GPT-6 Astra, GPT-6 Sol, and GPT-6.1 Sol models support low, medium, and high output verbosity through the Responses API. The tracked configs currently use low.
gpt-6-astra, gpt-6-sol, and gpt-6.1-sol IDs using openai-responses or openai-codex-responses in stow/pi/.pi/agent/extensions/gpt-verbosity.ts. Change the verbosity: "low" value, then run /reload in Pi.model_verbosity in stow/codex/.codex/config.toml. Change the value, then restart Codex.textVerbosity per provider and model in stow/opencode/.config/opencode/opencode.jsonc. Update each GPT-5 model entry you use under provider.openai.models or provider.opencode.models, then restart OpenCode.For example, an OpenCode model override uses this shape:
"providers": {
"openai": {
"models": {
"gpt-5.6-sol-fast": {
"settings": {
"textVerbosity": "low",
},
},
},
},
}
Run claude-log from your regular shell instead of claude to start Claude Code through an opt-in local Anthropic request logger:
claude-log
The command starts a local proxy on a temporary loopback port, launches Claude Code against it, and stops the proxy when Claude exits. Each /v1/messages request is written under ~/.claude/logs/requests/ as readable Markdown plus the raw JSON payload. The Markdown includes request sizes, ranked tool schemas, redacted request headers, the full payload, and the streamed provider response. This logger is inspired by Matt Pocock's agent proxy. It does not capture direct MCP network traffic.
The directory and files use owner-only permissions. The logs can contain sensitive source code, prompts, and connected-service data. Review them before sharing, and remove them when finished:
rm -rf ~/.claude/logs/requests
Set CLAUDE_REQUEST_LOG_DIR to store logs somewhere else. Normal claude sessions do not write request logs.
Claude Code uses the Tokyo Night night theme from stow/claude/.claude/themes/tokyonight-night.json, taken from folke's tokyonight.nvim Claude Code extras. Claude Code watches ~/.claude/themes/, so theme edits apply to running sessions.
Mods live in stow/claude/.claude/mods/ and are not stowed. CLAUDE_CODE_PLUGIN_DIRS in settings.json loads them from the repository, and interactive sessions reload a mod when its files are saved.
optojr-slack registers optojr_slack_send, which posts as @OptoJr through the relay credentials in the macOS Keychain.skill-toast shows a toast when a skill loads.session-relaunch adds /update, which runs claude update and resumes the session on the new version, and /restart, which resumes the session on whatever version is installed. The relaunch comes from the claude function in stow/bash/.config/bash/functions.bash. Sessions started any other way print the claude --resume command instead of exiting.jev-pipeline registers mcp__jev-pipeline__run, which fetches a list from one MCP tool, optionally enriches each item with a second MCP call, and classifies or reranks every item with Jev. The items never enter the model's context: the reply holds counts and the top items, and the full results go to $TMPDIR/jev-pipeline-<session id>-<timestamp>.json.jev-coding adds a jev_decide reminder to the first Edit or Write of each turn, and records every edit the turn makes. When the agent tries to finish, Jev classifies each edit against the recent prompts as requested, scope creep, speculative, or leftover. A confidently flagged edit blocks the stop once with the list, so the agent reverts or justifies it; a Jev failure shows a toast and lets the turn finish.jev-screen runs jev_screen on every Exa and Ref fetch result, in 20,000-character chunks. A review or block verdict, or a failed screen, adds a warning the model reads after the result, and block also shows a toast.model-cost writes the session's cost per model to $TMPDIR/claude-model-cost-<session id>.json, and statusline.sh shows it after the context usage. The engine reports only a session total, so each response is charged the total's growth since the previous response.Add a new mod folder to CLAUDE_CODE_PLUGIN_DIRS to load it. Check a mod with claude plugin validate <folder> and claude plugin test <folder>.
Run pi-log from your regular shell instead of pi when you need to inspect the exact payload Pi sends to its model provider:
pi-log
The command enables the tracked request-logger.ts extension for that Pi process only. Each request is written as a readable Markdown file under ~/.pi/agent/logs/requests/, including a size audit, ranked tool schemas, the complete provider payload, the normalized assistant response, and response status metadata when the active provider exposes it. The directory and files use owner-only permissions. These logs can contain source code, prompts, tool results, Gmail, Slack, or Drive data, so do not commit or share them without reviewing the contents. Remove captured requests when finished:
rm -rf ~/.pi/agent/logs/requests
You can also run pi --request-log directly, or set PI_REQUEST_LOG_DIR to store logs somewhere else. Normal pi sessions do not write request logs.
# 1. Clone the repo
git clone https://github.com/manifoldfrs/dotfiles.git ~/dotfiles
# 2. Run the installer
cd ~/dotfiles
./scripts/bootstrap.sh
# 3. Fully quit and reopen your terminal
# 4. Verify Node.js works
node --version
# 5. Verify OpenCode 2 (installed by bootstrap)
opencode --version
Use the daily Stow wrapper when the repo is already on the machine and you just want the latest dotfiles applied.
# 1. Get the latest committed dotfiles
cd ~/dotfiles
git pull
# 2. Validate, then reapply all tracked shell/editor/terminal and Herdr config
./scripts/validate-dotfiles.sh
./scripts/stow.sh apply
# First time on this machine? Install Bash, Starship, and the supporting tools:
brew bundle --file=Brewfile
# or only the packaged shell stack:
# brew install bash starship zoxide fzf mise ripgrep fd gawk
# 3. Fully quit and reopen your terminal
# 4. Install/update declared Herdr plugins and reload a running server
./scripts/sync_herdr_plugins.sh
# 5. Verify the basics
node --version
herdr --version
Use ./scripts/bootstrap.sh instead when you also want to install or refresh Homebrew packages, Node.js, and Neovim plugins.
What this already handles for you:
Ctrl-a bindings, persistence, and agent-aware workspacesWhat ./scripts/bootstrap.sh additionally handles for you:
BrewfileWhat is still separate:
./mcp_setup.sh install for the optional Claude Desktop MCP config~/.claude.json, which stays local because it contains credentials and account-specific stateThe herdr Stow package also manages ~/.config/herdr/plugins.txt and ~/.config/plannotator-tui/config.toml. The default profile includes it. Stow only applies configuration, it does not install or update plugins.
# Existing machines: install the terminal review tool, then sync plugins.
# Requires herdr >= 0.8.0, Bun, and jq on PATH.
brew tap plannotator/tap
# On Homebrew versions that support trust:
# brew trust plannotator/tap
brew install plannotator/tap/plannotator-tui
./scripts/stow.sh apply
./scripts/sync_herdr_plugins.sh
The sync command installs or updates plannotator/herdr-annotate and paulbkim-dev/vim-herdr-navigation, checks the config, and reloads a running Herdr server. Plannotator TUI opens in a full-tab overlay. The agent sidebar prioritizes agents needing attention, uses distinct status symbols, and asks before closing workspaces. Pane-history persistence remains disabled.
| Shortcut | Action |
|---|---|
Ctrl-h/j/k/l | Navigate Neovim splits, then adjacent Herdr panes at the edge |
Ctrl-a a | Annotate selected terminal text |
Ctrl-a Shift-a | Copy annotations as agent context |
Ctrl-a m | Manage annotations |
Ctrl-a Shift-o | Review documents in the current folder |
Ctrl-a Shift-l | Review the agent's last reply |
Neovim visual <leader>a | Send the selection to Herdr Annotate |
Press Ctrl-a, release it, then press the shortcut's second key.
Launch Plannotator TUI directly from a shell:
plannotator-tui README.md # Review a file
plannotator-tui docs/ # Browse a folder
plannotator-tui herdr open . # Review this folder in a Herdr overlay
plannotator-tui herdr last # Annotate the agent's last reply in Herdr
plannotator-tui last --host pi # Review the latest Pi reply outside Herdr
plannotator-tui last --host claude # Review the latest Claude reply outside Herdr
The plannotator-tui commands open the terminal interface. The browser-based plannotator integration is installed alongside it.
Existing Ctrl-a o pane cycling and Ctrl-a z zoom bindings are unchanged. Global Ctrl-k and Ctrl-l navigation takes precedence over shell line deletion and screen clearing inside Herdr. Neovim outside Herdr retains ordinary split navigation. The annotation handoff uses a private temporary file that the plugin consumes and deletes.
In Ghostty on macOS, hold Shift + Cmd and click a link to open it, including inside Herdr. This bypasses application mouse capture and lets Ghostty handle the link. This gesture was verified in this setup, while Herdr's documented Ctrl-click gesture did not work. Keep mouse capture enabled to preserve Herdr's mouse UI. See Herdr's mouse guide.
Pi renders Markdown links as terminal hyperlinks, but relative targets such as docs/plan.md remain unresolved relative paths. The global Pi rules request absolute file:/// URLs for local files in chat and full https:// URLs for web links. File labels can still show readable repository-relative paths and line numbers. Line numbers are informational, not editor jump targets. Links written inside repository documentation remain relative for portability.
Shift-Cmd-click uses the system opener rather than the Herdr Annotate plugin. To review Markdown in Plannotator TUI, use Ctrl-a Shift-o or plannotator-tui herdr open <file.md>. Existing messages are not rewritten by the rule change. Start a new Pi session or use /reload to refresh the global instructions in an existing session.
cloak.nvim visually masks values in .env, .dev.vars, selected shell configuration files, and TOML token assignments. Use <leader>uC to toggle masking. This only affects display, not file contents, clipboard access, or agent access.:TSC runs the project's TypeScript compiler with --noEmit and opens errors in quickfix. Install TypeScript in the project first.ts-error-translator.nvim improves the readability of TypeScript diagnostics.:TwoslashQueriesEnable enables inline type queries and `:Thooks/register.ts 209 lines1import type { EngineInterface, Register } from 'claude-code'
2
3import {
4 CLASSIFY_BATCH,
5 JEV_TEXT_LIMIT,
6 RERANK_BATCH,
7 chunks,
8 fillArgs,
9 listAt,
10 mapLimit,
11 mcpJson,
12 textOf,
13 toItem,
14 type Classify,
15 type Item,
16 type PipelineInput,
17 type Rerank,
18} from './pipeline.ts'
19
20const TOOL_NAME = 'run'
21// Matches Pi's codemode demo: four concurrent per-item fetches keep MCP servers from rate limiting.
22const ENRICH_CONCURRENCY = 4
23
24const mcpStep = {
25 server: { type: 'string', description: 'MCP server name as /mcp lists it, e.g. "claude.ai Linear"' },
26 tool: { type: 'string', description: 'Tool name on that server, without the mcp__ prefix' },
27 args: { type: 'object', description: 'Tool arguments' },
28}
29
30const pathList = { type: 'array', items: { type: 'string' }, minItems: 1 }
31
32const inputSchema = {
33 type: 'object',
34 additionalProperties: false,
35 required: ['source', 'fields', 'judge', 'limit'],
36 properties: {
37 source: {
38 type: 'object',
39 required: ['server', 'tool'],
40 properties: {
41 ...mcpStep,
42 items: { type: 'string', description: 'Dot path to the item array in the JSON result; omit when the result is the array' },
43 },
44 },
45 fields: {
46 type: 'object',
47 required: ['text'],
48 properties: {
49 id: { type: 'string', description: 'Dot path to each item id' },
50 title: { type: 'string', description: 'Dot path to a short label shown in the reply' },
51 text: { ...pathList, description: 'Dot paths whose values make up the text Jev judges' },
52 },
53 },
54 enrich: {
55 type: 'object',
56 description: 'Optional per-item MCP call whose results are appended to the item text. A string arg that is exactly "{{path}}" takes the item value at that path.',
57 required: ['server', 'tool', 'text'],
58 properties: {
59 ...mcpStep,
60 items: { type: 'string', description: 'Dot path to an array in the per-item result, e.g. "comments"' },
61 text: { ...pathList, description: 'Dot paths within each result (or each array entry) to append' },
62 },
63 },
64 judge: {
65 type: 'object',
66 required: ['kind'],
67 description:
68 'kind "classify": classes (2+, strong descriptions), optional purpose, context, auto_accept, minimum_margin. kind "rerank": query, optional top_k.',
69 properties: {
70 kind: { enum: ['classify', 'rerank'] },
71 classes: {
72 type: 'array',
73 items: { type: 'object', required: ['description'], properties: { id: { type: 'string' }, description: { type: 'string' } } },
74 },
75 purpose: { type: 'string' },
76 context: {},
77 auto_accept: { type: 'number' },
78 minimum_margin: { type: 'number' },
79 query: { type: 'string' },
80 top_k: { type: 'integer', minimum: 1 },
81 },
82 },
83 limit: { type: 'integer', minimum: 1, maximum: 100, description: 'How many items to list per class (classify) or in total (rerank)' },
84 },
85}
86
87type Verdict = { classification: string; top_probability: number; decision: string }
88
89function failed(text: string) {
90 return { isError: true as const, result: text, text }
91}
92
93async function callJson($: EngineInterface, step: { server: string; tool: string }, args: unknown) {
94 // A JSON round trip drops undefined optional fields before they reach the server.
95 return mcpJson(await $.mcp.call(step.server, step.tool, JSON.parse(JSON.stringify(args ?? {}))))
96}
97
98async function enrich($: EngineInterface, input: PipelineInput, raws: unknown[], items: Item[]) {
99 const step = input.enrich
100 if (!step) return items
101
102 return mapLimit(items, ENRICH_CONCURRENCY, async item => {
103 const raw = raws[Number(item.key)]
104 const result = await callJson($, step, fillArgs(step.args, raw))
105 const entries = step.items ? listAt(result, step.items) : [result]
106 const extra = entries.map(entry => textOf(entry, step.text)).filter(Boolean)
107 return { ...item, text: [item.text, ...extra].join('\n') }
108 })
109}
110
111function jevItems(items: Item[]) {
112 return items.map(item => ({ id: item.key, text: item.text.slice(0, JEV_TEXT_LIMIT) }))
113}
114
115async function classify($: EngineInterface, judge: Classify, items: Item[], limit: number) {
116 const verdicts = new Map<string, Verdict>()
117 for (const batch of chunks(items, CLASSIFY_BATCH)) {
118 const { kind: _, ...options } = judge
119 const result = (await callJson($, { server: 'jev', tool: 'jev_classify' }, { ...options, items: jevItems(batch) })) as {
120 results?: (Verdict & { id: string })[]
121 }
122 for (const verdict of result.results ?? []) verdicts.set(verdict.id, verdict)
123 }
124
125 const judged = items.map(item => ({ ...item, verdict: verdicts.get(item.key) }))
126 const byClass: Record<string, { count: number; review: number; top: string[] }> = {}
127 const ranked = [...judged].sort((a, b) => (b.verdict?.top_probability ?? 0) - (a.verdict?.top_probability ?? 0))
128 for (const item of ranked) {
129 const name = item.verdict?.classification ?? 'unjudged'
130 const group = (byClass[name] ??= { count: 0, review: 0, top: [] })
131 group.count++
132 if (item.verdict?.decision === 'review') group.review++
133 if (group.top.length < limit) {
134 const review = item.verdict?.decision === 'review' ? ', review' : ''
135 group.top.push(`${item.id} ${item.title} (p=${item.verdict?.top_probability ?? '?'}${review})`)
136 }
137 }
138
139 return { judged, summary: { by_class: byClass } }
140}
141
142async function rerank($: EngineInterface, judge: Rerank, items: Item[], limit: number) {
143 const scores = new Map<string, number>()
144 for (const batch of chunks(items, RERANK_BATCH)) {
145 const result = (await callJson($, { server: 'jev', tool: 'jev_rerank' }, { query: judge.query, candidates: jevItems(batch) })) as {
146 ranked?: { id: string; relevance: number }[]
147 }
148 for (const entry of result.ranked ?? []) scores.set(entry.id, entry.relevance)
149 }
150
151 // Jev scores each candidate independently, so scores from separate batches compare directly.
152 const judged = items
153 .map(item => ({ ...item, verdict: { relevance: scores.get(item.key) } }))
154 .sort((a, b) => (b.verdict.relevance ?? 0) - (a.verdict.relevance ?? 0))
155 const top = judged.slice(0, Math.min(limit, judge.top_k ?? limit))
156 return { judged, summary: { top: top.map(item => `${item.id} ${item.title} (relevance=${item.verdict.relevance ?? '?'})`) } }
157}
158
159async function resultsFile($: EngineInterface) {
160 const dir = (await $.env.get('TMPDIR')) || '/tmp/'
161 return `${dir.replace(/\/?$/, '/')}jev-pipeline-${await $.session.id()}-${await $.clock.now()}.json`
162}
163
164function judgeProblem(judge: Classify | Rerank) {
165 if (judge.kind === 'classify' && !(judge.classes?.length >= 2)) return 'judge.classes needs at least two classes'
166 if (judge.kind === 'rerank' && !judge.query) return 'judge.query is required for rerank'
167 return undefined
168}
169
170async function run($: EngineInterface, input: PipelineInput) {
171 const problem = judgeProblem(input.judge)
172 if (problem) throw new Error(problem)
173
174 const raws = listAt(await callJson($, input.source, input.source.args), input.source.items)
175 const items = await enrich($, input, raws, raws.map((raw, index) => toItem(raw, index, input.fields)))
176 const judged = input.judge.kind === 'classify'
177 ? await classify($, input.judge, items, input.limit)
178 : await rerank($, input.judge, items, input.limit)
179
180 const path = await resultsFile($)
181 await $.fs.write(path, JSON.stringify(judged.judged.map(({ key: _, ...item }) => item), null, 2))
182 return { total: items.length, results_file: path, ...judged.summary }
183}
184
185export const register: Register = on => {
186 on('session.start', async ($, e, next) => {
187 await $.tool.register({
188 name: TOOL_NAME,
189 description: [
190 'Fetch a list of items from one MCP tool, optionally enrich each item with a second MCP call, and judge every item with TypeSafe Jev (classify or rerank).',
191 'The fetched data never enters your context: the reply holds counts and the top items only, and every item with its verdict is saved to results_file for Read or jq.',
192 'Use it instead of reading many records and pasting them into jev_classify or jev_rerank yourself.',
193 ].join(' '),
194 inputSchema,
195 })
196
197 return next(e)
198 })
199
200 on('tool.call', { tool: 'mcp__jev-pipeline__run' }, async ($, e) => {
201 const { tool: _, ...input } = e as unknown as PipelineInput & { tool: string }
202 try {
203 return { result: await run($, input) }
204 } catch (error) {
205 return failed(`jev-pipeline failed: ${error instanceof Error ? error.message : String(error)}`)
206 }
207 })
208}
209hooks/pipeline.ts 97 lines1import type { McpToolResult } from 'claude-code'
2
3export type McpStep = { server: string; tool: string; args?: Record<string, unknown> }
4export type Fields = { id?: string; title?: string; text: string[] }
5export type Classify = {
6 kind: 'classify'
7 classes: { id?: string; description: string }[]
8 purpose?: string
9 context?: unknown
10 auto_accept?: number
11 minimum_margin?: number
12}
13export type Rerank = { kind: 'rerank'; query: string; top_k?: number }
14
15export type PipelineInput = {
16 source: McpStep & { items?: string }
17 fields: Fields
18 enrich?: McpStep & { items?: string; text: string[] }
19 judge: Classify | Rerank
20 limit: number
21}
22
23export type Item = { key: string; id: string; title: string; text: string }
24
25// Jev truncates item text at 2,000 characters; trimming here keeps request payloads small.
26export const JEV_TEXT_LIMIT = 2_000
27export const CLASSIFY_BATCH = 64
28export const RERANK_BATCH = 250
29
30export function mcpJson(result: McpToolResult): unknown {
31 const text = result.content.find(block => block.type === 'text')?.text
32 if (result.isError) throw new Error(text ?? 'MCP call failed')
33 if (result.structuredContent !== undefined) return result.structuredContent
34 if (text === undefined) throw new Error('MCP call returned no text')
35 return JSON.parse(text)
36}
37
38export function at(value: unknown, path: string | undefined): unknown {
39 if (!path) return value
40 return path.split('.').reduce<unknown>((node, key) => (node && typeof node === 'object' ? (node as Record<string, unknown>)[key] : undefined), value)
41}
42
43export function listAt(value: unknown, path: string | undefined): unknown[] {
44 const list = at(value, path)
45 if (!Array.isArray(list)) throw new Error(`No array at "${path ?? ''}" in the MCP result`)
46 return list
47}
48
49function asText(value: unknown) {
50 if (value === undefined || value === null) return ''
51 return typeof value === 'string' ? value : JSON.stringify(value)
52}
53
54export function textOf(value: unknown, paths: string[]) {
55 return paths
56 .map(path => asText(at(value, path)))
57 .filter(Boolean)
58 .join('\n')
59}
60
61// A string that is exactly "{{path}}" becomes the item's value at that path, keeping its type.
62export function fillArgs(args: unknown, item: unknown): unknown {
63 if (typeof args === 'string') {
64 const match = /^\{\{(.+)\}\}$/.exec(args)
65 return match ? at(item, match[1]) : args
66 }
67 if (Array.isArray(args)) return args.map(arg => fillArgs(arg, item))
68 if (args && typeof args === 'object') {
69 return Object.fromEntries(Object.entries(args).map(([key, arg]) => [key, fillArgs(arg, item)]))
70 }
71 return args
72}
73
74export function toItem(raw: unknown, index: number, fields: Fields): Item {
75 const text = textOf(raw, fields.text)
76 const id = fields.id ? asText(at(raw, fields.id)) : ''
77 const title = fields.title ? asText(at(raw, fields.title)) : ''
78 return { key: String(index), id: id || String(index), title: title || text.slice(0, 100), text }
79}
80
81export function chunks<T>(list: T[], size: number) {
82 return Array.from({ length: Math.ceil(list.length / size) }, (_, index) => list.slice(index * size, (index + 1) * size))
83}
84
85export async function mapLimit<T, R>(list: T[], limit: number, fn: (item: T) => Promise<R>) {
86 const results: R[] = new Array(list.length)
87 let next = 0
88 async function worker() {
89 while (next < list.length) {
90 const index = next++
91 results[index] = await fn(list[index] as T)
92 }
93 }
94 await Promise.all(Array.from({ length: Math.min(limit, list.length) }, worker))
95 return results
96}
97