SLOPSHOPPER

jev-pipeline

Fetch items over MCP and judge them with Jev without passing the data through context

newguardtool
A shopper browsing a rack in a slop shop
README

dotfiles

Configuration files for Bash, Starship, Homebrew, Ghostty, Herdr, Neovim, OpenCode, Claude Code, Codex, and Pi. GNU Stow manages symlinks from stow/* into $HOME. Ghostty, Herdr, Neovim, Pi, OpenCode, Claude Code, and Codex use Tokyo Night with MonoLisaCode 14 pt.

Theme and font status

  • Ghostty, Herdr, Neovim, Pi, OpenCode, and Codex use Tokyo Night.
  • Neovim uses the Tokyo Night night style with transparent editor, sidebar, float, statusline, and tab-fill surfaces.
  • Ghostty uses MonoLisaCode at 14 pt with explicit regular, italic, bold, and bold-italic styles.
  • Starship uses the intended Nerd Font glyphs through Ghostty's built-in Symbols Nerd Font fallback. See plans/theme-font-glyph-followups.md.
  • Starship and FZF inherit the Tokyo Night terminal palette from Ghostty.
  • Cursor and Zed are archived under old/ and are not restored or rethemed.

Requirements

ToolMinimum VersionNotes
Neovim>= 0.11.0Required for mason-lspconfig v2 and vim.lsp.config()
Git>= 2.19.0Required for lazy.nvim partial clones
GNU Stow>= 2.4.0Symlink manager for tracked dotfiles
GhosttyLatestUses macos-option-as-alt syntax
HerdrLatestAgent-aware terminal workspace manager
Node.jsLTSFor LSP servers via Mason
tree-sitter-cli>= 0.26.1Required for nvim-treesitter main branch parser compilation
lazygit>= 0.40Required for snacks.lazygit keymap (<leader>gg)
ripgrep>= 13.0Required for nvim-spectre search backend
gnu-sedLatestRecommended on macOS for nvim-spectre replace engine (brew install gnu-sed)
imagemagick>= 7.0Required for snacks.image preview support

Spotify Terminal Visualizer

spotify-visualizer is a standalone TypeScript command managed by the bin Stow package. It renders a procedural terminal dot matrix based on the website music visualizer colors, then uses Spotify only for the current track, artist, play state, and track-specific animation seed.

Setup:

# 1. Create or reuse a Spotify developer app.
# 2. Add this redirect URI to that app:
#    http://127.0.0.1:8974/callback
# 3. Export the client id before launching the visualizer:
export SPOTIFY_CLIENT_ID=your_spotify_client_id

spotify-visualizer

The command stores OAuth tokens under ~/.cache/dotfiles/spotify-visualizer/. Run it in any Herdr pane or tab when you want a dedicated visualizer screen.

Controls:

KeyAction
SpaceToggle Spotify play or pause
nSkip to the next track
pSkip to the previous track
sToggle shuffle
rCycle repeat off, context, and current track
q / Ctrl-CQuit and restore the terminal

The visualizer shows this key legend in the header. Short notices, such as pressing a playback key before Spotify has an active track, replace the legend for about 3 seconds.

Shuffle and repeat state use compact status tokens in the header. Shuffle uses [S:-] when inactive and yellow [S:*] when active. Repeat uses gray [R:-] when inactive, red [R:all] for repeat context, and red [R:1] for repeat current track.

If Spotify returns 401 after scopes change, remove the cached token and authorize again:

rm ~/.cache/dotfiles/spotify-visualizer/tokens.json
spotify-visualizer

TypeSafe Jev

The TypeSafe skill defines when and how to request a Jev judgment. Pi 1.0 uses native Codemode classifiers; it no longer installs a custom TypeSafe SDK adapter. OpenCode retains its typesafe_evaluate adapter under stow/opencode/.config/opencode/plugins/typesafe-ai/. Both send the supplied state and questions to TypeSafe and return typed judgments.

In Pi Codemode:

const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");
if (!jev) throw new Error("Jev classifier is unavailable");
const result = await models.classify(jev, {
  state: { message: "A customer asks to cancel today" },
  questions: {
    urgent: {
      type: "bool",
      instructions: "Does this need a reply today?",
      criteria: { true: "Needs a reply today", false: "Can wait" },
    },
  },
});
if (result.stopReason !== "stop") throw new Error(result.errorMessage || "Jev classification failed");
return result.answers.urgent;

Native questions use choice, score, or bool; a bool answer supplies a probability, not a boolean decision. Specialized Jev MCP tools remain available for screening, verification, and gates. Do not include credentials, secrets, or unrelated private data.

OpenCode pins the TypeSafe tool in its Code Mode catalog and adds tool invocation guidance to outgoing contexts where the catalog lists it. Shared global rules and the Jev skill own judgment policy; the adapter explains only the Code Mode calling convention and result shape. Use exact search, parsing, arithmetic, and tests for deterministic facts. Inside execute, the tool returns a validated object with answers, model, and usage; no JSON.parse is needed. For example, when the catalog lists tools.typesafe_evaluate:

const result = await tools.typesafe_evaluate({
  state: "A customer asks to cancel today",
  noul_questions: [{ id: "urgent", instructions: "Does this need a reply today?" }],
});
return result.answers.urgent;

The Noul answer is { type: "noul", noul: probability }, not a boolean. Choice and Score answers include confidence and probability distributions. Missing or malformed provider answers fail explicitly rather than returning an empty success. Tool results retain readable text and model/token metadata for other consumers. Judgments advise; they do not replace permission checks or prove correctness.

Keep TYPESAFE_API_KEY as machine-local state in ~/.config/bash/local.bash:

export TYPESAFE_API_KEY="YOUR_API_KEY"

scripts/bootstrap.sh installs the OpenCode adapters' pinned runtime dependencies. Pi's native classifier needs no separate SDK installation. Apply and reload live configuration only after reviewing and approving the tracked changes.

OpenCode Configuration

The opencode Stow package owns the tracked sources under stow/opencode/.config/opencode/.

Tracked sourcePurpose
opencode.jsoncModel defaults, permissions, MCP servers, providers, skills, and global instructions
cli.jsonTokyo Night, TUI layout, permission handling, and keybindings
commands/Slash commands such as /lg
plugins/typesafe-ai/TypeSafe Jev tool
plugins/tui-conveniences//copy-all, /restart, /update, skill-load confirmations, and the Git status footer
plugins/request-logger/Opt-in private HTTP request capture

New sessions use openai/gpt-6.1-sol-fast with medium reasoning effort and low response verbosity. Its pinned limits match the running ChatGPT catalog checked on 2026-10-04: 400,000 context, 272,000 input, and 128,000 output tokens. This catalog alias sends gpt-6.1-sol with the priority service tier. The limits preserve the current compaction budget rather than assuming the public API's larger window applies to this connection. The TUI hides the session sidebar and persistent tab strip. The TUI also provides Pi-style navigation shortcuts.

OpenCode can use the current user's filesystem, processes, and network without a permission prompt. Review the tracked configuration before you apply it.

REF_API_KEY, EXA_API_KEY, and TYPESAFE_API_KEY are machine-local state. Do not put real credentials in tracked sources.

Apply only the OpenCode Stow package with:

cd ~/dotfiles
./scripts/stow.sh apply opencode

Bootstrap installs plugin dependencies automatically. Restart OpenCode after an apply.

OpenCode HTTP request logger

The last local plugin is request-logger, disabled by default with options.enabled: false. After reviewing the tracked changes and approving live activation, enable it through the plugin's options in opencode.jsonc:

{
  "package": "../../code/personal/dotfiles/stow/opencode/.config/opencode/plugins/request-logger",
  "options": {
    "enabled": true,
    "maxFiles": 100
  }
}

This is an entry in the existing plugins array, not a replacement configuration. Bootstrap installs its dependencies along with the other local plugins. The logger runs in the background service, so a flag on a new CLI process is not used to enable it. It writes to ~/.local/state/opencode/requests, or an absolute options.directory, with directory mode 0700 and file mode 0600. It refuses a symlink at the log directory and retains the newest 100 logger-owned files by default; maxFiles must be a positive integer. Retention limits file count, not total bytes, and leaves unrelated files alone.

Each JSON file records the timestamp, session, agent, catalog model, request kind, raw body string, and byte counts for instructions, input, and tools when those fields exist. The body shows the actual wire model and provider options; the catalog model can be an alias. Byte counts are not token counts, and the instruction count excludes system messages embedded in the input array. For non-UTF-8 bodies, bodyBase64 preserves the original bytes. The original request remains unchanged and readable; logging failures emit a generic server warning without error details and do not block dispatch.

Logs contain full prompts, tool schemas, and file contents, including any secrets already present in the body. The logger deliberately omits URLs and request headers, but does not redact bodies because that would hide what was sent. Keep logs outside version control, review them before sharing, and disable logging when the investigation ends. Capture covers HTTP requests for primary turns, compaction, titles, transient generation, and retries, not responses or WebSocket frames. Keep the logger after request-mutating plugins; hooks registered later can still change the request after capture. Automatic updates, MCP package updates, and automatic compaction remain unchanged.

Codex Chrome Extension Bridge

OpenCode's codex-chrome MCP server uses codex-control-chrome-mcp to control the existing Chrome profile through the Codex Chrome extension. Unlike the isolated chrome-devtools server, it can use existing signed-in tabs and capture screenshots of localhost apps. It grants access to page contents and browser actions; the community bridge does not enforce per-site permissions.

After approving live configuration changes, install the bridge and register it for Google Chrome only:

npm install -g codex-control-chrome-mcp@1.4.1
codex-control-chrome-mcp install-native-host --browser chrome
codex-control-chrome-mcp status --browser chrome

The tracked MCP entry starts codex-control-chrome-mcp from PATH; on Linux the binary is absent, so only that server fails to start. The installer backs up Chrome's existing native-host manifest and records its original host for proxy mode. Reload the Codex Chrome extension after installation, then reconnect codex-chrome through OpenCode's /mcps menu or restart OpenCode. Automatic registration repair remains enabled: after a Codex update restores its own host registration, starting the bridge re-registers it, and the extension may need another reload.

To restore the previous Chrome native-host registration, disconnect codex-chrome in OpenCode and run:

codex-control-chrome-mcp uninstall-native-host --browser chrome

Also remove the codex-chrome MCP entry if you no longer want OpenCode to start the bridge.

GPT Response Verbosity

OpenAI GPT-5 models and the verified GPT-6 Astra, GPT-6 Sol, and GPT-6.1 Sol models support low, medium, and high output verbosity through the Responses API. The tracked configs currently use low.

  • Pi sets verbosity for every GPT-5 model and the verified gpt-6-astra, gpt-6-sol, and gpt-6.1-sol IDs using openai-responses or openai-codex-responses in stow/pi/.pi/agent/extensions/gpt-verbosity.ts. Change the verbosity: "low" value, then run /reload in Pi.
  • Codex sets verbosity with model_verbosity in stow/codex/.codex/config.toml. Change the value, then restart Codex.
  • OpenCode sets textVerbosity per provider and model in stow/opencode/.config/opencode/opencode.jsonc. Update each GPT-5 model entry you use under provider.openai.models or provider.opencode.models, then restart OpenCode.

For example, an OpenCode model override uses this shape:

"providers": {
  "openai": {
    "models": {
      "gpt-5.6-sol-fast": {
        "settings": {
          "textVerbosity": "low",
        },
      },
    },
  },
}

Claude provider request logger

Run claude-log from your regular shell instead of claude to start Claude Code through an opt-in local Anthropic request logger:

claude-log

The command starts a local proxy on a temporary loopback port, launches Claude Code against it, and stops the proxy when Claude exits. Each /v1/messages request is written under ~/.claude/logs/requests/ as readable Markdown plus the raw JSON payload. The Markdown includes request sizes, ranked tool schemas, redacted request headers, the full payload, and the streamed provider response. This logger is inspired by Matt Pocock's agent proxy. It does not capture direct MCP network traffic.

The directory and files use owner-only permissions. The logs can contain sensitive source code, prompts, and connected-service data. Review them before sharing, and remove them when finished:

rm -rf ~/.claude/logs/requests

Set CLAUDE_REQUEST_LOG_DIR to store logs somewhere else. Normal claude sessions do not write request logs.

Claude Code theme and mods

Claude Code uses the Tokyo Night night theme from stow/claude/.claude/themes/tokyonight-night.json, taken from folke's tokyonight.nvim Claude Code extras. Claude Code watches ~/.claude/themes/, so theme edits apply to running sessions.

Mods live in stow/claude/.claude/mods/ and are not stowed. CLAUDE_CODE_PLUGIN_DIRS in settings.json loads them from the repository, and interactive sessions reload a mod when its files are saved.

  • optojr-slack registers optojr_slack_send, which posts as @OptoJr through the relay credentials in the macOS Keychain.
  • skill-toast shows a toast when a skill loads.
  • session-relaunch adds /update, which runs claude update and resumes the session on the new version, and /restart, which resumes the session on whatever version is installed. The relaunch comes from the claude function in stow/bash/.config/bash/functions.bash. Sessions started any other way print the claude --resume command instead of exiting.
  • jev-pipeline registers mcp__jev-pipeline__run, which fetches a list from one MCP tool, optionally enriches each item with a second MCP call, and classifies or reranks every item with Jev. The items never enter the model's context: the reply holds counts and the top items, and the full results go to $TMPDIR/jev-pipeline-<session id>-<timestamp>.json.
  • jev-coding adds a jev_decide reminder to the first Edit or Write of each turn, and records every edit the turn makes. When the agent tries to finish, Jev classifies each edit against the recent prompts as requested, scope creep, speculative, or leftover. A confidently flagged edit blocks the stop once with the list, so the agent reverts or justifies it; a Jev failure shows a toast and lets the turn finish.
  • jev-screen runs jev_screen on every Exa and Ref fetch result, in 20,000-character chunks. A review or block verdict, or a failed screen, adds a warning the model reads after the result, and block also shows a toast.
  • model-cost writes the session's cost per model to $TMPDIR/claude-model-cost-<session id>.json, and statusline.sh shows it after the context usage. The engine reports only a session total, so each response is charged the total's growth since the previous response.

Add a new mod folder to CLAUDE_CODE_PLUGIN_DIRS to load it. Check a mod with claude plugin validate <folder> and claude plugin test <folder>.

Pi provider request logger

Run pi-log from your regular shell instead of pi when you need to inspect the exact payload Pi sends to its model provider:

pi-log

The command enables the tracked request-logger.ts extension for that Pi process only. Each request is written as a readable Markdown file under ~/.pi/agent/logs/requests/, including a size audit, ranked tool schemas, the complete provider payload, the normalized assistant response, and response status metadata when the active provider exposes it. The directory and files use owner-only permissions. These logs can contain source code, prompts, tool results, Gmail, Slack, or Drive data, so do not commit or share them without reviewing the contents. Remove captured requests when finished:

rm -rf ~/.pi/agent/logs/requests

You can also run pi --request-log directly, or set PI_REQUEST_LOG_DIR to store logs somewhere else. Normal pi sessions do not write request logs.

Test Dotfiles

Quick Start (New Mac)

# 1. Clone the repo
git clone https://github.com/manifoldfrs/dotfiles.git ~/dotfiles

# 2. Run the installer
cd ~/dotfiles
./scripts/bootstrap.sh

# 3. Fully quit and reopen your terminal

# 4. Verify Node.js works
node --version

# 5. Verify OpenCode 2 (installed by bootstrap)
opencode --version

Update an Existing Mac / Work Laptop

Use the daily Stow wrapper when the repo is already on the machine and you just want the latest dotfiles applied.

# 1. Get the latest committed dotfiles
cd ~/dotfiles
git pull

# 2. Validate, then reapply all tracked shell/editor/terminal and Herdr config
./scripts/validate-dotfiles.sh
./scripts/stow.sh apply

# First time on this machine? Install Bash, Starship, and the supporting tools:
brew bundle --file=Brewfile
# or only the packaged shell stack:
# brew install bash starship zoxide fzf mise ripgrep fd gawk

# 3. Fully quit and reopen your terminal

# 4. Install/update declared Herdr plugins and reload a running server
./scripts/sync_herdr_plugins.sh

# 5. Verify the basics
node --version
herdr --version

Use ./scripts/bootstrap.sh instead when you also want to install or refresh Homebrew packages, Node.js, and Neovim plugins.

What this already handles for you:

  • stows Bash, Starship, Git, Ghostty, Herdr, Neovim, OpenCode, Claude Code, Codex, Pi settings, and local bin config
  • configures Herdr with Tokyo Night, Bash, tmux-style Ctrl-a bindings, persistence, and agent-aware workspaces
  • avoids rerunning full-machine bootstrap tasks during normal dotfile updates

What ./scripts/bootstrap.sh additionally handles for you:

  • installs Homebrew packages from Brewfile
  • runs Neovim headless plugin sync automatically
  • installs Bun and Plannotator TUI, then syncs the declared Herdr plugins

What is still separate:

  • ./mcp_setup.sh install for the optional Claude Desktop MCP config
  • Claude Code's user-scoped MCP config in ~/.claude.json, which stays local because it contains credentials and account-specific state
  • OpenCode install if you use it on that machine

Herdr and Neovim integrations

The herdr Stow package also manages ~/.config/herdr/plugins.txt and ~/.config/plannotator-tui/config.toml. The default profile includes it. Stow only applies configuration, it does not install or update plugins.

# Existing machines: install the terminal review tool, then sync plugins.
# Requires herdr >= 0.8.0, Bun, and jq on PATH.
brew tap plannotator/tap
# On Homebrew versions that support trust:
# brew trust plannotator/tap
brew install plannotator/tap/plannotator-tui
./scripts/stow.sh apply
./scripts/sync_herdr_plugins.sh

The sync command installs or updates plannotator/herdr-annotate and paulbkim-dev/vim-herdr-navigation, checks the config, and reloads a running Herdr server. Plannotator TUI opens in a full-tab overlay. The agent sidebar prioritizes agents needing attention, uses distinct status symbols, and asks before closing workspaces. Pane-history persistence remains disabled.

ShortcutAction
Ctrl-h/j/k/lNavigate Neovim splits, then adjacent Herdr panes at the edge
Ctrl-a aAnnotate selected terminal text
Ctrl-a Shift-aCopy annotations as agent context
Ctrl-a mManage annotations
Ctrl-a Shift-oReview documents in the current folder
Ctrl-a Shift-lReview the agent's last reply
Neovim visual <leader>aSend the selection to Herdr Annotate

Press Ctrl-a, release it, then press the shortcut's second key.

Launch Plannotator TUI directly from a shell:

plannotator-tui README.md       # Review a file
plannotator-tui docs/           # Browse a folder
plannotator-tui herdr open .    # Review this folder in a Herdr overlay
plannotator-tui herdr last      # Annotate the agent's last reply in Herdr
plannotator-tui last --host pi      # Review the latest Pi reply outside Herdr
plannotator-tui last --host claude  # Review the latest Claude reply outside Herdr

The plannotator-tui commands open the terminal interface. The browser-based plannotator integration is installed alongside it.

Existing Ctrl-a o pane cycling and Ctrl-a z zoom bindings are unchanged. Global Ctrl-k and Ctrl-l navigation takes precedence over shell line deletion and screen clearing inside Herdr. Neovim outside Herdr retains ordinary split navigation. The annotation handoff uses a private temporary file that the plugin consumes and deletes.

Clickable links in terminal chat

In Ghostty on macOS, hold Shift + Cmd and click a link to open it, including inside Herdr. This bypasses application mouse capture and lets Ghostty handle the link. This gesture was verified in this setup, while Herdr's documented Ctrl-click gesture did not work. Keep mouse capture enabled to preserve Herdr's mouse UI. See Herdr's mouse guide.

Pi renders Markdown links as terminal hyperlinks, but relative targets such as docs/plan.md remain unresolved relative paths. The global Pi rules request absolute file:/// URLs for local files in chat and full https:// URLs for web links. File labels can still show readable repository-relative paths and line numbers. Line numbers are informational, not editor jump targets. Links written inside repository documentation remain relative for portability.

Shift-Cmd-click uses the system opener rather than the Herdr Annotate plugin. To review Markdown in Plannotator TUI, use Ctrl-a Shift-o or plannotator-tui herdr open <file.md>. Existing messages are not rewritten by the rule change. Start a new Pi session or use /reload to refresh the global instructions in an existing session.

Neovim secret masking and TypeScript tools

  • cloak.nvim visually masks values in .env, .dev.vars, selected shell configuration files, and TOML token assignments. Use <leader>uC to toggle masking. This only affects display, not file contents, clipboard access, or agent access.
  • :TSC runs the project's TypeScript compiler with --noEmit and opens errors in quickfix. Install TypeScript in the project first.
  • ts-error-translator.nvim improves the readability of TypeScript diagnostics.
  • In TypeScript buffers, :TwoslashQueriesEnable enables inline type queries and `:T
Source 2 files
hooks/register.ts 209 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3import {
4  CLASSIFY_BATCH,
5  JEV_TEXT_LIMIT,
6  RERANK_BATCH,
7  chunks,
8  fillArgs,
9  listAt,
10  mapLimit,
11  mcpJson,
12  textOf,
13  toItem,
14  type Classify,
15  type Item,
16  type PipelineInput,
17  type Rerank,
18} from './pipeline.ts'
19
20const TOOL_NAME = 'run'
21// Matches Pi's codemode demo: four concurrent per-item fetches keep MCP servers from rate limiting.
22const ENRICH_CONCURRENCY = 4
23
24const mcpStep = {
25  server: { type: 'string', description: 'MCP server name as /mcp lists it, e.g. "claude.ai Linear"' },
26  tool: { type: 'string', description: 'Tool name on that server, without the mcp__ prefix' },
27  args: { type: 'object', description: 'Tool arguments' },
28}
29
30const pathList = { type: 'array', items: { type: 'string' }, minItems: 1 }
31
32const inputSchema = {
33  type: 'object',
34  additionalProperties: false,
35  required: ['source', 'fields', 'judge', 'limit'],
36  properties: {
37    source: {
38      type: 'object',
39      required: ['server', 'tool'],
40      properties: {
41        ...mcpStep,
42        items: { type: 'string', description: 'Dot path to the item array in the JSON result; omit when the result is the array' },
43      },
44    },
45    fields: {
46      type: 'object',
47      required: ['text'],
48      properties: {
49        id: { type: 'string', description: 'Dot path to each item id' },
50        title: { type: 'string', description: 'Dot path to a short label shown in the reply' },
51        text: { ...pathList, description: 'Dot paths whose values make up the text Jev judges' },
52      },
53    },
54    enrich: {
55      type: 'object',
56      description: 'Optional per-item MCP call whose results are appended to the item text. A string arg that is exactly "{{path}}" takes the item value at that path.',
57      required: ['server', 'tool', 'text'],
58      properties: {
59        ...mcpStep,
60        items: { type: 'string', description: 'Dot path to an array in the per-item result, e.g. "comments"' },
61        text: { ...pathList, description: 'Dot paths within each result (or each array entry) to append' },
62      },
63    },
64    judge: {
65      type: 'object',
66      required: ['kind'],
67      description:
68        'kind "classify": classes (2+, strong descriptions), optional purpose, context, auto_accept, minimum_margin. kind "rerank": query, optional top_k.',
69      properties: {
70        kind: { enum: ['classify', 'rerank'] },
71        classes: {
72          type: 'array',
73          items: { type: 'object', required: ['description'], properties: { id: { type: 'string' }, description: { type: 'string' } } },
74        },
75        purpose: { type: 'string' },
76        context: {},
77        auto_accept: { type: 'number' },
78        minimum_margin: { type: 'number' },
79        query: { type: 'string' },
80        top_k: { type: 'integer', minimum: 1 },
81      },
82    },
83    limit: { type: 'integer', minimum: 1, maximum: 100, description: 'How many items to list per class (classify) or in total (rerank)' },
84  },
85}
86
87type Verdict = { classification: string; top_probability: number; decision: string }
88
89function failed(text: string) {
90  return { isError: true as const, result: text, text }
91}
92
93async function callJson($: EngineInterface, step: { server: string; tool: string }, args: unknown) {
94  // A JSON round trip drops undefined optional fields before they reach the server.
95  return mcpJson(await $.mcp.call(step.server, step.tool, JSON.parse(JSON.stringify(args ?? {}))))
96}
97
98async function enrich($: EngineInterface, input: PipelineInput, raws: unknown[], items: Item[]) {
99  const step = input.enrich
100  if (!step) return items
101
102  return mapLimit(items, ENRICH_CONCURRENCY, async item => {
103    const raw = raws[Number(item.key)]
104    const result = await callJson($, step, fillArgs(step.args, raw))
105    const entries = step.items ? listAt(result, step.items) : [result]
106    const extra = entries.map(entry => textOf(entry, step.text)).filter(Boolean)
107    return { ...item, text: [item.text, ...extra].join('\n') }
108  })
109}
110
111function jevItems(items: Item[]) {
112  return items.map(item => ({ id: item.key, text: item.text.slice(0, JEV_TEXT_LIMIT) }))
113}
114
115async function classify($: EngineInterface, judge: Classify, items: Item[], limit: number) {
116  const verdicts = new Map<string, Verdict>()
117  for (const batch of chunks(items, CLASSIFY_BATCH)) {
118    const { kind: _, ...options } = judge
119    const result = (await callJson($, { server: 'jev', tool: 'jev_classify' }, { ...options, items: jevItems(batch) })) as {
120      results?: (Verdict & { id: string })[]
121    }
122    for (const verdict of result.results ?? []) verdicts.set(verdict.id, verdict)
123  }
124
125  const judged = items.map(item => ({ ...item, verdict: verdicts.get(item.key) }))
126  const byClass: Record<string, { count: number; review: number; top: string[] }> = {}
127  const ranked = [...judged].sort((a, b) => (b.verdict?.top_probability ?? 0) - (a.verdict?.top_probability ?? 0))
128  for (const item of ranked) {
129    const name = item.verdict?.classification ?? 'unjudged'
130    const group = (byClass[name] ??= { count: 0, review: 0, top: [] })
131    group.count++
132    if (item.verdict?.decision === 'review') group.review++
133    if (group.top.length < limit) {
134      const review = item.verdict?.decision === 'review' ? ', review' : ''
135      group.top.push(`${item.id} ${item.title} (p=${item.verdict?.top_probability ?? '?'}${review})`)
136    }
137  }
138
139  return { judged, summary: { by_class: byClass } }
140}
141
142async function rerank($: EngineInterface, judge: Rerank, items: Item[], limit: number) {
143  const scores = new Map<string, number>()
144  for (const batch of chunks(items, RERANK_BATCH)) {
145    const result = (await callJson($, { server: 'jev', tool: 'jev_rerank' }, { query: judge.query, candidates: jevItems(batch) })) as {
146      ranked?: { id: string; relevance: number }[]
147    }
148    for (const entry of result.ranked ?? []) scores.set(entry.id, entry.relevance)
149  }
150
151  // Jev scores each candidate independently, so scores from separate batches compare directly.
152  const judged = items
153    .map(item => ({ ...item, verdict: { relevance: scores.get(item.key) } }))
154    .sort((a, b) => (b.verdict.relevance ?? 0) - (a.verdict.relevance ?? 0))
155  const top = judged.slice(0, Math.min(limit, judge.top_k ?? limit))
156  return { judged, summary: { top: top.map(item => `${item.id} ${item.title} (relevance=${item.verdict.relevance ?? '?'})`) } }
157}
158
159async function resultsFile($: EngineInterface) {
160  const dir = (await $.env.get('TMPDIR')) || '/tmp/'
161  return `${dir.replace(/\/?$/, '/')}jev-pipeline-${await $.session.id()}-${await $.clock.now()}.json`
162}
163
164function judgeProblem(judge: Classify | Rerank) {
165  if (judge.kind === 'classify' && !(judge.classes?.length >= 2)) return 'judge.classes needs at least two classes'
166  if (judge.kind === 'rerank' && !judge.query) return 'judge.query is required for rerank'
167  return undefined
168}
169
170async function run($: EngineInterface, input: PipelineInput) {
171  const problem = judgeProblem(input.judge)
172  if (problem) throw new Error(problem)
173
174  const raws = listAt(await callJson($, input.source, input.source.args), input.source.items)
175  const items = await enrich($, input, raws, raws.map((raw, index) => toItem(raw, index, input.fields)))
176  const judged = input.judge.kind === 'classify'
177    ? await classify($, input.judge, items, input.limit)
178    : await rerank($, input.judge, items, input.limit)
179
180  const path = await resultsFile($)
181  await $.fs.write(path, JSON.stringify(judged.judged.map(({ key: _, ...item }) => item), null, 2))
182  return { total: items.length, results_file: path, ...judged.summary }
183}
184
185export const register: Register = on => {
186  on('session.start', async ($, e, next) => {
187    await $.tool.register({
188      name: TOOL_NAME,
189      description: [
190        'Fetch a list of items from one MCP tool, optionally enrich each item with a second MCP call, and judge every item with TypeSafe Jev (classify or rerank).',
191        'The fetched data never enters your context: the reply holds counts and the top items only, and every item with its verdict is saved to results_file for Read or jq.',
192        'Use it instead of reading many records and pasting them into jev_classify or jev_rerank yourself.',
193      ].join(' '),
194      inputSchema,
195    })
196
197    return next(e)
198  })
199
200  on('tool.call', { tool: 'mcp__jev-pipeline__run' }, async ($, e) => {
201    const { tool: _, ...input } = e as unknown as PipelineInput & { tool: string }
202    try {
203      return { result: await run($, input) }
204    } catch (error) {
205      return failed(`jev-pipeline failed: ${error instanceof Error ? error.message : String(error)}`)
206    }
207  })
208}
209
hooks/pipeline.ts 97 lines
1import type { McpToolResult } from 'claude-code'
2
3export type McpStep = { server: string; tool: string; args?: Record<string, unknown> }
4export type Fields = { id?: string; title?: string; text: string[] }
5export type Classify = {
6  kind: 'classify'
7  classes: { id?: string; description: string }[]
8  purpose?: string
9  context?: unknown
10  auto_accept?: number
11  minimum_margin?: number
12}
13export type Rerank = { kind: 'rerank'; query: string; top_k?: number }
14
15export type PipelineInput = {
16  source: McpStep & { items?: string }
17  fields: Fields
18  enrich?: McpStep & { items?: string; text: string[] }
19  judge: Classify | Rerank
20  limit: number
21}
22
23export type Item = { key: string; id: string; title: string; text: string }
24
25// Jev truncates item text at 2,000 characters; trimming here keeps request payloads small.
26export const JEV_TEXT_LIMIT = 2_000
27export const CLASSIFY_BATCH = 64
28export const RERANK_BATCH = 250
29
30export function mcpJson(result: McpToolResult): unknown {
31  const text = result.content.find(block => block.type === 'text')?.text
32  if (result.isError) throw new Error(text ?? 'MCP call failed')
33  if (result.structuredContent !== undefined) return result.structuredContent
34  if (text === undefined) throw new Error('MCP call returned no text')
35  return JSON.parse(text)
36}
37
38export function at(value: unknown, path: string | undefined): unknown {
39  if (!path) return value
40  return path.split('.').reduce<unknown>((node, key) => (node && typeof node === 'object' ? (node as Record<string, unknown>)[key] : undefined), value)
41}
42
43export function listAt(value: unknown, path: string | undefined): unknown[] {
44  const list = at(value, path)
45  if (!Array.isArray(list)) throw new Error(`No array at "${path ?? ''}" in the MCP result`)
46  return list
47}
48
49function asText(value: unknown) {
50  if (value === undefined || value === null) return ''
51  return typeof value === 'string' ? value : JSON.stringify(value)
52}
53
54export function textOf(value: unknown, paths: string[]) {
55  return paths
56    .map(path => asText(at(value, path)))
57    .filter(Boolean)
58    .join('\n')
59}
60
61// A string that is exactly "{{path}}" becomes the item's value at that path, keeping its type.
62export function fillArgs(args: unknown, item: unknown): unknown {
63  if (typeof args === 'string') {
64    const match = /^\{\{(.+)\}\}$/.exec(args)
65    return match ? at(item, match[1]) : args
66  }
67  if (Array.isArray(args)) return args.map(arg => fillArgs(arg, item))
68  if (args && typeof args === 'object') {
69    return Object.fromEntries(Object.entries(args).map(([key, arg]) => [key, fillArgs(arg, item)]))
70  }
71  return args
72}
73
74export function toItem(raw: unknown, index: number, fields: Fields): Item {
75  const text = textOf(raw, fields.text)
76  const id = fields.id ? asText(at(raw, fields.id)) : ''
77  const title = fields.title ? asText(at(raw, fields.title)) : ''
78  return { key: String(index), id: id || String(index), title: title || text.slice(0, 100), text }
79}
80
81export function chunks<T>(list: T[], size: number) {
82  return Array.from({ length: Math.ceil(list.length / size) }, (_, index) => list.slice(index * size, (index + 1) * size))
83}
84
85export async function mapLimit<T, R>(list: T[], limit: number, fn: (item: T) => Promise<R>) {
86  const results: R[] = new Array(list.length)
87  let next = 0
88  async function worker() {
89    while (next < list.length) {
90      const index = next++
91      results[index] = await fn(list[index] as T)
92    }
93  }
94  await Promise.all(Array.from({ length: Math.min(limit, list.length) }, worker))
95  return results
96}
97