SLOPSHOPPER

jev-search

Web search through Jev Search. Answers Claude Code's built-in WebSearch with Jev (sources and time window chosen by Jev, every result ranked by relevance) and…

newguardtoaststatuspromptnetwork
★ 6v0.3.1MITupdated 2026-10-02mukiwu/jev-search-mcp
A shopper browsing a rack in a slop shop
README

jev-search-mcp

Jev Search for Claude Code and Codex: a function-hook plugin that answers the built-in WebSearch with Jev, plus the same search as an MCP server, a CLI and a skill. Zero runtime dependencies

把 Jev Search 接進 Claude Code 和 Codex。裝成 Claude Code plugin 時,內建的 WebSearch 會直接由 Jev 回答,Jev 答不出來才退回內建;裝成 npm 套件時,是一個 MCP server 加 CLI,Claude Code、Codex 或任何能跑 shell 的 agent 都能用

Jev Search 的流程是:Jev 模型先讀懂你的一句話,決定要查哪些來源、哪段時間、用什麼關鍵字,再透過 Search1API 同時打 Google、DuckDuckGo、Yandex,必要時加上 Hacker News、Reddit、GitHub、X、arXiv、YouTube、Wikipedia、IMDb、WeChat,最後每一筆結果都由 Jev 打相關度分數。回來的是排好序的連結和摘要,不是生成的答案

這個 repo 同時是 npm 套件(src/)和 Claude Code plugin(hooks/、.claude-plugin/、.mcp.json、skills/),plugin 直接用套件裡的純函式,兩邊行為一致

安裝方式一:Claude Code plugin

這是給 Claude Code 使用者的建議路徑,裝完不用改任何指令或習慣,WebSearch 照常呼叫,答案換成 Jev 的

Function hooks 是 Claude Code 的 early access 功能,需要 2.1.271 以上,並且在 Claude Code 讀得到的地方開旗標,例如 ~/.claude/settings.json:

{ "env": { "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" } }

然後加入 muki-ai-plugins 這個 marketplace 並安裝,shell 或 session 裡的斜線指令都可以:

claude plugin marketplace add mukiwu/muki-ai-plugins
claude plugin install jev-search@muki-ai-plugins

安裝時會問三個設定:Jev Search 實例網址、要不要攔截 WebSearch、每次回幾筆,全部維持預設就是打官方的 jev.s1.dev。裝完重啟 Claude Code 或執行 /reload-plugins。不需要在 CLAUDE.md 寫任何東西,引導都由 plugin 自帶

之後你會得到:

  • 模型知道 WebSearch 現在是 Jev。plugin 會在對話開頭的 context 加一小段說明,跟 CLAUDE.md 同一層:一般網路問題先用 WebSearch,幾秒回來、一次跨多站;但問的是精確數字、排名、留言數這類站台資料,而站台又有正規 API 時,模型直接打 API 仍然是更好的選擇,plugin 不會擋
  • WebSearch 由 Jev 回答。模型讀到的格式跟內建一樣,多一行 Answered by Jev Search 和帶相關度百分比的排序清單。搜完通知列會跳一行 Jev Search: N results via 哪些來源 in 幾秒。Jev 回錯誤、被限流、網路不通、或網域過濾後一筆都不剩,就自動退回內建 WebSearch,對話裡會留一行暗色提示
  • WebSearch 的描述多一段提醒,讓模型把查詢寫成一句話,需要時用文字點名站台或時間範圍
  • 一個 jev_search MCP 工具,要明確指定 sources 或 window 時用
  • 一個 jev-search skill,教模型什麼時候該用、怎麼下請求

設定之後在 /config 裡改,改完 plugin 會重新載入。hook 的細節、退回條件、網域對應表在 hooks/README.md

不想安裝、只想從 checkout 試:

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir .

marketplace 那邊只是一筆指向這個 repo 的紀錄,plugin 本體就是這裡的 .claude-plugin/plugin.json、hooks/、.mcp.json 和 skills/,安裝時會 clone 整個 repo

安裝方式二:npm

給 Codex、其他 MCP host、或不想開 function hooks 的人。套件零依賴,npx 直接跑,需要 Node 22

當 MCP server

# Claude Code
claude mcp add --scope user jev-search -- npx -y jev-search-mcp

# Codex
codex mcp add jev-search -- npx -y jev-search-mcp

Codex 也可以直接寫 ~/.codex/config.toml,桌面版不一定帶著你的 shell PATH,command 建議寫絕對路徑:

[mcp_servers.jev-search]
command = "npx"
args = ["-y", "jev-search-mcp"]
startup_timeout_sec = 20
tool_timeout_sec = 60

[mcp_servers.jev-search.env]
JEV_SEARCH_BASE_URL = "https://jev.s1.dev"

這條路不會攔截內建搜尋,要讓模型優先用 jev_search,在全域指令加一段,Claude Code 放 ~/.claude/CLAUDE.md,Codex 放 ~/.codex/AGENTS.md:

## 網頁搜尋先走 jev_search

- 要查網路資料時,先用 jev_search,不要先用內建的 web search
- 需求用一句話寫,Jev 會自己挑來源和時間範圍,要鎖來源填 sources,要鎖時間填 window
- jev_search 回錯誤或被限流時,才退回內建 web search
- 拿到連結後要讀全文,照平常的方式抓網頁,Jev 只做搜尋不抓頁面

想徹底關掉內建搜尋:Claude Code 在 permissions.deny 加 WebSearch,Codex 在 config.toml 頂層加 web_search = "disabled"。但這樣 Jev 被限流時就沒有備援

當 CLI

npx -y jev-search-mcp search "what do Reddit users think of the Framework laptop this month"
npx -y jev-search-mcp search "new papers on speculative decoding" --window 30d --sources arxiv --max 8
npx -y jev-search-mcp search "bun 1.3 release notes" --json

只裝 skill

npx skills add mukiwu/jev-search-mcp

skills CLI 會把 skills/jev-search/SKILL.md 裝進 Claude Code、Codex、Cursor 等工具的 skill 目錄。這條路不接 MCP,模型看到 skill 之後會改用上面的 CLI 從 shell 查

運作方式

  1. 請求送到 POST /api/ask,帶同源的 Origin 標頭,超過 300 字在字邊界截斷
  2. 上游以 NDJSON 串流回 intent、每個引擎的 lane、最後 done,這裡收齊後照上游的規則以 URL 去重合併,同一個 URL 被多個引擎命中會合併成一列
  3. 排序跟網頁版一致:Jev 的相關度百分比優先,同分看幾個引擎命中,再看原始名次,還沒被打分的排最後
  4. 攔截 WebSearch 時,allowed_domains 對得上 Jev 來源就直接限制來源,對不上就寫進請求文字並在結果端過濾,blocked_domains 只在結果端過濾

它不做什麼

  • 不做本機檔案搜尋,Grep、Glob 那類工具跟它無關
  • 不抓網頁全文,拿到連結後還是用 WebFetch 或原本的方式讀頁面
  • 不儲存查詢紀錄,所有請求直接送到你設定的 Jev Search 實例

設定

Plugin 的三個欄位:

欄位預設說明
baseUrlhttps://jev.s1.devJev Search 實例,hook 和 MCP server 共用
intercepttrue關掉就什麼都不攔、不引導,只留 jev_search 工具
maxResults10攔截 WebSearch 時回幾筆

MCP server 與 CLI 的環境變數:

變數預設說明
JEV_SEARCH_BASE_URLhttps://jev.s1.devJev Search 實例的網址,只取 origin
JEV_SEARCH_TIMEOUT_MS35000單次搜尋的逾時,伺服器端本身是 30 秒
JEV_SEARCH_MAX_RESULTS10沒填 max_results 時的預設筆數

jev_search 工具與 CLI 的參數:

參數CLI說明
query位置參數一句話描述要找什麼
window--windowany、24h、7d、30d,不填讓 Jev 判斷
sources--sources a,b來源清單,不填讓 Jev 判斷
max_results--max回傳幾筆,最多 40

自架 Jev Search

官方實例 jev.s1.dev 是別人的帳單,也有每個 IP 每分鐘約 10 次的限制,量大或想穩定就自己架。照上游 README 部署到 Cloudflare Workers,需要 Search1API 的 key 和至少一個 Jev provider 的憑證,架好後把 plugin 的 baseUrl 或環境變數 JEV_SEARCH_BASE_URL 指過去即可

本機開發時上游跑在 http://localhost:3030,同樣可以直接指過去

開發

npm install                # 只裝 dev 依賴,執行時零依賴
npm test                   # node:test,含真實 stdio 協定測試與 hook 測試,不需要網路
npm run typecheck:hooks    # 用 types/claude-code.d.ts 檢查 hook
npm run validate           # claude plugin validate
npm run smoke -- "Rust async runtimes on Hacker News this month"
  • src/client.js 打 POST /api/ask,串流或整段文字都能收,純函式部分給 hook 共用
  • src/rank.js 從上游移植 URL 去重和排序規則
  • src/format.js 把結果排成給模型讀的文字
  • src/websearch-bridge.js WebSearch 輸入與 Jev 請求、Jev 結果與 WebSearch 輸出之間的轉換
  • src/tool.js 工具定義、參數驗證、設定讀取,server 和 CLI 共用
  • src/mcp.js 手寫的 JSON-RPC over stdio,只實作 tools 相關方法
  • src/server.js MCP 進入點,src/cli.js npx 進入點
  • hooks/jev.js function hook,見 hooks/README.md
  • types/claude-code.d.ts Claude Code 的型別快照,用 /plugin-types 重新產生
  • .claude-plugin/plugin.json plugin manifest,marketplace 紀錄在 muki-ai-plugins

授權

MIT,Jev Search 本身也是 MIT,TypeSafe 與 Jev 的名稱和商標屬於各自的持有者

Source 5 files
hooks/jev.js 109 lines
1// Claude Code function hooks: answer the built-in WebSearch with Jev Search, and
2// fall back to the built-in tool whenever Jev cannot answer.
3//
4// Runs in the plugin environment (no Node, no DOM): everything outside comes
5// through `$`. Reads the request body as one string via $.http.fetch.
6import { foldEvents, httpError, parseNdjsonText, requestFor } from '../src/client.js';
7import { describeWebSearch, planWebSearch, toWebSearchResult } from '../src/websearch-bridge.js';
8
9export const DEFAULT_BASE_URL = 'https://jev.s1.dev';
10
11/** The block appended to the conversation's first message, so no CLAUDE.md is needed. */
12export const CONTEXT_BLOCK_NAME = 'jevSearch';
13export const CONTEXT_GUIDANCE = [
14  'The jev-search plugin is active in this session.',
15  'WebSearch is answered by Jev Search: Jev reads the request, picks the sources (the open web, plus Hacker News, Reddit, GitHub, X, arXiv, YouTube, Wikipedia, IMDb or WeChat when the request calls for them) and a time window, and ranks every result by relevance. It falls back to the built-in search only when Jev cannot answer.',
16  'For questions that need the web, such as news, docs, releases, opinions, or what people are saying on Hacker News, Reddit or X, call WebSearch first, with the request as one plain-language sentence that names the site and the time span in words (for example "what Hacker News says about Bun this month"). It answers in a few seconds and covers several sites in one call.',
17  'When the question is really about a site\'s data (exact counts, points, ids, strict date ranges, full comment threads) and the site has a proper API, querying that API directly is the better tool: WebSearch ranks by relevance, not by engagement.',
18  'Use the jev_search MCP tool only when you must force sources or window.',
19].join('\n');
20
21function flag(value, fallback) {
22  if (value === undefined) return fallback;
23  return value === true || value === 'true';
24}
25
26/**
27 * Normalise the plugin's userConfig values, tolerating strings from older settings.
28 * @param {import('claude-code').PluginOptions} [options]
29 */
30export function readOptions(options = {}) {
31  const baseUrl = typeof options.baseUrl === 'string' && options.baseUrl.trim() ? options.baseUrl.trim() : DEFAULT_BASE_URL;
32  const intercept = flag(options.intercept, true);
33  const n = Number(options.maxResults);
34  const maxResults = Number.isInteger(n) && n >= 1 && n <= 40 ? n : 10;
35  return { baseUrl, intercept, maxResults };
36}
37
38/**
39 * Answer one WebSearch call through Jev. Resolves to `{ result }` for the engine,
40 * or `null` when the built-in tool should run instead (no hits after filtering).
41 * Throws on transport or upstream failure; the caller decides to fall back.
42 *
43 * @param {import('claude-code').EngineInterface} $
44 * @param {{ tool_use_id: string, query: string, allowed_domains?: string[], blocked_domains?: string[] }} e
45 * @param {ReturnType<typeof readOptions>} settings
46 */
47export async function answerWithJev($, e, settings) {
48  const started = Date.now();
49  const plan = planWebSearch(e);
50  const { url, init } = requestFor({ baseUrl: settings.baseUrl, query: plan.request, sources: plan.sources });
51  const response = await $.http.fetch(url, init);
52  if (!response.ok) throw httpError(response.status, response.text);
53  const folded = foldEvents(parseNdjsonText(response.text));
54  const { result, hits } = toWebSearchResult(folded, {
55    toolUseId: e.tool_use_id,
56    query: e.query,
57    allowed: plan.allowed,
58    blocked: plan.blocked,
59    maxResults: settings.maxResults,
60    durationSeconds: Math.round((Date.now() - started) / 100) / 10,
61  });
62  return hits > 0 ? { result, folded, hits } : null;
63}
64
65/**
66 * The hooks module's entry: Claude Code calls it once per activation with the plugin's options.
67 * @param {import('claude-code').On} on
68 * @param {import('claude-code').PluginOptions} options
69 */
70export function register(on, options) {
71  const settings = readOptions(options);
72
73  // Guidance travels with the plugin: appended to the first message's context blocks,
74  // beside CLAUDE.md, so every install gets it without the user writing anything.
75  on('prompt.context', ($, e, next) => {
76    if (!settings.intercept) return next(e);
77    const blocks = e.blocks.filter((b) => b.name !== CONTEXT_BLOCK_NAME);
78    return next({ ...e, blocks: [...blocks, { name: CONTEXT_BLOCK_NAME, text: CONTEXT_GUIDANCE }] });
79  });
80
81  on('tool.describe', { tool: 'WebSearch' }, ($, e, next) =>
82    settings.intercept ? { ...e, description: describeWebSearch(e.description) } : next(e)
83  );
84
85  on('tool.call', { tool: 'WebSearch' }, async ($, e, next) => {
86    if (!settings.intercept) return next(e);
87    $.ui.status('Jev Search…');
88    try {
89      const answer = await answerWithJev($, e, settings);
90      if (!answer) {
91        $.ui.log(`Jev Search found nothing for "${e.query}"; using the built-in WebSearch`, { to: 'debug' });
92        return next(e);
93      }
94      const { intent, totalMs } = answer.folded;
95      const summary = `${answer.hits} result${answer.hits === 1 ? '' : 's'} via ${intent.sources.join(', ')} in ${((totalMs ?? 0) / 1000).toFixed(1)}s`;
96      $.ui.toast(`Jev Search: ${summary}`);
97      $.ui.log(`Jev Search answered WebSearch: ${summary}`, { to: 'debug' });
98      return { result: answer.result };
99    } catch (error) {
100      if (next.signal.aborted) throw error;
101      const reason = error instanceof Error ? error.message : String(error);
102      $.ui.log(`Jev Search unavailable (${reason}); using the built-in WebSearch`);
103      return next(e);
104    } finally {
105      $.ui.status(undefined);
106    }
107  });
108}
109
src/client.js 195 lines
1// Client for jev-search's POST /api/ask, which streams newline-delimited JSON:
2// one `intent`, then `found` and `lane` events as engines answer, then `done`.
3//
4// The pure pieces (prepareQuery, requestFor, parseNdjsonText, foldEvents) are shared with
5// the Claude Code function hook, which runs without Node and reads the body as one string.
6import { mergeItems } from './rank.js';
7
8export const QUERY_MAX_CHARS = 300;
9export const WINDOWS = ['any', '24h', '7d', '30d'];
10export const SOURCES = [
11  'google',
12  'duckduckgo',
13  'yandex',
14  'hackernews',
15  'reddit',
16  'github',
17  'x',
18  'arxiv',
19  'youtube',
20  'wikipedia',
21  'imdb',
22  'wechat',
23];
24
25/** Cut at the last word boundary that fits, so the engines still get whole words. */
26function trimToWord(text, max) {
27  if (text.length <= max) return text;
28  const cut = text.slice(0, max);
29  const space = cut.lastIndexOf(' ');
30  return (space > max / 2 ? cut.slice(0, space) : cut).trim();
31}
32
33/** Trim and bound the request; throws on an empty one. */
34export function prepareQuery(query) {
35  const requested = (query ?? '').trim();
36  if (!requested) throw new Error('query must be a non-empty string');
37  if (requested.length <= QUERY_MAX_CHARS) return { q: requested, truncatedQuery: false };
38  return { q: trimToWord(requested, QUERY_MAX_CHARS), truncatedQuery: true };
39}
40
41/** The origin of a base URL, without relying on URL being present. */
42export function originOf(baseUrl) {
43  try {
44    return new URL(baseUrl).origin;
45  } catch {
46    const origin = /^(https?:\/\/[^/?#]+)/i.exec(baseUrl.trim())?.[1];
47    if (!origin) throw new Error(`Invalid Jev Search base URL: ${baseUrl}`);
48    return origin;
49  }
50}
51
52/**
53 * Build the HTTP request for one search. The endpoint accepts only same-origin
54 * callers, so a non-browser client states the origin itself.
55 *
56 * @param {{ baseUrl: string, query: string, window?: string, sources?: string[] }} input
57 */
58export function requestFor({ baseUrl, query, window, sources }) {
59  const { q, truncatedQuery } = prepareQuery(query);
60  const origin = originOf(baseUrl);
61  /** @type {{ q: string, w?: string, s?: string[] }} */
62  const body = { q };
63  if (window) body.w = window;
64  if (sources && sources.length > 0) body.s = sources;
65  return {
66    url: `${origin}/api/ask`,
67    init: {
68      method: 'POST',
69      headers: { 'Content-Type': 'application/json', Accept: 'application/x-ndjson', Origin: origin },
70      body: JSON.stringify(body),
71    },
72    q,
73    truncatedQuery,
74  };
75}
76
77/** Parse a whole NDJSON body. Blank lines are skipped; a broken line throws. */
78export function parseNdjsonText(text) {
79  return text
80    .split('\n')
81    .map((line) => line.trim())
82    .filter(Boolean)
83    .map((line) => JSON.parse(line));
84}
85
86/** Yield one parsed JSON value per line, tolerating chunk boundaries anywhere. */
87export async function* readNdjson(stream) {
88  const reader = stream.getReader();
89  // Node's TextDecoder takes { stream: true }; the hook environment's declaration does not, and never runs this path.
90  const decoder = /** @type {any} */ (new TextDecoder());
91  let buffer = '';
92  const parse = (line) => (line.trim() ? JSON.parse(line) : undefined);
93  try {
94    for (;;) {
95      const { value, done } = await reader.read();
96      if (done) break;
97      buffer += decoder.decode(value, { stream: true });
98      const lines = buffer.split('\n');
99      buffer = lines.pop() ?? '';
100      for (const line of lines) {
101        const event = parse(line);
102        if (event) yield event;
103      }
104    }
105    buffer += decoder.decode();
106    const last = parse(buffer);
107    if (last) yield last;
108  } finally {
109    reader.releaseLock();
110  }
111}
112
113/**
114 * Fold the event sequence into one result: lanes merged by URL, lane errors
115 * collected, timing from `done`. Throws when no `intent` ever arrived.
116 */
117export function foldEvents(events) {
118  let intent = null;
119  let items = [];
120  const lanes = [];
121  const errors = [];
122  let totalMs = null;
123  let tokens = null;
124  let streamError = null;
125
126  for (const event of events) {
127    switch (event?.type) {
128      case 'intent':
129        intent = event;
130        break;
131      case 'found':
132        break; // progress only; the scored lane follows
133      case 'lane':
134        lanes.push(event);
135        items = mergeItems(items, event.items ?? []);
136        if (event.error) errors.push({ source: event.source, engine: event.engine, message: event.error });
137        break;
138      case 'done':
139        totalMs = event.totalMs ?? null;
140        tokens = event.tokens ?? null;
141        break;
142      case 'error':
143        streamError = event.message ?? 'unknown stream error';
144        break;
145      default:
146        break;
147    }
148  }
149
150  if (!intent) throw new Error(streamError ?? 'Jev Search stream ended without an intent event');
151  return { intent, items, lanes, errors, totalMs, tokens, streamError };
152}
153
154/** Turn a non-2xx response into a readable error. */
155export function httpError(status, bodyText) {
156  let detail = '';
157  try {
158    detail = JSON.parse(bodyText).error ?? bodyText;
159  } catch {
160    detail = bodyText ?? '';
161  }
162  return new Error(`Jev Search HTTP ${status}${detail ? `: ${String(detail).slice(0, 300)}` : ''}`);
163}
164
165/**
166 * Run one search and collect the stream into a single result.
167 *
168 * @param {{ baseUrl: string, query: string, window?: string, sources?: string[] }} input
169 * @param {{ fetch?: (url: string, init?: Record<string, unknown>) => Promise<any>, signal?: AbortSignal }} [deps]
170 */
171export async function askJev(input, { fetch = /** @type {any} */ (globalThis).fetch, signal } = {}) {
172  const { url, init, q, truncatedQuery } = requestFor(input);
173  const response = await fetch(url, { ...init, signal });
174
175  if (!response.ok) {
176    let text = '';
177    try {
178      text = await response.text();
179    } catch {
180      /* body unreadable */
181    }
182    throw httpError(response.status, text);
183  }
184
185  let events;
186  if (response.body && typeof response.body.getReader === 'function') {
187    events = [];
188    for await (const event of readNdjson(response.body)) events.push(event);
189  } else {
190    events = parseNdjsonText(await response.text());
191  }
192
193  return { query: q, truncatedQuery, ...foldEvents(events) };
194}
195
src/websearch-bridge.js 120 lines
1// Translates between Claude Code's built-in WebSearch tool and Jev Search, in both
2// directions: the WebSearch input becomes a Jev request, the Jev result becomes the
3// record WebSearch's output schema expects. Pure functions, no I/O, so the function
4// hook stays a thin wrapper and this file is tested with node:test.
5import { formatSearch } from './format.js';
6import { rankItems } from './rank.js';
7
8/** Domains the model tends to pass as `allowed_domains` that map onto a Jev source. */
9export const DOMAIN_SOURCES = Object.freeze({
10  'reddit.com': 'reddit',
11  'news.ycombinator.com': 'hackernews',
12  'github.com': 'github',
13  'arxiv.org': 'arxiv',
14  'youtube.com': 'youtube',
15  'youtu.be': 'youtube',
16  'wikipedia.org': 'wikipedia',
17  'imdb.com': 'imdb',
18  'x.com': 'x',
19  'twitter.com': 'x',
20  'mp.weixin.qq.com': 'wechat',
21});
22
23function normaliseDomain(domain) {
24  return String(domain ?? '')
25    .trim()
26    .toLowerCase()
27    .replace(/^https?:\/\//, '')
28    .replace(/^www\./, '')
29    .replace(/\/.*$/, '');
30}
31
32function sourceForDomain(domain) {
33  const d = normaliseDomain(domain);
34  if (DOMAIN_SOURCES[d]) return DOMAIN_SOURCES[d];
35  for (const [known, source] of Object.entries(DOMAIN_SOURCES)) {
36    if (d.endsWith(`.${known}`)) return source;
37  }
38  return undefined;
39}
40
41/** True when `url` is on `domain` or one of its subdomains. */
42export function hostMatches(url, domain) {
43  const d = normaliseDomain(domain);
44  if (!d) return false;
45  let host;
46  try {
47    host = new URL(url).hostname.toLowerCase();
48  } catch {
49    return false;
50  }
51  host = host.replace(/^www\./, '');
52  return host === d || host.endsWith(`.${d}`);
53}
54
55/**
56 * Turn a WebSearch call into a Jev request. Allowed domains that all map onto Jev
57 * sources become `sources`; otherwise they are named in the sentence and enforced
58 * afterwards by filterByDomains. Blocked domains are only enforced afterwards.
59 *
60 * @param {{ query: string, allowed_domains?: string[], blocked_domains?: string[] }} call
61 */
62export function planWebSearch({ query, allowed_domains, blocked_domains }) {
63  const allowed = (allowed_domains ?? []).map(normaliseDomain).filter(Boolean);
64  const blocked = (blocked_domains ?? []).map(normaliseDomain).filter(Boolean);
65  let request = String(query ?? '').trim();
66  let sources;
67
68  if (allowed.length > 0) {
69    const mapped = allowed.map(sourceForDomain);
70    if (mapped.every(Boolean)) {
71      sources = [...new Set(mapped)];
72    } else {
73      request = `${request} from ${allowed.join(' or ')}`;
74    }
75  }
76
77  return { request, sources, allowed, blocked };
78}
79
80/** Keep only rows inside the allowed domains and outside the blocked ones. */
81export function filterByDomains(items, { allowed = [], blocked = [] } = {}) {
82  return items.filter((item) => {
83    if (blocked.some((d) => hostMatches(item.url, d))) return false;
84    if (allowed.length > 0 && !allowed.some((d) => hostMatches(item.url, d))) return false;
85    return true;
86  });
87}
88
89/**
90 * Build the record WebSearch's output schema describes from a folded Jev result.
91 *
92 * `results` carries one hit list (what the transcript renders as Links) and one text
93 * block with Jev's ranking, so the model reads the same shape the built-in tool gives.
94 */
95export function toWebSearchResult(folded, { toolUseId, query, allowed, blocked, maxResults = 10, durationSeconds = 0 }) {
96  const kept = filterByDomains(folded.items ?? [], { allowed, blocked });
97  const ranked = rankItems(kept).slice(0, Math.max(1, maxResults));
98  const text = formatSearch({ ...folded, items: kept }, { maxResults });
99  const result = {
100    query,
101    results: [
102      { tool_use_id: toolUseId, content: ranked.map((item) => ({ title: item.title, url: item.url })) },
103      `Answered by Jev Search.\n${text}`,
104    ],
105    durationSeconds,
106    searchCount: 1,
107  };
108  return { result, hits: ranked.length };
109}
110
111const DESCRIPTION_SUFFIX =
112  'This tool is answered by Jev Search: phrase the query as one plain-language sentence and name a site or a time span in words when it matters (for example "what Hacker News says about Bun this month"). Results come back ranked with a relevance percentage and a snippet each, in a few seconds, across the open web plus Hacker News, Reddit, GitHub, arXiv, YouTube, Wikipedia, IMDb and WeChat. Reach for it first for web questions; a site\'s own API remains the better tool when you need exact counts, ids or strict date ranges. Use the jev_search MCP tool only when you must force sources or window.';
113
114/** Append the Jev guidance to WebSearch's own description, once. */
115export function describeWebSearch(description) {
116  const base = String(description ?? '').trimEnd();
117  if (base.includes('Jev Search')) return base;
118  return `${base}\n\n${DESCRIPTION_SUFFIX}`;
119}
120
src/rank.js 79 lines
1// Ordering and URL folding, ported from jev-search's src/lib/rank.ts and merge.ts so
2// the MCP tool ranks the same way the web UI does.
3
4const TRACKING_PARAM = /^(utm_|ref$|ref_|fbclid|gclid|igshid|share_id|rdt|si$|feature$|lang$|s$|t$)/i;
5
6/** Host + path + the query params that identify content, minus tracking noise. */
7export function canonicalUrl(url) {
8  try {
9    const u = new URL(url);
10    let host = u.hostname.toLowerCase().replace(/^www\./, '');
11    if (host === 'twitter.com') host = 'x.com';
12    const path = u.pathname.replace(/\/+$/, '').toLowerCase();
13    // forEach rather than entries(): it is the one iteration method every runtime here declares.
14    /** @type {[string, string][]} */
15    const pairs = [];
16    u.searchParams.forEach((v, k) => pairs.push([k, v]));
17    const params = pairs
18      .filter(([k]) => !TRACKING_PARAM.test(k))
19      .sort(([a], [b]) => a.localeCompare(b))
20      .map(([k, v]) => `${k}=${v}`)
21      .join('&');
22    return `${host}${path}${params ? `?${params}` : ''}`;
23  } catch {
24    return url;
25  }
26}
27
28/**
29 * Fold a new lane into the rows already collected. The same URL from a second
30 * engine is one row: engines are unioned, the higher relevance wins, the better
31 * rank and the longer snippet are kept.
32 */
33export function mergeItems(existing, incoming) {
34  const byUrl = new Map();
35  const out = [];
36  for (const item of existing) {
37    byUrl.set(canonicalUrl(item.url), item);
38    out.push(item);
39  }
40  for (const item of incoming) {
41    const key = canonicalUrl(item.url);
42    const found = byUrl.get(key);
43    if (!found) {
44      byUrl.set(key, item);
45      out.push(item);
46      continue;
47    }
48    const publication = (!found.publishedDate && item.publishedDate) || found.ageHours === null ? item : found;
49    const merged = {
50      ...found,
51      engines: [...new Set([...found.engines, ...item.engines])],
52      relevance: Math.max(found.relevance, item.relevance),
53      ranked: found.ranked || item.ranked,
54      position: Math.min(found.position, item.position),
55      publishedDate: publication.publishedDate,
56      ageHours: publication.ageHours,
57      freshness: publication.freshness,
58      snippet: found.snippet.length >= item.snippet.length ? found.snippet : item.snippet,
59    };
60    byUrl.set(key, merged);
61    out[out.indexOf(found)] = merged;
62  }
63  return out;
64}
65
66/** Best match: judge percentage, ties by engine agreement, then engine rank. Unscored rows last. */
67export function compareItems(a, b) {
68  if (a.ranked !== b.ranked) return a.ranked ? -1 : 1;
69  const ra = Math.round(a.relevance * 100);
70  const rb = Math.round(b.relevance * 100);
71  if (ra !== rb) return rb - ra;
72  if (a.engines.length !== b.engines.length) return b.engines.length - a.engines.length;
73  return a.position - b.position;
74}
75
76export function rankItems(items) {
77  return [...items].sort(compareItems);
78}
79
src/format.js 73 lines
1// Turn one search result into the text block the model reads.
2import { rankItems } from './rank.js';
3
4function age(item) {
5  if (item.publishedDate) return item.publishedDate;
6  const h = item.ageHours;
7  if (h === null || h === undefined || !Number.isFinite(h)) return null;
8  if (h < 1) return 'just now';
9  if (h < 48) return `${Math.round(h)} hours ago`;
10  const days = Math.round(h / 24);
11  if (days < 60) return `${days} days ago`;
12  return `${Math.round(days / 30)} months ago`;
13}
14
15function clip(text, max) {
16  const t = (text ?? '').replace(/\s+/g, ' ').trim();
17  return t.length > max ? `${t.slice(0, max - 1)}…` : t;
18}
19
20function seconds(ms) {
21  return ms === null || ms === undefined ? null : `${(ms / 1000).toFixed(1)}s`;
22}
23
24/**
25 * @param {import('./client.js').askJev extends (...a: any) => Promise<infer R> ? R : never} result
26 * @param {{ maxResults?: number, snippetChars?: number }} [options]
27 */
28export function formatSearch(result, { maxResults = 10, snippetChars = 300 } = {}) {
29  const { intent, errors = [], totalMs, truncatedQuery, streamError } = result;
30  const ranked = rankItems(result.items ?? []);
31  const shown = ranked.slice(0, Math.max(1, maxResults));
32  const hidden = ranked.length - shown.length;
33
34  const head = [
35    `Query: ${intent.query}`,
36    `sources: ${intent.sources.join(', ')}`,
37    `window: ${intent.window}`,
38    `judge: ${intent.judge}`,
39  ];
40  const took = seconds(totalMs);
41  if (took) head.push(took);
42
43  const lines = [head.join(' | ')];
44
45  if (truncatedQuery) lines.push(`Note: the request was trimmed to ${300} characters, the Jev Search limit.`);
46  if (streamError) lines.push(`Note: the stream ended early (${streamError}); results below may be partial.`);
47
48  if (shown.length === 0) {
49    lines.push('', 'No results.');
50  } else {
51    lines.push('');
52    shown.forEach((item, i) => {
53      const score = item.ranked ? `${Math.round(item.relevance * 100)}%` : 'unscored';
54      lines.push(`${i + 1}. [${score}] ${clip(item.title, 200)}`);
55      lines.push(`   ${item.url}`);
56      const snippet = clip(item.snippet, snippetChars);
57      if (snippet) lines.push(`   ${snippet}`);
58      const meta = [`via ${item.engines.join(', ')}`];
59      const when = age(item);
60      if (when) meta.push(when);
61      lines.push(`   ${meta.join(' · ')}`);
62    });
63    if (hidden > 0) lines.push('', `… ${hidden} more result${hidden === 1 ? '' : 's'} not shown; raise max_results to see them.`);
64  }
65
66  if (errors.length > 0) {
67    lines.push('', `Lane errors: ${errors.map((e) => `${e.source}/${e.engine}: ${e.message}`).join('; ')}`);
68  }
69
70  lines.push('', 'Snippets are engine excerpts; fetch a URL when you need the page itself.');
71  return lines.join('\n');
72}
73