SLOPSHOPPER

ctxroute-wrap-up-sensor

ctxroute wrapUp (experimental): tells ctxroute how full the context window is after each turn

newprocess
v0.1.0MITupdated 2026-10-07zenonlab/ctxroute/mods/wrap-up-sensor
A shopper browsing a rack in a slop shop
README

ctxroute

test

Declarative context routing for coding agents. ctxroute is a small, deliberately non-Turing-complete language: you declare when a piece of knowledge (an invariant, a pitfall, a project skill) must reach an agent, and the engine injects it into the agent's context at the exact gesture — the tool call — where it matters. Predictable, explainable, harness-agnostic.

  • Language reference: LANGUAGE.md (derived from the engine's constants — a gate fails if it drifts).
  • Harness contract & conformity test: HARNESS-CONTRACT.md.

Platform status — what is PROVEN, and what is not

The engine is platform-agnostic by construction — a CI gate forbids any source from knowing a harness or an OS dialect. What differs is the plumbing around it, and only a run on the real machine proves that. This table is what a real supervisor said on a real runner, and nothing else.

PlatformService unitsWhat that means
Linuxgreeninstalled, loaded by systemd, and it ANSWERS a real request
macOS, eager modegreensame, via the direct-bind plist
macOS, socket activationfailsthe daemon installs and answers, but the cell proving the OUTAGE WINDOW IS CLOSED does not pass
Windowsfails on CIthe installer cannot register the scheduled task on a GitHub runner; it works on the maintainer's machine daily

⚠️ The macOS failure, stated plainly — its CAUSE is NOT established. Socket activation exists so that a request arriving while no daemon runs is QUEUED by the kernel and served by the next instance; that is the whole point, and the cell that proves it does not pass. A cause read off a machine that does not reproduce it would be a guess, and this project does not ship those. What is known: the same platform in eager mode is green, so macOS is usable today — it simply keeps the small outage window socket activation was meant to remove.

⚠️ The Windows failure is about the CI ENVIRONMENT, and that is a claim, not a proof. A GitHub runner has no interactive logon session, which is what the task's trigger needs; the same installer runs daily on the maintainer's machine. Until someone separates "the runner cannot host this" from "the installer is wrong", it is written here as unresolved rather than explained away.

🛑 So this repository does not claim macOS or Windows are proven. Linux is. A framework whose entire purpose is to refuse silent defects cannot begin by hiding one of its own.


Why

Prose instructions ("remember to…") decay: they rely on the agent's vigilance. ctxroute replaces them with a mechanical guarantee — knowledge is delivered when a decidable fact occurs (a file touched, a shell command run, an MCP tool called, a project perimeter entered), never by guessing intent.

The four sources

SourceTriggerExample
File docsfrontmatter match/rules on paths & shell commandsmatch: deploy.sh
MCP docsthe doc's path: docs/mcp/{server}[/{tool}].mddocs/mcp/stripe.md
Tool docsfrontmatter tool: — exact tool name, * wildcardtool: [WebSearch]
Skillsregistry entry (skills in the config): files ∪ MCP servers ∪ toolsproject knowledge, auto-loaded

All sources share one closed boolean base — match (∃) · scope (∃, AND of ORs) · exclude (∀¬) — plus a global filterMode/filterList target filter. Details and proofs: LANGUAGE.md.

Choosing a harness

Routing is deterministic on every harness. Transport is not, and the gap between harnesses is structural, not a question of maturity. The conditions a harness must satisfy to be deployed as a FLEET — and the verdict on each known harness — are published in HARNESS-CONTRACT.md. In one line: HTTP is the only industrial transport, and an HTTP handler is necessary but not sufficient — a harness that caps the size of a hook's output forces the knowledge across N declarations, and fires them as N simultaneous connections it owns.

What is measured here, and the honest answer is that the CAUSE is still open:

  • Rate, not red lines. On the current deployment the loss sits between 0.5 % and 1.6 % of POSTs, stationary over weeks and going back down on its own (20,896 POSTs over six live sessions on 2026-09-20: 1.13 %). An earlier, far worse regime — 16 % and never recovering — belonged to a loopback address the project left on 2026-09-03, and must not be read as this one.
  • Nothing on the server side explains it, and the accept queue does NOT settle it either way. The daemon's connection high-water mark has read 32 against an accept queue measured at 232 on one calm minute, and 254 against that same 232 twelve seconds before a live burst. That counter increments on ACCEPT, so a queue that is full while the loop is starved accepts nothing and is counted as nothing: a low reading is not evidence of a healthy queue. A packet capture at the moment of a failure shows no connection attempt on the wire at all — no SYN leaves, no RST answers. Several plausible stories (queue overflow, a blocked event loop, a filtering driver) were each built and each refuted.
  • Three clients, one loses. A Node client and a .NET client against the very same daemon, on the very same address, under the documented reproducer load (4,800 connections, six concurrent writers): zero failures. Only the harness's own client fails.
  • Upstream will not fix it: anthropics/claude-code#29963 describes a matching failure and is closed as not planned (re-verified 2026-09-20). We are the server; no line of our code can retry a connection that was never attempted.
  • Nothing is lost silently: content promised to a frame that never connects is harvested and carried by the next invocation (src/carryover-pure.js). A lost frame is still an occasion lost for that tool call — a consolation, never a repair.
  • And it has never been measured on anything but a developer workstation, a machine also running browsers, other services and full test suites. Under a saturating local run the rate reaches ~2 %; at rest, the same session lost zero.

⇒ Reliability here is a property of the DEPLOYMENT, not of the framework. Where an error is costly, pick the configuration with no burst; where speed matters more, take the http lane and accept a bounded, measured, non-silent loss. A client you write yourself removes the class entirely: HTTP/2 carries N frames as N streams over ONE connection, so the accept queue is touched once and a refusal cannot occur.

Gemini CLI is not a candidate today: its PreToolUse does not expose the injection channel at all — a capability hole, not a size one.

Install (Claude Code)

  1. Clone this folder anywhere.
  2. Wire the hooks in ~/.claude/settings.json (absolute paths). The gate is declared N times — that is the per-gesture BANDWIDTH, checked by node tools/doctor.js --settings against frames in the config.

🛑 N IS NOT A TUNING KNOB, IT IS THE CAPACITY OF ONE ACTION. The harness caps each hook's OUTPUT, so one declaration carries roughly 7,700 characters. A 50,000-character skill declared at frames: 2 therefore spreads over seven tool calls, and the agent acts six times without knowledge it was owed — which is the exact defect this project exists to remove. The example below is minimal on purpose; size frames against the LARGEST thing you will inject, never against a round number. The live deployment runs 32.

⚠️ Claude Code also implements type: "http", which replaces ~330 ms of node startup per declaration with one local POST to a resident daemon (measured 5,300 ms → 182 ms per action). Read HARNESS-CONTRACT.md before choosing: it is faster, and it is the lane where the transport loss described above exists at all.

⏻ The daemon runs only while a harness uses it, by default — http.lifecycle in the config: auto (default) · on-demand · login. on-demand starts with the first session (Windows: the src/hooks/daemon-ensure.js hook asks Task Scheduler; Linux/macOS: the socket-activated unit starts it at the first connection) and leaves by itself after http.idleSeconds (default 1800) with no request. login keeps it up from login. auto picks on-demand on Linux and macOS (the OS holds the socket, so a restart is guaranteed) and login on Windows, where security suites' "do not disturb" modes can pause Task Scheduler and leave an idle-exited daemon unable to restart — Windows users opt into on-demand knowingly. An idle daemon costs no CPU; what on-demand gives back is its memory. Wire daemon-ensure.js on SessionStart and UserPromptSubmit (the generated wiring does).

{
  "hooks": {
    "PreToolUse": [
      { "matcher": "*", "hooks": [
        { "type": "command", "command": "node /path/to/ctxroute/src/hooks/doc-inject.js --frame 1 --frames 2", "timeout": 10 },
        { "type": "command", "command": "node /path/to/ctxroute/src/hooks/doc-inject.js --frame 2 --frames 2", "timeout": 10 }
      ]}
    ],
    "SessionStart": [
      { "hooks": [{ "type": "command", "command": "node /path/to/ctxroute/src/hooks/session-inject.js --harness claudeCode", "timeout": 10 }] }
    ],
    "SubagentStart": [
      { "hooks": [{ "type": "command", "command": "node /path/to/ctxroute/src/hooks/session-inject.js --harness claudeCode", "timeout": 10 }] }
    ],
    "PostToolUse": [
      { "matcher": "Write|Edit", "hooks": [{ "type": "command", "command": "node /path/to/ctxroute/src/hooks/doc-write-guard.js", "timeout": 10 }] }
    ],
    "PreCompact": [
      { "hooks": [{ "type": "command", "command": "node /path/to/ctxroute/src/hooks/ctxroute-reset.js", "timeout": 5 }] }
    ],
    "UserPromptSubmit": [
      { "hooks": [
        { "type": "command", "command": "node /path/to/ctxroute/src/hooks/turn-count.js", "timeout": 5 },
        { "type": "command", "command": "node /path/to/ctxroute/src/hooks/canary-check.js", "timeout": 5 }
      ]}
    ]
  }
}
  1. Drop docs: docs/mcp/{server}.md for MCP servers, or any .md with a match: frontmatter in your file-docs folder. That's all — no code.
  1. Optional — tune ctxroute-config.json (everything has safe defaults).

Put it outside the clone. ctxroute looks for a per-user configuration at the location your operating system reserves for one, and uses it as soon as the file exists:

PlatformLocation
Linux / BSD$XDG_CONFIG_HOME/ctxroute/ctxroute-config.json, or ~/.config/ctxroute/ctxroute-config.json when that variable is unset
Windows%APPDATA%\ctxroute\ctxroute-config.json (i.e. %USERPROFILE%\AppData\Roaming\…)
macOS~/Library/Application Support/ctxroute/ctxroute-config.json

That file survives a git pull and a re-clone; a config left inside the clone does not. Precedence, highest first: CTXROUTE_CONFIG_PATH (reserved for tests and doctor.js) → --ctxroute-config <absolute path> on the hook command line → the per-user file above → ctxroute-config.json next to the code. With no per-user file present, nothing changes.

{
  "enabled": true,
  "showNotification": true,
  "mode": "smart",
  "defaultThreshold": 4,
  "frames": 2,
  "filterMode": "none",
  "filterList": [],
  "servers": { "odoo": { "subToolParam": "args.tool" } },
  "defaults": { "file": { "mode": "smart" } },
  "skills": { "myproject": { "match": ["myproject"], "mode": "once" } }
}

Codex CLI is supported with thin shells (src/hooks/codex-doc-inject.js, src/hooks/codex-doc-write-guard.js) — declare additionalContextLimit = 0 on the emitters (checked by doctor.js --codex-hooks). Wire session-inject.js --harness codex on SessionStart AND SubagentStart: the shared session shell learns which harness it serves from that flag, and without it a doc restricted by category to the main agent or to sub-agents reaches nobody.

Porting to another harness

The engine is portable by construction (CI gate: no source may know a harness dialect; the dialect lives in harness-profile.js, as data).

  1. Read HARNESS-CONTRACT.md.
  2. Capture one real hook payload from your harness.
  3. node tools/doctor.js --harness payload.json → supported / degraded (each point named with its consequence) / incompatible.

Guarantees (how this is not on faith)

  • Independent executable spec confronted to the engine exhaustively (~400k cases per npm test) — the judge that catches semantic bugs the engine's own tests cannot see.
  • Atoms table: every source × projection × operator cell probed by behavior; blind cells carry a written justification or the build is red.
  • Mutation testing at 100 % with a per-file floor; property-based laws; differential parity against the previous engine on real gestures.
  • Delivery of any size (RFC 2046/6455-style framing + queue): a doc is never dropped for being large; truncation by a harness is loud (seal), never silent.
  • Dead-man switches: doctor.js (engine + wiring) and a canary that watches the other end of the pipe (state/canary.json).

Diagnostics

  • node tools/explain.js --doc <name> --tool X --input '{...}' — why a doc did or did not inject (exact reason, from the real engine).
  • node tools/doctor.js [--settings …] [--codex-hooks …] [--harness …] — is the wiring alive, does the harness conform.
  • node tools/lint-corpus.js — audit of the whole doc corpus.

Journals — where failures are written, and how much disk they may use

Two files under the state directory (stateDir, state/ by default):

FileWhat it records
ctxroute-daemon.logThe daemon's life: start, exits and their cause, stalls, refused connections.
ctxroute-hooks.logA hook or a shared module that failed and stayed fail-open (the agent never sees it; the journal does), including a state write that was lost.

One line per record: <ISO instant> event=<name> key=value…. A failure the daemon survives is event=daemon-error site=<where>; a hook's is event=hook-error hook=<name>; each distinct failure is written once per process, so a failure repeating on every request cannot flood the journal. Errors are always written, whatever the level: a failure at night with debug off still leaves its line, and costs nothing while nothing breaks. The debug level adds the verbose trace: one line per daemon request and each hook's decision. Switch it on in the config to diagnose, off when done; the next record obeys, no restart.

"logging": { "level": "debug", "maxBytes": 1048576, "keptFiles": 5 }
KeyDefaultBoundsMeaning
levelerrorerror · debugdebug adds the trace on top of errors.
maxBytes26214416384 – 8388608Size at which a file rotates: it becomes .1, older ones shift.
keptFiles21 – 10Files kept per journal, the current one included; lowering it frees the extra generations at the next rotation.

Every journal is bounded for life by construction: at most keptFiles × maxBytes each, and at the widest setting 2 × 10 × 8 MB = 160 MB in total, whatever the uptime or the traffic. A value outside the bounds is refused by name: the journal writes event=logging-refused reason=… and keeps running on the defaults. Measured cost on the daemon: 49 µs per request with debug off, about +0.1 ms per request with it on.

Experimental: wrapUp — write the session's knowledge down before the context wall

When an agent's context window fills up, what it learned in the session is summarised away. With wrapUp on, once the fill crosses your threshold the agent may not END ITS TURN until your judges say its knowledge is written down — injectable docs, every skill covering the change, memory, regression tests, whatever your judges check. Off by default; switched off, the generated wiring is byte-identical.

  • atPercent is a share of the session's OWN window (default 70), as Claude Code computes it for the model in use — a 200k and a 1M window both fire at 70 % of themselves, so one setting fits every user. Keep it below the point where your harness compacts on its own. The figure arrives after each whole turn, so the refusal comes at the end of the turn AFTER the crossing.
"wrapUp": {
  "enabled": true,
  "atPercent": 70,
  "maxNudges": 3,
  "judgeTimeoutSeconds": 120,
  "message": "optional — your own words, any language, {percent} is replaced",
  "judges": {
    "docs": { "command": ["node", "scripts/check-docs.js"], "match": ["my-project"] }
  }
}
  • A judge is any program (a script, a suite of judges, a call to a model), given as an argument vector and run WITHOUT a shell, so it behaves the same on Windows, Linux and macOS. It receives on stdin { "version": 2, "sessionId", "cwd", "context": { "tokens", "window", "percent" }, "atPercent", "nudges", "touched" }, exits 0 when the work is done, anything else when it is not; its stdout (plain text, or { "ok": false, "findings": [{ "file", "message", "severity" }] }) is handed to the agent.
  • touched = the files THIS session wrote inside the judge's perimeter ({ "files": [absolute paths], "complete": true }), its sub-agents included, recorded after each file write while the option is on. Several agents working in one repository are therefore never judged on each other's files. complete turns false past 1,000 files: never trust a cut list, widen instead. A file written by a shell command is not in it (the harness does not report what a command writes).
  • A ready example: examples/judges/undocumented-changes.js fails while a file the session wrote has no injectable doc, and names each one. It asks ctxroute's own engine, so "covered" means exactly what the injection would deliver. Without a complete touched list it judges the git repository's changes instead. Wire it with "command": ["node", "<ctxroute>/examples/judges/undocumented-changes.js"]; optional arguments in the same array: "--ignore", ".md" to skip paths, "--since", "main" to include committed work in the git mode. Copy it as the starting point of your own judges.
  • Perimeters use the language's own words (match / scope / exclude / rules / keys), matched on the session directory AND on every file the session wrote: every judge whose perimeter covers either runs, in parallel — an agent started in another folder still meets the judge of the project it wrote in. No judge for a project = one nudge with the message, then the context is settled.
  • Bounded, never a trap: at most maxNudges refusals per context, then the session is released loudly; a judge that cannot start or overruns judgeTimeoutSeconds is stopped (its whole process tree) and never holds the session. It never blocks a compaction.
  • Harnesses: Claude Code ≥ 2.1.287 (sensor = the mod in mods/wrap-up-sensor, loaded with --plugin-dir, CLAUDE_CODE_PLUGIN_DIRS or a folder marketplace). Codex: not supported — its hooks receive no token count. See HARNESS-CONTRACT.md.

Known issues

Windows: a security suite's "do not disturb" / game mode pauses the daemon's task

Symptom. At a prompt, Claude Code shows ctxroute: the daemon could not be started — the scheduled task "ctxroute-http" is DISABLED …, or node tools/doctor.js --settings … reports the OS supervisor can start the daemon as failed. Agents then act without the knowledge ctxroute delivers, until the daemon is back.

Cause, measured (2026-09-29). Some security suites pause Windows scheduled tasks while an application runs full screen — a browser in full screen, the screenshot tool, a game. On Avast the rule is "Pause system background tasks" in Do Not Disturb Mode; its own log (C:\ProgramData\Avast Software\Avast\log\GamingMode.log) shows the rule turning on and off at the exact second Windows records the task as disabled then updated. A daemon that is already running is NOT affected: a disabled task never stops its running instance. What is affected is a START that falls inside a pause — typically the logon trigger, right after a reboot, while a full-screen application restores itself.

What ctxroute does on its own.

  • On Windows the daemon stays up from login (http.lifecycle resolves to login), so a pause mid-session costs nothing.
  • Every harness session and every prompt asks Task Scheduler to start the daemon again (src/hooks/daemon-ensure.js), so a start that was paused succeeds at the next prompt once the pause ends.
  • You are told, live, only when it matters: the notice above appears when the task is disabled AND the daemon does not answer. A disabled task under a daemon that still answers stays silent.

What you can do. In your security suite, exclude ctxroute-http from the full-screen / game-mode rules, or turn off the option that pauses background or scheduled tasks. On Avast: Performance → Do Not Disturb Mode → settings (gear icon) → untick "Pause system background tasks" (path reported by Avast users, not by Avast's own documentation — the wording may differ between versions). Nothing else needs changing, and you never need to disable your antivirus.

Linux and macOS are not affected: the OS holds the listening socket and starts the daemon on the first connection, whatever any other program does to scheduled jobs.

Source 1 files
hooks/register.ts 41 lines
1// ctxroute `wrapUp` SENSOR for Claude Code (EXPERIMENTAL).
2//
3// What it does: after each main-thread turn the engine raises `session.measure`
4// with the live context fill (`context.percent`, documented in the mod API as
5// "compare them with your own threshold here"). When the fill moved, this mod
6// hands it to ctxroute's sensor door, `src/hooks/wrap-up-observe.js`, which
7// records it for the end-of-turn judge (`src/hooks/wrap-up-stop.js`).
8//
9// Rules this file keeps:
10// - It OBSERVES and never changes the event: `next(e)` is always returned.
11// - It never waits for ctxroute: the hand-off runs unawaited, so a slow or
12//   absent ctxroute costs one observation, never a turn.
13// - It decides nothing. The threshold, the judges and the option's switch live
14//   in ctxroute; this file only carries the fact. A threshold read here would be
15//   a second copy of a setting, and two copies drift.
16// - The door is found from this plugin's own folder (`$.plugin.root`), never
17//   from a path typed elsewhere: the plugin lives at `<ctxroute>/mods/wrap-up-sensor`.
18import type { Register } from 'claude-code'
19
20export const register: Register = (on) => {
21  on('session.measure', async ($, e, next) => {
22    const result = await next(e)
23    const context = e.context
24    if (!e.changed.includes('context') || context.percent === undefined) return result
25    const door = `${$.plugin.root}/../../src/hooks/wrap-up-observe.js`
26    void (async () => {
27      const session_id = await $.session.id()
28      const cwd = await $.session.cwd()
29      await $.process.run(['node', door], {
30        stdin: JSON.stringify({
31          session_id,
32          cwd,
33          context: { tokens: context.tokens, window: context.window, percent: context.percent },
34        }),
35        timeoutMs: 30000,
36      })
37    })().catch(() => undefined)
38    return result
39  })
40}
41