SLOPSHOPPER

Orchestration Skills

Dual-host development orchestration for Claude Code and Codex: super-plan creates wave-ready tasks with machine-checkable contracts, multi-model executes…

newprompt
★ 5v4.16.0no licenseupdated 2026-10-08TemMax/agent-skills/plugins/orchestration
A shopper browsing a rack in a slop shop
README

agent-skills

A dual Claude Code and Codex plugin marketplace (temmax) with two plugins covering the full development pipeline: plan → supervised execution → review. Model-specific guidance is scoped to cited first-party Anthropic and OpenAI sources; shared safety rules stay provider-neutral. Load-bearing logic ships as tested code, not prose — a Claude wave runner, a Codex state/verifier helper, a plan linter, and the offline release suite (see tests/README.md).

  • orchestration — the full pipeline: ship conducts planning (super-plan) and execution (multi-model) into a reviewed PR. The orchestrator model researches, plans into contract-carrying waves, and launches executor subagents through the selected host adapter. Claude and exact GPT-5.6 and GPT-6 Astra/Sol/Luna profiles plus a conservative generic fallback guide routing; each executor is isolated in its own worktree and judged against its contract by a different model.
  • code-review — critical, evidence-based review of uncommitted changes or a GitHub PR, performed by the session's own model.

Each skill works on whatever model the session runs on: it reads its own model identity and loads the matching profile from references/. There is no -opus variant to pick between any more.

PluginSkillWhat it does
orchestrationsuper-planWave-native planning: research to decomposition depth, questions only about product forks and contradictions, one start approval, tasks carrying machine-checkable contracts grouped into waves by file-independence, validated by the shipped plan-lint.mjs before the start approval. Planning discipline adapted from Jesse Vincent's superpowers (MIT, attribution shipped).
orchestrationshipThe pipeline conductor: one command from request to reviewed PR — super-plan → supervised waves on a feature branch → critical-review of the PR and its threads. Adds no machinery or gate of its own: the branch and the PR are stated in planning's one start approval, fixes routed by behavior change, and the merge always stays with the user.
orchestrationmulti-modelModel routing, effort selection, task-prompt template, review checklist, and supervised waves executed by native Claude/Codex drivers using the shipped policy — isolated executors judged against a machine-checkable contract by a different model, with the escalation ladder as tested code — plus an orchestrator-drift advisory hook that watches the orchestrator session itself.
code-reviewcritical-reviewScope detection, PR description+threads protocol, tiered findings table (Blocker → Nit), and a post-review fix phase that answers and resolves the PR threads its findings came from.

How the model routing works

Claude Code states the session model in the system prompt ("You are powered by the model named X. The exact model ID is Y"). Step 0 of each skill maps that ID to exactly one profile file and forbids reading the others:

Model IDorchestration profilecode-review profile
claude-opus-5-5 (any context suffix)references/orchestrator-opus-5-5.mdreferences/reviewer-opus-5-5.md
claude-fable-5-1references/orchestrator-fable-5-1.mdreferences/reviewer-fable-5-1.md
claude-fable-5references/orchestrator-fable-5.mdreferences/reviewer-fable-5.md
claude-opus-5 (any context suffix)references/orchestrator-opus-5.mdreferences/reviewer-opus-5.md
claude-opus-4-8 (any context suffix, e.g. [1m])references/orchestrator-opus-4-8.mdreferences/reviewer-opus-4-8.md
anything elsenone — model-agnostic rules only, and the skill says sosame

Opus 5.5 (claude-opus-5-5) is the default heavy executor, verifier, and open-research route: an upgrade to Opus 5 on every evaluation in its summary table at a lower list price ($4 / $20 per million input/output tokens vs Opus 5's $5 / $25), and it matches Fable 5.1 as the most injection-robust route through tool results (IPI 0.1% at k=1). Untrusted text must still be handed to it by path, not pasted: compliance with instructions planted in pasted text rises from 2.1% at default effort to 7.4% at max. Grounded in the Claude Opus 5.5 system card (230 pp., September 2026).

Opus 5 (claude-opus-5) is the previous default heavy executor and verifier, retained as the supervisor fallback; Opus 4.8 is retained only for compiled-binary reverse-engineering (Opus 5's Fable-class cyber classifier blocks it) and as the cyber-refusal fallback. Opus 5's effort rule inverts Opus 4.8's — higher effort makes it worse on long-horizon work (documented overthinking / self-verification loops), so its profile runs at high, not xhigh, and its effort self-check flags too-*high*, not too-low. Its card also names an unverified-subagent-relay failure mode, so the Opus 5 orchestrator profile doubles down on verifying subagent claims. Grounded in the Claude Opus 5 system card (193 pp., July 2026).

Fable 5.1 (claude-fable-5-1) has its own profiles, grounded in the Claude Fable 5.1 & Mythos 5.1 system card (212 pp., September 2026). It beat Opus 5 on long-horizon coding (FrontierSWE v2 0.57 vs Opus 5's 0.52, pp. 170–171) at roughly half Fable 5's cost per task (p. 5), but Opus 5.5 now leads the lineup on that same metric (FrontierSWE v2 62.3 vs Fable 5.1's 56.3, p. 179), so Fable 5.1 is no longer the strongest long-horizon coder overall — it is a premium route, used only on the user's word (said in the session or written in their instruction files) and recorded in approvals.premium. Three of its measurements still change the rules: as a judge it is the first model since Opus 4.7 with a measured self-recognition bias (0.1 points out of 10, lenient when told the author is Claude, p. 124) — the runner's judge prompt never names the executor and the bias is bounded by the contract's mechanical half, so Fable 5.1 (claude-fable-5-1) still judges Opus 5, and the prompt rule is now a contract test; on scoped coding its score peaks at medium because higher effort adds unrequested out-of-scope edits (p. 169), so every Fable 5.1 executor prompt carries a scope line; and it is the most injection-robust model to date (IPI 0.1% at k=1, p. 83), the executor for untrusted content whose compromise would reach secrets or actions — and, like every Fable 5.1 executor route, only used with approvals.premium recorded from the user's word. Its card also documents an orchestrator failure the profile guards against: distorting user intent to subagents, including a fabricated user authorization and a bypassPermissions launch (pp. 95–96). Plans address it as claude-fable-5-1 (fable is only its Agent-tool alias); Fable 5's profiles and dossier sections stay for history.

The profile carries everything that is genuinely model-specific: the session's reasoning-effort guidance, amendments to the numbered process steps, and the model's own documented failure modes. The shared body carries everything else.

Why the split matters. Merging the variants naively — leaving an unconditional "run this session at xhigh reasoning effort" in the shared overview — made a Fable 5 orchestrator adopt Opus 4.8's effort directive in 3 of 3 test runs, reasoning that "the imperative is phrased generically", and import Opus-specific process amendments along with it. A model with no profile at all hedged instead of falling back cleanly. With the effort directive scoped inside the profile, a model-ID gate on each profile, and an explicit fallback row, 14 of 14 runs across Fable 5, Opus 4.8 and Sonnet 5 loaded the right profile, refused the wrong one, and applied the right effort. Repeated for Fable 5.1 on 2026-09-01 with headless claude -p --plugin-dir runs: 6 of 6 on claude-fable-5-1 (four multi-model, two critical-review) loaded exactly orchestrator-fable-5-1.md / reviewer-fable-5-1.md — one run captured with --output-format stream-json shows a single profile read and nothing else — and a Sonnet 5 control reported no matching profile.

The Opus profile also self-checks the session effort: Step 0 surfaces the live value via the ${CLAUDE_EFFORT} substitution, and the profile halts an orchestration started at medium or below with a request to restart higher (review notes the shortfall rather than halting). If the substitution ever fails to expand, the skill treats effort as unknown and proceeds — verified across 10 runs (medium halts, high notes the floor, xhigh proceeds silently, an unexpanded placeholder degrades gracefully, Fable never false-warns).

Recommendation for Opus 4.8 sessions: run the orchestrator at xhigh reasoning effort; high is the floor when latency-bound. Grounding from the Opus 4.8 system card: SWE-bench Pro peaks at xhigh (69.8, p. 196), deep-research agentic scores rise monotonically through max (DRACO 80.4, p. 208), Anthropic's own multi-agent harnesses ran the orchestrator at max effort (p. 214), and higher effort roughly halves prompt-injection susceptibility (p. 80). Low/medium effort on Opus 4.8 is executor territory (its minimum effort already matches Opus 4.7's maximum). No equivalent level is pinned for Fable 5 — that measurement does not exist for it, and the Fable profile says so explicitly. The same holds for Fable 5.1, whose profile names xhigh as the documented long-horizon sweet spot (xhigh matches max at 19–25% fewer tokens, pp. 193–194) without pinning it.

Both skills also ship a dossier (references/model-dossiers.md, references/reviewer-dossier.md) with benchmark numbers, documented failure modes, and page references to the system cards — loaded on demand for contested calls.

Full model IDs. Plans and runner args name full Claude IDs, never aliases: claude-haiku-5-5, claude-haiku-4-5-20251001 and claude-sonnet-5 (retired routes that stay valid for approved plans), claude-sonnet-5-5, claude-opus-5-5, claude-opus-5, claude-opus-4-8, claude-fable-5-1. Aliases are rejected by name because they re-point silently — on 2026-09-22 opus moved from Opus 5 to Opus 5.5, so every route still written as opus would have changed model without an edit. By 2026-10-08 sonnet and haiku had moved the same way. The one alias-only surface is the Claude Code Agent tool, whose schema accepts only aliases; that exception is covered by the probe-dated alias mapping in multi-model's Model identifiers table, which is re-verified whenever a new Claude model ships.

All skills always reply to the user in the language the user writes in.

Hosts, models, and lifecycle limits

Both marketplaces expose the same plugin folders. Claude Code reads skills/, Codex reads skills-codex/, and the shared references live under skills/. Claude Code and Codex discover the three orchestration skills (super-plan, ship, and multi-model) plus the critical-review skill. The exact-profile roster is deliberately narrower than a claim that every profile is a production route:

HostExact model IDs with a profileRole / effort conclusion
Claude Codeclaude-opus-5-5, claude-fable-5-1, claude-fable-5, claude-opus-5, claude-opus-4-8Existing Claude routes retain each profile's documented role and effort guidance.
Codexgpt-5.6-solExact profile exists; no executor, orchestrator, reviewer, or supervisor role/effort is production-supported by the 2026-09-04–05 UTC calibration.
Codexgpt-5.6-terraExact profile exists; no executor, orchestrator, reviewer, or supervisor role/effort is production-supported by the 2026-09-04–05 UTC calibration.
Codexgpt-5.6-lunaExact profile exists; no executor, orchestrator, reviewer, or supervisor role/effort is production-supported by the 2026-09-04–05 UTC calibration.
Codexgpt-6-astraActive-session orchestration and review profiles; GPT-5.6 executors with a separate Astra supervisor are calibration candidates, not production-qualified routes. A separately approved Astra initial executor or final rung is uncalibrated and requires a fresh Astra supervisor.
Codexgpt-6-solExact profiles and dossiers; retired as an executor route (default executors moved to gpt-6.1-sol, decision 011), still valid for approved plans, but no longer the lower-cost review option; 2026-09-23 calibration (tests/eval/gpt-6-results-2026-09-23.md) — review unsupported; supervisor only as the standard all-Luna supervisor (policy, uncalibrated).
Codexgpt-6.1-solExact profiles and dossiers; default ordinary/difficult executor, Luna ladder rung, research and seam-audit route, and the standard supervisor of all-Luna waves (fixture 9/9, 2026-09-29); review route measured-supported (strict gate 10/10, 10/10, PR 3/4, 2026-09-30). Record: tests/eval/gpt-6-1-sol-results-2026-09-29.md.
Codexgpt-6-lunaExact profiles and dossiers; default Codex executors per shared routing; 2026-09-23 calibration (tests/eval/gpt-6-results-2026-09-23.md) — review and supervisor routes unsupported.
Eitherany other model IDThe generic profile applies; missing identity/effort stay unknown, and no model-specific reliability claim follows.

A bare family label such as GPT-6 does not select any exact profile by itself: Codex CLI 0.155.1 gives Astra, Sol, and Luna the identical host instruction "You are Codex, an agent based on GPT-6" (verified 2026-09-23), so that phrase cannot distinguish between them and the skills load the generic profile instead. An exact ID takes priority; unsupported IDs or unresolved conflicts select generic. Quotes and available child-model lists are not session identity.

The gpt-5.6 alias normalizes only to gpt-5.6-sol; it is not a plan model ID. The dated record is tests/eval/gpt-5-6-results-2026-09-04.md: the final post-fix regression recorded 87 default rows (63 pass, 24 fail) with 3/24 required skill cells passing; its critical run recorded 204 rows (162 pass, 42 fail). Therefore every GPT-5.6 seed route is unsupported and must delegate the routing decision upward to a separately supported provider route or an authorized calibration. The result does not turn invalid or failed cells into support.

When a skill starts on Astra, Astra remains in the active seat: it plans, coordinates, and reviews. A supervised implementation wave ordinarily uses only Luna, Terra, or Sol, with a separate Astra supervisor. A separately approved Astra initial executor or final rung requires astra_executor_reason and a fresh Astra supervisor; Sol exhaustion never promotes, resets, or raises effort automatically. A fresh Astra reviewer provides context separation, not a different-model check. The drift hook judges Astra, GPT-6 Sol and GPT-6 Luna orchestrators with gpt-6.1-sol at high, and a GPT-6.1 Sol orchestrator with gpt-5.6-sol (decision 012). On Claude Code the judge is claude-haiku-5-5 at low, started without MCP servers (decision 014). Other profiles retain their existing rules. See the role decision and the Astra dossier. The bounded Astra pilot records offline checks, bounded live cases, preserved failures and scorer disagreement, and the remaining end-to-end calibration gaps. Per decision 006, Codex routing for new plans moves to gpt-6-sol/gpt-6-luna executors with the fixed gpt-6-astra/high supervisor: GPT-5.6 IDs are no longer chosen for new plans, but an already-approved plan carrying GPT-5.6 fields still runs to completion.

Claude Code 2.1.287+ loads a UI-free lifecycle mod with each plugin. It obtains the main session model directly from the host, including headless sessions, and refreshes runtime context at lifecycle starts or a changed-model prompt. Child identity is explicitly unknown rather than borrowed from the parent. Effort fallback and profile/calibration guards remain in force. Codex loads the separate classic-only hook configuration. See the live adapter checks for the tested scenarios and remaining limits. The mod uses the host memory loader for applicable AGENTS.md files, preserving existing instructions and read rules. The quiet-review follow-up records the final dual-host checks and all retained failed approaches.

Both Codex manifests intentionally retain their hooks fields, including the orchestration advisory drift hook. Lifecycle behavior is host-dependent; ChatGPT surfaces do not run Codex lifecycle hooks. Treat hooks as advisory where supported, never as a substitute for the plan, worktree, contract, and independent-supervisor safety model. The generic plugin validator bundled with some tooling is not authoritative for these manifests because it rejects the approved hooks field when PyYAML is unavailable; repository contracts and a real disposable Codex install rehearsal are the release checks.

Codex waves default to codex-wave-runner.mjs, a deterministic, model-free driver that runs the native protocol's own state machine (codex-wave-state.mjs) one task per worktree under a shared --jobs limit, so an orchestrator model no longer spends its wall time on the protocol's tool-call round trips. It shells out to codex exec, which needs the repository's .git writable and network access to reach the model API — in a sandboxed Codex session, grant .git as a writable root (--add-dir <repo>/.git) and network access, or use full access — and never bypasses the state machine it drives. The native codex-wave-protocol.md action loop — the orchestrator model driving codex-wave-state.mjs directly, one tool call at a time — remains the fallback for a host or session that cannot run the runner script.

Installation

Claude Code installation

/plugin marketplace add TemMax/agent-skills
/plugin install orchestration@temmax
/plugin install code-review@temmax

Codex installation

Clone this repository, then register that checkout as the repository/team marketplace (replace the path, but keep the selector name):

codex plugin marketplace add /absolute/path/to/agent-skills
codex plugin list --marketplace temmax --available --json
codex plugin add orchestration@temmax --json
codex plugin add code-review@temmax --json
codex plugin list --marketplace temmax --json

Do not use a personal marketplace for this repository.

Local validation before publishing

For local development, register this checkout as the temmax marketplace using codex plugin marketplace add /absolute/path/to/agent-skills --json or claude plugin marketplace add /absolute/path/to/agent-skills --json, then install both plugins with that host's commands above. This changes the source for this marketplace; use its original GitHub source again when returning to released versions. Check the installed versions and open a new session. Do not copy files into a plugin cache manually. Installing and validating packages calls no model; behavioral evaluations are separate and require explicit authorization.

Migrating from the old names

The repository was TemMax/claude-skills and the marketplace temmax-skills. GitHub redirects the old repository URL, but the marketplace rename is not redirected: ...@temmax-skills selectors stop resolving, so re-register once.

Claude Code:

/plugin marketplace remove temmax-skills
/plugin marketplace add TemMax/agent-skills
/plugin install orchestration@temmax
/plugin install code-review@temmax

Codex (after git pull in your checkout, or a fresh clone):

codex plugin remove orchestration@temmax-skills --json
codex plugin remove code-review@temmax-skills --json
codex plugin marketplace remove temmax-skills --json
codex plugin marketplace add /absolute/path/to/agent-skills
codex plugin add orchestration@temmax --json
codex plugin add code-review@temmax --json

Usage

There are two ways to invoke the skills.

Automatic (primary). The skills trigger on their own: each skill's description is always in Claude's context, and a matching request loads the skill automatically — just ask in plain text:

Decompose this into agents and run in parallel: <task>
Orchestrate this task across subagents: <task>
Разбей на агентов и запусти параллельно: <задача>
Review my uncommitted changes critically
Сделай критическое ревью ПР #42

Explicit slash command. Guarantees the skill loads. Everything after the skill name is passed as the task description:

/orchestration:ship Add multi-currency support to the pricing module
/orchestration:super-plan Plan multi-currency support for the pricing module
/orchestration:multi-model Add multi-currency support to the pricing module
/code-review:critical-review <PR number optional>

ship runs the whole chain as one command: super-plan produces the plan file whose machine half feeds the wave-runner directly (each task: json entry

  • its prose section as the description), multi-model executes it in

supervised waves on a pushed feature branch, critical-review closes the loop on the PR — and the merge stays with the user. Each link also runs standalone.

Type /orch or /code and let autocomplete fill in the namespaced name.

What to expect from planning. super-plan researches the codebase and asks you only about genuine product forks — collected in one batch — and contradictions in the feature; with none, it asks nothing. It chooses the supervisors and the review models itself and reports the design without waiting for an answer. Then it stops once: one message with a summary of the design and of the finished plan — its shape (the waves, the tasks that run in parallel in each, the critical path in waves) and the routes, plus, inside ship, the branch and the pull request — and waits for your approval to start. The plan must pass the shipped linter (same-wave file overlap, contract completeness) before you ever see it. No message, table or report carries a time or cost estimate.

What to expect from orchestration. The orchestrator loads its profile, shows you a table (task | model | effort | rationale), then launches the waves through the shipped runner: every executor works in its own worktree, and an independent judge model checks out the branch, re-runs the contract's commands itself and issues a verdict — executor self-reports are never trusted. Rework, model escalation and the unsatisfiable-contract stop are code, not judgment calls. A Codex wave no longer stops when only the sandbox blocks a check: the runner runs that check outside the sandbox and the wave goes on; it stops for the environment only when the machine itself is broken. After each wave the orchestrator removes the merged tasks' worktrees and wave/<id> branches; the final report lists what is left under Left behind: together with the command to clean up after the merge. Once you confirm the merge, it removes the run records and deletes the feature branch, locally and on the remote — only what the run created, clean and already merged.

During fixes the agent decides by itself whether to fix directly or through agents, and your direct instruction ("fix it yourself", "use agents") wins; it does not ask you to approve a route or a model. It asks only about a contradiction in the feature, a change of scope, weakening or removing a test, an irreversible action on something the run did not create, replies in colleagues' thre

Source 1 files
hooks/runtime-context.mjs 63 lines
1// Claude Code 2.1.287+: lifecycle context, no UI or additional model calls.
2const plugin = 'orchestration';
3const policy = 'User-facing updates state the task, checks and results. Do not announce skills, instruction/profile filenames or loading, runtime model/effort or selection metadata.';
4const prefix = `PLUGIN_RUNTIME_CONTEXT_V1 plugin=${plugin} `;
5const modelPattern = /^claude-[a-zA-Z0-9._-]+(?:\[[a-zA-Z0-9]+\])?$/;
6
7function context(model) {
8  // Lifecycle events have no trustworthy effort. Existing skill fallback stays.
9  return `${prefix}host=claude model=${model} effort=unknown\n${policy}`;
10}
11function replaceOwnContext(result, model) {
12  return { ...result, additionalContext: [
13    ...(result.additionalContext ?? []).filter(text => text !== policy && !text.startsWith(prefix)),
14    context(model),
15  ] };
16}
17
18export function register(on) {
19  let deliveredModel;
20  on('classic.SessionStart', async ($, e, next) => {
21    deliveredModel = undefined;
22    const result = await next(e);
23    let model;
24    try { model = await $.session.model(); } catch { return result; }
25    if (typeof model !== 'string' || !modelPattern.test(model)) return result;
26    deliveredModel = model;
27    return replaceOwnContext(result, model);
28  });
29  on('prompt.submit', async ($, e, next) => {
30    let model;
31    try { model = await $.session.model(); } catch { return next(e); }
32    if (typeof model !== 'string' || !modelPattern.test(model) || model === deliveredModel) return next(e);
33    const result = await next({ ...e, context: [...(e.context ?? []), context(model)] });
34    if (result.context?.includes(context(model))) deliveredModel = model;
35    return result;
36  });
37  on('classic.SubagentStart', async ($, e, next) => {
38    const result = await next(e);
39    // session.model() describes main, not this child. Explicit unknown prevents
40    // a fork inheriting its parent's identity and loading the wrong profile.
41    return replaceOwnContext(result, 'unknown');
42  });
43  on('prompt.context', async ($, e, next) => {
44    // Another mod rewrote the memory text: its source files are unknown.
45    if (e.instructionFiles === undefined) return next(e);
46    let ancestors;
47    try { ancestors = await $.fs.ancestors({ names: ['AGENTS.md'] }); }
48    catch { return next(e); } // Existing skill discovery remains the fallback.
49    const paths = new Set(e.instructionFiles.map(file => file.path));
50    const added = [];
51    for (const ancestor of ancestors) {
52      for (const part of ancestor.parts) {
53        if (paths.has(part.path)) continue;
54        paths.add(part.path);
55        // The loader flattens imports without exposing their direct importer.
56        added.push({ path: part.path, kind: 'project', content: part.content });
57      }
58    }
59    if (!added.length) return next(e);
60    return next({ ...e, instructionFiles: [...e.instructionFiles, ...added] });
61  });
62}
63