SLOPSHOPPER

claude-image-generation

Generate and edit images using Google Gemini, OpenAI GPT Image, xAI Grok Image, and OpenRouter APIs

newguardtoolprocess
★ 10v2026.9.1MITupdated 2026-09-25hex/claude-image-generation
A shopper browsing a rack in a slop shop
README

claude-image-generation

Claude Code plugin for generating and editing images using Google Gemini, OpenAI GPT Image, xAI Grok Image, and OpenRouter APIs.

Features

  • Text-to-image generation with Google Gemini, OpenAI GPT Image 2, xAI Grok Image, or any image model on OpenRouter
  • OpenRouter gateway — reach any OpenRouter image model (Gemini, GPT Image, and more) through one key via the chat-completions API
  • Image editing with text instructions (all providers)
  • Multi-image input — repeatable --input-image for multi-image edits (all providers) and Gemini reference-based generation
  • Parallel generation across all providers via scripts/run-all.sh — one shared streaming pane, council-style colored banners, and a waiting on line naming the pending providers (animated until the first image renders, then written once under each block)
  • Interactive provider selection via AskUserQuestion at runtime
  • Inline image preview -- generated images display directly in the terminal (iTerm2, Kitty, Ghostty, WezTerm, Sixel terminals)
  • Tmux pane display -- opens a split pane for image preview when running inside tmux (works with Claude Code). Providers running at the same time share one pane, however they were launched
  • Streaming display -- images appear progressively in a shared pane during parallel generation, accumulating as each provider finishes
  • Open in Finder/Preview -- press 'f' for Finder or 'p' for Preview in the display pane

Installation

From marketplace (recommended)

# Add the hex-plugins marketplace (once)
/plugin marketplace add hex/claude-marketplace

# Install the plugin
/plugin install claude-image-generation

From GitHub

/plugin install hex/claude-image-generation

Manual

git clone https://github.com/hex/claude-image-generation.git
claude --plugin-dir /path/to/claude-image-generation

Configuration

API Keys

Set any of these as environment variables:

VariableProviderGet a key
GEMINI_API_KEYGoogle GeminiGoogle AI Studio
OPENAI_API_KEYOpenAIOpenAI Platform
XAI_API_KEY or GROK_API_KEYxAIxAI Console
OPENROUTER_API_KEYOpenRouterOpenRouter Keys

At least one key is required.

Model Selection

Override the default model per provider via environment variables:

VariableDefaultPurpose
GEMINI_IMAGE_MODELgemini-3-pro-imageGemini model used for generation and editing
OPENAI_IMAGE_MODELgpt-image-2OpenAI model used for generation and editing
XAI_IMAGE_MODELgrok-imagine-image-2.0xAI model used for generation and editing
OPENROUTER_IMAGE_MODELgoogle/gemini-3.1-flash-imageOpenRouter model slug used for generation and editing

Command-line --model flag on the scripts takes precedence over environment variables.

Display Size

Control the terminal image display dimensions (in pixels):

VariableDefaultPurpose
DISPLAY_IMAGE_WIDTH512Max image width in pixels for terminal display
DISPLAY_IMAGE_HEIGHT512Max image height in pixels for iTerm2 display

These apply to inline display (iTerm2, Sixel) and tmux pane display.

Available Gemini Models

ModelCharacteristics
gemini-3-pro-imagePro tier, premium quality, 10 aspect ratios, up to 14 reference images (default, "Nano Banana Pro", GA since 2026-05-28)
gemini-3.1-flash-image14 aspect ratios (incl. extreme 1:4, 8:1), 512-4K resolution, thinking, Google Search grounding ("Nano Banana 2", GA since 2026-05-28)
gemini-3.1-flash-lite-imageCheapest tier ("Nano Banana 2 Lite")
gemini-2.5-flash-imagePrevious generation, 1K only (scheduled shutdown 2026-10-02)

The -preview IDs of the two GA models still answer but passed Google's earliest shutdown date (2026-06-25) and are gone from its model tables; pass the GA IDs.

Available OpenAI Models

ModelCharacteristics
gpt-image-2Latest flagship, snapshot gpt-image-2-2026-04-21 (default)
gpt-image-1.5Previous flagship, superior text rendering, transparent backgrounds, quality tiers (shutdown 2026-12-01)
gpt-image-1-mini3-4x cheaper, cost-efficient for drafts and previews (shutdown 2026-12-01)
gpt-image-1Older generation (shutdown 2026-10-23)

Available xAI Models

ModelCharacteristics
grok-imagine-image-2.0Flagship since 2026-08-07: --quality low/medium/auto, up to 5 reference images, 21:9 and 5:2 ratios (default)
grok-imagine-image-qualityQuality mode from 2026-05-06; grok-imagine-image-pro has redirected here since its 2026-05-15 retirement
grok-imagine-imageStandard tier, 1K/2K resolution, 300 RPM, same endpoint and parameters

Available OpenRouter Models

OpenRouter is a gateway, so --model (or OPENROUTER_IMAGE_MODEL) accepts any OpenRouter slug that supports image output. A few:

ModelCharacteristics
google/gemini-3.1-flash-imageFast Gemini image model, generation + editing (default)
google/gemini-3-pro-imagePro-tier Gemini image model, premium quality
x-ai/grok-imagine-image-2.0xAI's flagship image model via OpenRouter
openai/gpt-image-2OpenAI's flagship image model via OpenRouter (openai/gpt-5-image and openai/gpt-5-image-mini are the older chat-image models)

Browse the full list at openrouter.ai/models.

Usage

Slash Command

/generate-image a golden retriever in a field of sunflowers
/generate-image --edit ./photo.png remove the background and make it transparent

Without the generate tool (see Generate Tool under Usage), the command prompts you to select a provider (Gemini, OpenAI, xAI, OpenRouter, or all in parallel) and an output path.

Agent (Automatic)

The image-generator agent triggers automatically when conversation context involves image creation. It handles provider selection, parallel generation, and result delivery without requiring the slash command.

Generate Tool (function hooks, early access)

On Claude Code builds with function hooks, the plugin also registers a tool, mcp__claude-image-generation__generate. The model calls it with a typed input instead of writing a bash scripts/run-all.sh ... line:

FieldRequiredMeaning
promptyesWhat to draw, or how to change the input images
providersnoAny of gemini, openai, xai, openrouter; omitted, the Default providers setting
outputBasenoPath without extension; each provider saves <outputBase>-<provider>.png. Omitted, <Output directory>/image-<timestamp>
inputImagesnoImages to edit; any entry switches to edit mode
aspectRationoWhole-number W:H such as 16:9; passed to gemini and xai only

The tool runs scripts/run-all.sh, so the streaming pane and the retry offer work as they do for the slash command. It answers with a saved or missing line per expected file, the exit code, and whatever the providers printed. When no provider saved a file it refuses the call, with the same text. It also refuses a call that fails a field check before anything runs: a blank prompt, a field of the wrong type, an unknown provider, an outputBase ending in .png, .jpg, .jpeg or .webp (any case), or an aspect ratio that is not whole-number W:H (so xAI's auto, 19.5:9 and 9:19.5 need the scripts).

A call runs for at most ten minutes, the most the engine allows a process. That covers the slowest provider plus the 45-second retry offer.

Outside tmux the tool shows no inline preview: it sends the providers' terminal image output to /dev/null (through DISPLAY_IMAGE_TARGET), so the images do not land on Claude Code's own screen. Inside tmux they stream into the pane as usual.

Two rows in /config set its defaults:

SettingValuesDefault
Default providersall (gemini, openai, xai), gemini, openai, xai, openrouterall
Output directoryany path, relative to the session's working directory.

The slash command, the agent and the skill use the tool whenever the session offers it, and skip the provider and path questions. They fall back to running the scripts through Bash on builds without it, or when a request needs an option the tool does not take (image size, quality, transparent background, a specific model, or another per-provider flag).

Direct Script Usage

Scripts are located in scripts/ and can be invoked directly.

gemini.sh
# Generate
bash scripts/gemini.sh \
  --mode generate \
  --prompt "a mountain at sunset" \
  --output ./mountain.png

# Generate with aspect ratio
bash scripts/gemini.sh \
  --mode generate \
  --prompt "a wide landscape" \
  --output ./landscape.png \
  --aspect-ratio 16:9

# Edit
bash scripts/gemini.sh \
  --mode edit \
  --prompt "add snow to the peaks" \
  --input-image ./mountain.png \
  --output ./snowy.png

# Generate at 4K with thinking mode
bash scripts/gemini.sh \
  --mode generate \
  --prompt "a detailed sci-fi cityscape" \
  --output ./city.png \
  --image-size 4K \
  --thinking-level High

# Generate with Google Search grounding
bash scripts/gemini.sh \
  --mode generate \
  --prompt "Search for the latest SpaceX Starship and draw it at sunset on the launch pad" \
  --output ./starship.png \
  --search-grounding

# Use a specific model
bash scripts/gemini.sh \
  --mode generate \
  --prompt "quick sketch" \
  --output ./sketch.png \
  --model gemini-3-pro-image

Flags:

FlagValuesDefaultRequired
--modegenerate, edit--Yes
--prompttext--Yes
--outputfile path--Yes
--input-imagefile path, repeatable (max 14)--Edit mode; optional in generate mode as references
--aspect-ratio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9 on Pro (default); add 1:4, 4:1, 1:8, 8:1 on gemini-3.1-flash-image1:1No
--image-size512, 1K, 2K, 4K (UPPERCASE); 512 requires gemini-3.1-flash-image(API default 1K)No
--thinking-levelminimal, Highunset (API default minimal)No
--image-only(flag, no value)offNo
--search-grounding(flag, no value)offNo
--modelGemini model namegemini-3-pro-imageNo
openai.sh
# Generate
bash scripts/openai.sh \
  --mode generate \
  --prompt "a mountain at sunset" \
  --output ./mountain.png

# Generate with options
bash scripts/openai.sh \
  --mode generate \
  --prompt "company logo on transparent background" \
  --output ./logo.png \
  --size 1024x1024 \
  --quality high \
  --background transparent

# Edit
bash scripts/openai.sh \
  --mode edit \
  --prompt "add snow to the peaks" \
  --input-image ./mountain.png \
  --output ./snowy.png

Flags:

FlagValuesDefaultRequired
--modegenerate, edit--Yes
--prompttext--Yes
--outputfile path--Yes
--input-imagefile path, repeatable (max 16; dall-e-2 allows 1)--Edit mode only
--sizeauto or WxH. On gpt-image-2, any size with both edges multiples of 16, longest edge up to 3840, ratio at most 3:1 and 655,360 to 8,294,400 total pixels (such as 2048x1152, 3840x2160); older models take 1024x1024, 1536x1024, 1024x15361024x1024No
--qualityauto, low, medium, highhighNo
--backgroundauto, transparent, opaqueautoNo
--output-formatpng, jpeg, webppngNo
--output-compressioninteger 0-100 (jpeg/webp only)--No
--moderationauto, lowautoNo
--input-fidelitylow, high (edit only); not accepted with gpt-image-2 (that model always uses high fidelity; openai.sh refuses the flag)unset (API default low)No
--modelOpenAI model namegpt-image-2No
xai.sh
# Generate
bash scripts/xai.sh \
  --mode generate \
  --prompt "a mountain at sunset" \
  --output ./mountain.png

# Generate with aspect ratio
bash scripts/xai.sh \
  --mode generate \
  --prompt "a wide landscape" \
  --output ./landscape.png \
  --aspect-ratio 16:9

# Edit
bash scripts/xai.sh \
  --mode edit \
  --prompt "add snow to the peaks" \
  --input-image ./mountain.png \
  --output ./snowy.png

# Generate at 2K resolution
bash scripts/xai.sh \
  --mode generate \
  --prompt "a cat in a tree" \
  --output ./cat.png \
  --resolution 2k

# Use the May 2026 quality-mode model instead of the 2.0 default
bash scripts/xai.sh \
  --mode generate \
  --prompt "a cat in a tree" \
  --output ./cat.png \
  --model grok-imagine-image-quality

Flags:

FlagValuesDefaultRequired
--modegenerate, edit--Yes
--prompttext--Yes
--outputfile path--Yes
--input-imagefile path, repeatable (max 5 on grok-imagine-image-2.0; older models take 3)--Edit mode only
--aspect-ratio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 2:1, 1:2, 19.5:9, 9:19.5, 20:9, 9:20, 21:9, 5:2, auto(none)No
--resolution1k, 2k (LOWERCASE)(API default)No
--qualitylow, medium, auto (grok-imagine-image-2.0 only)unset (API auto: low for generation, medium for edits; billed as served)No
--modelxAI model namegrok-imagine-image-2.0No

Note: For single-image edits, xAI ignores --aspect-ratio and uses the input image's ratio. Multi-image edits allow aspect ratio override (up to 5 images).

openrouter.sh

OpenRouter is a gateway to many image models through a single key. It uses the chat-completions API, so --model accepts any OpenRouter model slug that supports image output.

# Generate (default model: google/gemini-3.1-flash-image)
bash scripts/openrouter.sh \
  --mode generate \
  --prompt "a mountain at sunset" \
  --output ./mountain.png

# Generate with a specific model
bash scripts/openrouter.sh \
  --mode generate \
  --prompt "a cat in a tree" \
  --output ./cat.png \
  --model openai/gpt-image-2

# Edit (single or multiple --input-image)
bash scripts/openrouter.sh \
  --mode edit \
  --prompt "add snow to the peaks" \
  --input-image ./mountain.png \
  --output ./snowy.png

Flags:

FlagValuesDefaultRequired
--modegenerate, edit--Yes
--prompttext--Yes
--outputfile path--Yes
--input-imagefile path, repeatable--Edit mode only
--modelany OpenRouter image model sluggoogle/gemini-3.1-flash-imageNo
--site-urlURL(none)No (sent as HTTP-Referer for OpenRouter attribution)
--site-nametext(none)No (sent as X-Title for OpenRouter attribution)

--site-url / --site-name also default from OPENROUTER_SITE_URL / OPENROUTER_SITE_NAME.

Reference Images and Multi-Image Composition

--input-image is repeatable on all four scripts. Passing more images than a provider supports exits with code 1 before any API call:

ProviderMax imagesModes
Gemini14generate (references for a fresh composition) and edit
OpenAI16edit only (its generation endpoint takes no images)
xAI5 on grok-imagine-image-2.0 (older models take 3)edit only
OpenRoutermodel-dependentedit only (input images attached as chat image parts)
# Gemini: compose a new image from reference images (generate mode)
bash scripts/gemini.sh \
  --mode generate \
  --prompt "a product shot combining the chair from the first image with the fabric of the second" \
  --input-image ./chair.png \
  --input-image ./fabric.png \
  --output ./composite.png

# OpenAI: multi-image edit
bash scripts/openai.sh \
  --mode edit \
  --prompt "place the logo from the second image onto the mug in the first" \
  --input-image ./mug.png \
  --input-image ./logo.png \
  --output ./branded.png

# xAI: multi-image edit
bash scripts/xai.sh \
  --mode edit \
  --prompt "blend both scenes into one panorama" \
  --input-image ./left.png \
  --input-image ./right.png \
  --output ./panorama.png

# All providers in parallel (edit mode only)
bash scripts/run-all.sh \
  --mode edit \
  --prompt "combine these" \
  --input-image ./ref-a.png \
  --input-image ./ref-b.png \
  --output-base ./combined

Gemini's flat 14-image budget is best composed as up to 6 object + 5 character-consistency + 3 style-reference images. There is no API field to tag an image's role — the model infers it from the prompt, so state which images are objects, characters, or style references.

Notes:

  • run-all.sh forwards every --input-image to each selected provider, but only in --mode edit. Gemini's generate-mode reference images are not forwarded through run-all — call scripts/gemini.sh directly for generate-with-references.
  • dall-e-2 edits are a known limitation: the script rejects multiple images for dall-e-2, but single-image dall-e-2 edits also do not work — the script sends form fields only the gpt-image models accept.

Retries

Each provider script retries a transient API error (429 or a 5xx status, or a network failure) up to three times, with a delay that doubles each attempt (1 second, 2 seconds, 4 seconds). IMAGE_MAX_RETRIES and IMAGE_RETRY_DELAY tune the count and the starting delay. Inside a streaming pane, a retry shows on the waiting on line as (retry 2/3); once an image is on the pane that line is only written when the next block lands, so a retry that starts between blocks shows only in the final banner or error text.

When a provider run through run-all.sh fails outright, its error shows under a red ✗ provider error heading, and if any provider failed the pane offers [r] retry failed (xai) · [esc/ctrl-d] close 45s. The offer counts down until the first image is on the pane; after that it shows the time budget once, as [r] retry failed (xai) · [esc/ctrl-d] close (up to 45s), because a rewritten line erases the pane's images. Pressing r re-runs only the failed providers inside the same run; Esc closes the pane and the run returns with what it has. DISPLAY_PANE_RETRY_WAIT sets how long the offer stays open (default 45 seconds; 0 disables it), and a run with a failure can take that much longer plus one more provider round.

Provider Comparison

FeatureGeminiOpenAIxAIOpenRouter
Default modelgemini-3-pro-imagegpt-image-2grok-imagine-image-2.0google/gemini-3.1-flash-image
Max resolution4K (via --image-size)3840 px long edge on gpt-image-2 (via --size); 1536x1024 on older models2K (via --resolution)Model-dependent
Text renderingVery good (under 25 chars)ExcellentGoodModel-dependent
Transparent BGNoYes (preview on gpt-image-2, png or webp only)NoModel-dependent
Aspect ratios10 on Pro / 14 on 3.1 FlashAny WxH up to 3:1 on gpt-image-2; 3 fixed sizes on older models16 options (incl. 21:9, 5:2, auto)Prompt-driven
Image editingMulti-turn, up to 14 refs (generate + edit)Up to 16 input images/v1/images/edits, up to 5 imagesChat image parts (edit)
Quality tiersN/Aauto / low / medium / highlow / medium / auto (2.0 only)Model-dependent
Thinking modeYes (--thinking-level)NoNoModel-dependent
Search groundingYes (Google Search)NoNoModel-dependent
PricingToken-basedToken-basedFlat per-imagePer OpenRouter model
Prompt revisionNoNoYes (by chat model)No

Plugin Components

ComponentFilePurpose
Plugin manifest.claude-plugin/plugin.jsonPlugin metadata, version and the /config rows (defaultProviders, outputDir)
Skillskills/image-generation/SKILL.mdAPI knowledge, prompting tips, script reference
Commandcommands/generate-image.md/generate-image slash command
Agentagents/image-generator.mdAutonomous image generation
Gemini scriptscripts/gemini.shGemini API call execution
OpenAI scriptscripts/openai.shOpenAI API call execution
xAI scriptscripts/xai.shxAI API call execution
OpenRouter scriptscripts/openrouter.shOpenRouter chat-completions image call execution
Retry helperscripts/retry.shcurl_with_retry for transient API failures
Parallel runnerscripts/run-all.shForks all providers in parallel under one streaming pane; holds a pane token for the batch
Display utilityscripts/display.shMulti-protocol terminal image display (iTerm2, Kitty, Sixel, tmux pane, shared streaming pane with colored banners + pending-provider waiting line)
API referenceskills/image-generation/references/api-details.mdEndpoint and payload documentation
Skill evalsskills/image-generation/evals/evals.jsonTrigger prompts for the skill's eval suite
Hooks modulehooks/index.tsRegisters the generate tool and serves its calls through run-all.sh
Tool input ruleshooks/argv.tsChecks a generate call's input and builds run-all.sh's argv
Hooks manifesthooks/hooks.jsonPoints the engine at hooks/index.ts
Tool input typehooks/generate-input.d.tsTypes the generate tool's input for the tool.call hook
TypeScript configtsconfig.jsonType-checks hooks/ and tests/ against .claude/types
Automated teststests/bats test suite for all scripts; argv.test.ts for the tool's input rules

Development

Versioning

This plugin uses calendar versioning in YYYY.M.PATCH format (e.g., 2026.7.1). The version is tracked in both .claude-plugin/plugin.json and skills/image-generation/SKILL.md.

Testing

# Run all automated tests (requires bats)
./tests/run_tests.sh

# Or run bats directly
bats tests/

The hooks module has its own checks. Run /plugin-types in a Claude Code session in this directory once, to write .claude/types, then:

tsc -p .                 # type-check hooks/ and tests/argv.test.ts
claude plugin test .     # run tests/argv.test.ts
claude plugin validate . # what the engine sees: the module's hooks and the calls it makes

See TESTING.md for the full testing guide, including manual test procedures.

Architecture

The plugin is organized into Claude Code extension points:

.claude-plugin/plugin.json    -- Plugin identity and metadata
commands/                      -- Slash command definitions
agents/                        -- Autonomous agent definitions
skills/                        -- Skill knowledge and references
scripts/                       -- Shell scripts for API calls
hooks/                         -- Function hooks module registering the generate tool
tests/                         -- Automated tests (bats, argv.test.ts)

The scripts (gemini.sh, openai.sh, xai.sh, openrouter.sh) are standalone bash programs that handle API communication, base64 encoding/decoding, and error reporting. They are invoked by the command, agent, and skill layers.

Source 2 files
hooks/index.ts 93 lines
1// ABOUTME: Registers the generate tool, which runs scripts/run-all.sh from a typed input
2// ABOUTME: and answers with the files it saved; the defaults come from the plugin's /config rows.
3
4import type { Register } from 'claude-code'
5
6import { buildRun } from './argv'
7
8const TOOL = 'mcp__claude-image-generation__generate'
9// run-all.sh can wait on a slow provider for minutes and then hold a 45 s retry offer open;
10// ten minutes is the most $.process.run allows.
11const RUN_TIMEOUT_MS = 600_000
12
13const DESCRIPTION = `Generate or edit images with Gemini, OpenAI, xAI and OpenRouter in parallel. \
14Each provider saves <outputBase>-<provider>.png; the result lists the files that exist. \
15Omit providers and outputBase to use the person's /config defaults. \
16Providers: gemini (gemini-3-pro-image, professional assets, many reference images), \
17openai (gpt-image-2, text rendering, transparent backgrounds), \
18xai (grok-imagine-image-2.0, prompt revision, flat per-image pricing), \
19openrouter (google/gemini-3.1-flash-image by default, opt-in). \
20Pass inputImages to edit instead of generate. aspectRatio (W:H) applies to gemini and xai only.`
21
22const INPUT_SCHEMA = {
23  type: 'object',
24  properties: {
25    prompt: { type: 'string', description: 'What to draw, or how to change the input images.' },
26    providers: {
27      type: 'array',
28      items: { type: 'string', enum: ['gemini', 'openai', 'xai', 'openrouter'] },
29      description: 'Which providers to run; omitted, the configured default.',
30    },
31    outputBase: {
32      type: 'string',
33      description: 'Path without extension, e.g. "art/hero"; omitted, <configured dir>/image-<timestamp>.',
34    },
35    inputImages: { type: 'array', items: { type: 'string' }, description: 'Images to edit; switches to edit mode.' },
36    aspectRatio: { type: 'string', description: 'W:H such as 16:9; gemini and xai only.' },
37  },
38  required: ['prompt'],
39  additionalProperties: false,
40}
41
42export const register: Register = (on, options) => {
43  const defaults = {
44    defaultProviders: stringOption(options, 'defaultProviders'),
45    outputDir: stringOption(options, 'outputDir'),
46  }
47
48  on('session.start', async ($, e, next) => {
49    await $.tool.register({ name: 'generate', description: DESCRIPTION, inputSchema: INPUT_SCHEMA })
50    return next(e)
51  })
52
53  on('tool.call', { tool: TOOL }, async ($, e) => {
54    const run = buildRun(e, defaults, $.plugin.root, timestamp(await $.clock.now()))
55    if ('error' in run) return { deny: run.error }
56
57    // Outside a tmux pane the providers draw inline images to /dev/tty, which is Claude Code's own screen.
58    const { exitCode, stdout, stderr } = await $.process.run(run.argv, {
59      timeoutMs: RUN_TIMEOUT_MS,
60      env: { DISPLAY_IMAGE_TARGET: '/dev/null' },
61    })
62
63    const saved: string[] = []
64    const missing: string[] = []
65    for (const path of run.outputs) {
66      ;((await $.fs.exists(path)) ? saved : missing).push(path)
67    }
68    const lines = [
69      ...saved.map(path => `saved ${path}`),
70      ...missing.map(path => `missing ${path}`),
71      `run-all.sh exited ${exitCode}`,
72    ]
73    if (stdout.trim() !== '') lines.push('', stdout.trim())
74    if (stderr.trim() !== '') lines.push('', 'stderr:', stderr.trim())
75    const text = lines.join('\n')
76    if (saved.length === 0) return { deny: text }
77    return { result: text }
78  })
79}
80
81// plugin.json declares both fields as strings with defaults, so the engine always fills them in.
82function stringOption(options: Parameters<Register>[1], field: string): string {
83  const value = options[field]
84  if (typeof value !== 'string') throw new Error(`option ${field} should be a string from plugin.json userConfig, got ${JSON.stringify(value)}`)
85  return value
86}
87
88function timestamp(ms: number): string {
89  const d = new Date(ms)
90  const pad = (n: number) => String(n).padStart(2, '0')
91  return `${d.getFullYear()}${pad(d.getMonth() + 1)}${pad(d.getDate())}-${pad(d.getHours())}${pad(d.getMinutes())}${pad(d.getSeconds())}`
92}
93
hooks/argv.ts 86 lines
1// ABOUTME: Turns a generate tool call plus the plugin's options into the argv for run-all.sh,
2// ABOUTME: or the reason the call is refused. Pure, so every rule is testable without a run.
3
4export type GenerateOptions = {
5  defaultProviders: string
6  outputDir: string
7}
8
9export type Run = { argv: string[]; outputs: string[] } | { error: string }
10
11type GenerateInput = {
12  prompt: string
13  providers: string[] | undefined
14  outputBase: string | undefined
15  inputImages: string[]
16  aspectRatio: string | undefined
17}
18
19// What run-all.sh runs when no provider is named; openrouter is opt-in there too.
20const ALL = ['gemini', 'openai', 'xai']
21const PROVIDERS = [...ALL, 'openrouter']
22// The providers whose scripts take --aspect-ratio; openai sizes by --size, openrouter by model.
23const TAKES_ASPECT_RATIO = ['gemini', 'xai']
24
25export function buildRun(raw: Record<string, unknown>, options: GenerateOptions, root: string, stamp: string): Run {
26  const input = parseInput(raw)
27  if ('error' in input) return input
28
29  const providers = input.providers?.length
30    ? input.providers
31    : options.defaultProviders === 'all' ? ALL : [options.defaultProviders]
32  const unknown = providers.find(p => !PROVIDERS.includes(p))
33  if (unknown !== undefined) return { error: `unknown provider "${unknown}"; choose from ${PROVIDERS.join(', ')}` }
34  const ratio = input.aspectRatio
35  if (ratio !== undefined && !/^\d+:\d+$/.test(ratio)) return { error: `aspectRatio must look like 16:9, got "${ratio}"` }
36  const base = input.outputBase ?? `${options.outputDir}/image-${stamp}`
37  if (/\.(png|jpe?g|webp)$/i.test(base)) {
38    return { error: `outputBase is a path without an extension; "${base}" would save as "${base}-${providers[0]}.png"` }
39  }
40  const images = input.inputImages
41  const argv = [
42    'bash', `${root}/scripts/run-all.sh`,
43    '--mode', images.length > 0 ? 'edit' : 'generate',
44    '--prompt', input.prompt,
45    '--output-base', base,
46    '--providers', providers.join(','),
47    ...images.flatMap(image => ['--input-image', image]),
48    ...(ratio === undefined ? [] : providers
49      .filter(p => TAKES_ASPECT_RATIO.includes(p))
50      .flatMap(p => [`--${p}-extra`, `--aspect-ratio ${ratio}`])),
51  ]
52  return { argv, outputs: providers.map(p => `${base}-${p}.png`) }
53}
54
55// The model writes the call's input, so every field is checked for its type here, once.
56function parseInput(raw: Record<string, unknown>): GenerateInput | { error: string } {
57  const prompt = raw.prompt
58  if (typeof prompt !== 'string' || prompt.trim() === '') return { error: 'prompt is required' }
59  const providers = stringList(raw, 'providers')
60  if (isError(providers)) return providers
61  const inputImages = stringList(raw, 'inputImages')
62  if (isError(inputImages)) return inputImages
63  const outputBase = optionalString(raw, 'outputBase')
64  if (isError(outputBase)) return outputBase
65  const aspectRatio = optionalString(raw, 'aspectRatio')
66  if (isError(aspectRatio)) return aspectRatio
67  return { prompt, providers, inputImages: inputImages ?? [], outputBase, aspectRatio }
68}
69
70function isError<T>(value: T | { error: string }): value is { error: string } {
71  return typeof value === 'object' && value !== null && 'error' in value
72}
73
74function stringList(raw: Record<string, unknown>, field: string): string[] | undefined | { error: string } {
75  const value = raw[field]
76  if (value === undefined) return undefined
77  if (Array.isArray(value) && value.every(v => typeof v === 'string')) return value
78  return { error: `${field} must be a list of strings, got ${JSON.stringify(value)}` }
79}
80
81function optionalString(raw: Record<string, unknown>, field: string): string | undefined | { error: string } {
82  const value = raw[field]
83  if (value === undefined || typeof value === 'string') return value
84  return { error: `${field} must be a string, got ${JSON.stringify(value)}` }
85}
86