Generate and edit images using Google Gemini, OpenAI GPT Image, xAI Grok Image, and OpenRouter APIs

Claude Code plugin for generating and editing images using Google Gemini, OpenAI GPT Image, xAI Grok Image, and OpenRouter APIs.
--input-image for multi-image edits (all providers) and Gemini reference-based generationscripts/run-all.sh — one shared streaming pane, council-style colored banners, and a waiting on line naming the pending providers (animated until the first image renders, then written once under each block)# Add the hex-plugins marketplace (once)
/plugin marketplace add hex/claude-marketplace
# Install the plugin
/plugin install claude-image-generation
/plugin install hex/claude-image-generation
git clone https://github.com/hex/claude-image-generation.git
claude --plugin-dir /path/to/claude-image-generation
Set any of these as environment variables:
| Variable | Provider | Get a key |
|---|---|---|
GEMINI_API_KEY | Google Gemini | Google AI Studio |
OPENAI_API_KEY | OpenAI | OpenAI Platform |
XAI_API_KEY or GROK_API_KEY | xAI | xAI Console |
OPENROUTER_API_KEY | OpenRouter | OpenRouter Keys |
At least one key is required.
Override the default model per provider via environment variables:
| Variable | Default | Purpose |
|---|---|---|
GEMINI_IMAGE_MODEL | gemini-3-pro-image | Gemini model used for generation and editing |
OPENAI_IMAGE_MODEL | gpt-image-2 | OpenAI model used for generation and editing |
XAI_IMAGE_MODEL | grok-imagine-image-2.0 | xAI model used for generation and editing |
OPENROUTER_IMAGE_MODEL | google/gemini-3.1-flash-image | OpenRouter model slug used for generation and editing |
Command-line --model flag on the scripts takes precedence over environment variables.
Control the terminal image display dimensions (in pixels):
| Variable | Default | Purpose |
|---|---|---|
DISPLAY_IMAGE_WIDTH | 512 | Max image width in pixels for terminal display |
DISPLAY_IMAGE_HEIGHT | 512 | Max image height in pixels for iTerm2 display |
These apply to inline display (iTerm2, Sixel) and tmux pane display.
| Model | Characteristics |
|---|---|
gemini-3-pro-image | Pro tier, premium quality, 10 aspect ratios, up to 14 reference images (default, "Nano Banana Pro", GA since 2026-05-28) |
gemini-3.1-flash-image | 14 aspect ratios (incl. extreme 1:4, 8:1), 512-4K resolution, thinking, Google Search grounding ("Nano Banana 2", GA since 2026-05-28) |
gemini-3.1-flash-lite-image | Cheapest tier ("Nano Banana 2 Lite") |
gemini-2.5-flash-image | Previous generation, 1K only (scheduled shutdown 2026-10-02) |
The -preview IDs of the two GA models still answer but passed Google's earliest shutdown date (2026-06-25) and are gone from its model tables; pass the GA IDs.
| Model | Characteristics |
|---|---|
gpt-image-2 | Latest flagship, snapshot gpt-image-2-2026-04-21 (default) |
gpt-image-1.5 | Previous flagship, superior text rendering, transparent backgrounds, quality tiers (shutdown 2026-12-01) |
gpt-image-1-mini | 3-4x cheaper, cost-efficient for drafts and previews (shutdown 2026-12-01) |
gpt-image-1 | Older generation (shutdown 2026-10-23) |
| Model | Characteristics |
|---|---|
grok-imagine-image-2.0 | Flagship since 2026-08-07: --quality low/medium/auto, up to 5 reference images, 21:9 and 5:2 ratios (default) |
grok-imagine-image-quality | Quality mode from 2026-05-06; grok-imagine-image-pro has redirected here since its 2026-05-15 retirement |
grok-imagine-image | Standard tier, 1K/2K resolution, 300 RPM, same endpoint and parameters |
OpenRouter is a gateway, so --model (or OPENROUTER_IMAGE_MODEL) accepts any OpenRouter slug that supports image output. A few:
| Model | Characteristics |
|---|---|
google/gemini-3.1-flash-image | Fast Gemini image model, generation + editing (default) |
google/gemini-3-pro-image | Pro-tier Gemini image model, premium quality |
x-ai/grok-imagine-image-2.0 | xAI's flagship image model via OpenRouter |
openai/gpt-image-2 | OpenAI's flagship image model via OpenRouter (openai/gpt-5-image and openai/gpt-5-image-mini are the older chat-image models) |
Browse the full list at openrouter.ai/models.
/generate-image a golden retriever in a field of sunflowers
/generate-image --edit ./photo.png remove the background and make it transparent
Without the generate tool (see Generate Tool under Usage), the command prompts you to select a provider (Gemini, OpenAI, xAI, OpenRouter, or all in parallel) and an output path.
The image-generator agent triggers automatically when conversation context involves image creation. It handles provider selection, parallel generation, and result delivery without requiring the slash command.
On Claude Code builds with function hooks, the plugin also registers a tool, mcp__claude-image-generation__generate. The model calls it with a typed input instead of writing a bash scripts/run-all.sh ... line:
| Field | Required | Meaning |
|---|---|---|
prompt | yes | What to draw, or how to change the input images |
providers | no | Any of gemini, openai, xai, openrouter; omitted, the Default providers setting |
outputBase | no | Path without extension; each provider saves <outputBase>-<provider>.png. Omitted, <Output directory>/image-<timestamp> |
inputImages | no | Images to edit; any entry switches to edit mode |
aspectRatio | no | Whole-number W:H such as 16:9; passed to gemini and xai only |
The tool runs scripts/run-all.sh, so the streaming pane and the retry offer work as they do for the slash command. It answers with a saved or missing line per expected file, the exit code, and whatever the providers printed. When no provider saved a file it refuses the call, with the same text. It also refuses a call that fails a field check before anything runs: a blank prompt, a field of the wrong type, an unknown provider, an outputBase ending in .png, .jpg, .jpeg or .webp (any case), or an aspect ratio that is not whole-number W:H (so xAI's auto, 19.5:9 and 9:19.5 need the scripts).
A call runs for at most ten minutes, the most the engine allows a process. That covers the slowest provider plus the 45-second retry offer.
Outside tmux the tool shows no inline preview: it sends the providers' terminal image output to /dev/null (through DISPLAY_IMAGE_TARGET), so the images do not land on Claude Code's own screen. Inside tmux they stream into the pane as usual.
Two rows in /config set its defaults:
| Setting | Values | Default |
|---|---|---|
| Default providers | all (gemini, openai, xai), gemini, openai, xai, openrouter | all |
| Output directory | any path, relative to the session's working directory | . |
The slash command, the agent and the skill use the tool whenever the session offers it, and skip the provider and path questions. They fall back to running the scripts through Bash on builds without it, or when a request needs an option the tool does not take (image size, quality, transparent background, a specific model, or another per-provider flag).
Scripts are located in scripts/ and can be invoked directly.
# Generate
bash scripts/gemini.sh \
--mode generate \
--prompt "a mountain at sunset" \
--output ./mountain.png
# Generate with aspect ratio
bash scripts/gemini.sh \
--mode generate \
--prompt "a wide landscape" \
--output ./landscape.png \
--aspect-ratio 16:9
# Edit
bash scripts/gemini.sh \
--mode edit \
--prompt "add snow to the peaks" \
--input-image ./mountain.png \
--output ./snowy.png
# Generate at 4K with thinking mode
bash scripts/gemini.sh \
--mode generate \
--prompt "a detailed sci-fi cityscape" \
--output ./city.png \
--image-size 4K \
--thinking-level High
# Generate with Google Search grounding
bash scripts/gemini.sh \
--mode generate \
--prompt "Search for the latest SpaceX Starship and draw it at sunset on the launch pad" \
--output ./starship.png \
--search-grounding
# Use a specific model
bash scripts/gemini.sh \
--mode generate \
--prompt "quick sketch" \
--output ./sketch.png \
--model gemini-3-pro-image
Flags:
| Flag | Values | Default | Required |
|---|---|---|---|
--mode | generate, edit | -- | Yes |
--prompt | text | -- | Yes |
--output | file path | -- | Yes |
--input-image | file path, repeatable (max 14) | -- | Edit mode; optional in generate mode as references |
--aspect-ratio | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9 on Pro (default); add 1:4, 4:1, 1:8, 8:1 on gemini-3.1-flash-image | 1:1 | No |
--image-size | 512, 1K, 2K, 4K (UPPERCASE); 512 requires gemini-3.1-flash-image | (API default 1K) | No |
--thinking-level | minimal, High | unset (API default minimal) | No |
--image-only | (flag, no value) | off | No |
--search-grounding | (flag, no value) | off | No |
--model | Gemini model name | gemini-3-pro-image | No |
# Generate
bash scripts/openai.sh \
--mode generate \
--prompt "a mountain at sunset" \
--output ./mountain.png
# Generate with options
bash scripts/openai.sh \
--mode generate \
--prompt "company logo on transparent background" \
--output ./logo.png \
--size 1024x1024 \
--quality high \
--background transparent
# Edit
bash scripts/openai.sh \
--mode edit \
--prompt "add snow to the peaks" \
--input-image ./mountain.png \
--output ./snowy.png
Flags:
| Flag | Values | Default | Required |
|---|---|---|---|
--mode | generate, edit | -- | Yes |
--prompt | text | -- | Yes |
--output | file path | -- | Yes |
--input-image | file path, repeatable (max 16; dall-e-2 allows 1) | -- | Edit mode only |
--size | auto or WxH. On gpt-image-2, any size with both edges multiples of 16, longest edge up to 3840, ratio at most 3:1 and 655,360 to 8,294,400 total pixels (such as 2048x1152, 3840x2160); older models take 1024x1024, 1536x1024, 1024x1536 | 1024x1024 | No |
--quality | auto, low, medium, high | high | No |
--background | auto, transparent, opaque | auto | No |
--output-format | png, jpeg, webp | png | No |
--output-compression | integer 0-100 (jpeg/webp only) | -- | No |
--moderation | auto, low | auto | No |
--input-fidelity | low, high (edit only); not accepted with gpt-image-2 (that model always uses high fidelity; openai.sh refuses the flag) | unset (API default low) | No |
--model | OpenAI model name | gpt-image-2 | No |
# Generate
bash scripts/xai.sh \
--mode generate \
--prompt "a mountain at sunset" \
--output ./mountain.png
# Generate with aspect ratio
bash scripts/xai.sh \
--mode generate \
--prompt "a wide landscape" \
--output ./landscape.png \
--aspect-ratio 16:9
# Edit
bash scripts/xai.sh \
--mode edit \
--prompt "add snow to the peaks" \
--input-image ./mountain.png \
--output ./snowy.png
# Generate at 2K resolution
bash scripts/xai.sh \
--mode generate \
--prompt "a cat in a tree" \
--output ./cat.png \
--resolution 2k
# Use the May 2026 quality-mode model instead of the 2.0 default
bash scripts/xai.sh \
--mode generate \
--prompt "a cat in a tree" \
--output ./cat.png \
--model grok-imagine-image-quality
Flags:
| Flag | Values | Default | Required |
|---|---|---|---|
--mode | generate, edit | -- | Yes |
--prompt | text | -- | Yes |
--output | file path | -- | Yes |
--input-image | file path, repeatable (max 5 on grok-imagine-image-2.0; older models take 3) | -- | Edit mode only |
--aspect-ratio | 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 2:1, 1:2, 19.5:9, 9:19.5, 20:9, 9:20, 21:9, 5:2, auto | (none) | No |
--resolution | 1k, 2k (LOWERCASE) | (API default) | No |
--quality | low, medium, auto (grok-imagine-image-2.0 only) | unset (API auto: low for generation, medium for edits; billed as served) | No |
--model | xAI model name | grok-imagine-image-2.0 | No |
Note: For single-image edits, xAI ignores --aspect-ratio and uses the input image's ratio. Multi-image edits allow aspect ratio override (up to 5 images).
OpenRouter is a gateway to many image models through a single key. It uses the chat-completions API, so --model accepts any OpenRouter model slug that supports image output.
# Generate (default model: google/gemini-3.1-flash-image)
bash scripts/openrouter.sh \
--mode generate \
--prompt "a mountain at sunset" \
--output ./mountain.png
# Generate with a specific model
bash scripts/openrouter.sh \
--mode generate \
--prompt "a cat in a tree" \
--output ./cat.png \
--model openai/gpt-image-2
# Edit (single or multiple --input-image)
bash scripts/openrouter.sh \
--mode edit \
--prompt "add snow to the peaks" \
--input-image ./mountain.png \
--output ./snowy.png
Flags:
| Flag | Values | Default | Required |
|---|---|---|---|
--mode | generate, edit | -- | Yes |
--prompt | text | -- | Yes |
--output | file path | -- | Yes |
--input-image | file path, repeatable | -- | Edit mode only |
--model | any OpenRouter image model slug | google/gemini-3.1-flash-image | No |
--site-url | URL | (none) | No (sent as HTTP-Referer for OpenRouter attribution) |
--site-name | text | (none) | No (sent as X-Title for OpenRouter attribution) |
--site-url / --site-name also default from OPENROUTER_SITE_URL / OPENROUTER_SITE_NAME.
--input-image is repeatable on all four scripts. Passing more images than a provider supports exits with code 1 before any API call:
| Provider | Max images | Modes |
|---|---|---|
| Gemini | 14 | generate (references for a fresh composition) and edit |
| OpenAI | 16 | edit only (its generation endpoint takes no images) |
| xAI | 5 on grok-imagine-image-2.0 (older models take 3) | edit only |
| OpenRouter | model-dependent | edit only (input images attached as chat image parts) |
# Gemini: compose a new image from reference images (generate mode)
bash scripts/gemini.sh \
--mode generate \
--prompt "a product shot combining the chair from the first image with the fabric of the second" \
--input-image ./chair.png \
--input-image ./fabric.png \
--output ./composite.png
# OpenAI: multi-image edit
bash scripts/openai.sh \
--mode edit \
--prompt "place the logo from the second image onto the mug in the first" \
--input-image ./mug.png \
--input-image ./logo.png \
--output ./branded.png
# xAI: multi-image edit
bash scripts/xai.sh \
--mode edit \
--prompt "blend both scenes into one panorama" \
--input-image ./left.png \
--input-image ./right.png \
--output ./panorama.png
# All providers in parallel (edit mode only)
bash scripts/run-all.sh \
--mode edit \
--prompt "combine these" \
--input-image ./ref-a.png \
--input-image ./ref-b.png \
--output-base ./combined
Gemini's flat 14-image budget is best composed as up to 6 object + 5 character-consistency + 3 style-reference images. There is no API field to tag an image's role — the model infers it from the prompt, so state which images are objects, characters, or style references.
Notes:
run-all.sh forwards every --input-image to each selected provider, but only in --mode edit. Gemini's generate-mode reference images are not forwarded through run-all — call scripts/gemini.sh directly for generate-with-references.dall-e-2 edits are a known limitation: the script rejects multiple images for dall-e-2, but single-image dall-e-2 edits also do not work — the script sends form fields only the gpt-image models accept.Each provider script retries a transient API error (429 or a 5xx status, or a network failure) up to three times, with a delay that doubles each attempt (1 second, 2 seconds, 4 seconds). IMAGE_MAX_RETRIES and IMAGE_RETRY_DELAY tune the count and the starting delay. Inside a streaming pane, a retry shows on the waiting on line as (retry 2/3); once an image is on the pane that line is only written when the next block lands, so a retry that starts between blocks shows only in the final banner or error text.
When a provider run through run-all.sh fails outright, its error shows under a red ✗ provider error heading, and if any provider failed the pane offers [r] retry failed (xai) · [esc/ctrl-d] close 45s. The offer counts down until the first image is on the pane; after that it shows the time budget once, as [r] retry failed (xai) · [esc/ctrl-d] close (up to 45s), because a rewritten line erases the pane's images. Pressing r re-runs only the failed providers inside the same run; Esc closes the pane and the run returns with what it has. DISPLAY_PANE_RETRY_WAIT sets how long the offer stays open (default 45 seconds; 0 disables it), and a run with a failure can take that much longer plus one more provider round.
| Feature | Gemini | OpenAI | xAI | OpenRouter |
|---|---|---|---|---|
| Default model | gemini-3-pro-image | gpt-image-2 | grok-imagine-image-2.0 | google/gemini-3.1-flash-image |
| Max resolution | 4K (via --image-size) | 3840 px long edge on gpt-image-2 (via --size); 1536x1024 on older models | 2K (via --resolution) | Model-dependent |
| Text rendering | Very good (under 25 chars) | Excellent | Good | Model-dependent |
| Transparent BG | No | Yes (preview on gpt-image-2, png or webp only) | No | Model-dependent |
| Aspect ratios | 10 on Pro / 14 on 3.1 Flash | Any WxH up to 3:1 on gpt-image-2; 3 fixed sizes on older models | 16 options (incl. 21:9, 5:2, auto) | Prompt-driven |
| Image editing | Multi-turn, up to 14 refs (generate + edit) | Up to 16 input images | /v1/images/edits, up to 5 images | Chat image parts (edit) |
| Quality tiers | N/A | auto / low / medium / high | low / medium / auto (2.0 only) | Model-dependent |
| Thinking mode | Yes (--thinking-level) | No | No | Model-dependent |
| Search grounding | Yes (Google Search) | No | No | Model-dependent |
| Pricing | Token-based | Token-based | Flat per-image | Per OpenRouter model |
| Prompt revision | No | No | Yes (by chat model) | No |
| Component | File | Purpose |
|---|---|---|
| Plugin manifest | .claude-plugin/plugin.json | Plugin metadata, version and the /config rows (defaultProviders, outputDir) |
| Skill | skills/image-generation/SKILL.md | API knowledge, prompting tips, script reference |
| Command | commands/generate-image.md | /generate-image slash command |
| Agent | agents/image-generator.md | Autonomous image generation |
| Gemini script | scripts/gemini.sh | Gemini API call execution |
| OpenAI script | scripts/openai.sh | OpenAI API call execution |
| xAI script | scripts/xai.sh | xAI API call execution |
| OpenRouter script | scripts/openrouter.sh | OpenRouter chat-completions image call execution |
| Retry helper | scripts/retry.sh | curl_with_retry for transient API failures |
| Parallel runner | scripts/run-all.sh | Forks all providers in parallel under one streaming pane; holds a pane token for the batch |
| Display utility | scripts/display.sh | Multi-protocol terminal image display (iTerm2, Kitty, Sixel, tmux pane, shared streaming pane with colored banners + pending-provider waiting line) |
| API reference | skills/image-generation/references/api-details.md | Endpoint and payload documentation |
| Skill evals | skills/image-generation/evals/evals.json | Trigger prompts for the skill's eval suite |
| Hooks module | hooks/index.ts | Registers the generate tool and serves its calls through run-all.sh |
| Tool input rules | hooks/argv.ts | Checks a generate call's input and builds run-all.sh's argv |
| Hooks manifest | hooks/hooks.json | Points the engine at hooks/index.ts |
| Tool input type | hooks/generate-input.d.ts | Types the generate tool's input for the tool.call hook |
| TypeScript config | tsconfig.json | Type-checks hooks/ and tests/ against .claude/types |
| Automated tests | tests/ | bats test suite for all scripts; argv.test.ts for the tool's input rules |
This plugin uses calendar versioning in YYYY.M.PATCH format (e.g., 2026.7.1). The version is tracked in both .claude-plugin/plugin.json and skills/image-generation/SKILL.md.
# Run all automated tests (requires bats)
./tests/run_tests.sh
# Or run bats directly
bats tests/
The hooks module has its own checks. Run /plugin-types in a Claude Code session in this directory once, to write .claude/types, then:
tsc -p . # type-check hooks/ and tests/argv.test.ts
claude plugin test . # run tests/argv.test.ts
claude plugin validate . # what the engine sees: the module's hooks and the calls it makes
See TESTING.md for the full testing guide, including manual test procedures.
The plugin is organized into Claude Code extension points:
.claude-plugin/plugin.json -- Plugin identity and metadata
commands/ -- Slash command definitions
agents/ -- Autonomous agent definitions
skills/ -- Skill knowledge and references
scripts/ -- Shell scripts for API calls
hooks/ -- Function hooks module registering the generate tool
tests/ -- Automated tests (bats, argv.test.ts)
The scripts (gemini.sh, openai.sh, xai.sh, openrouter.sh) are standalone bash programs that handle API communication, base64 encoding/decoding, and error reporting. They are invoked by the command, agent, and skill layers.
hooks/index.ts 93 lines1// ABOUTME: Registers the generate tool, which runs scripts/run-all.sh from a typed input
2// ABOUTME: and answers with the files it saved; the defaults come from the plugin's /config rows.
3
4import type { Register } from 'claude-code'
5
6import { buildRun } from './argv'
7
8const TOOL = 'mcp__claude-image-generation__generate'
9// run-all.sh can wait on a slow provider for minutes and then hold a 45 s retry offer open;
10// ten minutes is the most $.process.run allows.
11const RUN_TIMEOUT_MS = 600_000
12
13const DESCRIPTION = `Generate or edit images with Gemini, OpenAI, xAI and OpenRouter in parallel. \
14Each provider saves <outputBase>-<provider>.png; the result lists the files that exist. \
15Omit providers and outputBase to use the person's /config defaults. \
16Providers: gemini (gemini-3-pro-image, professional assets, many reference images), \
17openai (gpt-image-2, text rendering, transparent backgrounds), \
18xai (grok-imagine-image-2.0, prompt revision, flat per-image pricing), \
19openrouter (google/gemini-3.1-flash-image by default, opt-in). \
20Pass inputImages to edit instead of generate. aspectRatio (W:H) applies to gemini and xai only.`
21
22const INPUT_SCHEMA = {
23 type: 'object',
24 properties: {
25 prompt: { type: 'string', description: 'What to draw, or how to change the input images.' },
26 providers: {
27 type: 'array',
28 items: { type: 'string', enum: ['gemini', 'openai', 'xai', 'openrouter'] },
29 description: 'Which providers to run; omitted, the configured default.',
30 },
31 outputBase: {
32 type: 'string',
33 description: 'Path without extension, e.g. "art/hero"; omitted, <configured dir>/image-<timestamp>.',
34 },
35 inputImages: { type: 'array', items: { type: 'string' }, description: 'Images to edit; switches to edit mode.' },
36 aspectRatio: { type: 'string', description: 'W:H such as 16:9; gemini and xai only.' },
37 },
38 required: ['prompt'],
39 additionalProperties: false,
40}
41
42export const register: Register = (on, options) => {
43 const defaults = {
44 defaultProviders: stringOption(options, 'defaultProviders'),
45 outputDir: stringOption(options, 'outputDir'),
46 }
47
48 on('session.start', async ($, e, next) => {
49 await $.tool.register({ name: 'generate', description: DESCRIPTION, inputSchema: INPUT_SCHEMA })
50 return next(e)
51 })
52
53 on('tool.call', { tool: TOOL }, async ($, e) => {
54 const run = buildRun(e, defaults, $.plugin.root, timestamp(await $.clock.now()))
55 if ('error' in run) return { deny: run.error }
56
57 // Outside a tmux pane the providers draw inline images to /dev/tty, which is Claude Code's own screen.
58 const { exitCode, stdout, stderr } = await $.process.run(run.argv, {
59 timeoutMs: RUN_TIMEOUT_MS,
60 env: { DISPLAY_IMAGE_TARGET: '/dev/null' },
61 })
62
63 const saved: string[] = []
64 const missing: string[] = []
65 for (const path of run.outputs) {
66 ;((await $.fs.exists(path)) ? saved : missing).push(path)
67 }
68 const lines = [
69 ...saved.map(path => `saved ${path}`),
70 ...missing.map(path => `missing ${path}`),
71 `run-all.sh exited ${exitCode}`,
72 ]
73 if (stdout.trim() !== '') lines.push('', stdout.trim())
74 if (stderr.trim() !== '') lines.push('', 'stderr:', stderr.trim())
75 const text = lines.join('\n')
76 if (saved.length === 0) return { deny: text }
77 return { result: text }
78 })
79}
80
81// plugin.json declares both fields as strings with defaults, so the engine always fills them in.
82function stringOption(options: Parameters<Register>[1], field: string): string {
83 const value = options[field]
84 if (typeof value !== 'string') throw new Error(`option ${field} should be a string from plugin.json userConfig, got ${JSON.stringify(value)}`)
85 return value
86}
87
88function timestamp(ms: number): string {
89 const d = new Date(ms)
90 const pad = (n: number) => String(n).padStart(2, '0')
91 return `${d.getFullYear()}${pad(d.getMonth() + 1)}${pad(d.getDate())}-${pad(d.getHours())}${pad(d.getMinutes())}${pad(d.getSeconds())}`
92}
93hooks/argv.ts 86 lines1// ABOUTME: Turns a generate tool call plus the plugin's options into the argv for run-all.sh,
2// ABOUTME: or the reason the call is refused. Pure, so every rule is testable without a run.
3
4export type GenerateOptions = {
5 defaultProviders: string
6 outputDir: string
7}
8
9export type Run = { argv: string[]; outputs: string[] } | { error: string }
10
11type GenerateInput = {
12 prompt: string
13 providers: string[] | undefined
14 outputBase: string | undefined
15 inputImages: string[]
16 aspectRatio: string | undefined
17}
18
19// What run-all.sh runs when no provider is named; openrouter is opt-in there too.
20const ALL = ['gemini', 'openai', 'xai']
21const PROVIDERS = [...ALL, 'openrouter']
22// The providers whose scripts take --aspect-ratio; openai sizes by --size, openrouter by model.
23const TAKES_ASPECT_RATIO = ['gemini', 'xai']
24
25export function buildRun(raw: Record<string, unknown>, options: GenerateOptions, root: string, stamp: string): Run {
26 const input = parseInput(raw)
27 if ('error' in input) return input
28
29 const providers = input.providers?.length
30 ? input.providers
31 : options.defaultProviders === 'all' ? ALL : [options.defaultProviders]
32 const unknown = providers.find(p => !PROVIDERS.includes(p))
33 if (unknown !== undefined) return { error: `unknown provider "${unknown}"; choose from ${PROVIDERS.join(', ')}` }
34 const ratio = input.aspectRatio
35 if (ratio !== undefined && !/^\d+:\d+$/.test(ratio)) return { error: `aspectRatio must look like 16:9, got "${ratio}"` }
36 const base = input.outputBase ?? `${options.outputDir}/image-${stamp}`
37 if (/\.(png|jpe?g|webp)$/i.test(base)) {
38 return { error: `outputBase is a path without an extension; "${base}" would save as "${base}-${providers[0]}.png"` }
39 }
40 const images = input.inputImages
41 const argv = [
42 'bash', `${root}/scripts/run-all.sh`,
43 '--mode', images.length > 0 ? 'edit' : 'generate',
44 '--prompt', input.prompt,
45 '--output-base', base,
46 '--providers', providers.join(','),
47 ...images.flatMap(image => ['--input-image', image]),
48 ...(ratio === undefined ? [] : providers
49 .filter(p => TAKES_ASPECT_RATIO.includes(p))
50 .flatMap(p => [`--${p}-extra`, `--aspect-ratio ${ratio}`])),
51 ]
52 return { argv, outputs: providers.map(p => `${base}-${p}.png`) }
53}
54
55// The model writes the call's input, so every field is checked for its type here, once.
56function parseInput(raw: Record<string, unknown>): GenerateInput | { error: string } {
57 const prompt = raw.prompt
58 if (typeof prompt !== 'string' || prompt.trim() === '') return { error: 'prompt is required' }
59 const providers = stringList(raw, 'providers')
60 if (isError(providers)) return providers
61 const inputImages = stringList(raw, 'inputImages')
62 if (isError(inputImages)) return inputImages
63 const outputBase = optionalString(raw, 'outputBase')
64 if (isError(outputBase)) return outputBase
65 const aspectRatio = optionalString(raw, 'aspectRatio')
66 if (isError(aspectRatio)) return aspectRatio
67 return { prompt, providers, inputImages: inputImages ?? [], outputBase, aspectRatio }
68}
69
70function isError<T>(value: T | { error: string }): value is { error: string } {
71 return typeof value === 'object' && value !== null && 'error' in value
72}
73
74function stringList(raw: Record<string, unknown>, field: string): string[] | undefined | { error: string } {
75 const value = raw[field]
76 if (value === undefined) return undefined
77 if (Array.isArray(value) && value.every(v => typeof v === 'string')) return value
78 return { error: `${field} must be a list of strings, got ${JSON.stringify(value)}` }
79}
80
81function optionalString(raw: Record<string, unknown>, field: string): string | undefined | { error: string } {
82 const value = raw[field]
83 if (value === undefined || typeof value === 'string') return value
84 return { error: `${field} must be a string, got ${JSON.stringify(value)}` }
85}
86