Grades each prompt you send with TypeSafe's Jev model and shows what is missing above the prompt box.

A Claude Code mod that grades the prompt you are writing with TypeSafe's Jev model, before you send it. Start typing and a Grade prompt button appears above the prompt box. Press it and the result shows in its place:
●●○○○ weak [feature] no success criteria (12%) · too little context (30%)
Works in the Claude Code CLI and in the Code tab of the Claude desktop app.
Mods are an early-access Claude Code feature and need a recent version. This mod is tested on Claude Code 2.1.293; older versions such as 2.1.241 cannot load it.
claude --version
If yours is older, update it:
claude update
The desktop app updates its own Claude Code; restart the app if it is out of date.
Sign up at TypeSafe and create an API key. Direct API access may still be invite-only, in which case you join the waitlist first.
Add it to the env block of ~/.claude/settings.json (create the file if it does not exist):
{
"env": {
"TYPESAFE_API_KEY": "your-key-here"
}
}
This works for both the CLI and the desktop app. Setting TYPESAFE_API_KEY in your shell profile works too, but only for sessions started from that shell.
Start Claude Code in a terminal and type this at the prompt:
/plugin install jev-prompt-grader --marketplace Syptern/jev-prompt-grader
Then:
y when asked to add the marketplace.You should see Installed jev-prompt-grader. Plugin is now active. It runs straight away, in that session and every session after, in the CLI and the desktop app alike.
Prefer the shell? The same in two commands:
claude plugin marketplace add Syptern/jev-prompt-grader
claude plugin install jev-prompt-grader@jev-prompt-grader
Type a prompt (do not send it yet). A Grade prompt button appears above the prompt box. Press it (or its number key in the terminal) and the grade appears after a second or two.
claude plugin update jev-prompt-grader
claude plugin uninstall jev-prompt-grader
To remove the marketplace as well:
claude plugin marketplace remove jev-prompt-grader
Grading takes two Jev calls. The first works out the kind of request: bug fix, feature, refactor, question, explore or other. The second asks for an overall score (poor / weak / okay / good / great) and only the checks that fit that kind, so a question is never marked down for missing acceptance criteria.
| Check | Shown when it fails | Bug fix | Feature | Refactor | Question | Explore | Other |
|---|---|---|---|---|---|---|---|
| States a clear goal | no clear goal | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Gives enough context | too little context | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Sticks to one task | mixes several tasks | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Unambiguous wording | ambiguous wording | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Specific scope | vague scope | ✓ | ✓ | ✓ | |||
| Success criteria | no success criteria | ✓ | ✓ | ✓ | |||
| Points to files, functions, URLs | no concrete pointers | ✓ | ✓ | ✓ | |||
| Error message or repro steps | no error or repro steps | ✓ | |||||
| Constraints | no constraints | ✓ | ✓ | ||||
| Explains why | missing the why | ✓ | ✓ | ||||
| Asks a specific question | question too broad | ✓ | ✓ | ||||
| Says what answer is wanted | no answer format | ✓ | ✓ |
Each check is a Jev yes/no probability. Anything under 50% is listed, worst first, up to 4 plus "+N more".
Your last 3 sent prompts go along as context, so a follow-up like "now do the same for the footer" is graded in light of what came before instead of being flagged as vague. Only the draft itself is graded.
Edit the draft after grading and the result dims with a Re-grade button. Sending the prompt or emptying the box clears it.
Text is sent to TypeSafe's API (api.typesafe.ai) only when you press the button, never while you type. Each press sends:
Both Jev calls of one grade send the same text.
Nothing else is sent: no files, no replies from Claude, no project details. The mod keeps the earlier prompts for the current session only.
Anything you pasted into one of those prompts, a secret included, is sent again with each grade. If you pasted something sensitive, start a new session before grading.
| What you see | What to do |
|---|---|
| No button when you type | Check claude --version (step 1), then run /plugin and make sure jev-prompt-grader is installed and enabled. |
Set TYPESAFE_API_KEY to grade prompts | The key is not reaching Claude Code. Check step 3, then start a new session. |
Jev returned HTTP 401 | The key is wrong or revoked. |
Jev returned HTTP 4xx (other) | Your account may not have API access yet, or TypeSafe changed its API. Open an issue. |
Could not reach Jev | No network, or a firewall or sandbox blocks api.typesafe.ai. |
Clone the repository and load your copy for one session:
claude --plugin-dir ./jev-prompt-grader
Validate and test:
claude plugin validate .
claude plugin test .
Add a check by adding an entry to CHECKS in hooks/register.tsx with an id, a label (shown when it fails), instructions (the yes/no question for Jev) and kinds (the kinds of request it applies to).
Files:
hooks/register.tsx: the hooks (draft tracking, grading on request, and the band above the prompt)hooks/band.test.tsx: tests with a mocked Jev responsetypes/index.d.ts: the mod's state contract.claude-plugin/plugin.json: the mod's manifest.claude-plugin/marketplace.json: makes this repository installable with /plugin installhooks/register.tsx 364 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface as Engine, Register } from 'claude-code'
3
4import type { Check, Grade } from '../types'
5
6const grade = atom({ plugin: 'jev-prompt-grader', key: 'grade' } as const, null)
7const hasDraft = atom({ plugin: 'jev-prompt-grader', key: 'hasDraft' } as const, false)
8const isStale = atom({ plugin: 'jev-prompt-grader', key: 'isStale' } as const, false)
9const isGrading = atom({ plugin: 'jev-prompt-grader', key: 'isGrading' } as const, false)
10const history = atom({ plugin: 'jev-prompt-grader', key: 'history' } as const, [] as string[])
11
12const ENDPOINT = 'https://api.typesafe.ai/v1/systemone'
13const LEVELS = ['poor', 'weak', 'okay', 'good', 'great']
14// Earlier prompts sent along as context, and how much of each.
15const HISTORY_SIZE = 3
16const HISTORY_CHARS = 2000
17
18// The kinds of request Jev picks from in the first call; each check lists the kinds it fits.
19const KINDS = {
20 'bug fix': 'fix broken behaviour',
21 feature: 'build something new or extend existing behaviour',
22 refactor: 'restructure existing code without changing behaviour',
23 question: 'explain or answer something, no code change asked',
24 explore: 'brainstorm, research, investigate or plan',
25 other: 'anything else',
26} as const
27type Kind = keyof typeof KINDS
28
29const CODE_CHANGE: Kind[] = ['bug fix', 'feature', 'refactor']
30const ALL = Object.keys(KINDS) as Kind[]
31
32// Yes/no checks Jev answers with a probability, asked only for the kinds they fit.
33// Label is what we show when it fails.
34const CHECKS: { id: string; label: string; instructions: string; kinds: Kind[] }[] = [
35 {
36 id: 'goal',
37 label: 'no clear goal',
38 instructions: 'Does the prompt state a clear goal or desired outcome?',
39 kinds: ALL,
40 },
41 {
42 id: 'context',
43 label: 'too little context',
44 instructions:
45 'Does the prompt give enough context (relevant files, background, current behaviour) to act without guessing?',
46 kinds: ALL,
47 },
48 {
49 id: 'scope',
50 label: 'vague scope',
51 instructions:
52 'Is the scope specific (what to change and where) rather than vague or open-ended?',
53 kinds: CODE_CHANGE,
54 },
55 {
56 id: 'done',
57 label: 'no success criteria',
58 instructions:
59 'Does the prompt say how to tell the task is done (expected result, tests to pass, acceptance criteria)?',
60 kinds: CODE_CHANGE,
61 },
62 {
63 id: 'focus',
64 label: 'mixes several tasks',
65 instructions: 'Is the prompt focused on a single task rather than several unrelated asks?',
66 kinds: ALL,
67 },
68 {
69 id: 'references',
70 label: 'no concrete pointers',
71 instructions:
72 'Does the prompt point to concrete places to look (file names, functions, components, URLs, commands)?',
73 kinds: CODE_CHANGE,
74 },
75 {
76 id: 'repro',
77 label: 'no error or repro steps',
78 instructions: 'Does the prompt include the error message, logs or steps to reproduce the bug?',
79 kinds: ['bug fix'],
80 },
81 {
82 id: 'constraints',
83 label: 'no constraints',
84 instructions:
85 'Does the prompt state constraints where they matter (what must not change, libraries or patterns to use or avoid, compatibility, style)? Answer yes if the task is simple enough to need none.',
86 kinds: ['feature', 'refactor'],
87 },
88 {
89 id: 'why',
90 label: 'missing the why',
91 instructions:
92 'Does the prompt explain why the change is wanted or what problem it solves, so trade-offs can be judged? Answer yes if the reason is obvious from the request.',
93 kinds: ['feature', 'refactor'],
94 },
95 {
96 id: 'specific',
97 label: 'question too broad',
98 instructions:
99 'Does the prompt ask a specific question rather than about a broad topic?',
100 kinds: ['question', 'explore'],
101 },
102 {
103 id: 'format',
104 label: 'no answer format',
105 instructions:
106 'Does the prompt say what kind of answer is wanted (a short answer, an explanation, a comparison, a recommendation, a plan, a code example)? Answer yes if it is obvious from the question.',
107 kinds: ['question', 'explore'],
108 },
109 {
110 id: 'clear',
111 label: 'ambiguous wording',
112 instructions:
113 'Is the wording unambiguous, with no unclear references such as "it", "this" or "that thing" whose meaning cannot be told from the prompt alone?',
114 kinds: ALL,
115 },
116]
117
118// Failed checks shown in the band; the rest are summed up as "+N more".
119const MAX_SHOWN = 4
120
121type JevAnswer = { noul?: number; score?: number; choice?: string }
122type JevResponse = { answers?: Record<string, JevAnswer> }
123
124// Every question grades the last prompt only; earlier ones are there to resolve follow-ups.
125const CONTEXT_NOTE =
126 ' Judge only the prompt under "Prompt to grade". Earlier prompts are context: a follow-up that is clear given them counts as clear.'
127
128// What Jev reads: the draft alone, or the draft after the earlier prompts.
129function buildState(draft: string, earlier: readonly string[]) {
130 if (earlier.length === 0) {
131 return draft
132 }
133 const lines = earlier.map((t, i) => `${i + 1}. ${t.slice(0, HISTORY_CHARS)}`)
134 return `Earlier prompts in this conversation, oldest first:\n${lines.join('\n')}\n\nPrompt to grade:\n${draft}`
135}
136
137function checksFor(kind: Kind) {
138 return CHECKS.filter(c => c.kinds.includes(kind))
139}
140
141// First call: what kind of request this is.
142function kindQuestion(note: string) {
143 return {
144 kind: { type: 'choice', instructions: 'What kind of request is this?' + note, criteria: KINDS },
145 }
146}
147
148// Second call: the overall score and the checks that fit the kind.
149function gradeQuestions(kind: Kind, note: string) {
150 const questions: Record<string, unknown> = {
151 overall: {
152 type: 'score',
153 instructions:
154 `This is a ${kind} request (${KINDS[kind]}). How good is the prompt for an AI coding assistant to act on without follow-up questions?` +
155 note,
156 criteria: [
157 'Poor: unclear what is wanted',
158 'Weak: rough idea, major gaps',
159 'Okay: workable but the assistant must guess',
160 'Good: clear, minor gaps',
161 `Great: everything a ${kind} request needs`,
162 ],
163 },
164 }
165 for (const c of checksFor(kind)) {
166 questions[c.id] = { type: 'noul', instructions: c.instructions + note }
167 }
168 return questions
169}
170
171function toGrade(answers: Record<string, JevAnswer>, kind: Kind): Grade {
172 const score = answers.overall?.score ?? 0
173 const level = Math.max(0, Math.min(LEVELS.length - 1, Math.round(score)))
174 const missing: Check[] = checksFor(kind)
175 .map(c => ({ id: c.id, label: c.label, p: answers[c.id]?.noul ?? 1 }))
176 .filter(c => c.p < 0.5)
177 .sort((a, b) => a.p - b.p)
178 return { status: 'done', level, label: LEVELS[level] ?? 'okay', kind, missing }
179}
180
181// One Jev call: the answers, or the error to show.
182async function askJev(
183 $: Engine,
184 apiKey: string,
185 state: string,
186 questions: Record<string, unknown>,
187): Promise<{ answers: Record<string, JevAnswer> } | { error: string }> {
188 try {
189 const res = await $.http.fetch(ENDPOINT, {
190 method: 'POST',
191 headers: {
192 Authorization: `Bearer ${apiKey}`,
193 'Content-Type': 'application/json',
194 },
195 body: JSON.stringify({ model: 'jev-latest', state, questions }),
196 })
197 if (!res.ok) {
198 return { error: `Jev returned HTTP ${res.status}` }
199 }
200 return { answers: (JSON.parse(res.text) as JevResponse).answers ?? {} }
201 } catch {
202 return { error: 'Could not reach Jev' }
203 }
204}
205
206// Counts requests so a slow answer for an older draft never overwrites a newer one.
207let latest = 0
208// The draft the shown grade belongs to, so cursor moves and no-op edits do not mark it stale.
209let lastGraded = ''
210
211// Sends the draft to Jev in two calls: first the kind of request, then the checks that fit
212// it. Only ever called from the Grade button, so nothing leaves the machine unless the
213// person asks for it.
214async function gradeDraft($: Engine) {
215 const draft = (await $.prompt.read()).text.trim()
216 if (!draft) {
217 return
218 }
219 lastGraded = draft
220
221 const apiKey = await $.env.get('TYPESAFE_API_KEY')
222 if (!apiKey) {
223 await update($, grade, (): Grade => ({
224 status: 'error',
225 message: 'Set TYPESAFE_API_KEY to grade prompts',
226 }))
227 return
228 }
229
230 const id = ++latest
231 await update($, isGrading, () => true)
232 const earlier = await read($, history)
233 const state = buildState(draft, earlier)
234 const note = earlier.length > 0 ? CONTEXT_NOTE : ''
235
236 let result: Grade
237 const first = await askJev($, apiKey, state, kindQuestion(note))
238 if ('error' in first) {
239 result = { status: 'error', message: first.error }
240 } else {
241 const choice = first.answers.kind?.choice
242 const kind: Kind = choice && choice in KINDS ? (choice as Kind) : 'other'
243 // Skip the second call when a newer request has already taken over.
244 if (id !== latest) {
245 return
246 }
247 const second = await askJev($, apiKey, state, gradeQuestions(kind, note))
248 result =
249 'error' in second
250 ? { status: 'error', message: second.error }
251 : toGrade(second.answers, kind)
252 }
253
254 if (id === latest) {
255 await update($, grade, () => result)
256 await update($, isGrading, () => false)
257 await update($, isStale, () => false)
258 }
259}
260
261async function reset($: Engine) {
262 latest += 1
263 lastGraded = ''
264 await update($, grade, () => null)
265 await update($, hasDraft, () => false)
266 await update($, isStale, () => false)
267 await update($, isGrading, () => false)
268}
269
270export const register: Register = on => {
271 // Tracks the draft locally (nothing is sent) so the band knows when to show and when a grade is stale.
272 on('prompt.edit', async ($, e, next) => {
273 const box = await next(e)
274 const draft = box.text.trim()
275
276 if (!draft) {
277 await reset($)
278 return box
279 }
280 if (!(await read($, hasDraft))) {
281 await update($, hasDraft, () => true)
282 }
283 // Stale exactly when a grade is shown for text other than what is in the box now.
284 const stale = (await read($, grade)) !== null && draft !== lastGraded
285 if ((await read($, isStale)) !== stale) {
286 await update($, isStale, () => stale)
287 }
288 return box
289 })
290
291 // Sending the prompt empties the box, so the grade goes with it. The prompt is kept,
292 // in this session only, as context for grading the next ones.
293 on('prompt.submit', async ($, e, next) => {
294 const result = await next(e)
295 if (e.origin.kind === 'composer') {
296 await reset($)
297 const text = e.text.trim()
298 if (text && !text.startsWith('/')) {
299 await update($, history, h => [...h, text].slice(-HISTORY_SIZE))
300 }
301 }
302 return result
303 })
304
305 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
306 const g = await read($, grade)
307 const has = await read($, hasDraft)
308 if (e.props.hasSurvey || !has) {
309 return next(e)
310 }
311
312 const { Box, Button, Text } = $.ui.resolve(e)
313 const grading = await read($, isGrading)
314 const stale = await read($, isStale)
315 const gradeButton = grading ? null : (
316 <Button
317 key="grade"
318 label={g?.status === 'done' ? 'Re-grade' : 'Grade prompt'}
319 onPress={() => gradeDraft($)}
320 />
321 )
322
323 if (g === null) {
324 return (
325 <Box>
326 {grading ? <Text dimColor>Jev is grading your prompt…</Text> : gradeButton}
327 </Box>
328 )
329 }
330
331 if (g.status === 'error') {
332 return (
333 <Box>
334 <Text dimColor>Prompt grader: {g.message} </Text>
335 {gradeButton}
336 </Box>
337 )
338 }
339
340 const color = g.level >= 3 ? 'green' : g.level === 2 ? 'yellow' : 'red'
341 const dots = '●'.repeat(g.level + 1) + '○'.repeat(LEVELS.length - g.level - 1)
342 const shown = g.missing.slice(0, MAX_SHOWN).map(m => `${m.label} (${Math.round(m.p * 100)}%)`)
343 const more = g.missing.length - shown.length
344 const notes = g.missing.length
345 ? shown.join(' · ') + (more > 0 ? ` · +${more} more` : '')
346 : 'nothing missing'
347 const status = grading ? ' · grading…' : stale ? ' · edited since ' : ''
348
349 return (
350 <Box>
351 <Text color={stale || grading ? undefined : color} dimColor={stale || grading}>
352 {dots} {g.label}
353 </Text>
354 <Text dimColor>
355 {' '}
356 [{g.kind}] {notes}
357 {status}
358 </Text>
359 {stale ? gradeButton : null}
360 </Box>
361 )
362 })
363}
364types/index.d.ts 24 lines1export type Check = { id: string; label: string; p: number }
2
3export type Grade =
4 | { status: 'error'; message: string }
5 | {
6 status: 'done'
7 level: number // 0..4, rounded expected score
8 label: string
9 kind: string
10 missing: Check[] // checks with p < 0.5, worst first
11 }
12
13declare module 'claude-code' {
14 interface PluginState {
15 'jev-prompt-grader': {
16 grade: Grade | null // the last result, kept on screen while a new one loads
17 hasDraft: boolean // the prompt box holds text
18 isStale: boolean // the draft changed since it was graded
19 isGrading: boolean // a Jev request is in flight
20 history: string[] // the person's last sent prompts, oldest first, sent along as context
21 }
22 }
23}
24