Audits a URL for AI answer engines (ChatGPT search, Perplexity, Google AI Overviews, Copilot, Claude): crawler access, firewall, indexing directives, no-JS…

A free Claude Code mod that checks whether ChatGPT search, Perplexity, Google AI Overviews, Microsoft Copilot and Claude can get into a web page and read it.
Type /aeo yoursite.com. A pane opens with every check marked PASS, WARN, FAIL or INFO, the fix for anything that failed, and the source each rule comes from. The full report is saved as markdown.
Built for the video If I Wanted ChatGPT to Recommend My Business, I'd Run This Audit First by Jake Bauman. [LINK TO VIDEO]
You need Claude Code 2.1.287 or newer. The mod API is early access, so a future Claude Code release may change it.
git clone https://github.com/jbauman-26/aeo-audit-mod.git ~/aeo-audit-mod
Load it for one session:
claude --plugin-dir ~/aeo-audit-mod
Load it in every session, including the Claude desktop app: add this to the env block of ~/.claude/settings.json, then restart Claude Code.
"CLAUDE_CODE_PLUGIN_DIRS": "~/aeo-audit-mod"
/aeo https://example.com/pricing audits that page and opens the results pane./aeo on its own reopens the last report.audit_site tool, reads the report and writes the fixes.f) drafts that request for you, and Rerun (r) checks again after you change something.Reports are saved to ~/output/aeo-audits/<host>-<date>.md.
| Check | Why it matters | Source |
|---|---|---|
| robots.txt allows each search crawler: OAI-SearchBot, PerplexityBot, Googlebot, Bingbot, Claude-SearchBot | Blocking a search crawler removes you from that engine's answers. Blocking only training crawlers (GPTBot, Google-Extended, ClaudeBot, CCBot) does not, so those are reported but not scored. | OpenAI, Perplexity, Google, Bing, Anthropic |
| Firewall treats AI crawler user agents the same as a browser | A CDN or firewall rule can block AI crawlers before robots.txt is read. | |
| No noindex, nosnippet or max-snippet:0 | Google uses these to keep a page out of AI Overviews and AI Mode. | |
| Text visible without JavaScript | OpenAI's, Anthropic's and Perplexity's crawlers read the raw HTML and do not run JavaScript. | Vercel, Dec 2024 |
| Page status, title, meta description, canonical URL, one H1 | The basics every engine reads first. | |
| Questions with short, direct answers under them (headings or FAQ accordions) | A heuristic: answer engines match pages to questions, and a short direct answer is easy to quote. | |
| Schema (JSON-LD), sitemap, llms.txt | Listed, not scored, except broken JSON-LD and a missing sitemap. Google says AI features need no special schema, and no major engine has said it reads llms.txt. | Google, llmstxt.org |
The score covers this audit's checks only. It is not a prediction that any engine will cite you. Passing means the doors are open and the page is readable; whether an engine recommends you depends on your content, your reputation and what other sites say about you.
Three honest limits:
claude plugin validate ~/aeo-audit-mod
claude plugin test ~/aeo-audit-mod
The checks live in hooks/audit.ts as plain functions with no Claude Code dependency, so you can run them from any TypeScript runtime. hooks/register.tsx wires them to the /aeo command, the audit_site tool, the pane and the status line.
MIT. Use it, change it, ship it to clients.
hooks/register.tsx 253 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { SavedReport } from '../types'
5import {
6 BOT_USER_AGENTS,
7 BROWSER_USER_AGENT,
8 audit,
9 sitemapUrl,
10 toMarkdown,
11 type Check,
12 type Fetched,
13 type Status,
14} from './audit'
15
16type $ = EngineInterface
17
18const PANE = 'aeo-audit'
19const last = atom({ plugin: 'aeo-audit', key: 'last' } as const, null as SavedReport | null)
20const running = atom({ plugin: 'aeo-audit', key: 'running' } as const, null as string | null)
21
22const COLOR: Record<Status, string | undefined> = { pass: 'green', warn: 'yellow', fail: 'red', info: undefined }
23const MARK: Record<Status, string> = { pass: 'PASS', warn: 'WARN', fail: 'FAIL', info: 'INFO' }
24
25export function normalizeUrl(raw: string): string | undefined {
26 const trimmed = raw.trim()
27 if (trimmed === '') return undefined
28 try {
29 const url = new URL(/^https?:\/\//i.test(trimmed) ? trimmed : `https://${trimmed}`)
30
31 return url.hostname.includes('.') ? url.toString() : undefined
32 } catch {
33 return undefined
34 }
35}
36
37// Follows redirects itself so the audit reads the page a crawler lands on.
38async function get($: $, url: string, userAgent = BROWSER_USER_AGENT): Promise<Fetched | undefined> {
39 let current = url
40 try {
41 for (let hop = 0; hop < 5; hop += 1) {
42 const res = await $.http.fetch(current, { headers: { 'user-agent': userAgent, accept: 'text/html,*/*' } })
43 const location = res.headers.location
44 if (res.status >= 300 && res.status < 400 && location !== undefined) {
45 current = new URL(location, current).toString()
46 continue
47 }
48
49 return { status: res.status, headers: res.headers, text: res.text }
50 }
51 } catch {
52 return undefined
53 }
54
55 return undefined
56}
57
58async function runAudit($: $, url: string): Promise<SavedReport> {
59 const origin = new URL(url).origin
60 await update($, running, () => url)
61 $.ui.status(`AEO: auditing ${new URL(url).hostname}...`)
62 try {
63 const robots = await get($, `${origin}/robots.txt`)
64 const robotsText = robots?.status === 200 ? robots.text : undefined
65 const probeNames = Object.keys(BOT_USER_AGENTS)
66 const [page, llms, sitemap, ...probes] = await Promise.all([
67 get($, url),
68 get($, `${origin}/llms.txt`),
69 get($, sitemapUrl(robotsText, origin)),
70 ...probeNames.map(name => get($, url, BOT_USER_AGENTS[name])),
71 ])
72 const report = audit({
73 url,
74 page: page ?? { status: 0, headers: {}, text: '' },
75 robots,
76 llms,
77 sitemap,
78 botProbes: Object.fromEntries(probeNames.map((name, i) => [name, probes[i]])),
79 })
80
81 const at = new Date(await $.clock.now()).toISOString()
82 const day = at.slice(0, 10)
83 const home = await $.env.get('HOME')
84 const file = `${new URL(url).hostname}-${day}.md`
85 let path: string | null = null
86 if (home !== undefined) {
87 path = `${home}/output/aeo-audits/${file}`
88 try {
89 await $.fs.write(path, `${toMarkdown(report, day)}\n`)
90 } catch {
91 path = null
92 }
93 }
94
95 const saved: SavedReport = { ...report, at, path }
96 await update($, last, () => saved)
97 await $.store.set('last', saved)
98 $.ui.status(`AEO: ${new URL(url).hostname} ${report.score}/100${report.isPartial ? ' (partial)' : ''}`)
99
100 return saved
101 } finally {
102 await update($, running, () => null)
103 }
104}
105
106function summary(report: SavedReport): string {
107 const top = report.checks.filter(c => c.status === 'fail').map(c => `- ${c.label}`)
108
109 return [
110 `${report.url}: ${report.score}/100${report.isPartial ? ' (partial, the page refused the audit)' : ''}.`,
111 `${report.counts.fail} fail, ${report.counts.warn} warn, ${report.counts.pass} pass.`,
112 ...(top.length > 0 ? ['Failing:', ...top] : []),
113 report.path === null ? 'Report not saved.' : `Report: ${report.path}`,
114 ].join('\n')
115}
116
117export const register: Register = on => {
118 on('session.start', async ($, e, next) => {
119 await $.command.register({
120 name: 'aeo',
121 description: 'Audit a URL for AI answer engines (ChatGPT, Perplexity, AI Overviews, Copilot, Claude)',
122 argumentHint: '<url>',
123 })
124 await $.tool.register({
125 name: 'audit_site',
126 description:
127 'Audits one URL for AI answer-engine readiness: robots.txt access for each AI search crawler, firewall behaviour toward AI user agents, noindex/nosnippet, text visible without JavaScript, title, description, canonical, H1, question-shaped sections, schema, sitemap, llms.txt. Returns a markdown report with sources and fixes, and saves it under ~/output/aeo-audits/.',
128 inputSchema: {
129 type: 'object',
130 properties: { url: { type: 'string', description: 'The page to audit, e.g. https://example.com/pricing' } },
131 required: ['url'],
132 },
133 })
134 if ((await read($, last)) === null) {
135 const stored = await $.store.get('last')
136 if (stored !== undefined && stored !== null) await update($, last, () => stored as SavedReport)
137 }
138
139 return next(e)
140 })
141
142 on('command.run', { command: 'aeo' }, async ($, e) => {
143 if (e.args.trim() === '') {
144 const opened = await $.ui.open({ id: PANE, title: 'AEO audit' })
145
146 return { text: (await read($, last)) === null ? 'Usage: /aeo <url>' : opened.isPlaced ? 'AEO pane opened.' : `AEO pane waiting: ${opened.reason}` }
147 }
148 const url = normalizeUrl(e.args)
149 if (url === undefined) return { text: `Not a URL: ${e.args.trim()}` }
150 void $.ui.open({ id: PANE, title: 'AEO audit' })
151 const report = await runAudit($, url)
152
153 return { text: summary(report) }
154 })
155
156 on('tool.call', { tool: 'mcp__aeo-audit__audit_site' }, async ($, e) => {
157 // A plugin tool's arguments arrive on the event itself, beside `tool` and `tool_use_id`.
158 const url = normalizeUrl(String((e as { url?: unknown }).url ?? ''))
159 if (url === undefined) return { deny: 'audit_site needs a full URL such as https://example.com/' }
160 const report = await runAudit($, url)
161
162 return { result: `${toMarkdown(report, report.at.slice(0, 10))}\n\n${report.path === null ? '' : `Saved: ${report.path}`}` }
163 })
164
165 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
166 const { Box, Text, Button } = $.ui.resolve(e)
167 const report = await read($, last)
168 const busy = await read($, running)
169
170 if (report === null) {
171 return (
172 <Box flexDirection="column">
173 <Text>{busy === null ? 'No audit yet. Run /aeo <url>.' : `Auditing ${busy}...`}</Text>
174 </Box>
175 )
176 }
177
178 const scoreColor = report.score >= 80 ? 'green' : report.score >= 50 ? 'yellow' : 'red'
179 const row = (c: Check) => (
180 <Box key={c.id} flexDirection="column">
181 <Text wrap="truncate-end">
182 <Text color={COLOR[c.status]} bold={c.status === 'fail'}>
183 {MARK[c.status]}
184 </Text>{' '}
185 {c.label}
186 </Text>
187 {(c.status === 'fail' || c.status === 'warn') && (
188 <Text dimColor wrap="wrap">
189 {' '}
190 {c.fix ?? c.detail}
191 </Text>
192 )}
193 </Box>
194 )
195
196 return (
197 <Box flexDirection="column">
198 <Text bold wrap="truncate-end">
199 {report.url}
200 </Text>
201 <Text>
202 <Text color={scoreColor} bold>
203 {report.score}/100
204 </Text>
205 <Text dimColor>
206 {' '}
207 {report.counts.fail} fail, {report.counts.warn} warn, {report.counts.pass} pass
208 {report.isPartial ? ', partial' : ''}
209 </Text>
210 </Text>
211 {busy !== null && <Text color="yellow">Re-auditing {busy}...</Text>}
212 {(['Crawler access', 'Page', 'Content', 'Extras'] as const).map(group => {
213 const rows = report.checks.filter(c => c.group === group)
214 if (rows.length === 0) return null
215
216 return (
217 <Box key={group} flexDirection="column" marginTop={1}>
218 <Text bold>{group}</Text>
219 {rows.map(row)}
220 </Box>
221 )
222 })}
223 <Box marginTop={1}>
224 <Button
225 key="fix"
226 label="Fix with Claude"
227 hotkey="f"
228 variant="primary"
229 onPress={async () => {
230 const live = await read($, last)
231 if (live === null) return
232 const where = live.path === null ? 'the AEO audit in this session' : `the AEO audit at ${live.path}`
233 await $.prompt.fill({
234 text: `Read ${where} for ${live.url} and write the exact fixes for every FAIL and WARN, most impact first. Quote the source for each.`,
235 })
236 }}
237 />
238 <Text> </Text>
239 <Button
240 key="rerun"
241 label="Rerun"
242 hotkey="r"
243 onPress={async () => {
244 const live = await read($, last)
245 if (live !== null) await runAudit($, live.url)
246 }}
247 />
248 </Box>
249 </Box>
250 )
251 })
252}
253hooks/audit.ts 552 lines1// Pure analysis: takes what was fetched, returns findings. No network, no engine,
2// so the same code runs in the mod, in tests, and from a plain `bun` script.
3
4import type { Check, Report, Status } from '../types'
5
6export type { Check, Report, Status }
7
8export type Fetched = { status: number; headers: Record<string, string>; text: string }
9
10export type AuditInput = {
11 url: string
12 page: Fetched
13 robots: Fetched | undefined
14 llms: Fetched | undefined
15 sitemap: Fetched | undefined
16 // The same page requested with AI crawler user agents, keyed by crawler name.
17 botProbes: Record<string, Fetched | undefined>
18}
19
20
21// Crawlers that decide whether a site can appear in an AI answer engine at all.
22export const SEARCH_BOTS = [
23 {
24 token: 'OAI-SearchBot',
25 engine: 'ChatGPT search',
26 source: 'https://developers.openai.com/api/docs/bots',
27 },
28 {
29 token: 'PerplexityBot',
30 engine: 'Perplexity',
31 source: 'https://docs.perplexity.ai/guides/bots',
32 },
33 {
34 token: 'Googlebot',
35 engine: 'Google AI Overviews and AI Mode',
36 source: 'https://developers.google.com/search/docs/appearance/ai-features',
37 },
38 { token: 'Bingbot', engine: 'Microsoft Copilot (Bing index)', source: 'https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0' },
39 { token: 'Claude-SearchBot', engine: 'Claude search', source: 'https://support.claude.com/en/articles/8896518' },
40] as const
41
42// Crawlers that only collect training data. Blocking them is a policy choice, not a visibility problem.
43export const TRAINING_BOTS = ['GPTBot', 'Google-Extended', 'ClaudeBot', 'CCBot'] as const
44
45export const BOT_USER_AGENTS: Record<string, string> = {
46 'OAI-SearchBot': 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot',
47 PerplexityBot: 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)',
48}
49
50export const BROWSER_USER_AGENT =
51 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36'
52
53// ---------- robots.txt (RFC 9309 group selection, longest-match rules) ----------
54
55type Rule = { allow: boolean; pattern: string }
56type Group = { agents: string[]; rules: Rule[] }
57
58export function parseRobots(text: string): { groups: Group[]; sitemaps: string[] } {
59 const groups: Group[] = []
60 const sitemaps: string[] = []
61 let current: Group | undefined
62 let lastWasAgent = false
63
64 for (const raw of text.split(/\r?\n/)) {
65 const line = raw.replace(/#.*$/, '').trim()
66 const colon = line.indexOf(':')
67 if (colon === -1) continue
68 const key = line.slice(0, colon).trim().toLowerCase()
69 const value = line.slice(colon + 1).trim()
70
71 if (key === 'sitemap') {
72 if (value !== '') sitemaps.push(value)
73 continue
74 }
75 if (key === 'user-agent') {
76 if (current === undefined || !lastWasAgent) {
77 current = { agents: [], rules: [] }
78 groups.push(current)
79 }
80 current.agents.push(value.toLowerCase())
81 lastWasAgent = true
82 continue
83 }
84 if ((key === 'allow' || key === 'disallow') && current !== undefined) {
85 // An empty Disallow blocks nothing, so it adds no rule.
86 if (value !== '') current.rules.push({ allow: key === 'allow', pattern: value })
87 lastWasAgent = false
88 }
89 }
90
91 return { groups, sitemaps }
92}
93
94function ruleMatches(pattern: string, path: string): boolean {
95 const anchored = pattern.endsWith('$')
96 const body = (anchored ? pattern.slice(0, -1) : pattern)
97 .split('*')
98 .map(part => part.replace(/[.+?^${}()|[\]\\]/g, '\\$&'))
99 .join('.*')
100
101 return new RegExp(`^${body}${anchored ? '$' : ''}`).test(path)
102}
103
104// Whether `token` may fetch `path`, and which group decided it.
105export function robotsAllows(
106 robots: ReturnType<typeof parseRobots>,
107 token: string,
108 path: string,
109): { isAllowed: boolean; matchedAgent: string | undefined; rule: Rule | undefined } {
110 const lower = token.toLowerCase()
111 let groups = robots.groups.filter(g => g.agents.includes(lower))
112 let matchedAgent: string | undefined = groups.length > 0 ? token : undefined
113 if (groups.length === 0) {
114 groups = robots.groups.filter(g => g.agents.includes('*'))
115 matchedAgent = groups.length > 0 ? '*' : undefined
116 }
117
118 let best: Rule | undefined
119 for (const rule of groups.flatMap(g => g.rules)) {
120 if (!ruleMatches(rule.pattern, path)) continue
121 const isLonger = best === undefined || rule.pattern.length > best.pattern.length
122 const isTieAllow = best !== undefined && rule.pattern.length === best.pattern.length && rule.allow
123 if (isLonger || isTieAllow) best = rule
124 }
125
126 return { isAllowed: best === undefined || best.allow, matchedAgent, rule: best }
127}
128
129// The sitemap to fetch: the first one robots.txt lists, else /sitemap.xml.
130export function sitemapUrl(robotsText: string | undefined, origin: string): string {
131 return parseRobots(robotsText ?? '').sitemaps[0] ?? `${origin}/sitemap.xml`
132}
133
134// ---------- HTML helpers (regex, no DOM in the mod's environment) ----------
135
136const ENTITIES: Record<string, string> = { amp: '&', lt: '<', gt: '>', quot: '"', apos: "'", nbsp: ' ' }
137
138export function decode(text: string): string {
139 return text.replace(/&(#x[0-9a-f]+|#\d+|[a-z]+);/gi, (whole, code: string) => {
140 if (code[0] === '#') {
141 const n = code[1] === 'x' || code[1] === 'X' ? parseInt(code.slice(2), 16) : parseInt(code.slice(1), 10)
142
143 return Number.isFinite(n) ? String.fromCodePoint(n) : whole
144 }
145
146 return ENTITIES[code.toLowerCase()] ?? whole
147 })
148}
149
150export function stripToText(html: string): string {
151 return decode(
152 html
153 .replace(/<!--[\s\S]*?-->/g, ' ')
154 .replace(/<(script|style|template|svg|head)\b[\s\S]*?<\/\1>/gi, ' ')
155 .replace(/<[^>]+>/g, ' '),
156 )
157 .replace(/\s+/g, ' ')
158 .trim()
159}
160
161export function countWords(text: string): number {
162 return text === '' ? 0 : text.split(' ').filter(w => /[\p{L}\p{N}]/u.test(w)).length
163}
164
165function attrs(tag: string): Record<string, string> {
166 const out: Record<string, string> = {}
167 for (const m of tag.matchAll(/([a-zA-Z:-]+)\s*=\s*("([^"]*)"|'([^']*)'|([^\s>]+))/g)) {
168 const name = m[1]
169 if (name !== undefined) out[name.toLowerCase()] = decode(m[3] ?? m[4] ?? m[5] ?? '')
170 }
171
172 return out
173}
174
175function metaTags(html: string): Record<string, string>[] {
176 return [...html.matchAll(/<meta\b[^>]*>/gi)].map(m => attrs(m[0]))
177}
178
179function metaContent(html: string, name: string): string | undefined {
180 const tag = metaTags(html).find(a => (a.name ?? a.property ?? '').toLowerCase() === name)
181
182 return tag?.content
183}
184
185export type Heading = { level: number; text: string; index: number; end: number }
186
187export function headings(html: string): Heading[] {
188 return [...html.matchAll(/<h([1-6])\b[^>]*>([\s\S]*?)<\/h\1>/gi)].map(m => ({
189 level: Number(m[1]),
190 text: stripToText(m[2] ?? ''),
191 index: m.index ?? 0,
192 end: (m.index ?? 0) + m[0].length,
193 }))
194}
195
196export function jsonLdTypes(html: string): { types: string[]; broken: number } {
197 const types = new Set<string>()
198 let broken = 0
199 const visit = (node: unknown): void => {
200 if (Array.isArray(node)) return node.forEach(visit)
201 if (node === null || typeof node !== 'object') return
202 const record = node as Record<string, unknown>
203 const t = record['@type']
204 if (typeof t === 'string') types.add(t)
205 if (Array.isArray(t)) t.filter(x => typeof x === 'string').forEach(x => types.add(x))
206 if (record['@graph'] !== undefined) visit(record['@graph'])
207 }
208 for (const m of html.matchAll(/<script\b[^>]*type\s*=\s*["']?application\/ld\+json["']?[^>]*>([\s\S]*?)<\/script>/gi)) {
209 try {
210 visit(JSON.parse(m[1] ?? ''))
211 } catch {
212 broken += 1
213 }
214 }
215
216 return { types: [...types], broken }
217}
218
219// Ends in a question mark, ignoring decoration after it (an accordion's "+" icon).
220function isQuestion(text: string): boolean {
221 return /\?[^\p{L}\p{N}]*$/u.test(text.trim())
222}
223
224// Removes elements marked aria-hidden: icons and decoration, not content.
225function dropHidden(html: string): string {
226 return html.replace(/<(span|i|svg)\b[^>]*aria-hidden\s*=\s*["']?true["']?[^>]*>[\s\S]*?<\/\1>/gi, ' ')
227}
228
229// Question headings, and whether the text right under each one answers it briefly.
230export function answerSections(html: string): { questions: number; answered: number } {
231 const list = headings(html).filter(h => h.level >= 2 && h.level <= 4)
232 let questions = 0
233 let answered = 0
234 list.forEach((h, i) => {
235 if (!isQuestion(stripToText(dropHidden(html.slice(h.index, h.end))))) return
236 questions += 1
237 const nextStart = list[i + 1]?.index ?? html.length
238 const firstParagraph = /<p\b[^>]*>([\s\S]*?)<\/p>/i.exec(html.slice(h.end, nextStart))
239 const words = countWords(stripToText(firstParagraph?.[1] ?? ''))
240 // A direct answer an engine can lift: one paragraph, roughly 20 to 90 words.
241 if (words >= 20 && words <= 90) answered += 1
242 })
243
244 // FAQ accordions: <details><summary>Question?</summary>answer</details>. Collapsed
245 // answers are still in the raw HTML, so crawlers read them.
246 for (const m of html.matchAll(/<details\b[^>]*>([\s\S]*?)<\/details>/gi)) {
247 const inner = m[1] ?? ''
248 const summary = /<summary\b[^>]*>([\s\S]*?)<\/summary>/i.exec(inner)
249 if (summary === null || !isQuestion(stripToText(dropHidden(summary[1] ?? '')))) continue
250 questions += 1
251 const words = countWords(stripToText(inner.replace(summary[0], ' ')))
252 if (words >= 20 && words <= 120) answered += 1
253 }
254
255 return { questions, answered }
256}
257
258function isChallengePage(page: Fetched): boolean {
259 return (
260 page.headers['cf-mitigated'] === 'challenge' ||
261 /<title>\s*(just a moment|attention required)/i.test(page.text) ||
262 /cf-chl-|challenge-platform/i.test(page.text)
263 )
264}
265
266// ---------- the audit ----------
267
268const WEIGHT = { critical: 3, normal: 1, none: 0 } as const
269
270export function audit(input: AuditInput): Report {
271 const checks: Check[] = []
272 const add = (check: Check) => checks.push(check)
273 const { page } = input
274 const html = page.text
275 const target = new URL(input.url)
276 const path = `${target.pathname}${target.search}`
277
278 // Crawler access: robots.txt, per engine.
279 // RFC 9309: a 4xx robots.txt means no restrictions; a 5xx (or 429) means crawl nothing.
280 const robotsStatus = input.robots?.status
281 const robotsDown = robotsStatus !== undefined && (robotsStatus >= 500 || robotsStatus === 429)
282 const robotsMissing = robotsStatus === undefined || (robotsStatus >= 400 && robotsStatus < 500 && robotsStatus !== 429)
283 const robots = parseRobots(robotsMissing || robotsDown || input.robots === undefined ? '' : input.robots.text)
284
285 if (robotsDown) {
286 add({
287 id: 'robots-5xx',
288 group: 'Crawler access',
289 label: 'robots.txt reachable',
290 status: 'fail',
291 detail: `robots.txt answered ${input.robots?.status}. Google treats a server error on robots.txt as "do not crawl anything" until it recovers.`,
292 fix: 'Make /robots.txt return 200 (or 404 if you have none).',
293 source: 'https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt',
294 weight: WEIGHT.critical,
295 })
296 }
297
298 for (const bot of SEARCH_BOTS) {
299 const verdict = robotsAllows(robots, bot.token, path)
300 const by = verdict.matchedAgent === undefined ? 'no matching group' : `group "${verdict.matchedAgent}"`
301 add({
302 id: `robots-${bot.token}`,
303 group: 'Crawler access',
304 label: `${bot.engine} can crawl (${bot.token})`,
305 status: robotsDown ? 'fail' : verdict.isAllowed ? 'pass' : 'fail',
306 detail: robotsMissing
307 ? 'No robots.txt, so nothing is blocked.'
308 : verdict.isAllowed
309 ? `Allowed by ${by}.`
310 : `Blocked by ${by}: "Disallow: ${verdict.rule?.pattern}" in robots.txt as served to this audit. ${bot.engine} cannot use this page unless the site serves verified crawlers a different file.`,
311 fix: verdict.isAllowed ? undefined : `Add "User-agent: ${bot.token}" with "Allow: /" to robots.txt, or remove the Disallow rule.`,
312 source: bot.source,
313 weight: WEIGHT.critical,
314 })
315 }
316
317 const trainingBlocked = TRAINING_BOTS.filter(t => !robotsAllows(robots, t, path).isAllowed)
318 add({
319 id: 'robots-training',
320 group: 'Crawler access',
321 label: 'Training-only crawlers',
322 status: 'info',
323 detail:
324 trainingBlocked.length === 0
325 ? `All allowed (${TRAINING_BOTS.join(', ')}).`
326 : `Blocked: ${trainingBlocked.join(', ')}. This only opts out of model training. It does not remove the site from AI search answers.`,
327 source: 'https://developers.openai.com/api/docs/bots',
328 weight: WEIGHT.none,
329 })
330
331 // Firewall: does the page answer an AI crawler user agent the way it answers a browser?
332 const browserRefused = page.status >= 400 || isChallengePage(page)
333 const blockedProbes = Object.entries(input.botProbes).filter(
334 ([, probe]) => probe !== undefined && page.status < 400 && (probe.status >= 400 || isChallengePage(probe)),
335 )
336 add({
337 id: 'firewall',
338 group: 'Crawler access',
339 label: 'Firewall lets AI crawlers through',
340 status: browserRefused ? 'info' : blockedProbes.length === 0 ? 'pass' : 'warn',
341 detail: browserRefused
342 ? 'Not comparable: the site refused the browser request too.'
343 : blockedProbes.length === 0
344 ? `Same response for ${Object.keys(input.botProbes).join(' and ')} user agents as for a browser.`
345 : `${blockedProbes.map(([name, p]) => `${name} got ${p?.status}${p !== undefined && isChallengePage(p) ? ' (challenge page)' : ''}`).join(', ')} while a browser got ${page.status}. A CDN or firewall rule may be blocking AI crawlers. This test sends the user agent only, and real crawlers are verified by IP, so confirm in your CDN's bot settings.`,
346 fix: blockedProbes.length === 0 ? undefined : 'Check your CDN or WAF AI-bot setting (Cloudflare: Security > Bots) and allow the search crawlers you want.',
347 weight: browserRefused ? WEIGHT.none : WEIGHT.normal,
348 })
349
350 // Page: status and indexing directives.
351 const isWalled = page.status === 401 || page.status === 403 || page.status === 429 || isChallengePage(page)
352 const isReadable = page.status >= 200 && page.status < 300 && !isWalled
353 add({
354 id: 'status',
355 group: 'Page',
356 label: 'Page loads',
357 status: isReadable ? 'pass' : isWalled ? 'warn' : 'fail',
358 detail: isReadable
359 ? `HTTP ${page.status}.`
360 : isWalled
361 ? `HTTP ${page.status}${isChallengePage(page) ? ', bot challenge page' : ''}: the site refused this audit's request. It may still serve real crawlers, so the content checks below were skipped rather than guessed.`
362 : `HTTP ${page.status}. Nothing on this URL can be cited.`,
363 fix: isReadable || isWalled ? undefined : 'Serve the page with a 200 status.',
364 weight: isWalled ? WEIGHT.none : WEIGHT.critical,
365 })
366
367 if (!isReadable) return finish(input.url, checks, robots, input, true)
368
369 const directives = [page.headers['x-robots-tag'] ?? '', metaContent(html, 'robots') ?? '', metaContent(html, 'googlebot') ?? '']
370 .join(',')
371 .toLowerCase()
372 const noindex = /\b(noindex|none)\b/.test(directives)
373 const nosnippet = /\bnosnippet\b/.test(directives)
374 const maxSnippet = /max-snippet\s*:\s*(-?\d+)/.exec(directives)
375 const snippetLimit = maxSnippet?.[1] === undefined ? undefined : Number(maxSnippet[1])
376 const snippetBlocked = nosnippet || snippetLimit === 0
377 add({
378 id: 'directives',
379 group: 'Page',
380 label: 'No noindex or nosnippet',
381 status: noindex || snippetBlocked ? 'fail' : snippetLimit !== undefined && snippetLimit > 0 && snippetLimit < 50 ? 'warn' : 'pass',
382 detail: noindex
383 ? 'The page says noindex. Search engines, and the AI answers built on them, will drop it.'
384 : snippetBlocked
385 ? 'nosnippet or max-snippet:0 is set. Google uses these to keep the page out of AI Overviews and AI Mode.'
386 : snippetLimit !== undefined && snippetLimit > 0 && snippetLimit < 50
387 ? `max-snippet:${snippetLimit} limits how much text an answer can quote.`
388 : 'No blocking robots directives in the meta tags or X-Robots-Tag header.',
389 fix: noindex || snippetBlocked ? 'Remove noindex / nosnippet / max-snippet:0 from the robots meta tag and X-Robots-Tag header.' : undefined,
390 source: 'https://developers.google.com/search/docs/appearance/ai-features',
391 weight: WEIGHT.critical,
392 })
393
394 const title = stripToText(/<title\b[^>]*>([\s\S]*?)<\/title>/i.exec(html)?.[1] ?? '')
395 add({
396 id: 'title',
397 group: 'Page',
398 label: 'Title',
399 status: title === '' ? 'warn' : 'pass',
400 detail: title === '' ? 'No <title>.' : `"${title}"`,
401 fix: title === '' ? 'Add a <title> that names what the page is about.' : undefined,
402 weight: WEIGHT.normal,
403 })
404
405 const description = metaContent(html, 'description') ?? ''
406 add({
407 id: 'description',
408 group: 'Page',
409 label: 'Meta description',
410 status: description === '' ? 'warn' : 'pass',
411 detail: description === '' ? 'No meta description.' : `${description.length} characters.`,
412 fix: description === '' ? 'Add a one-sentence description of what the page offers and who it is for.' : undefined,
413 weight: WEIGHT.normal,
414 })
415
416 const canonical = [...html.matchAll(/<link\b[^>]*>/gi)].map(m => attrs(m[0])).find(a => (a.rel ?? '').toLowerCase() === 'canonical')?.href
417 add({
418 id: 'canonical',
419 group: 'Page',
420 label: 'Canonical URL',
421 status: canonical === undefined ? 'warn' : 'pass',
422 detail: canonical === undefined ? 'No rel=canonical link.' : canonical,
423 fix: canonical === undefined ? 'Add <link rel="canonical"> pointing at the preferred URL.' : undefined,
424 weight: WEIGHT.normal,
425 })
426
427 // Content: what a crawler that does not run JavaScript sees.
428 const words = countWords(stripToText(html))
429 const scripts = (html.match(/<script\b/gi) ?? []).length
430 add({
431 id: 'raw-text',
432 group: 'Content',
433 label: 'Text visible without JavaScript',
434 status: words < 50 ? 'fail' : words < 150 ? 'warn' : 'pass',
435 detail: `${words} words in the raw HTML (${scripts} script tags). OpenAI's, Anthropic's and Perplexity's crawlers read the raw HTML and do not run JavaScript (Vercel, Dec 2024), so text that only appears after scripts run is invisible to them.`,
436 fix: words < 150 ? 'Server-render or pre-render the main content so it is in the HTML response.' : undefined,
437 source: 'https://vercel.com/blog/the-rise-of-the-ai-crawler',
438 weight: WEIGHT.critical,
439 })
440
441 const h1s = headings(html).filter(h => h.level === 1)
442 add({
443 id: 'h1',
444 group: 'Content',
445 label: 'One clear H1',
446 status: h1s.length === 1 ? 'pass' : 'warn',
447 detail: h1s.length === 0 ? 'No H1 in the raw HTML.' : h1s.length === 1 ? `"${h1s[0]?.text}"` : `${h1s.length} H1s.`,
448 fix: h1s.length === 1 ? undefined : 'Use exactly one H1 that states what the page is.',
449 weight: WEIGHT.normal,
450 })
451
452 const sections = answerSections(html)
453 add({
454 id: 'answers',
455 group: 'Content',
456 label: 'Answer-shaped sections (heuristic)',
457 status: sections.answered > 0 ? 'pass' : 'warn',
458 detail:
459 sections.questions === 0
460 ? 'No question headings or FAQ items. Engines match pages to questions; a question with a short direct answer under it is easy to quote.'
461 : `${sections.questions} questions (headings or FAQ items), ${sections.answered} followed by a short direct answer.`,
462 fix: sections.answered > 0 ? undefined : 'Add a few headings phrased as the questions buyers ask, each followed by a two to three sentence answer.',
463 weight: WEIGHT.normal,
464 })
465
466 return finish(input.url, checks, robots, input, false, html)
467}
468
469function finish(
470 url: string,
471 checks: Check[],
472 robots: ReturnType<typeof parseRobots>,
473 input: AuditInput,
474 isPartial: boolean,
475 html?: string,
476): Report {
477 const add = (check: Check) => checks.push(check)
478
479 // Extras: useful context, but not requirements.
480 if (html !== undefined) {
481 const schema = jsonLdTypes(html)
482 add({
483 id: 'schema',
484 group: 'Extras',
485 label: 'Structured data (JSON-LD)',
486 status: schema.broken > 0 ? 'warn' : 'info',
487 detail:
488 (schema.types.length === 0 ? 'None found.' : `Types: ${schema.types.join(', ')}.`) +
489 (schema.broken > 0 ? ` ${schema.broken} JSON-LD block(s) fail to parse.` : '') +
490 ' Google says AI Overviews need no special schema; it still tells engines who you are.',
491 fix: schema.broken > 0 ? 'Fix the JSON syntax in the broken JSON-LD block(s).' : undefined,
492 source: 'https://developers.google.com/search/docs/appearance/ai-features',
493 weight: schema.broken > 0 ? WEIGHT.normal : WEIGHT.none,
494 })
495 }
496
497 const sitemapListed = robots.sitemaps.length > 0
498 const sitemapFound = input.sitemap !== undefined && input.sitemap.status === 200
499 const lastmods = (input.sitemap?.text.match(/<lastmod>/gi) ?? []).length
500 add({
501 id: 'sitemap',
502 group: 'Extras',
503 label: 'Sitemap',
504 status: sitemapFound || sitemapListed ? 'pass' : 'warn',
505 detail: sitemapFound
506 ? `${sitemapListed ? 'Listed in robots.txt' : 'Found at /sitemap.xml'}; ${lastmods} lastmod dates in the first file.`
507 : sitemapListed
508 ? `Listed in robots.txt (${robots.sitemaps.length}); the first one refused this audit's request.`
509 : 'No sitemap found in robots.txt or at /sitemap.xml.',
510 fix: sitemapFound || sitemapListed ? undefined : 'Publish a sitemap.xml with lastmod dates and list it in robots.txt.',
511 weight: WEIGHT.normal,
512 })
513
514 add({
515 id: 'llms',
516 group: 'Extras',
517 label: 'llms.txt',
518 status: 'info',
519 detail: `${input.llms?.status === 200 ? 'Present' : 'Not present'}. A proposed standard; no major answer engine has said it reads it.`,
520 source: 'https://llmstxt.org',
521 weight: WEIGHT.none,
522 })
523
524 const scored = checks.filter(c => c.weight > 0)
525 const total = scored.reduce((sum, c) => sum + c.weight, 0)
526 const earned = scored.reduce((sum, c) => sum + (c.status === 'pass' ? c.weight : c.status === 'warn' ? c.weight / 2 : 0), 0)
527 const counts: Record<Status, number> = { pass: 0, warn: 0, fail: 0, info: 0 }
528 checks.forEach(c => (counts[c.status] += 1))
529
530 return { url, score: total === 0 ? 0 : Math.round((earned / total) * 100), isPartial, counts, checks }
531}
532
533export function toMarkdown(report: Report, at: string): string {
534 const mark: Record<Status, string> = { pass: 'PASS', warn: 'WARN', fail: 'FAIL', info: 'INFO' }
535 const lines = [
536 `# AEO audit: ${report.url}`,
537 '',
538 `Run ${at}. Score ${report.score}/100 (this audit's checks only; not a prediction of citations).${report.isPartial ? ' Partial: the page refused this audit, so content checks were skipped.' : ''}`,
539 `${report.counts.fail} fail, ${report.counts.warn} warn, ${report.counts.pass} pass, ${report.counts.info} info.`,
540 ]
541 for (const group of ['Crawler access', 'Page', 'Content', 'Extras'] as const) {
542 lines.push('', `## ${group}`, '')
543 for (const c of report.checks.filter(x => x.group === group)) {
544 lines.push(`- **${mark[c.status]}** ${c.label}: ${c.detail}`)
545 if (c.fix !== undefined) lines.push(` - Fix: ${c.fix}`)
546 if (c.source !== undefined) lines.push(` - Source: ${c.source}`)
547 }
548 }
549
550 return lines.join('\n')
551}
552types/index.d.ts 35 lines1// Self-contained: the analyzer and the hooks module import these from here.
2export type Status = 'pass' | 'warn' | 'fail' | 'info'
3
4export type Check = {
5 id: string
6 group: 'Crawler access' | 'Page' | 'Content' | 'Extras'
7 label: string
8 status: Status
9 detail: string
10 fix?: string
11 source?: string
12 // How much this check counts toward the score; 0 for info-only checks.
13 weight: number
14}
15
16export type Report = {
17 url: string
18 score: number
19 // True when the page itself could not be read, so content checks were skipped.
20 isPartial: boolean
21 counts: Record<Status, number>
22 checks: Check[]
23}
24
25export type SavedReport = Report & { at: string; path: string | null }
26
27declare module 'claude-code' {
28 interface PluginState {
29 'aeo-audit': {
30 last: SavedReport | null
31 running: string | null
32 }
33 }
34}
35