SLOPSHOPPER

aeo-audit

Audits a URL for AI answer engines (ChatGPT search, Perplexity, Google AI Overviews, Copilot, Claude): crawler access, firewall, indexing directives, no-JS…

newpaneguardcommandstatustool
v0.1.0MITupdated 2026-10-06jbauman-26/aeo-audit-mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · aeo-audit
│ ┃ AEO audit ✕ › fix the failing auth test and add an audit log call │ ┃ No audit yet. Run /aeo <url>. │ ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /aeo │ ⎿ aeo-audit: Usage: /aeo <url> │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · AEO audit
No audit yet. Run /aeo <url>.
README

aeo-audit

A free Claude Code mod that checks whether ChatGPT search, Perplexity, Google AI Overviews, Microsoft Copilot and Claude can get into a web page and read it.

Type /aeo yoursite.com. A pane opens with every check marked PASS, WARN, FAIL or INFO, the fix for anything that failed, and the source each rule comes from. The full report is saved as markdown.

Built for the video If I Wanted ChatGPT to Recommend My Business, I'd Run This Audit First by Jake Bauman. [LINK TO VIDEO]

Install

You need Claude Code 2.1.287 or newer. The mod API is early access, so a future Claude Code release may change it.

git clone https://github.com/jbauman-26/aeo-audit-mod.git ~/aeo-audit-mod

Load it for one session:

claude --plugin-dir ~/aeo-audit-mod

Load it in every session, including the Claude desktop app: add this to the env block of ~/.claude/settings.json, then restart Claude Code.

"CLAUDE_CODE_PLUGIN_DIRS": "~/aeo-audit-mod"

Use

  • /aeo https://example.com/pricing audits that page and opens the results pane.
  • /aeo on its own reopens the last report.
  • Or ask Claude in plain words: "audit example.com for AI search and fix what fails." Claude calls the mod's audit_site tool, reads the report and writes the fixes.
  • In the pane, Fix with Claude (f) drafts that request for you, and Rerun (r) checks again after you change something.

Reports are saved to ~/output/aeo-audits/<host>-<date>.md.

What it checks

CheckWhy it mattersSource
robots.txt allows each search crawler: OAI-SearchBot, PerplexityBot, Googlebot, Bingbot, Claude-SearchBotBlocking a search crawler removes you from that engine's answers. Blocking only training crawlers (GPTBot, Google-Extended, ClaudeBot, CCBot) does not, so those are reported but not scored.OpenAI, Perplexity, Google, Bing, Anthropic
Firewall treats AI crawler user agents the same as a browserA CDN or firewall rule can block AI crawlers before robots.txt is read.
No noindex, nosnippet or max-snippet:0Google uses these to keep a page out of AI Overviews and AI Mode.Google
Text visible without JavaScriptOpenAI's, Anthropic's and Perplexity's crawlers read the raw HTML and do not run JavaScript.Vercel, Dec 2024
Page status, title, meta description, canonical URL, one H1The basics every engine reads first.
Questions with short, direct answers under them (headings or FAQ accordions)A heuristic: answer engines match pages to questions, and a short direct answer is easy to quote.
Schema (JSON-LD), sitemap, llms.txtListed, not scored, except broken JSON-LD and a missing sitemap. Google says AI features need no special schema, and no major engine has said it reads llms.txt.Google, llmstxt.org

What the score is not

The score covers this audit's checks only. It is not a prediction that any engine will cite you. Passing means the doors are open and the page is readable; whether an engine recommends you depends on your content, your reputation and what other sites say about you.

Three honest limits:

  • robots.txt is checked as served to this audit. Some large sites serve verified crawlers a different file.
  • The firewall check sends the crawler's user agent, not its IP address. Real crawlers are verified by IP, so a failure means "check your CDN's bot settings," not proof.
  • If a site refuses the audit itself (a 403 or a bot challenge), the content checks are skipped and the report says it is partial, rather than scoring the block page.

Develop

claude plugin validate ~/aeo-audit-mod
claude plugin test ~/aeo-audit-mod

The checks live in hooks/audit.ts as plain functions with no Claude Code dependency, so you can run them from any TypeScript runtime. hooks/register.tsx wires them to the /aeo command, the audit_site tool, the pane and the status line.

License

MIT. Use it, change it, ship it to clients.

Source 3 files
hooks/register.tsx 253 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { SavedReport } from '../types'
5import {
6  BOT_USER_AGENTS,
7  BROWSER_USER_AGENT,
8  audit,
9  sitemapUrl,
10  toMarkdown,
11  type Check,
12  type Fetched,
13  type Status,
14} from './audit'
15
16type $ = EngineInterface
17
18const PANE = 'aeo-audit'
19const last = atom({ plugin: 'aeo-audit', key: 'last' } as const, null as SavedReport | null)
20const running = atom({ plugin: 'aeo-audit', key: 'running' } as const, null as string | null)
21
22const COLOR: Record<Status, string | undefined> = { pass: 'green', warn: 'yellow', fail: 'red', info: undefined }
23const MARK: Record<Status, string> = { pass: 'PASS', warn: 'WARN', fail: 'FAIL', info: 'INFO' }
24
25export function normalizeUrl(raw: string): string | undefined {
26  const trimmed = raw.trim()
27  if (trimmed === '') return undefined
28  try {
29    const url = new URL(/^https?:\/\//i.test(trimmed) ? trimmed : `https://${trimmed}`)
30
31    return url.hostname.includes('.') ? url.toString() : undefined
32  } catch {
33    return undefined
34  }
35}
36
37// Follows redirects itself so the audit reads the page a crawler lands on.
38async function get($: $, url: string, userAgent = BROWSER_USER_AGENT): Promise<Fetched | undefined> {
39  let current = url
40  try {
41    for (let hop = 0; hop < 5; hop += 1) {
42      const res = await $.http.fetch(current, { headers: { 'user-agent': userAgent, accept: 'text/html,*/*' } })
43      const location = res.headers.location
44      if (res.status >= 300 && res.status < 400 && location !== undefined) {
45        current = new URL(location, current).toString()
46        continue
47      }
48
49      return { status: res.status, headers: res.headers, text: res.text }
50    }
51  } catch {
52    return undefined
53  }
54
55  return undefined
56}
57
58async function runAudit($: $, url: string): Promise<SavedReport> {
59  const origin = new URL(url).origin
60  await update($, running, () => url)
61  $.ui.status(`AEO: auditing ${new URL(url).hostname}...`)
62  try {
63    const robots = await get($, `${origin}/robots.txt`)
64    const robotsText = robots?.status === 200 ? robots.text : undefined
65    const probeNames = Object.keys(BOT_USER_AGENTS)
66    const [page, llms, sitemap, ...probes] = await Promise.all([
67      get($, url),
68      get($, `${origin}/llms.txt`),
69      get($, sitemapUrl(robotsText, origin)),
70      ...probeNames.map(name => get($, url, BOT_USER_AGENTS[name])),
71    ])
72    const report = audit({
73      url,
74      page: page ?? { status: 0, headers: {}, text: '' },
75      robots,
76      llms,
77      sitemap,
78      botProbes: Object.fromEntries(probeNames.map((name, i) => [name, probes[i]])),
79    })
80
81    const at = new Date(await $.clock.now()).toISOString()
82    const day = at.slice(0, 10)
83    const home = await $.env.get('HOME')
84    const file = `${new URL(url).hostname}-${day}.md`
85    let path: string | null = null
86    if (home !== undefined) {
87      path = `${home}/output/aeo-audits/${file}`
88      try {
89        await $.fs.write(path, `${toMarkdown(report, day)}\n`)
90      } catch {
91        path = null
92      }
93    }
94
95    const saved: SavedReport = { ...report, at, path }
96    await update($, last, () => saved)
97    await $.store.set('last', saved)
98    $.ui.status(`AEO: ${new URL(url).hostname} ${report.score}/100${report.isPartial ? ' (partial)' : ''}`)
99
100    return saved
101  } finally {
102    await update($, running, () => null)
103  }
104}
105
106function summary(report: SavedReport): string {
107  const top = report.checks.filter(c => c.status === 'fail').map(c => `- ${c.label}`)
108
109  return [
110    `${report.url}: ${report.score}/100${report.isPartial ? ' (partial, the page refused the audit)' : ''}.`,
111    `${report.counts.fail} fail, ${report.counts.warn} warn, ${report.counts.pass} pass.`,
112    ...(top.length > 0 ? ['Failing:', ...top] : []),
113    report.path === null ? 'Report not saved.' : `Report: ${report.path}`,
114  ].join('\n')
115}
116
117export const register: Register = on => {
118  on('session.start', async ($, e, next) => {
119    await $.command.register({
120      name: 'aeo',
121      description: 'Audit a URL for AI answer engines (ChatGPT, Perplexity, AI Overviews, Copilot, Claude)',
122      argumentHint: '<url>',
123    })
124    await $.tool.register({
125      name: 'audit_site',
126      description:
127        'Audits one URL for AI answer-engine readiness: robots.txt access for each AI search crawler, firewall behaviour toward AI user agents, noindex/nosnippet, text visible without JavaScript, title, description, canonical, H1, question-shaped sections, schema, sitemap, llms.txt. Returns a markdown report with sources and fixes, and saves it under ~/output/aeo-audits/.',
128      inputSchema: {
129        type: 'object',
130        properties: { url: { type: 'string', description: 'The page to audit, e.g. https://example.com/pricing' } },
131        required: ['url'],
132      },
133    })
134    if ((await read($, last)) === null) {
135      const stored = await $.store.get('last')
136      if (stored !== undefined && stored !== null) await update($, last, () => stored as SavedReport)
137    }
138
139    return next(e)
140  })
141
142  on('command.run', { command: 'aeo' }, async ($, e) => {
143    if (e.args.trim() === '') {
144      const opened = await $.ui.open({ id: PANE, title: 'AEO audit' })
145
146      return { text: (await read($, last)) === null ? 'Usage: /aeo <url>' : opened.isPlaced ? 'AEO pane opened.' : `AEO pane waiting: ${opened.reason}` }
147    }
148    const url = normalizeUrl(e.args)
149    if (url === undefined) return { text: `Not a URL: ${e.args.trim()}` }
150    void $.ui.open({ id: PANE, title: 'AEO audit' })
151    const report = await runAudit($, url)
152
153    return { text: summary(report) }
154  })
155
156  on('tool.call', { tool: 'mcp__aeo-audit__audit_site' }, async ($, e) => {
157    // A plugin tool's arguments arrive on the event itself, beside `tool` and `tool_use_id`.
158    const url = normalizeUrl(String((e as { url?: unknown }).url ?? ''))
159    if (url === undefined) return { deny: 'audit_site needs a full URL such as https://example.com/' }
160    const report = await runAudit($, url)
161
162    return { result: `${toMarkdown(report, report.at.slice(0, 10))}\n\n${report.path === null ? '' : `Saved: ${report.path}`}` }
163  })
164
165  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
166    const { Box, Text, Button } = $.ui.resolve(e)
167    const report = await read($, last)
168    const busy = await read($, running)
169
170    if (report === null) {
171      return (
172        <Box flexDirection="column">
173          <Text>{busy === null ? 'No audit yet. Run /aeo <url>.' : `Auditing ${busy}...`}</Text>
174        </Box>
175      )
176    }
177
178    const scoreColor = report.score >= 80 ? 'green' : report.score >= 50 ? 'yellow' : 'red'
179    const row = (c: Check) => (
180      <Box key={c.id} flexDirection="column">
181        <Text wrap="truncate-end">
182          <Text color={COLOR[c.status]} bold={c.status === 'fail'}>
183            {MARK[c.status]}
184          </Text>{' '}
185          {c.label}
186        </Text>
187        {(c.status === 'fail' || c.status === 'warn') && (
188          <Text dimColor wrap="wrap">
189            {'     '}
190            {c.fix ?? c.detail}
191          </Text>
192        )}
193      </Box>
194    )
195
196    return (
197      <Box flexDirection="column">
198        <Text bold wrap="truncate-end">
199          {report.url}
200        </Text>
201        <Text>
202          <Text color={scoreColor} bold>
203            {report.score}/100
204          </Text>
205          <Text dimColor>
206            {'  '}
207            {report.counts.fail} fail, {report.counts.warn} warn, {report.counts.pass} pass
208            {report.isPartial ? ', partial' : ''}
209          </Text>
210        </Text>
211        {busy !== null && <Text color="yellow">Re-auditing {busy}...</Text>}
212        {(['Crawler access', 'Page', 'Content', 'Extras'] as const).map(group => {
213          const rows = report.checks.filter(c => c.group === group)
214          if (rows.length === 0) return null
215
216          return (
217            <Box key={group} flexDirection="column" marginTop={1}>
218              <Text bold>{group}</Text>
219              {rows.map(row)}
220            </Box>
221          )
222        })}
223        <Box marginTop={1}>
224          <Button
225            key="fix"
226            label="Fix with Claude"
227            hotkey="f"
228            variant="primary"
229            onPress={async () => {
230              const live = await read($, last)
231              if (live === null) return
232              const where = live.path === null ? 'the AEO audit in this session' : `the AEO audit at ${live.path}`
233              await $.prompt.fill({
234                text: `Read ${where} for ${live.url} and write the exact fixes for every FAIL and WARN, most impact first. Quote the source for each.`,
235              })
236            }}
237          />
238          <Text> </Text>
239          <Button
240            key="rerun"
241            label="Rerun"
242            hotkey="r"
243            onPress={async () => {
244              const live = await read($, last)
245              if (live !== null) await runAudit($, live.url)
246            }}
247          />
248        </Box>
249      </Box>
250    )
251  })
252}
253
hooks/audit.ts 552 lines
1// Pure analysis: takes what was fetched, returns findings. No network, no engine,
2// so the same code runs in the mod, in tests, and from a plain `bun` script.
3
4import type { Check, Report, Status } from '../types'
5
6export type { Check, Report, Status }
7
8export type Fetched = { status: number; headers: Record<string, string>; text: string }
9
10export type AuditInput = {
11  url: string
12  page: Fetched
13  robots: Fetched | undefined
14  llms: Fetched | undefined
15  sitemap: Fetched | undefined
16  // The same page requested with AI crawler user agents, keyed by crawler name.
17  botProbes: Record<string, Fetched | undefined>
18}
19
20
21// Crawlers that decide whether a site can appear in an AI answer engine at all.
22export const SEARCH_BOTS = [
23  {
24    token: 'OAI-SearchBot',
25    engine: 'ChatGPT search',
26    source: 'https://developers.openai.com/api/docs/bots',
27  },
28  {
29    token: 'PerplexityBot',
30    engine: 'Perplexity',
31    source: 'https://docs.perplexity.ai/guides/bots',
32  },
33  {
34    token: 'Googlebot',
35    engine: 'Google AI Overviews and AI Mode',
36    source: 'https://developers.google.com/search/docs/appearance/ai-features',
37  },
38  { token: 'Bingbot', engine: 'Microsoft Copilot (Bing index)', source: 'https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0' },
39  { token: 'Claude-SearchBot', engine: 'Claude search', source: 'https://support.claude.com/en/articles/8896518' },
40] as const
41
42// Crawlers that only collect training data. Blocking them is a policy choice, not a visibility problem.
43export const TRAINING_BOTS = ['GPTBot', 'Google-Extended', 'ClaudeBot', 'CCBot'] as const
44
45export const BOT_USER_AGENTS: Record<string, string> = {
46  'OAI-SearchBot': 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot',
47  PerplexityBot: 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)',
48}
49
50export const BROWSER_USER_AGENT =
51  'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36'
52
53// ---------- robots.txt (RFC 9309 group selection, longest-match rules) ----------
54
55type Rule = { allow: boolean; pattern: string }
56type Group = { agents: string[]; rules: Rule[] }
57
58export function parseRobots(text: string): { groups: Group[]; sitemaps: string[] } {
59  const groups: Group[] = []
60  const sitemaps: string[] = []
61  let current: Group | undefined
62  let lastWasAgent = false
63
64  for (const raw of text.split(/\r?\n/)) {
65    const line = raw.replace(/#.*$/, '').trim()
66    const colon = line.indexOf(':')
67    if (colon === -1) continue
68    const key = line.slice(0, colon).trim().toLowerCase()
69    const value = line.slice(colon + 1).trim()
70
71    if (key === 'sitemap') {
72      if (value !== '') sitemaps.push(value)
73      continue
74    }
75    if (key === 'user-agent') {
76      if (current === undefined || !lastWasAgent) {
77        current = { agents: [], rules: [] }
78        groups.push(current)
79      }
80      current.agents.push(value.toLowerCase())
81      lastWasAgent = true
82      continue
83    }
84    if ((key === 'allow' || key === 'disallow') && current !== undefined) {
85      // An empty Disallow blocks nothing, so it adds no rule.
86      if (value !== '') current.rules.push({ allow: key === 'allow', pattern: value })
87      lastWasAgent = false
88    }
89  }
90
91  return { groups, sitemaps }
92}
93
94function ruleMatches(pattern: string, path: string): boolean {
95  const anchored = pattern.endsWith('$')
96  const body = (anchored ? pattern.slice(0, -1) : pattern)
97    .split('*')
98    .map(part => part.replace(/[.+?^${}()|[\]\\]/g, '\\$&'))
99    .join('.*')
100
101  return new RegExp(`^${body}${anchored ? '$' : ''}`).test(path)
102}
103
104// Whether `token` may fetch `path`, and which group decided it.
105export function robotsAllows(
106  robots: ReturnType<typeof parseRobots>,
107  token: string,
108  path: string,
109): { isAllowed: boolean; matchedAgent: string | undefined; rule: Rule | undefined } {
110  const lower = token.toLowerCase()
111  let groups = robots.groups.filter(g => g.agents.includes(lower))
112  let matchedAgent: string | undefined = groups.length > 0 ? token : undefined
113  if (groups.length === 0) {
114    groups = robots.groups.filter(g => g.agents.includes('*'))
115    matchedAgent = groups.length > 0 ? '*' : undefined
116  }
117
118  let best: Rule | undefined
119  for (const rule of groups.flatMap(g => g.rules)) {
120    if (!ruleMatches(rule.pattern, path)) continue
121    const isLonger = best === undefined || rule.pattern.length > best.pattern.length
122    const isTieAllow = best !== undefined && rule.pattern.length === best.pattern.length && rule.allow
123    if (isLonger || isTieAllow) best = rule
124  }
125
126  return { isAllowed: best === undefined || best.allow, matchedAgent, rule: best }
127}
128
129// The sitemap to fetch: the first one robots.txt lists, else /sitemap.xml.
130export function sitemapUrl(robotsText: string | undefined, origin: string): string {
131  return parseRobots(robotsText ?? '').sitemaps[0] ?? `${origin}/sitemap.xml`
132}
133
134// ---------- HTML helpers (regex, no DOM in the mod's environment) ----------
135
136const ENTITIES: Record<string, string> = { amp: '&', lt: '<', gt: '>', quot: '"', apos: "'", nbsp: ' ' }
137
138export function decode(text: string): string {
139  return text.replace(/&(#x[0-9a-f]+|#\d+|[a-z]+);/gi, (whole, code: string) => {
140    if (code[0] === '#') {
141      const n = code[1] === 'x' || code[1] === 'X' ? parseInt(code.slice(2), 16) : parseInt(code.slice(1), 10)
142
143      return Number.isFinite(n) ? String.fromCodePoint(n) : whole
144    }
145
146    return ENTITIES[code.toLowerCase()] ?? whole
147  })
148}
149
150export function stripToText(html: string): string {
151  return decode(
152    html
153      .replace(/<!--[\s\S]*?-->/g, ' ')
154      .replace(/<(script|style|template|svg|head)\b[\s\S]*?<\/\1>/gi, ' ')
155      .replace(/<[^>]+>/g, ' '),
156  )
157    .replace(/\s+/g, ' ')
158    .trim()
159}
160
161export function countWords(text: string): number {
162  return text === '' ? 0 : text.split(' ').filter(w => /[\p{L}\p{N}]/u.test(w)).length
163}
164
165function attrs(tag: string): Record<string, string> {
166  const out: Record<string, string> = {}
167  for (const m of tag.matchAll(/([a-zA-Z:-]+)\s*=\s*("([^"]*)"|'([^']*)'|([^\s>]+))/g)) {
168    const name = m[1]
169    if (name !== undefined) out[name.toLowerCase()] = decode(m[3] ?? m[4] ?? m[5] ?? '')
170  }
171
172  return out
173}
174
175function metaTags(html: string): Record<string, string>[] {
176  return [...html.matchAll(/<meta\b[^>]*>/gi)].map(m => attrs(m[0]))
177}
178
179function metaContent(html: string, name: string): string | undefined {
180  const tag = metaTags(html).find(a => (a.name ?? a.property ?? '').toLowerCase() === name)
181
182  return tag?.content
183}
184
185export type Heading = { level: number; text: string; index: number; end: number }
186
187export function headings(html: string): Heading[] {
188  return [...html.matchAll(/<h([1-6])\b[^>]*>([\s\S]*?)<\/h\1>/gi)].map(m => ({
189    level: Number(m[1]),
190    text: stripToText(m[2] ?? ''),
191    index: m.index ?? 0,
192    end: (m.index ?? 0) + m[0].length,
193  }))
194}
195
196export function jsonLdTypes(html: string): { types: string[]; broken: number } {
197  const types = new Set<string>()
198  let broken = 0
199  const visit = (node: unknown): void => {
200    if (Array.isArray(node)) return node.forEach(visit)
201    if (node === null || typeof node !== 'object') return
202    const record = node as Record<string, unknown>
203    const t = record['@type']
204    if (typeof t === 'string') types.add(t)
205    if (Array.isArray(t)) t.filter(x => typeof x === 'string').forEach(x => types.add(x))
206    if (record['@graph'] !== undefined) visit(record['@graph'])
207  }
208  for (const m of html.matchAll(/<script\b[^>]*type\s*=\s*["']?application\/ld\+json["']?[^>]*>([\s\S]*?)<\/script>/gi)) {
209    try {
210      visit(JSON.parse(m[1] ?? ''))
211    } catch {
212      broken += 1
213    }
214  }
215
216  return { types: [...types], broken }
217}
218
219// Ends in a question mark, ignoring decoration after it (an accordion's "+" icon).
220function isQuestion(text: string): boolean {
221  return /\?[^\p{L}\p{N}]*$/u.test(text.trim())
222}
223
224// Removes elements marked aria-hidden: icons and decoration, not content.
225function dropHidden(html: string): string {
226  return html.replace(/<(span|i|svg)\b[^>]*aria-hidden\s*=\s*["']?true["']?[^>]*>[\s\S]*?<\/\1>/gi, ' ')
227}
228
229// Question headings, and whether the text right under each one answers it briefly.
230export function answerSections(html: string): { questions: number; answered: number } {
231  const list = headings(html).filter(h => h.level >= 2 && h.level <= 4)
232  let questions = 0
233  let answered = 0
234  list.forEach((h, i) => {
235    if (!isQuestion(stripToText(dropHidden(html.slice(h.index, h.end))))) return
236    questions += 1
237    const nextStart = list[i + 1]?.index ?? html.length
238    const firstParagraph = /<p\b[^>]*>([\s\S]*?)<\/p>/i.exec(html.slice(h.end, nextStart))
239    const words = countWords(stripToText(firstParagraph?.[1] ?? ''))
240    // A direct answer an engine can lift: one paragraph, roughly 20 to 90 words.
241    if (words >= 20 && words <= 90) answered += 1
242  })
243
244  // FAQ accordions: <details><summary>Question?</summary>answer</details>. Collapsed
245  // answers are still in the raw HTML, so crawlers read them.
246  for (const m of html.matchAll(/<details\b[^>]*>([\s\S]*?)<\/details>/gi)) {
247    const inner = m[1] ?? ''
248    const summary = /<summary\b[^>]*>([\s\S]*?)<\/summary>/i.exec(inner)
249    if (summary === null || !isQuestion(stripToText(dropHidden(summary[1] ?? '')))) continue
250    questions += 1
251    const words = countWords(stripToText(inner.replace(summary[0], ' ')))
252    if (words >= 20 && words <= 120) answered += 1
253  }
254
255  return { questions, answered }
256}
257
258function isChallengePage(page: Fetched): boolean {
259  return (
260    page.headers['cf-mitigated'] === 'challenge' ||
261    /<title>\s*(just a moment|attention required)/i.test(page.text) ||
262    /cf-chl-|challenge-platform/i.test(page.text)
263  )
264}
265
266// ---------- the audit ----------
267
268const WEIGHT = { critical: 3, normal: 1, none: 0 } as const
269
270export function audit(input: AuditInput): Report {
271  const checks: Check[] = []
272  const add = (check: Check) => checks.push(check)
273  const { page } = input
274  const html = page.text
275  const target = new URL(input.url)
276  const path = `${target.pathname}${target.search}`
277
278  // Crawler access: robots.txt, per engine.
279  // RFC 9309: a 4xx robots.txt means no restrictions; a 5xx (or 429) means crawl nothing.
280  const robotsStatus = input.robots?.status
281  const robotsDown = robotsStatus !== undefined && (robotsStatus >= 500 || robotsStatus === 429)
282  const robotsMissing = robotsStatus === undefined || (robotsStatus >= 400 && robotsStatus < 500 && robotsStatus !== 429)
283  const robots = parseRobots(robotsMissing || robotsDown || input.robots === undefined ? '' : input.robots.text)
284
285  if (robotsDown) {
286    add({
287      id: 'robots-5xx',
288      group: 'Crawler access',
289      label: 'robots.txt reachable',
290      status: 'fail',
291      detail: `robots.txt answered ${input.robots?.status}. Google treats a server error on robots.txt as "do not crawl anything" until it recovers.`,
292      fix: 'Make /robots.txt return 200 (or 404 if you have none).',
293      source: 'https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt',
294      weight: WEIGHT.critical,
295    })
296  }
297
298  for (const bot of SEARCH_BOTS) {
299    const verdict = robotsAllows(robots, bot.token, path)
300    const by = verdict.matchedAgent === undefined ? 'no matching group' : `group "${verdict.matchedAgent}"`
301    add({
302      id: `robots-${bot.token}`,
303      group: 'Crawler access',
304      label: `${bot.engine} can crawl (${bot.token})`,
305      status: robotsDown ? 'fail' : verdict.isAllowed ? 'pass' : 'fail',
306      detail: robotsMissing
307        ? 'No robots.txt, so nothing is blocked.'
308        : verdict.isAllowed
309          ? `Allowed by ${by}.`
310          : `Blocked by ${by}: "Disallow: ${verdict.rule?.pattern}" in robots.txt as served to this audit. ${bot.engine} cannot use this page unless the site serves verified crawlers a different file.`,
311      fix: verdict.isAllowed ? undefined : `Add "User-agent: ${bot.token}" with "Allow: /" to robots.txt, or remove the Disallow rule.`,
312      source: bot.source,
313      weight: WEIGHT.critical,
314    })
315  }
316
317  const trainingBlocked = TRAINING_BOTS.filter(t => !robotsAllows(robots, t, path).isAllowed)
318  add({
319    id: 'robots-training',
320    group: 'Crawler access',
321    label: 'Training-only crawlers',
322    status: 'info',
323    detail:
324      trainingBlocked.length === 0
325        ? `All allowed (${TRAINING_BOTS.join(', ')}).`
326        : `Blocked: ${trainingBlocked.join(', ')}. This only opts out of model training. It does not remove the site from AI search answers.`,
327    source: 'https://developers.openai.com/api/docs/bots',
328    weight: WEIGHT.none,
329  })
330
331  // Firewall: does the page answer an AI crawler user agent the way it answers a browser?
332  const browserRefused = page.status >= 400 || isChallengePage(page)
333  const blockedProbes = Object.entries(input.botProbes).filter(
334    ([, probe]) => probe !== undefined && page.status < 400 && (probe.status >= 400 || isChallengePage(probe)),
335  )
336  add({
337    id: 'firewall',
338    group: 'Crawler access',
339    label: 'Firewall lets AI crawlers through',
340    status: browserRefused ? 'info' : blockedProbes.length === 0 ? 'pass' : 'warn',
341    detail: browserRefused
342      ? 'Not comparable: the site refused the browser request too.'
343      : blockedProbes.length === 0
344        ? `Same response for ${Object.keys(input.botProbes).join(' and ')} user agents as for a browser.`
345        : `${blockedProbes.map(([name, p]) => `${name} got ${p?.status}${p !== undefined && isChallengePage(p) ? ' (challenge page)' : ''}`).join(', ')} while a browser got ${page.status}. A CDN or firewall rule may be blocking AI crawlers. This test sends the user agent only, and real crawlers are verified by IP, so confirm in your CDN's bot settings.`,
346    fix: blockedProbes.length === 0 ? undefined : 'Check your CDN or WAF AI-bot setting (Cloudflare: Security > Bots) and allow the search crawlers you want.',
347    weight: browserRefused ? WEIGHT.none : WEIGHT.normal,
348  })
349
350  // Page: status and indexing directives.
351  const isWalled = page.status === 401 || page.status === 403 || page.status === 429 || isChallengePage(page)
352  const isReadable = page.status >= 200 && page.status < 300 && !isWalled
353  add({
354    id: 'status',
355    group: 'Page',
356    label: 'Page loads',
357    status: isReadable ? 'pass' : isWalled ? 'warn' : 'fail',
358    detail: isReadable
359      ? `HTTP ${page.status}.`
360      : isWalled
361        ? `HTTP ${page.status}${isChallengePage(page) ? ', bot challenge page' : ''}: the site refused this audit's request. It may still serve real crawlers, so the content checks below were skipped rather than guessed.`
362        : `HTTP ${page.status}. Nothing on this URL can be cited.`,
363    fix: isReadable || isWalled ? undefined : 'Serve the page with a 200 status.',
364    weight: isWalled ? WEIGHT.none : WEIGHT.critical,
365  })
366
367  if (!isReadable) return finish(input.url, checks, robots, input, true)
368
369  const directives = [page.headers['x-robots-tag'] ?? '', metaContent(html, 'robots') ?? '', metaContent(html, 'googlebot') ?? '']
370    .join(',')
371    .toLowerCase()
372  const noindex = /\b(noindex|none)\b/.test(directives)
373  const nosnippet = /\bnosnippet\b/.test(directives)
374  const maxSnippet = /max-snippet\s*:\s*(-?\d+)/.exec(directives)
375  const snippetLimit = maxSnippet?.[1] === undefined ? undefined : Number(maxSnippet[1])
376  const snippetBlocked = nosnippet || snippetLimit === 0
377  add({
378    id: 'directives',
379    group: 'Page',
380    label: 'No noindex or nosnippet',
381    status: noindex || snippetBlocked ? 'fail' : snippetLimit !== undefined && snippetLimit > 0 && snippetLimit < 50 ? 'warn' : 'pass',
382    detail: noindex
383      ? 'The page says noindex. Search engines, and the AI answers built on them, will drop it.'
384      : snippetBlocked
385        ? 'nosnippet or max-snippet:0 is set. Google uses these to keep the page out of AI Overviews and AI Mode.'
386        : snippetLimit !== undefined && snippetLimit > 0 && snippetLimit < 50
387          ? `max-snippet:${snippetLimit} limits how much text an answer can quote.`
388          : 'No blocking robots directives in the meta tags or X-Robots-Tag header.',
389    fix: noindex || snippetBlocked ? 'Remove noindex / nosnippet / max-snippet:0 from the robots meta tag and X-Robots-Tag header.' : undefined,
390    source: 'https://developers.google.com/search/docs/appearance/ai-features',
391    weight: WEIGHT.critical,
392  })
393
394  const title = stripToText(/<title\b[^>]*>([\s\S]*?)<\/title>/i.exec(html)?.[1] ?? '')
395  add({
396    id: 'title',
397    group: 'Page',
398    label: 'Title',
399    status: title === '' ? 'warn' : 'pass',
400    detail: title === '' ? 'No <title>.' : `"${title}"`,
401    fix: title === '' ? 'Add a <title> that names what the page is about.' : undefined,
402    weight: WEIGHT.normal,
403  })
404
405  const description = metaContent(html, 'description') ?? ''
406  add({
407    id: 'description',
408    group: 'Page',
409    label: 'Meta description',
410    status: description === '' ? 'warn' : 'pass',
411    detail: description === '' ? 'No meta description.' : `${description.length} characters.`,
412    fix: description === '' ? 'Add a one-sentence description of what the page offers and who it is for.' : undefined,
413    weight: WEIGHT.normal,
414  })
415
416  const canonical = [...html.matchAll(/<link\b[^>]*>/gi)].map(m => attrs(m[0])).find(a => (a.rel ?? '').toLowerCase() === 'canonical')?.href
417  add({
418    id: 'canonical',
419    group: 'Page',
420    label: 'Canonical URL',
421    status: canonical === undefined ? 'warn' : 'pass',
422    detail: canonical === undefined ? 'No rel=canonical link.' : canonical,
423    fix: canonical === undefined ? 'Add <link rel="canonical"> pointing at the preferred URL.' : undefined,
424    weight: WEIGHT.normal,
425  })
426
427  // Content: what a crawler that does not run JavaScript sees.
428  const words = countWords(stripToText(html))
429  const scripts = (html.match(/<script\b/gi) ?? []).length
430  add({
431    id: 'raw-text',
432    group: 'Content',
433    label: 'Text visible without JavaScript',
434    status: words < 50 ? 'fail' : words < 150 ? 'warn' : 'pass',
435    detail: `${words} words in the raw HTML (${scripts} script tags). OpenAI's, Anthropic's and Perplexity's crawlers read the raw HTML and do not run JavaScript (Vercel, Dec 2024), so text that only appears after scripts run is invisible to them.`,
436    fix: words < 150 ? 'Server-render or pre-render the main content so it is in the HTML response.' : undefined,
437    source: 'https://vercel.com/blog/the-rise-of-the-ai-crawler',
438    weight: WEIGHT.critical,
439  })
440
441  const h1s = headings(html).filter(h => h.level === 1)
442  add({
443    id: 'h1',
444    group: 'Content',
445    label: 'One clear H1',
446    status: h1s.length === 1 ? 'pass' : 'warn',
447    detail: h1s.length === 0 ? 'No H1 in the raw HTML.' : h1s.length === 1 ? `"${h1s[0]?.text}"` : `${h1s.length} H1s.`,
448    fix: h1s.length === 1 ? undefined : 'Use exactly one H1 that states what the page is.',
449    weight: WEIGHT.normal,
450  })
451
452  const sections = answerSections(html)
453  add({
454    id: 'answers',
455    group: 'Content',
456    label: 'Answer-shaped sections (heuristic)',
457    status: sections.answered > 0 ? 'pass' : 'warn',
458    detail:
459      sections.questions === 0
460        ? 'No question headings or FAQ items. Engines match pages to questions; a question with a short direct answer under it is easy to quote.'
461        : `${sections.questions} questions (headings or FAQ items), ${sections.answered} followed by a short direct answer.`,
462    fix: sections.answered > 0 ? undefined : 'Add a few headings phrased as the questions buyers ask, each followed by a two to three sentence answer.',
463    weight: WEIGHT.normal,
464  })
465
466  return finish(input.url, checks, robots, input, false, html)
467}
468
469function finish(
470  url: string,
471  checks: Check[],
472  robots: ReturnType<typeof parseRobots>,
473  input: AuditInput,
474  isPartial: boolean,
475  html?: string,
476): Report {
477  const add = (check: Check) => checks.push(check)
478
479  // Extras: useful context, but not requirements.
480  if (html !== undefined) {
481  const schema = jsonLdTypes(html)
482  add({
483    id: 'schema',
484    group: 'Extras',
485    label: 'Structured data (JSON-LD)',
486    status: schema.broken > 0 ? 'warn' : 'info',
487    detail:
488      (schema.types.length === 0 ? 'None found.' : `Types: ${schema.types.join(', ')}.`) +
489      (schema.broken > 0 ? ` ${schema.broken} JSON-LD block(s) fail to parse.` : '') +
490      ' Google says AI Overviews need no special schema; it still tells engines who you are.',
491    fix: schema.broken > 0 ? 'Fix the JSON syntax in the broken JSON-LD block(s).' : undefined,
492    source: 'https://developers.google.com/search/docs/appearance/ai-features',
493    weight: schema.broken > 0 ? WEIGHT.normal : WEIGHT.none,
494  })
495  }
496
497  const sitemapListed = robots.sitemaps.length > 0
498  const sitemapFound = input.sitemap !== undefined && input.sitemap.status === 200
499  const lastmods = (input.sitemap?.text.match(/<lastmod>/gi) ?? []).length
500  add({
501    id: 'sitemap',
502    group: 'Extras',
503    label: 'Sitemap',
504    status: sitemapFound || sitemapListed ? 'pass' : 'warn',
505    detail: sitemapFound
506      ? `${sitemapListed ? 'Listed in robots.txt' : 'Found at /sitemap.xml'}; ${lastmods} lastmod dates in the first file.`
507      : sitemapListed
508        ? `Listed in robots.txt (${robots.sitemaps.length}); the first one refused this audit's request.`
509        : 'No sitemap found in robots.txt or at /sitemap.xml.',
510    fix: sitemapFound || sitemapListed ? undefined : 'Publish a sitemap.xml with lastmod dates and list it in robots.txt.',
511    weight: WEIGHT.normal,
512  })
513
514  add({
515    id: 'llms',
516    group: 'Extras',
517    label: 'llms.txt',
518    status: 'info',
519    detail: `${input.llms?.status === 200 ? 'Present' : 'Not present'}. A proposed standard; no major answer engine has said it reads it.`,
520    source: 'https://llmstxt.org',
521    weight: WEIGHT.none,
522  })
523
524  const scored = checks.filter(c => c.weight > 0)
525  const total = scored.reduce((sum, c) => sum + c.weight, 0)
526  const earned = scored.reduce((sum, c) => sum + (c.status === 'pass' ? c.weight : c.status === 'warn' ? c.weight / 2 : 0), 0)
527  const counts: Record<Status, number> = { pass: 0, warn: 0, fail: 0, info: 0 }
528  checks.forEach(c => (counts[c.status] += 1))
529
530  return { url, score: total === 0 ? 0 : Math.round((earned / total) * 100), isPartial, counts, checks }
531}
532
533export function toMarkdown(report: Report, at: string): string {
534  const mark: Record<Status, string> = { pass: 'PASS', warn: 'WARN', fail: 'FAIL', info: 'INFO' }
535  const lines = [
536    `# AEO audit: ${report.url}`,
537    '',
538    `Run ${at}. Score ${report.score}/100 (this audit's checks only; not a prediction of citations).${report.isPartial ? ' Partial: the page refused this audit, so content checks were skipped.' : ''}`,
539    `${report.counts.fail} fail, ${report.counts.warn} warn, ${report.counts.pass} pass, ${report.counts.info} info.`,
540  ]
541  for (const group of ['Crawler access', 'Page', 'Content', 'Extras'] as const) {
542    lines.push('', `## ${group}`, '')
543    for (const c of report.checks.filter(x => x.group === group)) {
544      lines.push(`- **${mark[c.status]}** ${c.label}: ${c.detail}`)
545      if (c.fix !== undefined) lines.push(`  - Fix: ${c.fix}`)
546      if (c.source !== undefined) lines.push(`  - Source: ${c.source}`)
547    }
548  }
549
550  return lines.join('\n')
551}
552
types/index.d.ts 35 lines
1// Self-contained: the analyzer and the hooks module import these from here.
2export type Status = 'pass' | 'warn' | 'fail' | 'info'
3
4export type Check = {
5  id: string
6  group: 'Crawler access' | 'Page' | 'Content' | 'Extras'
7  label: string
8  status: Status
9  detail: string
10  fix?: string
11  source?: string
12  // How much this check counts toward the score; 0 for info-only checks.
13  weight: number
14}
15
16export type Report = {
17  url: string
18  score: number
19  // True when the page itself could not be read, so content checks were skipped.
20  isPartial: boolean
21  counts: Record<Status, number>
22  checks: Check[]
23}
24
25export type SavedReport = Report & { at: string; path: string | null }
26
27declare module 'claude-code' {
28  interface PluginState {
29    'aeo-audit': {
30      last: SavedReport | null
31      running: string | null
32    }
33  }
34}
35