Field guide

The AEO audit: check your AI visibility in an afternoon

The complete manual method: build the question set, run it across the engines, record what comes back, and turn the findings into a ranked work queue. No tools required.

7 chapters · 6 min read · Updated Aug 25, 2026

What does an AEO audit tell you?

One afternoon of disciplined checking produces the document most brands have never seen: the true state of their category inside the engines their buyers use, with names attached. No score from a tool, no proxy metric, but the answers themselves: which brands each engine names for each buying question, which sources it leaned on, and where you are absent. Everything downstream in Answer Engine Optimisation, what to publish, which platforms to earn presence on, which facts to fix, is prioritised by this document, which is why the audit comes before any spend.

The method below needs no software beyond the engines and a spreadsheet. It is the manual version of the reading MentionOS runs continuously as an agent, and doing it by hand once is the fastest way to understand what the discipline is for. Budget two to three hours for a first pass on one brand.

Step one: build the question set

The five question shapes

The audit lives or dies on asking what buyers ask, and buyer questions come in five recurring shapes. Category-open questions: "best [category] for [situation]", the purest shortlist request. Problem-first questions: "[problem] recommendations", where the buyer has not named your category yet and the engine chooses it for them. Validation questions: "is [your brand] legit", "is [your brand] worth it", asked by buyers who found you elsewhere and are checking. Competitive questions: "[rival] alternatives", "[your brand] vs [rival]", the highest-intent shape of all. And instructional questions: "how to [job your product does]", where brands get named inside advice.

Write ten to fifteen questions across those shapes, in buyer language instead of yours. The test for buyer language is unforgiving: if the question contains a term your industry uses and your customer does not, rewrite it. Your support inbox, sales calls and review text are the primary sources; the questions buyers ask you in person are the questions they ask engines when you are not in the room.

Where to steal the phrasing

Three free sources sharpen the set. Google's own People-Also-Ask boxes for your category terms show question phrasing with proven demand. Your Search Console query report shows the words real searchers already use to reach you, with the standing caveat that its tables are incomplete by design, anonymised queries are omitted, so treat it as a sample, never a census. And Reddit threads in your category are a corpus of unfiltered buyer phrasing that the engines themselves demonstrably read.

Step two: run the questions properly

Hygiene that keeps the data honest

Method decides whether you measured the engines or measured your own shadow. Use fresh chats for every question, because context bleeds inside a session and contaminates the next answer. Run logged out where the engine allows it, or at minimum on an account with no history of searching your own brand, personalisation is real and it flatters you. Ask each question once per engine per pass, verbatim, resisting the urge to rephrase toward the answer you want. And date-stamp every run, because answers drift and an undated record is trivia.

One honest caveat belongs in the record itself: engine answers vary run to run, sometimes visibly. A single pass is a reading, and three passes spread across a week is data. Treat anything seen once as provisional until a second pass confirms it.

The engines and what to capture

As of late 2026 the core set is ChatGPT, Gemini and Perplexity, plus Google's AI Overviews checked through ordinary searches, and Claude where your buyers skew technical. For every question on every engine, capture the same six fields: the question, the engine, whether your brand was named, which brands were named instead and in what order, which sources the answer cited, and the date. Perplexity cites by default and Google's AI Overviews list their links; ChatGPT shows sources when it browses. When an answer cites nothing, record that too, an uncited answer is running on training memory, which tells you which pool you are contesting.

The audit record: one row per question per engine, capturing named or not, who was named instead, sources cited, and the date

Step three: read what you recorded

The four patterns that matter

Finished records repeat the same four findings in every category we have watched. Absence: the engine names rivals and not you, the default state for most brands and the clearest work order. Wrong-room presence: you are named, but third or fourth, behind brands you outsell in the real market, which usually signals thin corroboration rather than thin product. Mischaracterisation: you are named with stale prices, dead products or a description you retired, an entity-consistency debt being repaid with interest. And generalist slots: the cited sources are Reddit, YouTube and big publishers with no category specialist among them, which is what an unsettled category looks like and is the single most actionable pattern on the list, because unsettled slots are winnable.

Scoring without kidding yourself

Two numbers summarise a pass honestly. Named-rate: the share of question-engine pairs where your brand appears at all. Source-overlap: the share of cited sources where your brand has any presence. Resist inventing a weighted composite score; at audit scale the two raw numbers and the four patterns carry all the signal, and precision theatre on twelve questions is self-deception. What the numbers are for is the delta: the same two numbers next month, same questions, same method, is when the document becomes a trendline.

Step four: turn findings into a work queue

Each finding maps to a workstream, and the mapping is nearly mechanical. Absence on a question maps to content: a page answering that exact question, plus presence on whichever sources that answer cited. Wrong-room presence maps to third-party corroboration: reviews, comparisons and communities, since the engines already know you exist and want reasons to trust you more. Mischaracterisation maps to entity work: correcting facts everywhere they appear, your own site first, and structured data plus an llms.txt so machines get the facts from you directly. Generalist slots map to specialist content ambition: the definitive category piece those slots are waiting for, the play our own data says is open in most young categories.

Order the queue by intent, competitive and validation questions before category-open ones, because the buyer asking "[rival] alternatives" is closest to money. Cap the first cycle at five items; an audit that produces a forty-item backlog produces nothing.

Step five: make it repeat, or make it someone's job

The audit's value compounds only on repetition, and repetition is where manual breaks. Same questions, same hygiene, monthly: the drift between passes shows which work moved answers, which rival is climbing, and which source a model newly trusts. In practice the monthly re-run is the first casualty of a busy quarter, which is the honest case for automating the reading: this exact loop, question set, multi-engine runs, recorded answers, source maps, drift, is what MentionOS operates continuously as the operating Agent for AEO, with the work queue turned into proposals it executes on your approval. Hand-run the first audit regardless; you will read every tool and agent in this category more sharply once you have felt the manual version.

Frequently asked questions

Is there a tool that can check my AI visibility score?

Several products measure AI visibility, and the honest sequence is audit first, tool second. The manual afternoon shows you exactly what any tool must capture to be useful, questions, naming, sources, drift, and makes vendor claims easy to evaluate. The category's options are compared honestly, our own included, in the best rated AEO tools.

How do I check what ChatGPT says about my brand?

Open a fresh chat, ask the validation questions from your set, "is [brand] worth it", "[brand] review", "best [category]", and record whether you are named, what is said, and which sources are cited when browsing is on. Repeat on a second day before treating any single answer as the state of things; run-to-run variance is normal.

How many questions do I need for a useful audit?

Ten to fifteen, spread across the five shapes, beats fifty scattershot ones. The constraint is your reading discipline: every question multiplies across engines and passes, and a set you will genuinely re-run monthly is worth more than an exhaustive one you run once.

How often should I re-run the audit?

Monthly is the working default; weekly on the competitive questions if your category is contested. Answers drift with model updates, retrieval changes and competitor publishing, so the interval is a bet on how fast your category moves. The asset is the trendline across passes; a single pass is only its first point.

The window is now

Read the blueprint. Or run the agent.