What happens between the question and the answer?
A pipeline runs, in under a second, and every stage of it is contestable. The buyer types a question. The engine draws on what its model learned in training and, for most commercial questions, on what its retrieval system fetches from the live web at that moment. It weighs which fetched sources deserve trust, composes a response, and names the brands that survived the whole process, often with citations attached. AEO is the practice of showing up well at each stage: known to the model, fetchable by the retriever, trusted as a source, and clean enough to name.
This guide walks the pipeline stage by stage, because the work only makes sense once the machine does. Definitions and the wider discipline live in the complete AEO guide; this is the mechanism.
Stage one: which pool answers, training or retrieval?

The training pool
Part of every answer comes from the model itself: patterns absorbed from the corpus it was trained on, months before your buyer asked anything. This pool decides the engine's background sense of a category, which brands feel canonical, which claims feel settled. It cannot be edited, petitioned or purchased; it shifts only when providers ship updated models, and it rewards one thing over time: having been widely and consistently written about before the training cut.
The retrieval pool
The half you can move this quarter. For fresh, commercial or specific questions, engines search the live web, fetch candidate pages and feed them to the model as working material. Perplexity leans hardest on this pattern, and as of late 2026 ChatGPT's browsing behaviour and Google's AI surfaces run their own versions of it, each with different appetites; these behaviours change without notice, which is itself a planning fact. What matters strategically is the asymmetry of clocks: retrieval can reflect a well-corroborated new page within weeks, while the training pool moves on release cycles. New brands win in retrieval first, and sustained retrieval presence is how a brand eventually soaks into training.
Stage two: how do engines decide which sources to trust?
Engines corroborate before they compose, and the trust decision is where most brands silently lose. A fetched page becomes a cited source when it is parseable, current, specific to the question, and consistent with what other fetched pages say. In practice the citation slots skew heavily third-party. In our August 2026 category checks, Google's AI Overviews cited Reddit, YouTube, HubSpot, Coursera and Forbes again and again, established generalists occupying a young category's slots because no specialist had earned them yet.
Two implications follow. First, your brand's presence on the platforms engines already trust, review sites, comparison pages, active communities, is not optional garnish; it is the raw material of corroboration. Second, the slots are winnable: a default incumbent citing generalists is what an unsettled category looks like, and specialists that publish original, checkable material displace defaults as categories mature. Google states the standard from its own side: original information, first-hand experience and evidence over commodity summaries, for search and generative surfaces alike.
Stage three: how does a brand become a nameable entity?
Resolution
Before an engine can name you, it has to resolve you: connecting the mentions, pages, profiles and facts scattered across the web into one entity it is confident about. Every consistent appearance lowers the cost of that resolution; every contradiction raises it. A brand whose name, description, product facts and claims agree everywhere is cheap to corroborate. A brand with three descriptions and stale facts forces the engine to reconcile, and engines composing under time pressure skip what they cannot cheaply verify.
The machinery of clarity
The mechanical aids are unglamorous and documented. Valid structured data gives engines explicit clues about the entities on a page, an eligibility signal, never a guaranteed outcome. Visible authorship and sourcing support trust the same way. An llms.txt file hands engines your brand's facts in a form built for machine consumption. None of this is optimisation theatre; it is lowering the cost of being believed.
Stage four: why do technically-invisible pages never enter the race?
Because the pipeline can only weigh what it can fetch and parse, and official guidance is blunt about the failure modes. Search and answer systems can only use content they can find, crawl and understand. Not every AI crawler executes JavaScript, so client-rendered content is invisible to some fetchers regardless of quality; server-rendering what matters is the safe cross-crawler default. Conflicting canonical signals make engines guess which version of a page is real. Indexing follows the mobile version of a page, and Core Web Vitals thresholds, paint within 2.5 seconds, interaction under 200 milliseconds, layout shift under 0.1, define the experience bar. Every one of these is a gate before the trust contest even begins: the best answer on the web, unrendered, is no answer at all.
Stage five: how does shipped work become a changed answer?
The loop, honestly timed
Work enters the pipeline in a sequence you can plan around. A new page, once crawlable, can join the retrieval pool within days to weeks. Citations arrive when the page proves specific and corroborated; naming follows when the entity behind it resolves cleanly. Movement shows first on retrieval-heavy engines and freshest questions, later on training-heavy ones. Anyone promising a fixed timeline is guessing, but the ordering is dependable, which is why the discipline front-loads fetchability and third-party corroboration.
Measuring without fooling yourself
The scoreboard is the answers themselves, read on a schedule, beside first-party data. Search Console shows queries and impressions with known gaps, its tables omit anonymised queries, so an empty row is missing data, never proof of zero demand. Analytics ties movement to outcomes, with attribution settings deciding how credit lands; a single last-click report proves little alone. Third-party tool scores are modelled estimates and should be labelled that way. The honest loop reads answers, reads first-party data, ships the next piece of work, and repeats, which is precisely the loop MentionOS runs as an agent: reading ChatGPT, Claude, Gemini, Google AI and Perplexity continuously, shipping approved work to your own site, and reporting movement with a receipt for every action.
The whole journey, end to end
Follow one thread and the mechanism stops being abstract. A brand's buyers start asking assistants a buying question the brand has never answered anywhere. The engines compose from what exists: a rival's comparison page, a Reddit thread, a generalist explainer, and the brand is absent. The brand notices, because it reads the answers on a schedule. It publishes a page that answers the exact question with evidence and dates, server-rendered and marked up; it earns a presence on the review platform the answers keep citing; it aligns its product facts everywhere they appear. Weeks pass. The page enters retrieval, gets cited on the freshest engines first, and the entity resolves cleanly enough that a shortlist finally includes the name. The answer changed because every stage of the pipeline was fed deliberately. That is AEO working, and the practical method for running it yourself starts in how to get ChatGPT to recommend your brand.
Frequently asked questions
Does ChatGPT use live web data or only training data?
Both, and the mix depends on the question. Settled general knowledge often comes from training; fresh, commercial and specific questions trigger retrieval from the live web. The practical consequence for brands: new work reaches retrieval-backed answers first, and training-pool presence follows from sustained, widely-corroborated visibility over time.
Why does AI recommend some brands over others?
Because those brands are cheaper to verify at every stage of the pipeline: fetchable pages, specific answers, corroboration on trusted third-party sources, and a consistent entity that resolves without contradictions. Quality matters, but checkability decides, which is why excellent products with fragmented online identities go unnamed.
How long does it take for AEO work to show up in answers?
The ordering is dependable even though no fixed timeline exists: retrieval-backed answers can reflect well-corroborated new pages within weeks, citations precede naming, and training-pool change arrives on model-update cycles. Contested categories move slower than open ones.
Can I see which sources an AI engine used for an answer?
Often, and it is the most useful reading you can do. Perplexity cites by default, Google's AI Overviews list source links, and ChatGPT shows sources when browsing. Recording which domains recur across your category's answers produces the map of where your brand needs presence, the exact exercise the audit guide in this series walks through.