The pipeline Retrieval Selection Trust Evidence grid Test it yourself FAQ
AI Search · Pillar Guide

How AI Search Engines Choose Which Sources to Cite

Citation is not one decision. It is a pipeline — retrieve, select a passage, decide whether that passage can stand alone with attribution. Most advice targets the wrong stage. Here is each factor, sorted by how much evidence it actually has.

Many candidate web pages narrowing through a funnel into a shortlist, then into two cited results marked with ticks — the retrieval, selection and trust stages of AI citation
Citation is a three-stage pipeline. Most advice targets the middle stage while the first one is quietly failing.

Quick share

Useful for anyone planning a quarter of AI visibility work.

Get a summary from AI

Short on time? Open this article in an answer engine and have it summarised for you.

The short version
  • Citation is a pipeline, not a decision. Retrieval decides who is eligible; passage selection decides who is quoted; a trust check decides whose name goes on it.
  • Stage one is ordinary SEO. If you can't be retrieved, nothing downstream matters — and most "GEO tactics" target stage two while stage one is broken.
  • Citability is the underrated factor. A claim that survives being lifted out of your page, alone, is what gets lifted.
  • Original data is the strongest structural advantage available to a small site, because it cannot be sourced from a competitor.
  • Freshness is conditional, not universal — it dominates on volatile queries and barely registers on stable ones.

Ask why a page got cited and you'll get answers pitched at wildly different levels: "it ranks well," "it's authoritative," "it's structured for AI." All three can be true simultaneously and none explains the mechanism.

What helps is separating the stages. A page must first be retrieved, then have a passage selected, then survive a trust check before its name appears under an answer. The factors that govern each stage are different, and most published advice targets the middle one while the first is quietly failing.

I've sorted every factor below by how much evidence supports it — documented by a platform, observed consistently, or still assumption. Where I'm inferring, I say so.

1. The citation pipeline

Every system in this category — AI Overviews, AI Mode, ChatGPT search, Perplexity, Claude search — runs some version of retrieval-augmented generation. The shared shape is:

The pipeline

query → intent classification → fan-out into sub-queries → retrieval → passage extraction → synthesis → attribution

Different systems weight the stages differently, but no major one generates commercial answers purely from parametric memory. They fetch, then write.

Two consequences fall straight out of that shape:

  1. Retrieval is a gate, not a factor. Fail it and the remaining twelve factors are irrelevant. This is why crawler access and rendering outrank everything marketed as GEO.
  2. The unit of competition is the passage, not the page. A mediocre page containing one exceptionally clear paragraph will out-cite an excellent page whose every claim depends on three paragraphs of setup.

2. Stage 1 — retrieval

Query intent

Evidence: documented. Intent determines whether an AI answer appears at all, how many sources it draws on, and what kind of source it favours. Broadly:

IntentTypical citation behaviourFavours
InformationalSeveral sources, synthesisedDepth, structure, explainers
ComparativeMultiple sources, often tabulatedGenuine comparison content, specs
CommercialFewer sources, more brand and product dataProduct pages, feeds, reviews
LocalBusiness profiles and directoriesBusiness Profile, NAP consistency
NavigationalRarely generates an AI answer

The practical implication is that a single "AI visibility" number is close to meaningless. Visibility varies by intent class, and you should be measuring the classes that map to your revenue.

Existing rankings

Evidence: documented for Google, strongly implied elsewhere. Google describes its generative features as rooted in core Search ranking and quality systems. Retrieval runs on the same infrastructure — so classic ranking is the substrate.

This is the single most important claim in the article and the one most resisted, because it implies AI visibility is mostly earned by conventional SEO. The industry has a commercial interest in it being a new discipline. The documentation says it isn't.

The nuance that rescues the interesting part: because of query fan-out, the ranking that matters may not be for your target keyword. A system decomposing "best CRM for a 12-person agency" into eight sub-queries retrieves against all eight. You can be cited off a page that ranks for an adjacent sub-question and nothing else. That rewards topical coverage over exact-match optimisation — a real strategic difference, arrived at without pretending ranking is irrelevant.

Crawler access and rendering

Evidence: documented. A binary gate, and the most common own-goal I find. Block OAI-SearchBot and you are absent from ChatGPT search regardless of quality. Render your prices client-side and a fetcher-class crawler receives a page with no prices in it.

I've covered both in the crawler policy guide and the agent-readiness guide. If you check nothing else from this article, check these.

3. Stage 2 — passage selection

Once retrieved, your page competes to supply a specific passage. This is where content decisions start to matter, and where most of the durable advantage lives.

Passage relevance

Evidence: strongly observed. Systems extract at passage level, not document level. What is scored is a chunk of your page against a sub-query, which means a long page can win a citation on a narrow question if one section answers it cleanly, and a well-optimised page can lose if its answer is distributed across the whole document.

What follows: each section should be independently comprehensible. Headings phrased as real questions. Answer first, elaborate second. Avoid pronouns that reach back three paragraphs for their referent — an extracted passage carries none of that context.

Content clarity

Evidence: strongly observed. Ambiguity is expensive to resolve, and systems selecting between candidate passages have an incentive to prefer the unambiguous one.

Concretely: state units, dates and qualifiers inside the sentence that carries the claim. "It improved significantly last year" is unusable. "Organic sessions rose 30% year over year between 2025 and 2026" can be lifted verbatim.

Direct answers

Evidence: strongly observed. The inverted pyramid is not a stylistic preference here, it's a structural one. Give the answer in the first two sentences under the heading, then earn the reader's remaining attention with detail.

The commonest failure in commercial content is a 400-word preamble before the answer. That preamble exists to satisfy a word-count target and it actively costs citations.

Citability

Evidence: inferred, but with a clear mechanism. This is the factor I'd elevate above the fashionable ones. A citable claim is one that survives extraction — self-contained, specific, attributable, and not dependent on surrounding setup.

Not citableCitable
"This can significantly improve your results.""Adding for attributes to form labels lets an agent map a user's phone number to the right input."
"As mentioned above, it varies.""Google restricted FAQ rich results in August 2023 to a narrow set of authoritative sites."
"Many businesses see improvements.""In one GA4 funnel rebuild we identified a 25% checkout drop-off at the shipping step."

Read your own page one paragraph at a time, with the rest covered. Whatever still makes sense alone is what can be quoted.

Original statistics

Evidence: observed, and structurally sound. The strongest advantage available to a small site. A synthesis of public knowledge competes against every other synthesis of the same knowledge; a number only you have cannot be sourced from anyone else.

It does not need to be a large study. A support-ticket analysis, an audit sample, a before-and-after with the method stated — all qualify. State your n, your window and your method, because an unmethodologised statistic is a claim, not data.

Comparisons and structured information

Evidence: observed. Comparative queries are among the highest-value commercial intents, and tables are the extractable form. "Compare product specifications" is one of Google's two named agent use cases; a specification buried in a paragraph is not comparable.

Use real HTML tables with proper <th> headers and scope attributes. Not screenshots of tables, not CSS-grid divs that only look like a table.

4. Stage 3 — trust

Being relevant is not sufficient. Systems attach your name to an assertion, which means they carry reputational risk when they cite you.

Authority

Evidence: documented in principle, opaque in mechanism. Google's quality systems and E-E-A-T framing predate AI search and carry into it. What we can say usefully is that experience and identity are legible signals: named authors with real credentials and corroborating profiles; first-hand accounts that could only come from doing the work; a clear entity behind the site.

What I would not claim is any specific formula. Anyone quoting an "authority score" as an input to AI citation is describing a third-party model, not the system.

Third-party validation

Evidence: observed. What others say about you appears to matter more than what you say about yourself — which is intuitive, since self-description is free and corroboration isn't.

This includes mentions in industry press and roundups, discussion on forums and communities, reviews on platforms you don't control, and citations from other practitioners. Note the word earned. Google explicitly names chasing inauthentic mentions as ineffective and adjacent to spam systems, so link-buying and mass-mention schemes are not the play here.

Freshness

Evidence: observed, and conditional. This one is routinely overstated. Freshness dominates on volatile queries — pricing, product availability, anything about a fast-moving platform — and barely registers on stable ones. Nobody needs a 2026 explanation of how HTTP redirects work.

The counterproductive version is updating dateModified without changing anything. It's detectable, it's dishonest, and on stable topics a genuinely old page with a stable URL and accumulated corroboration frequently outperforms a fresher one.

Entity context

Evidence: documented in principle. Systems reason about entities. Being an unambiguous, consistent entity — the same name, the same identifiers, the same profiles across every surface — makes you easier to include with confidence.

This is where structured data earns its place, and note the framing: it aids understanding, not ranking. I've argued that at length in the schema markup evidence review, including why the correlation studies don't show what they're quoted as showing.

5. The evidence grid

Everything above, sorted honestly by evidence strength and by how much control you have.

FactorStageEvidenceYour control
Crawler accessRetrievalDocumentedTotal
Server-side renderingRetrievalDocumentedTotal
Existing rankingsRetrievalDocumented (Google)High, slow
Query intent matchRetrievalDocumentedHigh
Passage relevanceSelectionStrongly observedTotal
Direct answersSelectionStrongly observedTotal
Content claritySelectionStrongly observedTotal
CitabilitySelectionInferred, clear mechanismTotal
Structured comparisonsSelectionObservedTotal
Original statisticsSelectionObservedHigh, costly
Entity contextTrustDocumented in principleHigh
FreshnessTrustObserved, conditionalHigh
AuthorityTrustDocumented, opaqueLow, slow
Third-party validationTrustObservedLow, earned

Notice the shape: the factors you fully control cluster in retrieval and selection. The ones you don't cluster in trust. That is a reasonable order of work — fix the gate, then make your passages the best available answer, then let corroboration accumulate.

6. Testing it on your own market

Everything above is either platform documentation or observation. Neither tells you what happens in your vertical, in your market, for the queries that pay you — and that is a gap you can close yourself in about a day.

I'm running exactly this design across India and the UAE at 120 prompts; the methodology is published in full and the results are pending collection. Rather than have you wait for mine, here is the protocol so you can run a smaller version on your own market.

Why publish the method before the results

A benchmark whose inputs are secret is marketing, not research — and stating hypotheses before you look is the only thing that lets them be wrong. If my findings contradict what I've written above, that's the more interesting outcome and I'll publish it that way.

A minimum viable test

Setup

  • Pick 20 prompts a real buyer would type — not keywords. Split them across discovery, comparison and decision intent
  • Test on the surfaces that matter to you: AI Overviews or AI Mode, ChatGPT search, Perplexity, Claude
  • Log out, disable memory and personalisation, use a fresh session per prompt
  • Pin your location; run each market separately rather than averaging
  • Complete all platforms inside one 72-hour window — the web moves under you
  • Screenshot every response. It is the only defence against your own memory

Record per response

  • Brands mentioned, and sources cited with a live link — these are different outcomes
  • Domain type: brand, publisher, forum, directory, video
  • Whether the cited page ranks in Google's top 10 for a related query, and at what position
  • Content format: guide, listicle, product page, comparison, forum thread
  • Published and modified dates
  • Whether the cited passage contains original data
  • Word count, and whether the quoted passage is self-contained

What it will tell you

  • Whether citation tracks ranking in your vertical, or diverges from it
  • Whether owned pages compete at all, or aggregators dominate your discovery queries
  • How much the four platforms actually agree — my expectation is less than people assume
  • Which content formats get lifted in your niche, specifically

State the limits when you publish. Single-run sampling leaves model non-determinism uncontrolled — the same prompt can return different sources an hour later. Twenty prompts is a small n. Logged-out testing removes personalisation that real users have. None of that invalidates the exercise; it bounds what you may claim from it. A study that's honest about its bounds is more citable than one that isn't, which is a pleasant piece of recursion.

7. What I'd actually do

First — clear the gate

  • Confirm no AI search crawler you want is blocked in robots.txt or at the WAF
  • Load your key pages with JavaScript disabled; confirm the substance survives
  • Confirm indexation and snippet eligibility in Search Console

Then — win the passage

  • Rewrite headings as the questions people actually ask
  • Move the answer into the first two sentences of every section
  • Make each section survive being read alone — no orphaned pronouns, no back-references
  • Convert comparisons from prose into real tables
  • Put units, dates and qualifiers inside the claim sentence
  • Publish one thing per quarter containing data only you have

Then — earn the trust

  • One consistent entity: same name, same identifiers, same profiles everywhere
  • Named authors with genuine, corroborated credentials
  • Update on volatile topics; leave stable pages alone and let them accumulate
  • Earn mentions by being worth mentioning — the alternative is explicitly named as spam-adjacent

8. Frequently asked questions

Do I need to rank in Google to be cited in AI Overviews?

Effectively yes, by design. AI Overviews retrieve using Google's core ranking systems, so retrievability in classic Search is the prerequisite. The nuance is that query fan-out means the ranking that matters may be for an adjacent sub-question rather than your target keyword.

Why does my competitor get cited when their content is worse?

Usually one of three things: they're retrievable and you aren't (crawler or rendering), their content is more extractable at passage level even if weaker overall, or they have third-party corroboration you don't. Check them in that order.

Does content length affect citation?

Not directly. What matters is whether a specific passage answers a specific sub-query cleanly. Long pages can win on narrow questions; padding to hit a word count actively hurts by burying the answer.

How much does freshness matter?

It depends on the query. On volatile topics — pricing, availability, fast-moving platforms — it matters a great deal. On stable topics it barely registers, and an established page with accumulated corroboration often outperforms a newer one. Updating dates without changing content is detectable and counterproductive.

Do all AI search engines cite the same sources?

No, and assuming they do is a planning error. They use different indexes, different retrieval and different attribution behaviour. Measuring one platform and generalising is one of the commonest mistakes in AI visibility reporting.

What is the single highest-leverage change?

For most sites, moving the answer to the top of each section and making that answer self-contained. It costs an editing pass, requires no new technology, and it targets the stage where you have complete control.

Can I buy my way into AI citations?

No. Google explicitly names chasing inauthentic mentions as ineffective and adjacent to its spam systems. Paid mention schemes carry risk without documented benefit.

How do I track whether this is working?

Search Console's generative AI performance report for Google's surfaces, server logs for crawler activity, a fixed prompt set retested on a schedule, and self-reported attribution on lead forms. Treat any third-party "AI visibility score" as a model, not a measurement — no external tool has access to internal platform metrics.

Work with me

Find out which stage of the citation pipeline you're losing at

Jayant Solanki

I'm Jayant Solanki — an SEO, GEO and automation strategist working with eCommerce, local-service and global brands. I test on your prompts and your market rather than quoting someone else's US-centric study back at you.

A citation readiness engagement typically covers:

  • Retrieval audit — crawler access, rendering and indexation on revenue pages
  • A fixed prompt set built from your real buyer questions, tested across four surfaces
  • Passage-level rewrite guidance on your highest-intent templates
  • Entity and corroboration review across profiles you don't control
  • A repeatable measurement loop, so the next test is comparable to this one

Ranked #1 for "metal buildings" · +30% YoY organic traffic · 120-prompt benchmark across India and the UAE in progress

Sources

Jayant Solanki

Jayant Solanki

AI-Ready SEO, GEO & AIO strategist based in Indore, India, working with eCommerce, local service and global brands across India, the UAE and the US. Ranked a client #1 for a 45,000/month commercial keyword; builds the measurement and automation behind the work.

More about Jayant →