How AI Works

How Google's AI Mode Actually Chooses Which Sources to Cite

Google's AI Mode chooses the sources it cites by running your question through its live search index first, breaking that question into a spread of narrower sub-questions, retrieving and ranking real pages for each one, and then checking whether the answer it wants to write is actually supported by passages on those pages. It cites the pages that survive that last check. That is the short version, and it is genuinely different from how ChatGPT or Perplexity behave. If you have read my breakdowns of how ChatGPT chooses sources and how Perplexity chooses sources, this is the third leg of the stool — and the one most businesses get wrong, because they assume Google's AI works like the other two. It doesn't.

I want to walk through the actual pipeline as best any outside practitioner can reconstruct it, then tell you what I'd change on your site if getting named by Google's AI were the goal. I'll flag clearly where I'm describing mechanics Google has confirmed versus where I'm inferring from behavior. The inference parts are my read, not a guarantee.

Google is grounding an answer, not writing from memory

Here is the core distinction. A raw language model answers from what it absorbed during training — a compressed, lossy memory of the web frozen at some cutoff. Google's AI Mode does something narrower and, for our purposes, better: it grounds its answer in documents it retrieves live at the moment you ask. The model's job is less "recall the answer" and more "read these retrieved pages and synthesize an answer I can defend with citations."

That single design choice explains almost everything about which sources get cited. If the answer has to be defensible against retrieved documents, then the sources that get cited are the ones that (a) got retrieved in the first place and (b) contain a passage that plainly supports a sentence the model wanted to write. Miss either condition and you are invisible — no matter how good your content is in the abstract.

This is why I tell people that Google's AI Mode rewards the same thing classic search rewarded, just with a higher bar for specificity. You still have to be in the index. You still have to be retrievable for the query. But now you also have to be quotable at the passage level, because a citation is earned sentence by sentence, not page by page.

Step one: the query fan-out

When you type a real question into AI Mode — something like "what's the best way for a small accounting firm to handle client document intake" — Google doesn't run that one string. It fans the query out into a cluster of related sub-queries: the security angle, the software-comparison angle, the workflow angle, the compliance angle, the cost angle. Google has described this fan-out behavior openly; it's central to how AI Mode covers a topic instead of just matching keywords.

Each of those sub-queries gets its own retrieval. So the set of documents in play is not "pages about client document intake" — it's the union of the best pages for a dozen adjacent questions the user never explicitly typed. I wrote a whole piece on query fan-out because it is, in my opinion, the most underappreciated mechanic in all of AI search. The businesses that win Google AI citations are usually the ones that happen to answer three or four of the fanned-out sub-questions well, not the ones that ranked number one for the literal phrase.

The practical consequence: breadth of coverage on a topic beats a single perfectly optimized page. If you have one page that answers the head question and nothing addressing the natural follow-ups, you'll get retrieved for one branch of the fan-out and lose the rest to whoever covered them.

Step two: retrieval and ranking, on the live index

This is the part people forget: AI Mode sits on top of Google Search. The retrieval that feeds it is drawing from the same index, shaped by many of the same signals — relevance, quality, the maze of things that have always determined whether a page can surface. Being fundamentally un-findable in Google Search means being un-retrievable for AI Mode. There is no separate "AI index" you can sneak into while remaining absent from the main one.

That's actually good news, and I say that as someone who spends most of his time on the newer engines. It means the discipline you may already have — crawlability, clean structure, genuine topical depth, real links from places that matter — is not wasted. It carries forward. What changes is the selection that happens after retrieval, which I'll get to. But the entry ticket is still the index.

One nuance I watch closely: retrieval for a fanned-out sub-query can surface a page that would never rank on page one for the head term. Google is retrieving for the narrow intent, not the broad one. So a tightly focused page that thoroughly answers one specific sub-question can get pulled into an AI answer even if it's modest on traditional rankings. I've seen this repeatedly. Depth on a narrow question is a legitimate path in.

Step three: the grounding and synthesis pass

Once Google has a pool of retrieved passages across all the sub-queries, the model drafts an answer and then does what I think of as a defensibility check. For each claim it wants to make, is there a retrieved passage that supports it? The citations you see are, functionally, the receipts — the sources whose passages backed the sentences that made it into the final answer.

This is where writing style stops being cosmetic and starts being mechanical. A page that states a claim cleanly, in a self-contained sentence, near the evidence for it, is easy to ground against. A page that buries the same fact inside a rambling paragraph, or implies it without ever stating it, is hard to ground against — so it doesn't get the citation even if it technically "covers" the topic. My rule of thumb: if a smart human skimming your page couldn't lift one clean sentence that answers the question, the model can't either.

What actually makes a passage citable

How this differs from ChatGPT and Perplexity

People lump the three engines together and it costs them. The differences are real:

Perplexity is retrieval-first by design and shows its citations aggressively; it tends to reward pages that read like a clean, sourced answer to the exact question. It's the most "link-forward" of the three.

ChatGPT blends what it learned in training with live browsing depending on the query and the mode. That means being in the training data — being written about enough, consistently enough, over time — carries weight there in a way it does not on Google's live-grounded path. I've called this the difference between the two memories of AI, and it matters most on ChatGPT.

Google's AI Mode is the most index-dependent and the most fan-out-driven. Its citations are the tightest to "what could we retrieve and defend right now," which means recency, crawlability, and passage-level clarity punch above their weight. Training-data fame helps you less here and retrievability helps you more.

So the honest strategic answer, if you're choosing where to spend, is: you don't optimize for one. You build the underlying entity authority all three respect, then you tune the surface. For Google specifically, the tuning is retrieval and passage clarity. This is opinion shaped by what I watch day to day, not a law of physics — engines change their behavior constantly, and any of this can shift in a quarter.

What I'd change on your site this week

If a client came to me tomorrow and said "I want to get named in Google's AI answers," here's the order I'd work in, and why.

  1. Confirm you're actually in the index and crawlable. It sounds basic. It is the single most common blocker I find. If AI Mode can't retrieve you, nothing downstream matters.
  2. Map the fan-out for your money questions. Take the three or four questions that actually lead to revenue and write down the sub-questions a curious person would ask around each. That list is your content gap analysis. Cover the branches, not just the trunk.
  3. Rewrite for passage-level answers. Go through your key pages and make sure each important claim exists as a clean, standalone, specific sentence. Add real question-and-answer sections. This is the highest-leverage editing you can do for Google's AI.
  4. Feed corroboration. Make it easy for other credible sources to say the same true things about you — consistent facts across your profiles, directories, and third-party mentions. Google grounds more confidently when you're not the only one saying it.
  5. Track whether it's working. Test your priority questions in AI Mode on a schedule and log who gets cited. You cannot manage what you don't measure, and AI citations move week to week.

The mistake I see most often

The single most common error I watch businesses make with Google's AI is treating it as a brand-new channel that needs a brand-new playbook, and in the panic, either doing nothing or chasing rumored "AI ranking factors" that don't exist. Both are wrong. Google's AI Mode is not a separate machine bolted onto search — it's a new way of reading the same index you've been living in. The work isn't exotic. It's the unglamorous discipline of being findable, being clear, and being corroborated, applied with more precision than the old ten blue links ever demanded.

The second most common mistake is optimizing a page for a keyword when you should be answering a question. Keywords were a proxy for intent; the fan-out makes intent explicit by decomposing it into sub-questions. If you're still writing to a keyword density rather than to the cluster of things a real person actually wants to know, you're optimizing for a game that ended. Write to the questions, structure the answers so a machine can locate them, and let the retrieval do its job.

None of this is a trick. That's the part I want to leave you with. Google's AI Mode is, if anything, less gameable than the classic ten blue links, because the grounding step keeps punishing claims that aren't supported. The businesses that win are the ones that are genuinely the clearest, most retrievable, most corroborated answer to a real question. That's been the whole thesis of Answer Engine Optimization from the start. Google's AI just enforces it more literally than anything before it.

Key takeaways

  • Google's AI Mode grounds answers in pages it retrieves live from its search index — so being retrievable, not just being 'good,' is the entry ticket.
  • It fans your question into many sub-questions and retrieves for each; breadth of coverage on a topic beats one perfectly optimized page.
  • Citations are earned at the passage level: a clean, specific, standalone sentence that answers the question is what gets cited.
  • It differs from ChatGPT (which leans on training-data fame) and Perplexity (retrieval-first, link-forward) — Google is the most index- and fan-out-dependent.
  • Corroboration helps: Google grounds more confidently in a claim other credible sources also make.
  • Highest-leverage work: confirm crawlability, map the fan-out for revenue questions, and rewrite key claims as liftable sentences.

Frequently asked questions

Is Google's AI Mode the same as AI Overviews?
They're closely related surfaces built on the same underlying approach — live retrieval from Google's index plus a grounding step. AI Mode is the more conversational, fan-out-heavy experience, while AI Overviews appear atop traditional results, but the source-selection logic I describe applies to both. The practical advice doesn't change between them.
Do I need special 'AI schema' to get cited by Google's AI?
No. There's no separate AI index or magic markup that gets you in. Legitimate structured data helps a machine parse your page, but the real levers are being in the index, being retrievable for the fanned-out sub-questions, and writing claims as clean, groundable sentences. Treat schema as parsing help, not a shortcut.
Why does my page rank well in normal search but never get cited in AI Mode?
Usually because the page covers the topic broadly but doesn't state the specific answer as a clean, standalone sentence the model can ground against — or because it doesn't address the follow-up sub-questions the fan-out generates. Ranking gets you retrieved; passage-level clarity gets you cited. They're related but not the same bar.
Scott Tischler

About the author

Scott Tischler is the Founder & Chairman of AIrecommend.ai and a practitioner-authority on AI search and Answer Engine Optimization. With 20+ years in marketing technology — including American Express, MetLife, and UBS — and executive and professional study at Wharton, Harvard, and Oxford, he helps businesses become the ones AI recommends.

Want to be the business AI recommends?

See how AIrecommend.ai builds the entity authority answer engines reward.

Explore AIrecommend.ai