How answer engines retrieve and cite sources
Most answer engines use retrieval-augmented generation (RAG), also called grounding: the model rewrites the question into one or more searches, retrieves pages from a search index, reads the relevant passages and writes an answer that links to the pages it relied on. If your page is not in the index the engine uses, it cannot be retrieved or cited.
The pipeline in five steps
- Decide whether to search. Gemini’s grounding documentation says the model “analyzes the prompt and determines if a Google Search can improve the answer.”[1] ChatGPT “may search the web automatically when your question would benefit from current information.”[2]
- Rewrite and fan out. The system turns the question into one or more search queries. Google calls this query fan-out: “a set of concurrent, related queries generated by the model.”[3] OpenAI says ChatGPT “typically rewrites your query into one or more targeted queries” and may send more specific follow-up queries.[2]
- Retrieve. Each query goes to a search index: Google’s index for AI Overviews, AI Mode and Gemini grounding;[3][1] Bing’s index for Copilot;[4] Perplexity’s own crawler plus live search;[5] search partners for ChatGPT;[2] and, per TechCrunch’s reporting, Brave Search for Claude.[6]
- Read and synthesize. The model reads the retrieved passages and writes an answer. Google says its systems “review the specific information from those retrieved pages to generate a more reliable and helpful response.”[3]
- Cite. The answer links to supporting pages. The Gemini API returns inline annotations that map spans of text to source URLs.[1]
Grounding vs training data
| Grounding (retrieval at answer time) | Training data (learned before release) | |
|---|---|---|
| Freshness | Current, as fresh as the index | Frozen at the training cutoff |
| Citations | Yes, links to retrieved pages | No reliable citation |
| Controlled by | Search crawlers (Googlebot, Bingbot, OAI-SearchBot, PerplexityBot, Claude-SearchBot) | Training crawlers or tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended) |
| What you can do | Be indexed, rank, answer clearly | Be widely and accurately described in public sources |
Vendors separate these controls on purpose. OpenAI says each setting “is independent of the others”: you can allow OAI-SearchBot to appear in search while disallowing GPTBot for training.[7] See AI crawlers and robots.txt.
Query fan-out in practice
Google’s example: for “how to fix a lawn that’s full of weeds,” fan-out queries might include “best herbicides for lawns,” “remove weeds without chemicals” and “how to prevent weeds in lawn.”[3] Google says AI Mode issues “a multitude of queries simultaneously” across subtopics.[8] Two implications:
- A page can be cited for a sub-question even if it does not rank for the original query. Ahrefs found that 76.1% of pages cited in AI Overviews ranked in Google’s top 10 for the query, which also means about a quarter did not.[9] Independent study
- Covering the real sub-questions of a topic well helps, but Google warns that creating separate pages for every fan-out variation to manipulate results violates its scaled content abuse policy.[3] Vendor-documented Try the fan-out planner to map sub-questions for one strong page.
Why a source gets cited
No vendor publishes a citation formula. What is documented:
- Google: AI features are “rooted in our core Search ranking and quality systems.”[3] Vendor-documented
- Bing: URLs are more likely to be chosen for grounding when content “stands on its own,” facts are explicit, each URL is about a single topic and key information appears early.[4] Vendor-documented
- Research: in the GEO benchmark, adding citations, quotations and statistics increased visibility in generated answers, with effects varying by domain.[10] Independent study
Where it goes wrong
The Tow Center tested eight AI search tools with 1,600 queries in 2025 and found incorrect answers to more than 60% of them, with fabricated links and citations of syndicated copies.[11] Researchers have also shown that prompt-injection text in a page can reorder which products a conversational engine recommends;[12] Bing lists “prompt injection and AI manipulation” as abuse.[4] Clear, verifiable facts on the page make accurate citation more likely and manipulation by others easier to spot.
Frequently asked questions
What is grounding?
Grounding (retrieval-augmented generation) means the model retrieves current web pages and bases its answer on them, with citations, instead of relying only on training data.
What is query fan-out?
The engine splits one question into several related searches, retrieves results for each and combines them. Google documents it for AI Overviews and AI Mode.
See also
References
Pages accessed September 28, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "Grounding with Google Search (Gemini API)". Google AI for Developers.
- ^ "Searching the web with ChatGPT". OpenAI Help Center.
- ^ "Optimizing your website for generative AI features on Google Search". Google Search Central. Added May 15, 2026; last updated July 10, 2026.
- ^ "Bing Webmaster Guidelines". Microsoft Bing Webmaster Tools.
- ^ "Perplexity Crawlers". Perplexity documentation.
- ^ "Anthropic appears to be using Brave to power web search for its Claude chatbot". TechCrunch. Published March 21, 2025.
- ^ "Overview of OpenAI Crawlers". OpenAI Platform documentation.
- ^ "AI Mode in Google Search: Updates from Google I/O 2025". Google (The Keyword). Published May 20, 2025.
- ^ "76% of AI Overview Citations Pull From the Top 10". Ahrefs. Published July 21, 2025.
- ^ "GEO: Generative Engine Optimization (arXiv:2311.09735; Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande)". arXiv / KDD 2024. Submitted November 16, 2023; revised June 28, 2024.
- ^ "AI Search Has a Citation Problem (Jaźwińska and Chandrasekar)". Columbia Journalism Review, Tow Center for Digital Journalism. Published March 6, 2025.
- ^ "Ranking Manipulation for Conversational Search Engines (Pfrommer, Bai, Gautam, Sojoudi)". Proceedings of EMNLP 2024, ACL Anthology. Published November 2024.