A software company has one quarter to improve its AI-search visibility. The content team proposes 120 pages: “best CRM for agencies,” “best CRM for small agencies,” “best CRM for 12-person agencies,” “best EU CRM for agencies,” and dozens more. The pitch sounds technical: an AI system may fan one buyer question into many hidden searches, so the site should own every possible variation.
That plan can consume the quarter and leave the buyer with a maze of repeated claims. It can also move the site toward the patterns Google calls doorway abuse and scaled content abuse: substantially similar, query-targeted pages that exist closer to the search result than to a useful, browsable information architecture.5 The expensive mistake is confusing a retrieval system’s internal research process with a publishing instruction.
The better question is: what must a buyer learn, verify, compare, and decide before the original request can be answered well? That reframing preserves the value of fan-out without pretending we can see inside it. It leads to fewer, stronger pages; clearer evidence; and a measurement plan that can distinguish improved decision coverage from wishful “AI optimization.”
Query fan-out is a retrieval pattern, not a ranking factor
Google’s first public AI Mode description in March 2025 said the system could issue multiple related searches concurrently across subtopics and multiple data sources, combine the results, and adjust its plan based on what it found. Its example compared sleep-tracking devices: one question required the system to plan, search, compare, and then support follow-up questions.1 At Google I/O two months later, the company described AI Mode as breaking a question into subtopics and issuing a “multitude” of queries simultaneously. It described Deep Search as taking the same technique further, potentially issuing hundreds of searches for a cited research report.2
In March 2026, a Google Search engineering leader explained the same idea through visual search. A model could identify the hat, shoes, and jacket in an outfit, retrieve results for those objects in parallel, and weave the evidence into one answer. The interview summarized the experience as doing roughly a dozen searches in the time of one.3 That is a useful explanation, not a fixed operational count. A simple factual question may require no branch at all; a high-stakes comparison may require many; a follow-up can create a new branch from what the system just learned.
Google’s current site-owner documentation supplies the most important boundary. AI Overviews and AI Mode may use fan-out across subtopics and data sources, and the responses and links can vary because the features may use different models and techniques. But there are no special optimizations or additional technical requirements. A supporting page must be indexed and eligible to appear in ordinary Google Search with a snippet; meeting those conditions still does not guarantee crawling, indexing, serving, or inclusion.4
Therefore, “query fan-out” is not a disclosed ranking factor that can be added to a checklist. It is a description of how a system can investigate a complex information need. Publishers control the clarity, evidence, accessibility, and architecture of their pages. They do not control or observe the private query plan that a platform generates for each response.
Scroll horizontally to read the full graphic.
Use five lenses to model the hidden journey
Because the real query plan is private and variable, a useful editorial model should be small enough to repeat and broad enough to expose obvious gaps. Five lenses cover most decision journeys without pretending to reproduce a platform’s internal behavior: subtopics, constraints, comparisons, evidence, and follow-ups.
Subtopics define the objects of the decision
Subtopics answer “what parts must be understood?” A payroll-software decision may involve employee records, contractor payments, tax documents, approvals, accounting integration, and support. A solar-panel decision may involve roof suitability, production estimates, financing, installer quality, warranties, and grid rules. These are not keyword variants. They are different objects or stages whose facts may live in different sources.
Start with the minimum set that can change the verdict. If removing a branch would not change the recommendation, implementation, or risk, it may be background rather than a core subtopic. This check protects the brief from becoming an encyclopedia simply because the topic has many adjacent concepts.
Constraints turn a generic answer into a usable one
Constraints answer “for whom, where, when, and under what limit?” They include location, company size, budget, deadline, existing technology, accessibility needs, risk tolerance, contractual terms, and required compatibility. Constraints often explain why two people asking nearly the same question need different answers. They also create the greatest temptation to build duplicate landing pages.
Put a constraint on its own page only when it changes the task and requires unique evidence. A country-specific tax guide may deserve a separate owner, source set, and update cadence. A page for each neighboring employee-count modifier probably does not. When constraints change only a filter or conclusion, use a comparison table, selector, or clearly labeled subsection on the main guide.
Comparisons require one set of criteria
Comparison branches answer “what are the realistic alternatives, and what changes the verdict?” A useful comparison applies the same criteria, evidence standard, date, and market boundary to every option. It explains exclusions and negative trade-offs instead of selecting only the facts that favor the publisher’s preferred answer.
Separate pages are justified when each comparison serves a distinct decision and can remain fair on its own. Otherwise, one maintained table is usually stronger than a cluster of pairwise pages that repeat the same product facts. Normalize units and feature definitions before writing the recommendation; apparent differences often disappear when pricing periods, limits, and required tiers are made comparable.
Evidence needs determine whether a claim is answerable
Evidence branches answer “what must be checked before this claim can be trusted?” A feature comparison needs current product documentation and a consistent definition. A performance claim needs a test environment, denominator, and date. A legal or security claim needs an authoritative document and a jurisdiction boundary. A customer-experience claim may need a transparent review method rather than a handful of selected quotes.
This lens is where many otherwise polished content programs fail. The site covers every noun but supports none of the consequential statements. An AI system may retrieve several pages, yet the answer still depends on an uncited pricing assumption or an outdated integration description. Build the source register before drafting. If the evidence cannot be obtained, narrow the claim instead of filling the gap with confident prose.
Follow-ups reveal the next decision, not another synonym
Follow-ups answer “what will a reasonable reader ask after learning this?” They often arise from the evidence: “Does that price include every user?”, “What happens if migration fails?”, “Which alternative avoids this limitation?”, or “How do I verify the configuration after launch?” A strong guide anticipates the highest-value follow-ups and gives them a visible path.
Follow-ups can justify supporting pages because they may represent complete downstream tasks. But they should not be generated mechanically from “People also ask,” autocomplete, or an LLM brainstorm. Validate them with support tickets, sales calls, on-site search, community questions, and actual user interviews. The hidden journey is useful as a hypothesis. Customer evidence decides which branches deserve production effort.
One prompt can become a research journey
Consider a buyer asking: “What is the best CRM for a 12-person agency in the EU that needs Gmail integration, simple migration, and predictable pricing?” A classic keyword workflow may reduce that request to “best CRM for agencies.” A decision workflow sees at least five distinct jobs.
- Define the fit. What does a 12-person agency actually need: pipeline stages, permissions, shared inboxes, reporting, or automation?
- Test constraints. Which products meet EU data, contract, residency, and administration requirements?
- Verify the integration. Does “Gmail integration” mean contact sync, two-way email, calendar, shared mailboxes, or an extension?
- Estimate transition cost. What can be imported, what breaks, how long does migration take, and who owns cleanup?
- Normalize price. Which required features sit behind higher tiers, usage charges, implementation fees, or annual commitments?
A useful answer may also need evidence about vendor support, recent product changes, customer examples, and viable alternatives. Some branches are parallel. Others are sequential: you cannot compare total cost until you know which features are required; you cannot recommend migration until you know the source system and data shape. Follow-up questions can revise the research plan again.
Scroll horizontally to read the full graphic.
Research supports decomposition—and warns about noise
Peer-reviewed retrieval research helps explain why fan-out is attractive. Complex questions often require facts that do not coexist in one document. A 2025 ACL workshop paper tested a pipeline that decomposed a question, retrieved passages for each subquestion, merged the candidate set, and reranked it against the original request. On MultiHop-RAG, the combined system raised MRR@10 from 0.464 to 0.635. The authors summarize the improvement over their standard baseline as 36.7% for retrieval and 11.6% for answer F1.8
The mechanism matters more than the headline effect. Decomposition broadened evidence coverage; reranking removed some of the noise introduced by the larger pool. This is not evidence that a page ranking for five keyword variants will be cited by Google. It is evidence, inside two benchmark RAG systems, that focused subquestions can retrieve complementary documents when an answer genuinely needs multiple facts.
The same paper shows why “more queries” is a poor content strategy. The prompt allowed up to five subqueries, and the model generated exactly five for 93.3% of MultiHop-RAG questions and 98.6% of HotpotQA questions, even though the number of gold evidence items varied. The subquery count was effectively a prompt budget, not a discovered law of the information need. The complete pipeline also took 18.9 seconds per query in the authors’ 250-query latency test, compared with 0.03 seconds for naive retrieval. Decomposition improved the tested result, but it added cost, latency, and failure points.8
A 2026 EACL paper tackles the next problem: once a request has many branches, how should a system spend a limited retrieval budget? The authors treated each subquery as an option whose usefulness becomes clearer as documents are observed. Their rank- and relevance-aware selection produced a reported 35% gain in document-level precision and a 15% increase in alpha-nDCG, and selected subsets improved support and coverage in generated reports.7 The lesson is not that publishers should copy a bandit algorithm. It is that breadth and precision compete. A research process must stop weak branches, deepen promising ones, and remove redundant evidence.
Both papers have firm boundaries. They test specific retrievers, corpora, prompts, models, rerankers, and budgets. One uses an offline supervised setting; the other uses benchmark question-answering datasets rather than live web search. Their gains do not establish a Google citation effect, a universal branch count, or a site architecture. They do support a practical editorial principle: decompose to reveal missing evidence, then consolidate to protect relevance.
Fan-out changes the retrieval footprint, not the need for judgment
The most useful live-system evidence comes from a 2026 Findings of ACL paper comparing Google organic search with Google AI Overviews, Gemini, two OpenAI search configurations, and Perplexity Sonar. The main dataset contained 4,706 queries across general search, chatbot prompts, politics, science, and shopping, plus separate trend and static-fact analyses. Queries were issued in English from the United States and Germany in September 2025.6
Retrieval behavior varied sharply. The median number of cited links averaged across datasets was zero for GPT-Tool, four for GPT-Search, and nine for Google AI Overviews. About 2% of the AI Overviews in the study retrieved more than 30 pages, while 15% retrieved two or fewer. Within this study, that spread is a counterexample to any universal “fan-out equals N queries” rule. Systems and questions used very different retrieval footprints.6
Google AI Overviews also reached beyond classic rankings: 53% of consulted domains were absent from the top 10 organic domains and 27% were absent from the top 100. Yet broader sourcing did not consistently produce broader conceptual coverage. Across the study’s datasets, generative systems produced average topic coverage between 0.71 and 0.78, close to organic search at 0.78. More retrieved sources can support different synthesis; they do not automatically yield a more complete answer.6
Source selection was also unstable. Across snapshots roughly two months apart, only 18% of AI Overview pages overlapped, compared with 45% for organic search. Overall topic coverage remained steadier than the exact sources.6 That distinction matters for measurement: a site can improve its coverage of a decision journey and still see citation volatility from run to run. One missing citation is not proof of an architectural failure, and one appearance is not proof of durable ownership.
Build an intent-and-evidence map before a content calendar
A keyword list records language. A decision map records what changes the customer’s action. Start with one costly decision and create rows for the information needs that must be satisfied. For each row, name the required proof, the freshness boundary, the best page owner, and the action a reader can take after learning it.
Scroll horizontally to read the full graphic.
Use five practical evidence classes:
- Definitions and boundaries: precise terminology, scope, exclusions, and who the advice is for.
- Comparisons: consistent criteria, alternatives, trade-offs, and the conditions that change the verdict.
- Verification: primary documentation, measurements, test methods, screenshots, dates, and known limitations.
- Implementation: prerequisites, sequence, owners, templates, failure modes, and a completion check.
- Outcomes: what to measure, over what window, with which denominator, and what threshold would change the decision.
These classes are not a schema for AI engines. They are a quality-control device for humans. They make it harder to publish a page that repeats a conclusion but lacks the proof or next step a research journey needs.
Turn the matrix into a publishable brief
The matrix becomes operational when each row produces a claim block rather than a vague topic assignment. A claim block contains five parts: the question, the bounded answer, the supporting source or method, the limitation, and the reader’s next action. For example:
Question: Will Gmail integration work for a shared agency inbox?
Bounded answer: Product A connects individual Gmail accounts, but its current documentation does not describe delegated shared-mailbox behavior.
Evidence: Vendor integration documentation reviewed on the stated date, plus a test in a trial workspace.
Limitation: The test did not cover enterprise administration or future product changes.
Next action: Run the supplied five-step shared-inbox test before migration.
This format makes a section extractable without writing for machines. A reader can find the answer, see why it is credible, understand where it may fail, and act. A retrieval system also has a coherent passage to evaluate. The same block can be maintained when the product changes because the owner knows which source and test must be repeated.
Next, order claim blocks by decision dependency. Definitions and non-negotiable constraints come before comparisons. Comparison criteria come before a verdict. A verdict comes before implementation. Implementation comes before verification and monitoring. This order prevents the common structure in which a page announces “the best” product, then retrofits criteria to justify it.
Add navigation that exposes the structure without multiplying URLs. Descriptive headings, a table of contents, anchored sections, contextual links, and concise comparison tables help readers move directly to a branch. Supporting pages should link back to the decision they inform, and the guide should link to the exact supporting task—not to a generic resource center. The result is a browsable graph whose relationships are legible to people even if no AI answer ever cites it.
Finally, create an update contract. Mark price, feature, policy, and availability claims with an owner and review date. Separate durable analysis from volatile facts so routine updates do not require rewriting the whole guide. If the organization cannot keep a branch current, link to the authoritative source and state the boundary. Freshness is part of the answer, especially when a research system can issue a branch specifically to verify recent information.
Choose a page boundary with the independent-value test
The difficult editorial decision is whether a branch deserves a section, a supporting page, or no new content. Apply four tests.
- Distinct job: does the branch help the reader complete a meaningfully different task?
- Distinct evidence: does it require its own data, process, expert, examples, or update cadence?
- Independent arrival: could a reader land there directly and finish the task without being funneled elsewhere?
- Maintenance owner: can someone keep its claims accurate without copying updates across many near-duplicates?
If the answer is “no” to two or more tests, keep the branch as a section in the decision guide. If all four answers are “yes,” a supporting page may be justified. For the CRM example, a detailed migration checklist could stand alone because it has a distinct workflow, evidence, owner, and direct audience. “Best CRM for 11-person agencies” and “best CRM for 12-person agencies” almost certainly cannot.
This approach produces a hub-and-support architecture without forcing every site into the same template. A small consultancy might publish one comprehensive guide with jump links and a downloadable worksheet. A large software platform may need separate integration, security, pricing, migration, and comparison pages because different teams own changing source data. The page boundary follows independent value, not company size or a guessed hidden query.
Scroll horizontally to read the full graphic.
The strongest counterargument
The strongest reasonable counterposition is that hidden subqueries still create real discovery opportunities. If AI search investigates pricing, compliance, alternatives, migration, and integrations separately, a narrowly focused page may match one branch better than a broad guide. Refusing to create supporting pages could leave useful evidence buried inside a 6,000-word article.
That argument is correct up to the page boundary. Focused pages are valuable when they solve focused problems. An integration page can document exact permissions and sync behavior. A security page can own certifications, subprocessors, controls, and update dates. A migration page can own field mappings and rollback steps. Each is independently useful, linkable, maintainable, and capable of answering a direct visit.
The argument fails when “focused” means changing only a modifier while repeating the same evidence and sending every reader to the same product page. Google’s doorway policy specifically identifies substantially similar pages made for similar queries and pages that funnel visitors to the actual relevant portion of a site. Its scaled-content policy focuses on many unoriginal pages created primarily to manipulate rankings, regardless of whether AI generated them.5 The safe dividing line is not broad versus narrow. It is independent reader value versus query-shaped duplication.
What the evidence does not show
- It does not reveal the exact subqueries generated for a live Google AI Overview or AI Mode response.
- It does not establish a fixed number of fan-out searches; public examples and observed retrieval footprints vary widely.
- It does not show that matching an imagined subquery causes a page to be selected, cited, clicked, or converted.
- It does not prove that every AI search system uses Google’s architecture or source-selection behavior.
- It does not turn benchmark RAG gains into a Google ranking, citation, or revenue effect.
- It does not show that publishing many near-duplicate pages improves AI visibility; current spam policies create the opposite risk.
- It does not make one citation screenshot a stable performance measure; live generative-search sources can vary across runs and time.
These are not minor caveats. They determine the responsible scope of the strategy. We can use fan-out to audit whether our content supports a complete decision. We cannot use it to promise that a particular page will enter a platform’s hidden research plan.
A practical query fan-out worksheet
Use this template for one high-value customer question. Keep it in a spreadsheet, brief, or issue tracker so every proposed page has a documented reason to exist.
| Field | Question to answer | Output |
|---|---|---|
| Customer decision | What costly action is the reader trying to take? | One sentence with audience, action, and stakes |
| Constraints | What budget, market, risk, timing, or compatibility conditions change the answer? | Named decision filters |
| Comparison set | Which realistic alternatives must be considered on the same criteria? | Fair shortlist and exclusion rule |
| Evidence needs | What would a skeptical reader need to verify each branch? | Primary sources, data, test, or example |
| Freshness | Which facts can expire, and who owns the update? | Date, review interval, and owner |
| Page boundary | Is this a section or an independently valuable task? | Guide section, supporting page, or no new content |
| Internal path | How can users and crawlers discover the evidence naturally? | Contextual links and browsable hierarchy |
| Outcome | What observable behavior would justify more investment? | Baseline, metric, denominator, and threshold |
What to do this week
- Select one decision. Use a question tied to revenue, retention, risk, or a frequent sales objection—not a broad topic.
- Interview the journey. Ask sales, support, and customers what they must verify before acting. Record constraints and follow-ups.
- Build the matrix. Map every intent to evidence, freshness, owner, and action. Mark unsupported rows as gaps.
- Consolidate duplicates. Compare existing pages. Merge pages that repeat the same job and evidence; preserve redirects when released.
- Strengthen one guide. Add the missing comparison criteria, primary evidence, limitations, implementation steps, and measurement block.
- Create only justified support. Use the independent-value test before opening a new URL.
- Verify discoverability. Confirm crawl access, indexability, canonical signals, textual content, and contextual internal links.
- Record the baseline. Capture current organic queries, landing pages, conversions, and a repeated AI-answer sample before release.
What to measure for 30 days
Measurement should test the content decision, not claim access to hidden fan-out queries. Use a fixed panel of customer questions and paraphrases, repeat it on a documented schedule, and keep platform, market, account state, date, and retrieval mode attached to every observation. Separate four layers:
- Coverage: how many decision rows have current, primary evidence and a clear owner?
- Discovery: are the intended pages crawled, indexed, internally linked, and receiving relevant impressions?
- Answer presence: across repeated runs, how often is the brand or page mentioned, cited, accurately represented, or recommended?
- Business outcome: do relevant visits, assisted conversions, qualified leads, or sales objections change after the content release?
Use a simple cadence: record the page and prompt-panel baseline on day 0; check technical discovery and evidence freshness weekly; repeat the same answer sample at least weekly without changing the panel mid-test; and review business outcomes on day 30. Preserve every run, including misses. If a platform or prompt must change, start a new series instead of blending unlike observations.
Preserve denominators. “Cited in 6 of 30 documented runs” is useful. “AI visibility improved” is not. Compare the 30-day period with a prior baseline where seasonality and release timing permit. Annotate product launches, model changes, indexing changes, and major news that could alter the result. Do not attribute a revenue movement to the article without a design that supports that causal claim.
When to change course
Continue the architecture when the evidence rows are becoming more complete, relevant organic impressions spread across the intended decision terms, readers engage with the guide and supporting tasks, and repeated answer samples show a directional improvement without accuracy loss. Expand a supporting page only when users need that task and the page has distinct evidence to own.
Change course if the team is producing pages faster than it can verify them; if new URLs cannibalize each other; if the same claims, examples, and calls to action appear across many pages; if indexed pages receive no meaningful impressions or engagement; or if answer samples repeat an inaccurate claim. In those cases, pause production, consolidate, correct the source of truth, and improve internal paths before adding another branch.
After 30 days, require a predeclared threshold for further investment. A small site might require one improved sales workflow and clear discovery growth before commissioning another guide. An enterprise publisher may require statistically credible movement across a stable prompt panel and multiple markets. The scale differs; the discipline does not.
The durable strategy is decision coverage
Query fan-out explains why a single customer question can touch more of the web than its visible wording suggests. That matters. A page can be useful for a constraint, comparison, or verification step even when it does not mirror the original prompt. Research also shows why the tempting shortcut fails: expanded retrieval brings noise, selection trade-offs, computational cost, and unstable sources.
The defensible publishing strategy is simple to state and difficult to fake. Understand the decision. Decompose it into real information needs. Supply current evidence for each need. Give every page an independent job. Link the journey so people and crawlers can navigate it. Measure repeated outcomes with visible denominators. That will not expose the hidden searches behind an AI answer, but it will produce a site worth researching.
Sources
- Google, “Expanding AI Overviews and introducing AI Mode”, March 5, 2025. First-party product announcement.
- Google, “AI in Search: Going beyond information to intelligence”, May 20, 2025. First-party product announcement and internal usage claim.
- Molly McHugh-Johnson, Google, “Ask a Techspert: How does AI understand my visual searches?”, March 5, 2026. First-party expert interview.
- Google Search Central, “AI features and your website”, last updated December 10, 2025. First-party technical documentation.
- Google Search Central, “Spam policies for Google web search”, reviewed September 1, 2026. First-party policy.
- Elisabeth Kirsten et al., “Characterizing Web Search in the Age of Generative AI”, Findings of ACL 2026, pp. 10827–10848. Peer-reviewed conference paper.
- Roxana Petcu et al., “Query Decomposition for RAG: Balancing Exploration-Exploitation”, EACL 2026, pp. 6857–6871. Peer-reviewed long paper.
- Paul J. L. Ammann, Jonas Golde, and Alan Akbik, “Question Decomposition for Retrieval-Augmented Generation”, ACL 2025 Student Research Workshop, pp. 497–507. Peer-reviewed workshop paper.