Query fan-out technique: structuring content for AI Mode
The query fan-out technique explained in Google's own words, plus a tactical method for restructuring headings around sub-queries without fragmenting a post.
The query fan-out technique explained in Google's own words, plus a tactical method for restructuring headings around sub-queries without fragmenting a post.

The query fan-out technique is real, and Google names it in its own documentation. It is also the most misread concept in AI search right now, because the obvious reaction, spin up a separate page for every sub-query, is the exact move Google's own guidance warns against. This post explains what fan-out actually does in Google's words, then walks through restructuring a real page around it: headings rewritten as sub-queries, answered inside one consolidated post, not scattered across many.
Our deep dive on Google AI Mode versus AI Overviews covers the mechanism and the citation-overlap research behind why the two surfaces diverge. This post picks up where that one leaves off: the tactical restructuring work, on a page you already have.
Query fan-out is the retrieval step where Google's AI systems split one question into several related queries, run them in parallel, and combine what comes back into a single answer. It is a retrieval decision, not a content format, and Google has published the definition directly rather than leaving it to third-party guesswork.
Google's guide to optimizing for generative AI features defines query fan-out as "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query." Its AI features documentation adds the mechanism: "Both AI Overviews and AI Mode may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources, to develop a response." The stated payoff is a wider, more diverse set of helpful links than a single-query search would surface, because each sub-query can pull from a different source.
Aleyda Solis, founder of Orainti, cites the same Google explanation in her own breakdown of the technique: "AI Mode uses our query fan-out technique, breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf." Her own read adds the detail Google's quote leaves out: those parallel sub-queries don't all hit the same index. They run across the live web, the knowledge graph, and specialized data sources like Google Shopping, which is why an AI Mode answer can cite a product spec, a forum thread, and a knowledge panel fact in the same response.
The mechanism has a paper trail. iPullRank's analysis ties fan-out to Google patents including US20240289407A1 ("Search with Stateful Chat"), which documents synthetic query generation inside conversational search sessions, and US12158907B1, the "Thematic Search" patent. That second one runs in the opposite order you'd expect: it doesn't sort sub-queries into themes before retrieval, it gathers the initial results first, generates summary descriptions of the returned passages, then clusters those summaries into themes afterward. A separate Google patent, US11663201B2 ("Generating Query Variants Using a Trained Generative Model"), describes training a model on real query-document pairs to produce synthetic query variants at request time.
None of this is new invention dressed up as AI magic. It's query expansion, a decades-old information retrieval technique, running through a generative model instead of a synonym table, at a scale and speed that changes what "covering a topic" needs to mean for a page that wants to get cited.
iPullRank's analysis of the patent family names eight distinct types a model can generate from one original query, and the training mechanism it describes matches US11663201B2 above: equivalent (a reworded version of the same question), follow-up (the natural next question), generalization (a broader version of the question), specification (a narrower, more detailed version), canonicalization (a standardized phrasing), translation (a different language), entailment (a question the original logically implies), and clarification (a question the system asks back to narrow intent).
That taxonomy matters more than it looks. Most writers plan content around equivalent and specification queries already, different phrasings and narrower angles of the same topic. Follow-up, entailment, and clarification are the types that get missed, because they aren't rephrasings of your target keyword at all. They're the next question a reader would actually ask once they have the first answer, and a page that stops at the head term never reaches them.
A page written to answer one query well used to be enough. Fan-out changes the unit of competition from the page to the sub-query, and there are now enough sub-queries per prompt that a page answering only the head term is competing for a shrinking slice of the answer.
Seer Interactive ran 501 tracked prompts through the Gemini 3 API using forced grounding and measured an average of 10.7 fan-out sub-queries per prompt, ranging from 3 to 28. That's a 78% jump from Gemini 2.5's average of 6.01 queries per prompt. Seer's writeup also points to a separate study, run by Chris Long at Nextiv, putting Gemini 3's fan-out volume at roughly 5x ChatGPT's. One prompt from a user is now, on average, more than ten separate retrieval events on Google's side.
The shape of those sub-queries is the more useful finding for a writer. The average fan-out sub-query ran just 6.7 words, 95% had zero measurable search volume of their own, and 26.4% included a brand name. These aren't queries you'd find in a keyword tool, and they never will be, because nobody types them. They exist only inside the model's own reasoning chain. That's the practical case for planning content around sub-topics and entities instead of a keyword list: the queries that actually drive citations mostly don't show up in one.
Once fan-out issues a sub-query, the system doesn't retrieve your whole page for it. Lumar's explainer on content chunking notes that passage-level retrieval pulls sections of roughly 100-300 words that semantically match a sub-query, rather than ingesting the full article, because feeding an entire long document into the model for every query is expensive and mostly unnecessary. Our RAG chunking post goes deeper into that mechanic: the chunk, not the page, is the unit a retrieval pipeline actually scores.
That's the part that makes fan-out a structure problem, not just a coverage problem. A page can mention a sub-topic accurately and still lose the citation if the sentence answering it is split across two paragraphs, or buried under three sentences of setup before it resolves. Ten sub-queries hitting one page only produce ten citations if ten distinct passages on that page each stand on their own.
The obvious reaction to "there are ten sub-queries now" is to write ten pages. Google's own guidance says not to. Its generative AI optimization guide states it plainly: "There's no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users." The same guide goes further and names the failure mode directly: creating separate content for every anticipated query variation "violates Google's scaled content abuse spam policy," and a high quantity of pages doesn't make a site higher quality or more relevant, just harder to maintain.
This is the correction the rest of this post is built around. Fan-out rewards sub-query coverage. It does not reward one page per sub-query. Those are opposite instructions if you read "cover more sub-queries" as "publish more pages," and conflating them is how a well-intentioned SEO push turns into the exact spam pattern Google is warning about.
The fix is inside the page you already have, not a new one. Take the sub-query list you would have turned into new pages and turn it into headings instead, each one answered where it belongs on the page that already owns the topic.
Start from the eight sub-query types, not a keyword list, since 95% of real fan-out sub-queries carry no measurable search volume and won't show up in one anyway. For a given topic, write out the equivalent phrasing, the obvious follow-up, a broader and a narrower version, and, most often skipped, the entailment: what does answering this question imply the reader needs next. That list is your section map before you touch a single heading.
Replace a vague section label with the literal question a sub-query would ask, then answer it in the first two sentences under the heading, before any setup or context. This is the same discipline our answer-first content structure post lays out as a checklist: a self-contained block that survives being lifted out and quoted alone, because that's exactly what happens when a passage gets retrieved for a sub-query that isn't the page's main topic.
When two sub-queries are close enough that separating them would mean repeating half the same explanation twice, they belong in the same section, not two sections and definitely not two pages. The test: if a reader landed on one and skipped the other, would they be missing information or just missing a slightly different phrasing of what they already read? If it's the second, consolidate. This is the same judgment call behind content refresh work, where the fix for a page that's fallen behind is usually adding sections to the existing post, not launching a competing one that then cannibalizes it.
Take a hypothetical post titled "How to reduce SaaS churn." A version written for one query has a single section: "Reduce churn with better onboarding," three paragraphs of general advice. Mapped against the eight sub-query types, that one section is actually standing in for several distinct sub-queries: the equivalent ("how to lower SaaS churn"), the specification ("how to reduce churn in the first 90 days"), the follow-up ("what churn rate is normal for B2B SaaS"), the entailment ("how do you measure whether an onboarding fix worked"), and a comparison ("does onboarding or pricing drive more churn").
Restructured, that one section becomes five, each an H3 phrased as the question, each answered in its first two sentences: "What's a normal SaaS churn rate?" (answered with a cited benchmark), "How do you reduce churn in the first 90 days?" (the specific tactic), "Does onboarding or pricing cause more churn?" (a direct comparison with a stated verdict), and so on. Same post, same URL, same slug. Five retrievable passages where there was one, because the headings now match what fan-out actually asks instead of what a table of contents used to look like.
You know a restructure worked when the sub-queries it targeted start showing up as citations for passages inside the page, not when the page ranks higher for its head keyword, since fan-out citations and classic rankings move somewhat independently. Check Search Console's generative AI performance report for the page before and after, and if you're running a citation tracker, watch for new sub-query phrasings appearing in the tracked terms rather than just watching the head term's position.
Give it real time before judging it. A heading restructure changes what's retrievable, but re-crawling and re-indexing the passage still has to happen before it shows up in an answer, and comparing a short pre/post window against normal week-to-week noise in a small citation count will mislead you either way. The signal worth trusting is a widening set of distinct sub-queries citing the same URL over a few weeks, not a single answer appearing once.
This is also where semantic SEO and fan-out coverage point at the same underlying discipline: covering a topic's real sub-questions in depth, connected by internal links, is what both a Google ranking algorithm and an AI Mode retrieval pass reward, for the same reason. Neither is impressed by one page stretched thin across ten disconnected phrases; both reward one page (or a well-linked cluster of them) that actually answers each question it raises.
Restructuring headings around sub-queries is exactly the kind of repetitive, high-discipline editing work that's easy to state as a rule and tedious to run consistently across an entire blog. Lyra writes new posts with question-shaped headings and self-contained passages by default, and treats a stale or under-structured post the same way this restructure does: add the missing sub-query sections, verify every fact and link, and open the result as a pull request you review before anything ships.
Query fan-out rewards sub-query coverage on one page, not a page per sub-query. Lyra writes and refreshes posts with that structure by default and opens every change as a PR you merge.
Step by step
Map the sub-queries a topic actually generates
List the follow-ups, comparisons, specifications, and clarifications a reader (or a model) would ask next, using the eight sub-query types as a checklist rather than guessing.
Rewrite each heading as the sub-query itself
Turn vague section labels into the literal question, and answer it in the first two sentences under the heading.
Consolidate instead of cloning
Fold near-duplicate sub-topics into the same page as their own section rather than spinning up a new thin post for each one.
Check passage boundaries, not just headings
Confirm each section reads as a complete 100-300 word answer on its own, since that is the unit AI Mode actually retrieves.
Recheck citations after the restructure
Compare which sub-queries you now show up for before and after, using Search Console or a citation tracker, to confirm the restructure actually changed anything.
FAQ
Query fan-out is the retrieval method behind AI Mode and AI Overviews. Google's own documentation defines it as 'a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results.' Instead of retrieving for one query, the system issues many related sub-queries in parallel, across the live web, the knowledge graph, and specialized sources, then synthesizes one answer from whatever each sub-query returns.
Seer Interactive tracked 501 prompts through the Gemini 3 API and measured an average of 10.7 fan-out sub-queries per prompt, ranging from 3 to 28, a 78% jump from Gemini 2.5's average of 6.01. Seer's writeup also cites a separate study putting Gemini 3's fan-out volume at roughly 5x ChatGPT's. Most of those sub-queries were oddly specific: 95% had zero measurable search volume on their own.
No. Google's generative AI optimization guide says this directly: there's no requirement to break content into tiny pieces for AI, and creating separate content for every anticipated query variation violates its scaled content abuse policy. Google's systems can understand multiple subtopics on one page and surface the relevant part to users, so the fix is restructuring the headings inside one consolidated post, not spawning a thin page per sub-query.
Passage level. AI Mode's retrieval pulls sections of roughly 100-300 words that semantically match a given sub-query, not the whole page. A page can cover a sub-topic well in its full text and still lose the citation if the specific passage answering that sub-query is buried mid-paragraph or split across two sections.
Keyword variations are different phrasings of roughly the same search intent. Fan-out sub-queries are often a different intent entirely: a follow-up question, a narrower specification, a comparison, or a clarification the model infers the user needs before it can answer confidently. Google's query-variant patent names eight types: equivalent, follow-up, generalization, specification, canonicalization, translation, entailment, and clarification. Structuring for fan-out means answering those distinct sub-intents, not repeating one phrase in different words.
Built by the tool you're reading about
Lyra finds the topics worth ranking for, writes them in your repo's voice, fact-checks every claim, and opens a pull request scored and ready to merge. You review and hit merge. Want to see what she'd write for you? Start free with three posts, no card.
Keep reading

Schema markup for AI search, done right: the JSON-LD for Article, FAQ, HowTo, and Organization, validated and shipped without breaking existing SEO.

The llms.txt standard sits on roughly 1 in 20 top sites, but Google says no AI system uses it yet. Who reads it, a generator list, and a real file example.

ChatGPT Ads Manager rolled out through 2026: the ad timeline, how ads differ from citations, and whether SaaS blogs should buy ads, chase GEO, or both.