Wikipedia AI citations: ChatGPT's top source, Google's not
Wikipedia is ChatGPT's top-cited source, and barely shows up in Google's AI Overviews. The citation data, the notability bar, and what it means for your pages.
Wikipedia is ChatGPT's top-cited source, and barely shows up in Google's AI Overviews. The citation data, the notability bar, and what it means for your pages.

Ask ChatGPT to explain almost any company, technology, or historical event, and there's a good chance one of its sources is Wikipedia. That's not a guess. Across a 680-million-citation study spanning August 2024 to June 2025, Profound found Wikipedia was ChatGPT's single most-cited domain at 7.8% of total citations, while the same domain accounted for just 0.6% of Google AI Overviews' citations in the same window. That's not a small gap. It's the difference between a source ChatGPT treats as load-bearing and a source Google's AI barely touches.
This post breaks down what the data actually shows, where the studies agree and where they diverge, and why "Wikipedia AI citations" is really two separate questions: how often a source gets named in an answer, and how much of the total citation pie it accounts for. Those two metrics tell almost opposite stories for Google AI Mode, which is exactly the kind of thing that gets flattened into a bad takeaway if you only read the headline. We'll also cover the practical version: what an accurate Wikipedia presence, or the entity pages that substitute for one, are actually worth to you, and where in answer engine optimization that effort belongs.
Wikipedia dominates ChatGPT's source mix. In every independent study we could find, it's either the top domain outright or close enough to it that the gap comes down to methodology, not a real disagreement about the underlying pattern.
The exact rank moves by a spot depending on how a study samples queries, but Wikipedia is always near the top. Ahrefs' July 2026 tracker of the 50 most-cited domains in ChatGPT puts Wikipedia at 8.9% of citations, second behind Reddit's 16.7%. Profound's larger, longer-running study puts Wikipedia at #1 with 7.8% of total citations, ahead of Reddit's 1.8% in that sample. Different query mixes and time windows produce different rankings against Reddit specifically, but Wikipedia clears every other domain in both studies by a wide margin. No other reference site, media outlet, or review site comes close.
Zoom in on just the highest-frequency sources and the concentration gets more extreme, at least by this measure. Profound's 680-million-citation dataset found Wikipedia accounts for roughly 47.9% of citations among ChatGPT's top-10 cited sources. Discovered Labs cites that same 47.9% figure directly from Profound's research rather than as an independent count of its own, so it's one data point, not two studies converging, but it's a striking one on its own terms: within just the leading tier of ChatGPT's sources, Wikipedia owns close to half the share, with the next-largest domain, Reddit in Profound's sample, taking a fraction of that.
The concentration isn't an accident of retrieval, it's baked in before ChatGPT ever runs a search. "Every LLM is trained on Wikipedia content, and it is almost always the largest source of training data in their data sets," Selena Deckelmann, the Wikimedia Foundation's Chief Product and Technology Officer, wrote in July 2023. Wikipedia is a large, structured, constantly updated text corpus with citations built into nearly every claim, which makes it unusually good raw material for a model to learn facts from. That gives ChatGPT a strong prior toward Wikipedia before any live citation happens.
On top of that prior, ChatGPT's live web citations lean heavily on Bing's index, and Wikipedia ranks well there for the kind of entity-defining, "what is X" query ChatGPT gets asked constantly. Stack a training-data bias with a retrieval-time bias in the same direction and you get the concentration the studies keep finding. That preference for structured, third-party-verified references over single-author blogs shows up further down ChatGPT's most-cited list too, in dictionary and reference sites like Merriam-Webster and Cambridge Dictionary; entity-rich, independently verified pages travel further in ChatGPT's citations than promotional pages do, the same logic behind author schema for AI citations.
Google's two AI surfaces don't share Wikipedia's outsized role in ChatGPT, but they don't behave identically to each other either. AI Overviews barely cites it at all. AI Mode names it far more often, just not as a bigger slice of its total citations. That distinction matters more than either number alone.
By share of total citations, the gap between ChatGPT and AI Overviews isn't close. In Profound's 680-million-citation study, Wikipedia took 7.8% of ChatGPT's total citations against just 0.6% of AI Overviews' total citations, roughly a 13x difference on the same measure. AI Overviews draw from Google's own web index and ranking signals across a single-pass synthesis, described in more detail in our Google AI Mode vs. AI Overviews breakdown, and that process spreads citations across a much wider set of domains rather than concentrating on one reference site the way ChatGPT does.
AI Mode complicates the "Google ignores Wikipedia" read, but only if you're careful about which metric you're reading. Ahrefs' study of 540,000 query pairs found Wikipedia appears in 28.9% of AI Mode responses, versus 18.1% of AI Overviews responses for the same queries, a real and repeatable gap in how often the surface names Wikipedia at all. That's a response-presence number: how many answers mention Wikipedia at least once, not what fraction of the surface's total citations point to it. Semrush's separate study of 230,000 prompts across 13 weekly snapshots measured the same kind of number, response presence, and found Wikipedia appearing in only around 2% of AI Mode's weekly responses through the window, while the same measure for ChatGPT fell from roughly 55% in early August to under 20% by mid-September. The two studies don't agree on AI Mode's exact percentage, different samples and time windows rarely do, but neither shows Wikipedia's response presence in AI Mode approaching its presence in ChatGPT.
Here's the part that should reframe how you read any of these stats: "response presence" and "citation share" are different measurements, and the studies above don't all use the same one. Response presence asks: in how many answers does a source show up at all? Citation share asks: out of every link cited across all answers, what fraction points to that source? Profound's headline numbers, 7.8% for ChatGPT and 0.6% for Google AI Overviews, are citation share: a straight count of total citations. Ahrefs' 28.9%-vs-18.1% figure and Semrush's weekly figures are both response presence: how often a source gets named at least once, regardless of how many other sources share the answer. A surface that cites nine domains per answer instead of two can name Wikipedia in a lot of responses while still handing it a small share of the total citation pie, simply because there's more total volume to spread across. That's the likely reason AI Mode names Wikipedia in more responses than AI Overviews does, while Google's AI surfaces overall still cite it at a fraction of ChatGPT's rate on the metric, citation share, that counts the whole pie. Read the response-presence number and it looks like Google is closing the gap with ChatGPT. Read the citation-share number and the gap barely moves.
None of this is trivia. It changes where a limited amount of effort on brand and entity pages actually pays off, and it means "get on Wikipedia" is not the general AI-search advice it sounds like.
If your product category, your competitors, or your company by name is the kind of thing people ask ChatGPT about, an accurate Wikipedia article is disproportionately a ChatGPT-visibility lever, not a general AI-search one. Take Lyra's own case: we don't clear Wikipedia's notability bar yet, no company this young has the multiple rounds of independent press coverage the guideline demands, so there's no article for ChatGPT to draw on when someone asks it about us. That's not a gap we can shortcut. It's the reason the glossary and author-schema work below is where our own entity effort goes right now, instead of at a Wikipedia draft that would get deleted on sight. The honest test for your own company: open ChatGPT, ask three questions you'd expect a prospective buyer to ask about your category or your company, and note every source it names. If Wikipedia shows up in the answer and your own entry is thin, outdated, or missing, that's the gap this data predicts, and it's worth fixing directly rather than hoping better blog content will crowd it out. An inaccurate or absent Wikipedia article costs you more than a citation. It's baked into what the model already believes about you before it searches anything, and that belief is expensive to correct one prompt at a time.
You can't just write yourself an article, and you shouldn't try. Wikipedia's own guideline is specific: an organization "is presumed notable if it has been the subject of significant coverage in multiple reliable secondary sources that are independent of the subject.") That guideline explicitly excludes press releases, self-promotion, and paid or sponsored media as evidence of notability. In practice, that means: real journalism or independent analysis, from more than one outlet, that covers you in depth rather than a passing mention in a funding roundup. If you don't have that yet, no amount of on-page optimization gets you a Wikipedia article, and a page built on weak sourcing is likely to get deleted anyway.
Most companies, honestly, don't clear that bar, and that's fine. The move isn't to fake notability, it's to build the entity signals you actually control. A glossary page that defines your category term precisely and gets cited as the definitional source does some of the same job Wikipedia does for a broader topic, just at your scale. Author schema and a verified sameAs trail (LinkedIn, GitHub, Crunchbase, Wikidata) build the same kind of reconciled, trustworthy entity that Wikipedia represents at a smaller radius, and our author schema guide walks through the markup. Wikidata specifically is worth a separate look here: its own notability policy accepts an entity described with "serious and publicly available references," without Wikipedia's requirement for significant independent media coverage, so it's often the more realistic near-term target while it feeds the same knowledge graphs Wikipedia feeds. None of these substitute for a Wikipedia article's weight, but they're the levers available while you build toward one, or instead of one if you never will.
An entity page you can't measure is a page you're guessing about. Once you've shipped a glossary entry, updated a Wikipedia article, or added author schema, check whether it's actually showing up in AI citations rather than assuming it worked. Our AI citation tracking guide covers the free GA4 setup for catching AI-assistant referral traffic and the paid tools that go further, so you're not flying blind on whether the Wikipedia and entity work is paying off. Pair that with the how to rank in ChatGPT checklist for the on-page half of the work, since a strong Wikipedia presence and a well-structured site are additive, not either-or.
Lyra fact-checks every claim before it ships, so a post like this one cites Wikipedia, Ahrefs, Profound, and Semrush by name instead of asserting a number and hoping it holds up.
FAQ
Two reasons stack on top of each other. Every major LLM trains on Wikipedia and it's almost always the largest single source in the training set, per the Wikimedia Foundation, so the model already 'knows' Wikipedia deeply before it ever runs a live search. On top of that, ChatGPT's live citations lean on Bing's index, and Wikipedia ranks well there for the encyclopedic, entity-defining queries ChatGPT gets asked most. Google's AI Overviews draw from Google's own index and ranking signals, which spread citations across a wider set of domains instead of concentrating on one reference site.
By response presence, yes: Ahrefs found Wikipedia shows up in 28.9% of AI Mode responses versus 18.1% of AI Overviews responses for the same 540,000 query pairs, and Semrush's separate weekly tracking put Wikipedia's response presence in AI Mode at around 2% across a 13-week window, far below ChatGPT's, which ran from roughly 55% down to under 20% over the same stretch. But the study that actually measures share of total citations, not response presence, is Profound's, and it found Google's AI Overviews cite Wikipedia at just 0.6% of all citations against 7.8% for ChatGPT. AI Mode names Wikipedia in more individual answers than AI Overviews does, but that doesn't mean it hands Wikipedia a bigger slice of its total citations, because each answer cites many more sources.
Not directly, most brand mentions in ChatGPT answers don't come with a citation at all. But an accurate Wikipedia article changes what the model already believes about you before it searches anything, because that article is training data. If ChatGPT users ask about your category, your competitors, or your company by name, a correct, well-sourced Wikipedia entry is one of the highest-leverage pages you don't control directly, and an inaccurate or missing one is a gap competitors' entries can fill instead.
An organization is presumed notable if it has been the subject of significant coverage in multiple reliable secondary sources that are independent of the subject, per Wikipedia's own guideline. Press releases, sponsored posts, interviews where the subject is the only source, and your own site don't count. You need real, independent journalism or analysis, from more than one outlet, that covers the company in depth rather than mentioning it in passing.
Build the entity signals you do control. A glossary page that defines your category term precisely, sourced author bios with schema markup, and a Wikidata entry (which has a lower bar than a full Wikipedia article) all feed the same knowledge graphs and training corpora Wikipedia feeds, just at smaller scale. None of them substitute for a Wikipedia article's authority, but they're the levers available to a company that isn't there yet.
Built by the tool you're reading about
Lyra finds the topics worth ranking for, writes them in your repo's voice, fact-checks every claim, and opens a pull request scored and ready to merge. You review and hit merge. Want to see what she'd write for you? Start free with three posts, no card.
Keep reading

Perplexity SEO runs on a six-stage retrieval pipeline that scores relevance, freshness, extractability, and authority before deciding which sources get cited.

G2 and Capterra reviews now drive most AI citations for software recommendations. Here's the data on why ChatGPT recommends competitors, and how to fix it.

AI writing tool data training policies vary sharply by vendor: the exact contract language to check before your SaaS drafts train someone else's model.