AI Search Statistics You Can Actually Cite (Primary Sources Only)
Quick answer
Most AI search statistics circulate without a traceable source, because stats pages cite each other in circles. This page holds a smaller set to a stricter rule. Every number links to a named primary source with a date and a method. Each was re-checked in July 2026. Numbers we could not trace are flagged, not repeated.
Key takeaways
- Pew Research Center found Google users clicked a result on 8 percent of pages with an AI summary, versus 15 percent without one, in March 2025 browsing data.
- A study of 15 domains found 72.4 percent of posts cited by ChatGPT led with a short answer capsule, and over nine in ten capsules contained no links.
- About 65 percent of pages cited by Google AI Mode and 71 percent cited by ChatGPT include structured data, yet no schema type predicted citations on its own.
- The academic GEO paper measured up to 40 percent visibility gains in generative responses, in a lab benchmark, not the live web.
- Several widely repeated AI search numbers could not be traced to any named primary source as of July 2026. They are flagged below, not repeated as fact.
Why most AI search statistics cannot be trusted
Because stats pages cite each other, not the research. One roundup invents or garbles a number, five roundups quote the roundup, and by the third generation the number has no source at all. It just has momentum. Search for AI search statistics today and the top results are lists of 60, 70, even 250 data points. Many of those points link to other lists, or to nothing.
This page runs on a stricter rule. Every number below links to a named primary source, with its date, its method, and what it truly measured. Each one was re-checked against that source in July 2026. Numbers we could not trace are flagged in their own section, not laundered back into use. The set is smaller than the big roundups. That is the point.
Citation policy: every statistic on this page links to a named primary source with a date and method, re-verified July 2026. If a number cannot be traced, it is flagged as untraceable, not repeated.
Do AI summaries reduce clicks? The Pew data
The cleanest public evidence comes from Pew Research Center. It analyzed the real browsing behavior of roughly 900 US adults during March 2025, and published the results on July 22, 2025.
| Statistic | Number | Source and date |
|---|---|---|
| Clicked a result when the page had an AI summary | 8% of visits | Pew Research Center, July 2025 |
| Clicked a result with no AI summary present | 15% of visits | Pew Research Center, July 2025 |
| Users who got at least one AI summary that month | 58% | Pew Research Center, July 2025 |
| Median length of an AI summary | 67 words | Pew Research Center, July 2025 |
Read it plainly. In this panel, an AI summary cut the chance of a click roughly in half, and users rarely clicked the sources cited inside summaries. Two honest limits: the data is from one month, March 2025, and it measures one panel of US adults. It is still the strongest named-source click data in circulation, which is why it is the one worth citing.
What gets cited by AI engines? Two named studies
Two separate 2025 and 2026 studies put numbers on what AI-cited pages have in common. They are the closest thing this young field has to repeated findings, and they point the same direction: structure and original substance, not tricks.
| Statistic | Number | Source and date |
|---|---|---|
| ChatGPT-cited posts that led with a short answer capsule | 72.4% | Saltbox Solutions via Search Engine Land, Nov 2025 |
| Cited posts containing original data or owned insight | 52.2% | Saltbox Solutions via Search Engine Land, Nov 2025 |
| Answer capsules containing no links at all | About 91% | Saltbox Solutions via Search Engine Land, Nov 2025 |
| Google AI Mode-cited pages with structured data | About 65% | SE Ranking research, Jan 2026 |
| ChatGPT-cited pages with structured data | 71% | SE Ranking research, Jan 2026 |
| Schema types that predicted citations on their own | None; correlations -0.106 to +0.039 | SE Ranking research, Jan 2026 |
The method notes matter more than the headlines. The Saltbox audit covered 15 domains, close to 2 million monthly organic sessions, and about 7,500 ChatGPT referral sessions. Real, but modest in scale. The SE Ranking finding is the one vendors misquote most. They cite the 65 percent as proof that schema markup earns citations. The same research found no schema type raises citation odds by itself. Both halves belong together, as the full guide to getting cited by chatgpt breaks down.
Is GEO proven? What the research paper measured
The term generative engine optimization comes from one traceable place: a paper by Aggarwal and colleagues, first posted in November 2023 and accepted to KDD 2024. Its headline number: content changes like adding citations, quotes, and statistics boosted visibility "by up to 40%" in generative engine responses.
Cite it with its context attached. The 40 percent came from a bench of test queries in a lab, not from live traffic on the open web. It is proof that how you present things changes what AI engines quote. It is not a promise that any site gains 40 percent of anything. Numbers stripped of that caveat are how the graveyard below gets new residents.
Policy facts with dates: what platforms actually changed
These are not statistics, but they are the dated facts that AI search statistics get built on, and they are the ones writers most often get wrong.
| Fact | Date | Source |
|---|---|---|
| Google says no special requirements or AI files are needed for AI Overviews or AI Mode | Doc updated Dec 2025 | Google Search Central, AI features |
| FAQ rich results removed from Google Search entirely | Announced May 2026 | Google Search Central changelog |
| HowTo rich results retired | 2023 | Google Search Central |
| Google confirms the keywords meta tag is not used in web ranking | Sept 2009 | Google Search Central blog |
| Knowledge Graph launches: "things, not strings" | May 16, 2012 | Google official blog, Amit Singhal |
| llms.txt proposed as a convention for AI crawlers | Sept 2024 | llmstxt.org, Jeremy Howard |
One small original dataset, offered honestly
This site is a rebuilt domain that has run an Entity-First Method for about three months. Here is its own Search Console data, with the sample size stated plainly: one site. Between April 21 and July 20, 2026, daily impressions roughly doubled, from the low thirties to a 60 to 85 range, on zero link building. And among the logged queries are prompts clearly written by AI assistants searching on behalf of their users, complete with instructions to the search engine inside the query text.
We have not seen that last observation documented elsewhere. AI agents are already visible inside ordinary Search Console reports, before most tracking tools report them. One site proves nothing at scale. But unlike an untraceable percentage, you know exactly where this came from, and you are welcome to cite it as what it is.
The graveyard: numbers we could not trace
Each claim below circulates in current stats roundups. As of July 2026, we could not follow any of them to a named primary source with a stated method. That does not prove they are false. It does mean you should not cite them until someone shows the study. If you know the primary source for one, email us and we will move it upstairs with credit.
- The share of AI citations that come from earned media, commonly given as a suspiciously clean 84 percent, with no named study attached.
- The share of total search that AI tools now handle, quoted anywhere from a quarter to more than half, with methods that are never described.
- Precise AI referral market shares quoted to one decimal place, with no measurement panel named.
- Click-loss percentages attributed to unnamed "field experiments" with no author, venue, or date.
Here is the 60-second check that filled this section. Click the link under any statistic. If it lands on another roundup, keep clicking. If you never arrive at a named group with a date and a method, the number is folklore wearing a percent sign. That check is also step one of any honest content audit.
How to cite this page
Everything here is free to cite with a link. The page says exactly when each number was last checked, and we re-check the sources quarterly. A link here will not rot into the graveyard it warns about. Suggested citation: "AI Search Statistics You Can Actually Cite, JNAbear, updated July 2026." Corrections are welcome at info@jnabear.com, and confirmed corrections get credited on this page.
Want the strategy behind the numbers instead? Start with the ai seo checklist, which sorts the work by evidence strength. Or see how topical authority mapping turns findings like these into a content plan. Both follow the same rule as this page: claims trace to sources, or they get labeled as bets.
Sources & further reading
- Pew Research Center: Google users are less likely to click on links when an AI summary appears (July 22, 2025)
- Search Engine Land: How to get cited by ChatGPT, the content traits LLMs quote most (Adam Gnuse, Nov 2025)
- SE Ranking: Structured data for SEO and LLMs (research, January 2026)
- GEO: Generative Engine Optimization (Aggarwal et al., arXiv 2311.09735, KDD 2024)
- Google Search Central: AI features and your website
- Google: Introducing the Knowledge Graph, things not strings (May 2012)
- llms.txt: the proposal (llmstxt.org, September 2024)
Topics & entities in this article
Frequently asked questions
The best named-source evidence says yes. Pew Research Center analyzed March 2025 browsing from about 900 US adults. Users clicked a result on 8 percent of pages with an AI summary, versus 15 percent without one, roughly half the click rate.
No Google-published number exists, and tracker estimates vary widely with the query mix they sample. That is why this page does not quote one. In Pew's March 2025 panel, 58 percent of users got at least one AI summary during the month, which is a user-level view, not a query-level one.
Per a November 2025 audit of 15 domains, 72.4 percent of ChatGPT-cited posts led with a short self-contained answer. Another 52.2 percent contained original data. Over nine in ten of those answer blocks were link-free. Structure plus original substance, in short.
The data is two-sided. About 65 to 71 percent of AI-cited pages include structured data, per SE Ranking research from January 2026. Yet the same research found no schema type that predicts citations on its own. Schema declares your entities; it does not buy citations.
It is real and traceable: the GEO paper, accepted to KDD 2024, measured visibility gains of up to 40 percent in generative responses when content added citations, quotations, and statistics. It came from a lab benchmark, not live traffic, so cite it with that caveat.
Click the link under the number. If it lands on another roundup, keep clicking until you reach a named organization with a date and a stated method. If you never do, treat the number as folklore. Every statistic on this page passes that test or is flagged as untraceable.
Related service
Topical Authority Mapping
Topical authority mapping structures your entire topic space around entities. The map defines every pillar, cluster, and gap, so your site covers the subject comprehensively and search engines treat you as the authority.