How AI Answers Questions About Brands · Part Two of Four
Part Two: The Source Study
2,230 cited pages audited · 25,271 citations traced · 10 national brands
The series
Part One · The Retrieval Study — do the engines even search before answering?
Part Two · The Source Study — you’re reading it
Part Three · The Reddit Pipeline — the one source every engine reads, and what it says about you
Part Four · The Agency Playbook — what to fix first, and how much holds still from one day to the next
Where an AI answer actually comes from
When a buyer asks an AI assistant about a brand — is it worth it, which one’s best, is it legit — the assistant usually goes out to the web, picks a handful of pages, reads them, and writes its answer from what those pages say. The pages show up under the answer as citations. Most people never click one. Almost nobody asks the bigger question: what’s actually on them?
It matters because an AI answer isn’t a ranking. A Google ranking is a position on a list you can see — type the search, count down the page, there you are. An AI answer is composed, written from a stack of pages the engine chose and read on the buyer’s behalf. Your answer is only as good as the engine’s reading list, and nobody gets to see the reading list. Part One of this study measured how often the engines search at all; this part is about what they read when they do.
So I pulled the reading list. Across ten national brands — Gymshark, YETI, Hoka, Helix, Peak Design, AG1, Sonos, Patagonia, OLIPOP and Caraway — the four engines put 25,271 citations behind 2,000 live answers, pointing at 10,648 different pages. From those, the audit selected the fifty or so most load-bearing pages per question type for every brand: 2,230 pages. I sent the instrument after every one of them.
“Read” is a strong claim, so here is exactly what it means. Every page the instrument could open went through its reading layer: the full text, the entire comment tree when it was a forum thread, the complete transcript when it was a video — with up to thirty-one judgments recorded per page. Where does the brand appear, and how prominently? What does the page actually say about them? Which rivals share it? Who wrote it, and when? Did it even open? The machine read every word. I read the record it produced.
Here’s what a single entry in that record looks like — AG1’s cited YouTube videos, exactly as the reading layer logged them:
Finding 1 · One website is on every brand’s reading list: Reddit
Start with the simplest question you can ask about a reading list: which websites keep showing up? Ten brands, ten different industries — cookware to coolers to mattresses — so you’d expect ten different rosters. Mostly, you get one.
Exactly one website was cited for all ten brands by the engines: Reddit. Here are the sites that recur across the study, counted from all 25,271 citations:
- reddit.com1,606citations — the only site on all ten brands’ lists
- youtube.com402all ten brands
- forbes.com353nine of ten brands
- amazon.com292all ten brands
- trustpilot.com170all ten brands
- en.wikipedia.org95all ten brands
Look at what that list actually is. A forum. A video site. A store. A review site. An encyclopedia. The places your buyers talk to each other and the places that rank things. What’s not on it is any brand’s own website — the surface most marketing budgets are spent polishing. The engines’ reading list is overwhelmingly about you, and overwhelmingly not by you.
Finding 2 · Four engines, four completely different diets
The roster hides a second story: the four engines are not eating from the same table. Take Reddit, the one source they all touch. For Perplexity, Reddit is a tenth of its entire diet — 1,517 of its citations. For Gemini it’s about four percent. For ChatGPT, one citation in a thousand. And Claude? Claude cited Reddit exactly zero times — in 7,132 citations. Same questions, same day, same ten brands.
The engines don’t even agree on how much to show you. Here’s the median number of sources each engine displayed with an answer:
- Perplexity3030 is the most it ever shows on one answer — and its typical answer shows all 30
- Claude15
- Gemini4
- ChatGPT0the median ChatGPT answer shows no sources at all
For anyone doing this work professionally, this is the finding to sit with. There is no single “AI diet” to feed. The page your brand lives on for one engine can be invisible to another, and a source strategy built by looking at only one engine’s citations is a strategy for a quarter of the market. Here’s Reddit’s share of each engine’s diet, drawn out:
Reddit’s share of each engine’s citation diet
Bar chart of Reddit's share of each engine's total citations: Perplexity 10.2%, Gemini 4.2%, ChatGPT 0.1%, Claude a measured zero out of 7,132 citations.
| Label | Reddit share of citations (%) |
|---|---|
| Perplexity · 10.2% | 10.2 |
| Gemini · 4.2% | 4.2 |
| ChatGPT · 0.1% | 0.1 |
| Claude · 0 | 0 |
Share of each engine’s own citations that point at reddit.com, across all ten brands. Claude’s bar isn’t missing — it’s a measured zero: 0 Reddit citations out of 7,132. Shares are within-engine because the engines show different numbers of sources.
Finding 3 · A third of the reading list is somebody’s ranked list
Reading 1,828 pages tells you something counting them never could: what kind of pages the engines actually lean on. The reading layer classified every one, and the shape is lopsided. A third of everything the engines read is a ranked list — “best cookware brands,” “top 10 coolers,” somebody’s round-up with your brand at position four. Another quarter is forum threads: your buyers talking to each other. Traditional editorial — an article by a journalist, about you — is thirteen percent.
One brand breaks the pattern, and it’s worth naming: OLIPOP is the only brand in the study whose reading list has more editorial than ranked lists. That mix is earnable. Nine brands out of ten just haven’t earned it.
The full composition of the 1,828 pages the instrument read:
Every page the instrument read, classified by what it is. Two of every three pages are a ranked list or a forum thread — the places that rank you and the places that talk about you.
Being on the page is not being the page
Here’s where reading beats counting. A citation count says a page “mentions” your brand. But “mentioned” can mean the whole page is about you — or one line in a list of six. The engines cite both kinds, and only the reading layer can tell you which one you got. On every page where it could be measured — 784 of the read pages — the instrument recorded exactly how the brand appears. Three real entries from the record:
What “mentioned” actually looked like, from the reading layer’s record
| Brand | The cited page | What the reading layer recorded |
|---|---|---|
| Caraway | A top-10 cookware round-up | “Appears 3 times in top 10 list — ranked 3rd, 4th, and 10th positions.” |
| Sonos | A 12-product speaker round-up | “Three of 12 products featured; positions 5, 6, and 8 in round-up.” |
| Gymshark | A magazine gift guide | “One line in a quick links list of 6 products.” |
Finding 4 · Half the reading list is over a year old
Part One showed the engines writing years nobody asked for into their own searches. This is the other clock — not what they searched for, but how old the pages they cite turned out to be. Of the pages that carry a readable date — 1,373 of them; another 838 carry none, and that’s said plainly in the methodology — more than half are over a year old, and a third are over two.
And the age isn’t spread evenly. It lives in the forums. Nearly half the forum threads the engines cite are more than two years old; for ranked lists it’s one in six, because list publishers refresh for a living. The conversation layer of your reading list — the part where buyers talk to each other — is the part most likely to be describing a product you’ve since fixed, repriced, or discontinued.
How old the cited pages are
Bar chart of cited-page age: 55.7% older than one year, 34.9% older than two years, of 1,373 dated pages. Forum threads over two years old: 49.2%; ranked lists: 16.7%.
| Label | Share of dated pages (%) |
|---|---|
| Older than 1 year | 55.7 |
| Older than 2 years | 34.9 |
| Forum threads over 2 years | 49.2 |
| Ranked lists over 2 years | 16.7 |
Of the 1,373 cited pages with a parseable date (838 more carry no date at all). The two-year rate splits sharply by page type: forum threads 49.2%, ranked lists 16.7%.
Two entries from the record. Neither page is wrong to exist — the finding is that engines still hand them to buyers as if they were news.
Finding 5 · Nearly one page in five wouldn’t open — including the brands’ own
The last finding is the one I didn’t see coming. To read 2,230 pages you first have to open them, and 402 — 18.0% — wouldn’t open at all for our machine reader: blocked, dead, or gone. Six of the cited Reddit threads now exist only as deletion notices; three more have vanished entirely. That layer of the study belongs to Part Three.
But the group that failed hardest is the one nobody would guess. The brands’ own websites: 77 of the 132 cited own-site pages — 58.3% — wouldn’t open, mostly because the site’s bot protection turned our machine reader away. Dedicated review platforms were worse: 20 cited pages, 20 blocked, a perfect score. The pages brands control most tightly are the pages a machine reader is least able to read.
One page makes the point better than any rate. The fourth most-cited page in the entire study is a Forbes greens-powder round-up — 27 citations, all four engines lean on it for AG1 answers. Open it in a normal browser and it loads fine — it names AG1 as its top pick. But send any machine reader, ours included, and Forbes turns it away at the door. Humans can read it. Machines mostly can’t. The engines cite it anyway. Whatever route they take to it, an automated audit can’t follow them in.
The obvious next question — the one I asked myself — is whether the engines could read those pages either. That can’t be claimed, and I won’t. Engines fetch wearing different badges: a site that blocks one crawler may wave another through, some engines hold privileged access to publishers who block everyone else, and some of these pages were alive when the engine first read them and died later — the deleted Reddit threads prove that path. What can be said without qualification is narrower, and it still matters: a citation behind a machine wall can’t be verified by any automated reader — not an audit tool, not a monitoring service, not an AI assistant asked to double-check its own source. A person can still click through; everything built on clicking by hand doesn’t scale past a person. And in the deleted-thread cases, there’s nothing left for anyone.
And one brand proves it’s a choice, not physics: Caraway’s own site let our reader through on 10 of its 11 cited pages — the best own-site record in the study. Whether a machine reader gets into your site comes down to configuration — the bot protection in front of it and the robots.txt file behind it. Both are settings, both are checkable in minutes, and for a lot of brands both were set by a security vendor’s default rather than by anyone deciding how the site should treat an AI reader.
“Wouldn’t open” = our machine reader couldn’t fetch the page at read time. A person in a browser can often still get in — the wall is for automated readers. Not a claim about the engines’ own access.
Who blocks the machine reader, by source type
Bar chart of unreachable rates by source type: review platforms 100% of 20, brands' own sites 58.3% of 132, general web pages 19.0% of 1,296, YouTube 6.3% of 252, Reddit 0.6% of 482.
| Label | Unreachable (%) |
|---|---|
| Review platforms | 100 |
| Brands’ own sites | 58.3 |
| General web | 19 |
| YouTube | 6.3 |
| 0.6 |
Share of cited pages our machine reader couldn’t open, by source type. Denominators: review platforms 20 · brands’ own sites 132 · general web 1,296 · YouTube 252 · Reddit 482. The open web stays mostly open; the controlled, commercial layer is the wall.
The ten-second version of this study
Fetch your own homepage the way an AI crawler does — curl -I -A "PerplexityBot" yourdomain.com — and look at the status code. A surprising share of brands will watch their own front door answer with an error.
The full version is below: all ten brand reports, every cited page, what it is, what it says, and whether it opened. This is a dated snapshot of a moving system — if you run it on your own clients and the shape comes back different, tell me. How much the system moves is Part Four’s question.
A personal note
Forty years in the creative industry teaches you where the money goes: onto the surfaces a brand controls. The website, the campaign, the photography. What this part of the study says, as plainly as I can put it, is that the surfaces deciding your AI answers are mostly surfaces you don’t control and have never read — a reading list that is one-third ranked lists, half over a year old, and partly unreadable even to the people it’s about.
Part Three follows the roster to its strangest entry: Reddit — the only source all four engines share, the most negative layer of the whole study, and a pipeline carried almost entirely by a single engine.
Glenn · Founder, styleforge.io · September 2026
The full record, brand by brand
The complete audit for every brand in this article is published below — the same ten reports from Part One, generated by StyleForge’s AI Visibility Pulse, and each one runs 200 to 250 pages. The reading layer this article is built on fills its own chapter in every report: every cited page, what kind of page it is, how old it is, what it actually says about the brand, and whether it opened.
Click any brand to open its full report. If you run an agency, read one the way your client would — this is the deliverable, the thing you hand across a desk.
Each of these reports is white-labeled — generated under the agency’s own name and presentable straight to the client, and you control which sections every report includes. The agency name on these ten is a stand-in.
Appendix: how it was run
Every check in the study, written out in full.
| The check | How it was run |
|---|---|
| The pages | From 25,271 citations resolving to 10,648 distinct pages, the audit selected the top ~50 most load-bearing pages per question type per brand — chosen by the audit model for quality and relevance before any page was read. That’s 2,230 pages: 1,828 read in full, 402 unreachable. Every rate in this article names which of those denominators it uses. |
| The reading layer | Each page fetched and machine-read whole — full text, complete comment trees on forum threads, full transcripts on videos — with up to 31 judgments recorded per page: prominence, sentiment, rivals present, page type, author, date, status. The record was then reviewed by a person; the machine read every word, the person read the record. |
| Dates | 1,392 pages carry a publish date; 1,373 of those parse cleanly; 838 carry none. Every staleness rate uses the 1,373 — undated pages are never assumed old or new. |
| Reachability | Status recorded at read time from a US connection. “Unreachable” means our machine reader could not open the page. A person in a browser may still get through — the finding is about what automated readers can verify, never about what the engines can reach. |
| Citation counting | Citation counts are answer-level (25,271 total). The engines cap displayed sources differently — Perplexity shows up to 30, the median ChatGPT answer shows none — so all diet comparisons are within-engine shares, never raw counts across engines. |
| Internal cross-check | Source-layer totals were recomputed from the raw rows and matched the audit platform’s own totals before drafting. One row-count differs by one (1,828 vs 1,829 read rows, a healing-era artifact); it moves no headline figure at one decimal and is disclosed here rather than hidden. |
Limitations, stated plainly
- Single dated snapshot (September 1, 2026), US locale. The reading list changes between runs; how much is Part Four’s measurement.
- The 2,230 pages are the audit’s curated top layer, not a random sample of the web — rates describe the pages doing the most work in these answers, not all pages everywhere.
- This is the reading list the engines disclosed. Part One showed Gemini sometimes searches and shows nothing; whatever it read on those answers is in nobody’s data, including this study’s.
- Sentiment and prominence are recorded only where a page supports the judgment (1,172 and 784 pages respectively); every such rate quotes its own denominator.
This study was run on StyleForge
Every measurement in this article ran on StyleForge’s AI Visibility Pulse — and the Pulse is one tool of many. StyleForge is a creative operating system for agencies and brand teams: intelligence, design, production, publishing, and measurement in one connected system that helps small agencies work like teams five times their size.
Explore the AI Visibility Pulse