The AI Retrieval Study, part one of a four-part field study on how AI answers questions about brands: ChatGPT searched the web on 166 of 500 brand answers; Claude on 445, Gemini on 494, Perplexity on 500 of 500; Google’s AI Overview appeared on 367 of 500 questions.

Do AI assistants actually search the web before answering about your brand?

I asked the four leading AI assistants — and Google’s AI Overview — 2,500 live questions about ten well-known national brands, and recorded what each one actually did. This is part one of a four-part field study on how AI answers questions about brands.

How AI Answers Questions About Brands · Part One of Four

Part One: The AI Retrieval Study

10 national brands · 50 buyer questions × 5 AI engines · 2,500 live answers · 69,361 data points

First published September 2, 2026 · updated September 3 with a fresh full run — the study now measures five engines, adding Google’s AI Overview.

The series
Part One · The Retrieval Study — you’re reading it
Part Two · The Source Study — the pages behind the answers, how old they are, and how many won’t open
Part Three · The Reddit Pipeline — the one source every engine reads, and what it says about you
Part Four · The Agency Playbook — what to fix first, and how much holds still from one day to the next

SEO vs. GEO: both camps are arguing from the answer

I’ve spent about forty years in the creative industry, and the last few building software for the people who do this work for brands. For most of this year I’ve watched the same argument loop through every feed I read. The SEO camp says AI search is just search with a new coat of paint. The GEO camp says the rules have been rewritten. Strong opinions everywhere. Measurements almost nowhere. And both camps argue from the same evidence: the answer the engine finally gave. Nobody could say what the engine actually did to produce it.

That’s the part I decided to record.

When you ask ChatGPT a question, you see the answer. What you don’t see is what it did to get there. Behind every answer there’s a hidden layer of activity: did the engine go out to the web at all, what searches did it write for itself, which pages did it lean on. The app shows you almost none of it. But the engines themselves disclose that layer through their developer channels, answer by answer, if you build something to catch it. So I built something to catch it.

Then I pointed it at ten national brands: Gymshark, YETI, Hoka, Helix, Peak Design, AG1, Sonos, Patagonia, OLIPOP and Caraway. Fifty buyer questions per brand, asked live to ChatGPT, Claude, Gemini and Perplexity — and, new in this wave, checked against Google’s AI Overview on the live search page. That’s 2,500 answers, each one kept whole, along with the hidden layer underneath it. All ten brand reports are published in full at the bottom of this piece.

What we found was hard to ignore. Presented with the same question about a Helix mattress, on the same day, ChatGPT ran the web search “Helix mattress price comparison models 2023” — while Claude searched for the 2026 version. Nobody asked either engine for a year. Each one added its own.

One definition before any number, because the study lives or dies on precision. When I say an answer showed no web search, I mean the engine’s own record of what it did shows no search. I am not claiming to know what happens inside the model, and neither should anyone else.

Finding 1 · Three engines looked. ChatGPT mostly didn’t.

Start with the simplest question in that hidden layer, because it turns out to be the loudest: when a buyer asks about your brand, does the engine even look at the live web before it answers?

Same ten brands, same fifty questions, 500 answers per engine. Here’s how often each one went out to the web:

  • ChatGPT166 of 50033.2% — didn’t search on 334
  • Claude445 of 50089.0% — didn’t search on 55
  • Gemini494 of 50098.8% — didn’t search on 6
  • Perplexity500 of 500100% — searched on every answer

Sit with that spread for a second, because it’s the whole finding. Those percentages are the share of each engine’s 500 answers where it actually went out and read the live web before answering. Three of the four engines searched on nearly nine answers out of ten, or better. ChatGPT, the engine most of your buyers actually use, searched on just one answer in three. Identical questions. Same day. These four products get talked about as one thing — “AI search” — and they are not one thing.

One thing to name before anyone else does. Every engine was asked through its API in its search-enabled mode, with the choice to search left to the engine — the same choice the consumer apps make. Whether it searched was read from each engine’s own disclosures, not guessed from the answer text.

Why should an agency care? Because this split is the first budget decision in AI visibility, and today almost nobody makes it. Content and links can reach an engine that reads the live web now. Whatever ChatGPT is doing on the other 334 answers, your new page is not part of it.

Here is that same count drawn out square by square, so you can see what 33.2% looks like sitting next to 100%:

Did the engine search the web? · each square = 1% of that engine’s 500 answers
ChatGPT
166 of 500 searched
Claude
445 of 500 searched
Gemini
494 of 500 searched
Perplexity
500 of 500 searched

How this was recorded

Every question was asked live, five times over — once per AI assistant through its API, plus Google’s AI Overview read off the live search page, and every answer kept whole. When an engine searched, I kept the searches it wrote for itself: 2,466 of them, verbatim. And I captured every source each engine showed with its answer: 21,213 citations, pointing at 8,247 different pages.

Then came the slow part, and it’s the part I care most about. Knowing an engine leaned on a page tells you very little by itself. So I opened the cited pages — 2,357 of them, every one sent through a full machine read — and recorded what each page actually is. The machine read every word; I read the record it produced. Where does the brand appear on it, headline or footnote? What does the page really say about them: praise, complaint, or a passing mention? Which competitors share that page, and how do they come off? What kind of page is it — a ranked list, a forum thread, a news story, a review? Who wrote it, and when? A citation that trashes you still counts as a citation, and this is the layer that tells the difference.

The videos got the same treatment. When an engine cited a YouTube video, I didn’t log the link and move on. The whole transcript came with it: what was said about the brand, who said it, at what timestamp, how many times the brand came up, how many times a rival did, and in what light. That’s 368 full transcripts, 54.8 hours of video, never clipped. And when a source was a Reddit thread, the entire comment tree came along too.

Stack all of that together and the count comes to 69,361 recorded observations. I went that deep because that’s where the story actually lives. The answer a buyer gets isn’t conjured from nothing. It’s assembled from this layer — the pages, the threads, the videos, and what every one of them says about you and the brands you compete with. The pages and their age carry Part Two. The Reddit layer carries Part Three. The work order carries Part Four. This part is about the first thing recorded: whether the AI looked at the web at all.

It isn’t the brand. It’s the engine.

The natural objection: maybe famous brands don’t need looking up, and these numbers just reflect who’s in the sample. The data says no. Across the three engines that report their own searches, the ten brands run from 70.0% (OLIPOP) to 80.7% (Helix). That’s an 11-point spread across brands, against a 66-point spread across those same three engines. Add Perplexity’s 100% and the four-engine spread is 66.8.

One honest caveat, because a careful reader will find it anyway. Each brand’s rate pools three engines together, and pooling flattens differences. So here is the same test inside a single engine: on ChatGPT alone, the ten brands run from 24% (Caraway and OLIPOP) to 44% (Hoka). A 20-point spread — still a third of the engine gap. The conclusion holds either way.

Whether an AI engine reads the web before answering about you is mostly a property of the engine, not of your brand, your fame, or your content. Be precise about what that does and doesn’t say. It doesn’t mean content is irrelevant — when an engine does search, what it finds is entirely about your pages. It means you don’t get to earn the search. That decision was made before your question arrived.

Finding 2 · ChatGPT searched least exactly where buyers decide

The standing theory — mine included, going in — says engines search when they don’t already know. Rare questions force lookups; familiar ground doesn’t. The data inverts that, and only on ChatGPT.

The fifty questions per brand come in five types: category (“best cookware brands”), brand (“is Helix worth it”), versus (“YETI vs Pelican”), long-tail (“best mattress if I have fibromyalgia”), and reputation (“is AG1 legit”). I’ll call these lanes; you’ll see them again in every part of this series.

On the ground it should know best, category and brand questions, ChatGPT searched about half the time. On the questions buyers decide with, it stopped. “X vs Y”: 4 searches in 80. “Best X for my situation”: 17 in 120. The harder the buying question, the less it looked anything up.

Claude, Gemini and Perplexity don’t have this dip. Ask them any of the five question types and their search rate barely moves: never below 85%, usually close to 100%. The collapse on versus and long-tail questions is ChatGPT’s alone, and it lands exactly where a buyer is comparing you to someone else. Whatever content you publish to win “your brand vs rival,” ChatGPT read the live web for that answer 4 times in 80.

Here’s the full grid, every engine against every question type. The two orange cells are the finding:

How often each engine searched, by question type
ChatGPT
Claude
Gemini
Perplexity
Categoryn=180
52.2%
90.6%
99.4%
100%
Brandn=80
45%
87.5%
100%
100%
Versusn=80
5%
91.2%
98.8%
100%
Long-tailn=120
14.2%
86.7%
99.2%
100%
Reputationn=40
37.5%
87.5%
92.5%
100%

Each cell: how often that engine searched the web for that type of question, across all ten brands. The “engines search when they don’t already know” thesis inverts on ChatGPT — and only on ChatGPT.

The question types: category “best cookware brands” · brand “is Helix worth it” · versus “YETI vs Pelican” · long-tail “best mattress if I have fibromyalgia” · reputation “is AG1 legit”

Finding 3 · The engines wrote their own keyword list

How often an engine searches doesn’t track with how hard the question is. How hard it searches does. Versus questions drew 3.40 searches per searching answer; brand questions drew 1.91. Gemini writes 5.49 separate searches for a single “X vs Y” answer. That answer is assembled from five different result sets, and you are fighting for space in all of them. ChatGPT writes exactly one search, every time. (Reading search count as “difficulty” is my interpretation; the counts are the measurement.)

And the searches themselves are the find of the study. When an engine searches, it doesn’t forward the buyer’s question. It writes its own. I’ll be honest: I build software in this space and I didn’t know that until the data made it impossible to miss. But it isn’t a secret, and you don’t have to take it from me — it sits in plain sight in the engines’ own developer channels. When an engine searches, it reports back the exact query strings it wrote, answer by answer, and anyone with API access can read them. Gemini even shows some of its searches right in the app, inside the sources panel.

Three patterns I didn’t expect

One search in seventeen was dated to a year the buyer never mentioned. 141 of the 2,466 queries carried a year older than 2026 that appeared nowhere in the question, and it happened for every one of the ten brands. One real pair from the record: a buyer asks for the best down jackets for cold weather, no year anywhere in sight, and ChatGPT’s search reads “best down jackets for cold weather 2023.” ChatGPT wrote 93 of those stale-year searches, Gemini 29, Claude 19. (Another 61 queries echoed a year that was in the question’s own text; those are counted separately.)

They compare without being asked. 206 queries were comparison-phrased, and 50 of those were written for questions that never asked for a comparison at all. Someone asks Gemini for a technical jacket for alpine mountaineering — a question with no brand in it anywhere — and Gemini’s searches read “gore-tex pro vs active alpine mountaineering” and “pertex shield vs gore-tex.” The buyer asked a plain question. The answer got built out of a materials fight. You don’t get to opt out of the versus page. You only decide whether yours exists.

And they bring your rivals into your own answers. The engines typed a competitor’s name into 220 of their own searches, across 44 brand-and-rival pairs. YETI drew Pelican 12 times. Sonos drew Bose 8. Gymshark drew Lululemon 8 — and Nike 8 more. The pair that stopped me: a buyer asks for top trail running shoes for beginners — no brand in the question, squarely in Hoka’s category — and Gemini’s searches come back “Nike ACG Pegasus Trail beginner review” and “Altra Lone Peak 9 beginner review.” The buyer asked for beginner shoes. The engine decided whose reviews to read. Run this on a client and you get their real rival list — not the one from the pitch deck, the one the engines reach for. It’s usually not the list the client expects.

The harder the question, the harder they search

Bar chart of search queries written per searching answer by question type, across the three engines that report their queries: versus 3.40, long-tail 2.27, reputation 2.24, category 1.92, brand 1.91, from 2,466 captured queries.

Bar chart of search queries written per searching answer by question type, across the three engines that report their queries: versus 3.40, long-tail 2.27, reputation 2.24, category 1.92, brand 1.91, from 2,466 captured queries.
LabelSearches written per searching answer
Versus3.4
Long-tail2.27
Reputation2.24
Category1.92
Brand1.91

Average number of searches written per searching answer, from the three engines that report their queries — 2,466 captured searches in all. Gemini alone runs 5.49 searches for a single versus answer and 3.05 for a brand answer.

What the buyer asked vs. what the engine searched

Brand · laneThe buyer’s questionEngineThe search the engine actually ran
Helix · brandHelix mattress price comparison modelsChatGPT“Helix mattress price comparison models 2023”
Helix · brandHelix mattress price comparison modelsClaude“Helix mattress price comparison models 2026”
YETI · reputationIs YETI worth the money?Gemini“YETI coolers vs competitors value” · “YETI tumblers vs competitors value” · “YETI bags vs competitors value”
OLIPOP · categoryProbiotic drinks similar to kombuchaGemini“kefir vs kombucha” · “kvass vs kombucha” · “rejuvelac vs kombucha”

Read that first pair twice: the same question, on the same day, sent one engine searching in 2023 and another in 2026. If your comparison page is dated, one of those engines is reading your rival’s older one. And the YETI row is a reputation question — “is it worth the money” — that Gemini answered by going shopping across three product lines. The brand never chose that comparison. The engine did.

Gemini searched on 90 answers and showed you nothing

If you check AI answers by hand, the only clue you can see is the row of source links under the answer. No links, no search — that’s the natural read, and for ChatGPT and Claude it happens to be roughly true.

For Gemini, the clue lies. On 90 of its 500 answers, Gemini searched the web and then showed no links at all. It looked, it read, and it kept its sources to itself. Judge Gemini by its links and you’d call it the blindest engine in the group. Measured by its own record, it’s the second-heaviest searcher in the study.

So before you tell a client “this engine never looked at your site”: counting links only works on some engines. On Gemini it will send you to exactly the wrong conclusion.

The links you can see vs. the searches that actually happened

Grouped bar chart per engine comparing answers where no web search happened (ChatGPT 334, Claude 55, Gemini 6, Perplexity 0 of 500 each) against answers that displayed zero source links (ChatGPT 339, Claude 55, Gemini 96, Perplexity 0). Gemini's two bars differ by 90 answers: it searched the web but showed no links.

Grouped bar chart per engine comparing answers where no web search happened (ChatGPT 334, Claude 55, Gemini 6, Perplexity 0 of 500 each) against answers that displayed zero source links (ChatGPT 339, Claude 55, Gemini 96, Perplexity 0). Gemini's two bars differ by 90 answers: it searched the web but showed no links.
LabelAnswers where no search happenedAnswers showing zero source links
ChatGPT334339
Claude5555
Gemini696
Perplexity00

Out of 500 answers per engine. On ChatGPT and Claude the two bars all but match — the visible links tell the truth. Gemini’s bars are 90 answers apart: it searched, then showed nothing. Perplexity always searches and always shows its sources — though never the queries it wrote to find them.

The fifth engine doesn’t get a choice. Google does.

This wave adds a fifth engine, and it plays by different rules. Google’s AI Overview — the AI answer sitting above the results on the search page a billion people use — never decides whether to search. An Overview is retrieval. The choice Google makes is a different one: whether to answer at all.

Across the same 500 questions, Google showed an Overview on 367 — 73.4%. And the pattern in the other 133 is the finding:

  • Versus78 of 8097.5% — “X vs Y” almost always gets an Overview
  • Reputation39 of 4097.5% — “is it legit” almost always gets one too
  • Brand73 of 8091.2%
  • Long-tail71 of 12059.2% — two in five got no Overview
  • Category106 of 18058.9% — “best X” questions, same gap

Google declines to answer exactly where discovery happens. On the questions where a buyer has already named your brand or your rival, the Overview is nearly universal. On the broad “best X” and specific-situation questions — the ones a brand is trying to get found in — it sits out roughly two times in five. Reading that as editorial caution is interpretation; the split itself is the measurement.

When it does answer, it’s a strong witness. On shown Overviews it named the brand in question 73.0% of the time — the highest mention rate of the five engines — and cited the brand’s own site on 23.7%, second only to Perplexity and about six times ChatGPT’s rate. What it reads to build those answers — and how much of that is Reddit and YouTube — is Part Two’s territory, and it changes Part Three’s story outright.

One more property, and it matters for anyone doing this work: like Perplexity, Google discloses none of the searches behind its Overviews. Two of the five engines now run retrieval whose shape you cannot see. The 2,466 captured queries in this study all come from the three that show their work.

We tested whether fame changes it. It doesn’t.

One part of this study was designed before a single question was asked. The theory: famous brands live in the models’ memory, so ChatGPT should skip searching on famous names more often than on smaller ones. The brand slate was built to test exactly that — Patagonia and Sonos as the famous pair, OLIPOP and Caraway as the mid-tier pair. If the theory held, the two pairs would come back far apart.

They came back seven points apart, in the wrong direction. ChatGPT skipped searching on 69.0% of the famous-brand answers and 76.0% of the mid-tier answers (100 each). Fame doesn’t decide whether ChatGPT looks you up; ChatGPT decides. (The one directional signal ran on Claude, and opposite to the prediction: it searched more for the mid-tier brands.)

I’m publishing this one prediction that didn’t pan out right next to the ones that did, because designed tests you only report when they flatter you aren’t tests. And the result earns its place: it closes the last easy escape from Finding 1. You can’t say “well, these are famous brands, of course it doesn’t search.” Famous or not, the engine behaves the same.

Does fame change whether ChatGPT searches? No.
69.0%
MEGA-BRANDS
Patagonia · Sonos · n=100
76.0%
MID-TIER
OLIPOP · Caraway · n=100

The share of answers where ChatGPT never searched the web — famous pair vs. mid-tier pair, 100 answers each. Seven points apart — and in the wrong direction for the theory: fame made no difference in the direction predicted.

Don’t take my word for any of this

If you do this work for a living, you don’t need instructions from me — and you’re exactly the reader I expect to challenge these numbers. Good. Everything in this article reproduces on any brand, in minutes, with no special tooling: ask the same buyer question in all four assistants, google it and see whether an AI Overview answers, read the search strings Gemini discloses in its sources panel, count what each engine shows you.

And if you run this on your own clients and see different proportions, tell me — genuinely. This is a dated snapshot of systems that move. How much they move, measured a day apart, is Part Four.

A personal note

After forty years in the creative industry I’ve watched a few of these platform shifts from the inside, and the pattern repeats: the loudest arguments happen before anyone has measured anything. This series is my attempt to skip the argument stage. And if one number here changes how you plan, let it be the first one: these five engines are not one channel. A plan that treats “AI visibility” as a single job is doing work that one of the engines, most of the time, never sees. Five engines, ten brands, every answer on the record — what they searched, what they cited, what those pages really say, and how much of it holds still from one day to the next.

Part Two opens the reading list: the 2,357 cited pages I fetched and read, how old they are, and how many wouldn’t open at all. If you’re curious what your own brand’s answers are built on, the checklist above is the same one I’d run.

Glenn · Founder, styleforge.io · September 2026

Want to see what an AI actually says about a brand? Read one.

For every brand in the study, the complete audit is published below — all fifty questions, every answer all five engines gave, the grading, every citation, and the page-by-page audit of those citations with the observations recorded along the way. A typical report runs 200 to 250 pages. It’s the whole file, not highlights.

These reports are what the instrument produces. The tool is called the AI Visibility Pulse — it’s the StyleForge audit that agencies run on their own clients, and it generated every page below. The pace is the part that still surprises me: the initial read of a brand takes under two minutes, asking all fifty questions live across the five engines takes about four, and the full citation audit about ten more. All ten brands in this study, including researching each brand and its competitor set, took me about two hours end to end.

Click any brand to open its full report. If you run an agency, read one the way your client would — this is the deliverable, the thing you hand across a desk.

Each of these reports is white-labeled — generated under the agency’s own name and presentable straight to the client, and you control which sections every report includes. The agency name on these ten is a stand-in.

Appendix: how it was run

Every check in the study, written out in full.

The checkHow it was run
SampleTen national brands, chosen deliberately, not sampled — mid-tier DTC plus two mega-brands (that split was itself a designed experiment; null result reported above). 50 questions per brand: category 18 · long-tail 12 · brand 8 · versus 8 · reputation 4 — identical for all ten.
Engines & modeChatGPT, Claude, Gemini, Perplexity — each queried live through its API in the default consumer mode with web access on, and the choice to search left to the engine — plus Google’s AI Overview, read from the live US search page for the same fifty questions. 2,500 answers; four capture calls failed (all on the Overview) and are excluded from every rate.
Retrieval capturePer answer, from each engine’s own record of what it did: Gemini’s grounding metadata, Claude’s search tool use, OpenAI’s web-search calls. Every one of the 2,000 assistant answers came back a clear yes or no; none were inferred. For the Overview the recorded state is different in kind — whether Google showed one at all. It did on 367 of 500; the 129 no-Overview questions are excluded from every rate, never scored as absences. One earlier run predates this logging — it has no search record at all, so it can’t be scored either way. It’s excluded in full, never merged in.
Fan-out queriesThe engines’ own search strings, verbatim — 2,466 across 1,105 retrieving assistant answers. Perplexity and Google’s Overview disclose no queries — a property of those products, not missing data. Engine-invented stale years counted separately from years echoed out of question text (61 echoes excluded from the 141).
Citation captureEvery source shown with every answer — 21,213, resolving to 8,247 distinct pages. The cited-page reading layer (2,357 pages, up to 31 judgments each) is the top ~50 per lane per brand, chosen by the audit model for quality and relevance before any page was read. It powers Parts Two and Three.
Internal cross-checkEvery headline number was recalculated from the raw answer records and had to match what the platform reported — 420 of 420 matched. Every lane rate multiplies back to its engine total (166, 445, 494, 500), and the per-answer search rates multiply back to 2,466. Check it. This catches arithmetic and plumbing errors; it does not independently verify the engines’ own disclosures. Nobody outside the engines can.

This study was run on StyleForge

Every measurement in this article ran on StyleForge’s AI Visibility Pulse — and the Pulse is one tool of many. StyleForge is a creative operating system for agencies and brand teams: intelligence, design, production, publishing, and measurement in one connected system that helps small agencies work like teams five times their size.

Explore the AI Visibility Pulse

Video Transcripts