How AI Answers Questions About Brands · Part Four of Four
Part Four: The Agency Playbook
557 pages that never name the brand · 35 pages all four engines trust · 2,000 questions asked twice
The series
Part One · The Retrieval Study — do the engines even search before answering?
Part Two · The Source Study — what the engines actually read when they do
Part Three · The Reddit Pipeline — the one source every engine reads
Part Four · The Agency Playbook — you’re reading it
The meeting every agency is sitting in now
It starts with a screenshot from the client. An AI answer that recommends a rival, or repeats a complaint, or simply never mentions them. Then the question every agency has learned to brace for: “What do we do about this?”
Data isn’t the problem. Good tools expose real GEO and AEO numbers now — Semrush and Ahrefs among them — and anyone can see who’s mentioned where. What’s been missing is the part the client is actually asking for: the order of operations. What to fix first. What each fix is worth. And what number will move next month to prove it worked.
Parts One through Three of this study measured the terrain: which engines search, what they read, and what Reddit does inside their answers. This part is about what those measurements mean for that order of operations — what kind of work the data actually supports, and what to measure after it.
Finding 1 · Mentions and citations are two different campaigns
Start with the split most plans miss. Across the 1,500 answers from the three engines that report their own searching, answers where the engine searched the web got cited pages behind them — and a citation went to the brand 17.1% of the time. Answers where it didn’t search? Zero citations. Not few — zero, in all 406 answers. Searching is the only road to a citation.
But here’s the twist: those no-search answers still mentioned the brand just as often — 69.2%, against 67.9% when the engine searched. Retrieval buys links, not love. Being named lives in a different place than being linked — and a plan that funds “AI visibility” as one line item is quietly buying only one of the two.
One confound, said out loud
The mention comparison above pools all question types, and the engines skip searching most on exactly the question types where mention rates run low — so the clean version of the claim is the citation number, which needs no adjustment: no search, no citation, 0 for 406. On the long-tail lane alone the mention comparison actually runs the other way (32.1% searched vs 41.5% not). We print the confound rather than the convenient version.
Finding 2 · The fix-first list already exists, and it’s short
Part Two read 1,828 cited pages. Sort them by how many engines lean on them and a species split appears. Pages cited by three or four engines — 19.1% of the read set, just 425 pages, only 35 of them carrying all four — are overwhelmingly ranked lists, almost never negative (4.8%), and usually name the brand. Single-engine pages are the opposite animal: forum-heavy, 24.2% negative, and brand-absent a third of the time.
That split is an ordering principle. The consensus pages are the shared load-bearing wall of every engine’s answer — fix those first. For Caraway, one page — a cookware review on theroundup.org, mixed verdict, cited 15 times by all four engines — does more work in AI answers than dozens of others combined. Peak Design’s own warranty page is a four-engine consensus source, proof that an owned page can make the list. And its Trustpilot page is a four-engine consensus source that no machine reader can open — a person in a browser sees it fine, but every automated audit hits a wall there, ours included.
Consensus pages and single-engine pages are different species
Grouped bar chart comparing pages cited by three or more engines against single-engine pages: negative sentiment 4.8% vs 24.2%; never names the brand 12.7% vs 34.8%.
| Label | Consensus pages (3–4 engines) | Single-engine pages |
|---|---|---|
| Negative about the brand | 4.8 | 24.2 |
| Never names the brand | 12.7 | 34.8 |
The 425 pages cited by 3–4 engines vs. single-engine pages, of the 1,828 read. The pages every engine shares are safer, brand-present, and few — a workable priority list.
Finding 3 · A third of the reading list never says your name
Here is the study’s biggest opportunity number. Of the 1,828 cited pages the instrument read, 557 — 30.5% — never mention the brand whose buyers they’re answering. These pages already rank, already get cited, already sit inside your category’s answers. They simply don’t include you.
The spread across brands is the proof it’s fixable: AG1’s reading list is brand-absent on 48.2% of pages, OLIPOP’s on 43.5%, Caraway’s on 42.4% — while Patagonia’s sits at 16.6%. Same engines, same kinds of pages. The difference is presence on pages someone earned. For an agency, each of those 557 pages has a known URL, a known author, and a known reason to update. That isn’t a statistic — it’s a call list.
Share of each brand’s cited pages that never name it
Bar chart of brand-absent share of read cited pages: AG1 48.2%, OLIPOP 43.5%, Caraway 42.4%, Gymshark 39.2%, study average 30.5%, Patagonia 16.6%.
| Label | Cited pages that never name the brand (%) |
|---|---|
| AG1 | 48.2 |
| OLIPOP | 43.5 |
| Caraway | 42.4 |
| Gymshark | 39.2 |
| Study average | 30.5 |
| Patagonia | 16.6 |
Of each brand’s read cited pages. The ten-brand average is 30.5%; the spread from 48.2% down to 16.6% is the evidence that this number responds to work.
Finding 4 · Your rival list is short, specific, and already written
When the reading layer recorded which competitors appear on cited pages, the sets came back category-locked. Only two names — Nike and Adidas — recur across multiple brands’ reading lists. Everyone else’s competitive field is a handful of names that never leave the category. The engines have already decided who your client actually competes with — and it’s a shorter list than the one the client worries about, and often a different one.
Caraway’s entire rival field, as the engines’ reading list carries it
| Rival named on Caraway’s cited pages | Pages |
|---|---|
| GreenPan | 51 |
| Calphalon | 41 |
| Made In | 33 |
| All-Clad | 31 |
| Lodge | 30 |
| Le Creuset | 10 |
Finding 5 · Your own site has a per-engine ceiling
How often does an engine cite the brand’s own website in that brand’s answers? The rates are low everywhere, and they’re an engine property like everything else in this series:
- Perplexity29.2%of its 500 answers cite the brand’s own site
- Gemini18.4%
- Claude14.8%
- ChatGPT4.2%and for one brand — Hoka — 0 of 50
Pair that with Part Two’s finding that 58.3% of cited own-site pages wouldn’t even open for a machine reader, and the shape of the own-site job is clear: modest ceiling, mostly unclaimed. And on the demand side, the biggest hole in the study: the long-tail lane — “best X for my situation” questions — where brands get mentioned in just 35.8% of answers overall, and AG1 in 2%. The questions closest to a purchase are the questions brands are most absent from.
The closing experiment · We asked all 2,000 questions twice
None of this matters if the target citation won’t hold still. So the study ends with the measurement everything else depends on: we re-asked every one of the 2,000 questions, same engines, same wording, within a twenty-four-hour period — and compared everything. The answer mention rates barely moved: 69.0% became 68.3%. But underneath, the engines rewrote roughly half their own search queries — and what happened to the cited sources splits cleanly by engine:
Asked again within 24 hours: how much held still
| Engine | Cited the same pages | Same websites | Rewrote its own searches |
|---|---|---|---|
| Perplexity | 90.3% of the time | 92.1% | — (doesn’t disclose queries) |
| Claude | 66.8% | 70.1% | 54% rewritten |
| ChatGPT | 51.7% | 56.0% | 48% rewritten |
| Gemini | 33.9% | 38.4% | 74% rewritten |
Like searching, like Reddit, like everything else in this series: consistency is an engine property. This measures a single day — how the picture decays over weeks is the follow-up measurement, and it isn’t published until it’s run.
For an agency this number is the license to work. The reading list holds still enough to be worth fixing — and moves enough that every fix needs a re-measurement behind it. Where recurrence is high, a page you win stays won. Where it’s low, visibility is a monthly reading, not a trophy. The instrument this series ran on treats that as a feature: StyleForge’s AI Visibility Pulse versions every run, so a second pulse reports its variance against the previous one — every engine rate, share-of-voice position, and question type carrying its change:
Our approach, disclosed
Everything in this series ran on StyleForge’s AI Visibility Pulse, and this is the part where I say plainly what it is and why we built it this way. Plenty of tools now report GEO and AEO data, and some of it is excellent. Our premise is narrower: arm the agency with findings it can put in front of a client and act on together. The reports exist so agency and client can work from the same document toward the same goal — growing the brand’s visibility with AI assistants — with every finding pointing at a fix, and every fix pointing at the number that will show it worked.
In practice: the question pack is drafted for you by AI — or drag in your own list from a template — with the competitors, question types, and weighting under your control, so an audit can be an investigation, not a preset. The citation audit reads the sources channel by channel, at the depth you choose, down to full YouTube transcripts. Every pulse is versioned, and a second run reports its movement against the last one at the question-type level. And the client report — 200 to 250 pages, white-labeled to the agency, section-selectable — closes on a recommended-actions page where every action names an owner and the number that will move next pulse — the same discipline this series has practiced throughout: no claim without a number, no action without a re-measurement behind it.
The complete audit for every brand in this series is published below — every question, every answer, every cited page read. Click any brand to open its full report.
Each of these reports is white-labeled — generated under the agency’s own name and presentable straight to the client, and you control which sections every report includes. The agency name on these ten is a stand-in.
A personal note, to close the series
I’ve been building websites since 1995 — hundreds of them, hand-coded, optimized, fussed over — and I watched the whole SEO industry grow up around that work: the early tools, the guesswork, the slow arrival of real measurement. That experience is why I added an AI visibility tool to StyleForge: agencies deserve to see what’s actually happening under the hood of these answers, and to have a way to act on it with their clients — so no brand finds out secondhand how the assistants talk about them.
Just as the analytics tools of early SEO evolved and got better year over year, I think that’s exactly what’s happening with GEO right now. This series is my contribution to it: measurements on the table, the full record published.
If you’d like to see your own brand — or your clients — through the same lens, you’re warmly invited to try the AI Visibility Pulse, and whether you do or not, I genuinely welcome your feedback.
Glenn · Founder, styleforge.io · September 2026
Appendix: how it was run
Every check in the study, written out in full.
| The check | How it was run |
|---|---|
| Citations vs. mentions | Computed on the 1,500 answers from the three engines that disclose searching. “Cited” means the brand’s own site appears in the answer’s sources; subdomains count, third-party pages don’t. The 0-of-406 is unconditional; the mention comparison carries the question-type confound stated in the body. |
| Consensus attribution | A page’s engine count is the number of distinct engines citing it anywhere in the study: 425 pages at three-plus, 35 at four. Species rates use the 1,828 read pages. |
| Gap inventory | “Never names the brand” is the reading layer’s judgment on the full page text — 557 of 1,828 read pages. Per-brand rates use each brand’s own read set. |
| The repeat experiment | Every question re-asked with identical wording, same engines, same US locale, within a twenty-four-hour period (individual gaps ranged from roughly 4 to 16 hours). 500 paired answers per engine. Page-recurrence is computed per question at exact-URL and website level; stripping tracking parameters changed no figure by more than 0.1 points, so the raw numbers are published. |
| What the repeat can’t say | One day of spacing measures short-run stability only. Decay over weeks is a separate measurement, currently unpublished because it hasn’t finished being run. |
| Internal cross-check | All totals recomputed from raw rows before drafting; every rate in this part traces to the same records behind Parts One through Three. |
Limitations, stated plainly
- The stability findings cover a twenty-four-hour interval. Nothing here claims a fix lasts a quarter; it claims the target holds still long enough to measure honestly.
- These priorities are derived from ten consumer brands’ data; a category with different source dynamics (B2B, local services) deserves its own audit before inheriting these priorities.
- Own-site citation rates are conservative floors — the matching is deliberately strict and undercounts by a few percent rather than overcounting.
- These are the sources the engines disclosed; searches that showed no sources are in nobody’s data, including this study’s.
This study was run on StyleForge
Every measurement in this series ran on StyleForge’s AI Visibility Pulse — and the Pulse is one tool of many. StyleForge is a creative operating system for agencies and brand teams: intelligence, design, production, publishing, and measurement in one connected system that helps small agencies work like teams five times their size.
Explore the AI Visibility Pulse