The Agency Playbook, part four of a four-part field study on how AI answers questions about brands: the gap inventory, the consensus pages, and the repeat experiment — and what they mean for agencies.

AI visibility: what should an agency actually fix first?

The last measurements of the study — the pages that never name the brand, the pages all four engines trust, and the experiment where we asked all 2,000 questions twice — and what they mean for the agencies doing this work. This is part four of a four-part field study on how AI answers questions about brands.

How AI Answers Questions About Brands · Part Four of Four

Part Four: The Agency Playbook

557 pages that never name the brand · 35 pages all four engines trust · 2,000 questions asked twice

The series
Part One · The Retrieval Study — do the engines even search before answering?
Part Two · The Source Study — what the engines actually read when they do
Part Three · The Reddit Pipeline — the one source every engine reads
Part Four · The Agency Playbook — you’re reading it

The meeting every agency is sitting in now

It starts with a screenshot from the client. An AI answer that recommends a rival, or repeats a complaint, or simply never mentions them. Then the question every agency has learned to brace for: “What do we do about this?”

Data isn’t the problem. Good tools expose real GEO and AEO numbers now — Semrush and Ahrefs among them — and anyone can see who’s mentioned where. What’s been missing is the part the client is actually asking for: the order of operations. What to fix first. What each fix is worth. And what number will move next month to prove it worked.

Parts One through Three of this study measured the terrain: which engines search, what they read, and what Reddit does inside their answers. This part is about what those measurements mean for that order of operations — what kind of work the data actually supports, and what to measure after it.

Finding 1 · Mentions and citations are two different campaigns

Start with the split most plans miss. Across the 1,500 answers from the three engines that report their own searching, answers where the engine searched the web got cited pages behind them — and a citation went to the brand 17.1% of the time. Answers where it didn’t search? Zero citations. Not few — zero, in all 406 answers. Searching is the only road to a citation.

But here’s the twist: those no-search answers still mentioned the brand just as often — 69.2%, against 67.9% when the engine searched. Retrieval buys links, not love. Being named lives in a different place than being linked — and a plan that funds “AI visibility” as one line item is quietly buying only one of the two.

Finding 2 · The fix-first list already exists, and it’s short

Part Two read 1,828 cited pages. Sort them by how many engines lean on them and a species split appears. Pages cited by three or four engines — 19.1% of the read set, just 425 pages, only 35 of them carrying all four — are overwhelmingly ranked lists, almost never negative (4.8%), and usually name the brand. Single-engine pages are the opposite animal: forum-heavy, 24.2% negative, and brand-absent a third of the time.

That split is an ordering principle. The consensus pages are the shared load-bearing wall of every engine’s answer — fix those first. For Caraway, one page — a cookware review on theroundup.org, mixed verdict, cited 15 times by all four engines — does more work in AI answers than dozens of others combined. Peak Design’s own warranty page is a four-engine consensus source, proof that an owned page can make the list. And its Trustpilot page is a four-engine consensus source that no machine reader can open — a person in a browser sees it fine, but every automated audit hits a wall there, ours included.

Consensus pages and single-engine pages are different species

Grouped bar chart comparing pages cited by three or more engines against single-engine pages: negative sentiment 4.8% vs 24.2%; never names the brand 12.7% vs 34.8%.

Grouped bar chart comparing pages cited by three or more engines against single-engine pages: negative sentiment 4.8% vs 24.2%; never names the brand 12.7% vs 34.8%.
LabelConsensus pages (3–4 engines)Single-engine pages
Negative about the brand4.824.2
Never names the brand12.734.8

The 425 pages cited by 3–4 engines vs. single-engine pages, of the 1,828 read. The pages every engine shares are safer, brand-present, and few — a workable priority list.

Finding 3 · A third of the reading list never says your name

Here is the study’s biggest opportunity number. Of the 1,828 cited pages the instrument read, 557 — 30.5% — never mention the brand whose buyers they’re answering. These pages already rank, already get cited, already sit inside your category’s answers. They simply don’t include you.

The spread across brands is the proof it’s fixable: AG1’s reading list is brand-absent on 48.2% of pages, OLIPOP’s on 43.5%, Caraway’s on 42.4% — while Patagonia’s sits at 16.6%. Same engines, same kinds of pages. The difference is presence on pages someone earned. For an agency, each of those 557 pages has a known URL, a known author, and a known reason to update. That isn’t a statistic — it’s a call list.

Share of each brand’s cited pages that never name it

Bar chart of brand-absent share of read cited pages: AG1 48.2%, OLIPOP 43.5%, Caraway 42.4%, Gymshark 39.2%, study average 30.5%, Patagonia 16.6%.

Bar chart of brand-absent share of read cited pages: AG1 48.2%, OLIPOP 43.5%, Caraway 42.4%, Gymshark 39.2%, study average 30.5%, Patagonia 16.6%.
LabelCited pages that never name the brand (%)
AG148.2
OLIPOP43.5
Caraway42.4
Gymshark39.2
Study average30.5
Patagonia16.6

Of each brand’s read cited pages. The ten-brand average is 30.5%; the spread from 48.2% down to 16.6% is the evidence that this number responds to work.

Finding 4 · Your rival list is short, specific, and already written

When the reading layer recorded which competitors appear on cited pages, the sets came back category-locked. Only two names — Nike and Adidas — recur across multiple brands’ reading lists. Everyone else’s competitive field is a handful of names that never leave the category. The engines have already decided who your client actually competes with — and it’s a shorter list than the one the client worries about, and often a different one.

Caraway’s entire rival field, as the engines’ reading list carries it

Rival named on Caraway’s cited pagesPages
GreenPan51
Calphalon41
Made In33
All-Clad31
Lodge30
Le Creuset10

Finding 5 · Your own site has a per-engine ceiling

How often does an engine cite the brand’s own website in that brand’s answers? The rates are low everywhere, and they’re an engine property like everything else in this series:

  • Perplexity29.2%of its 500 answers cite the brand’s own site
  • Gemini18.4%
  • Claude14.8%
  • ChatGPT4.2%and for one brand — Hoka — 0 of 50

Pair that with Part Two’s finding that 58.3% of cited own-site pages wouldn’t even open for a machine reader, and the shape of the own-site job is clear: modest ceiling, mostly unclaimed. And on the demand side, the biggest hole in the study: the long-tail lane — “best X for my situation” questions — where brands get mentioned in just 35.8% of answers overall, and AG1 in 2%. The questions closest to a purchase are the questions brands are most absent from.

The closing experiment · We asked all 2,000 questions twice

None of this matters if the target citation won’t hold still. So the study ends with the measurement everything else depends on: we re-asked every one of the 2,000 questions, same engines, same wording, within a twenty-four-hour period — and compared everything. The answer mention rates barely moved: 69.0% became 68.3%. But underneath, the engines rewrote roughly half their own search queries — and what happened to the cited sources splits cleanly by engine:

Asked again within 24 hours: how much held still

EngineCited the same pagesSame websitesRewrote its own searches
Perplexity90.3% of the time92.1%— (doesn’t disclose queries)
Claude66.8%70.1%54% rewritten
ChatGPT51.7%56.0%48% rewritten
Gemini33.9%38.4%74% rewritten
The stability gap
90.3%
PERPLEXITY RETURNED TO THE SAME PAGES
Asked again within a day, it cited the same URLs nine times out of ten.
33.9%
GEMINI DID — ONE TIME IN THREE
Ask it twice, get an answer built from a different reading list.

Like searching, like Reddit, like everything else in this series: consistency is an engine property. This measures a single day — how the picture decays over weeks is the follow-up measurement, and it isn’t published until it’s run.

For an agency this number is the license to work. The reading list holds still enough to be worth fixing — and moves enough that every fix needs a re-measurement behind it. Where recurrence is high, a page you win stays won. Where it’s low, visibility is a monthly reading, not a trophy. The instrument this series ran on treats that as a feature: StyleForge’s AI Visibility Pulse versions every run, so a second pulse reports its variance against the previous one — every engine rate, share-of-voice position, and question type carrying its change:

Our approach, disclosed

Everything in this series ran on StyleForge’s AI Visibility Pulse, and this is the part where I say plainly what it is and why we built it this way. Plenty of tools now report GEO and AEO data, and some of it is excellent. Our premise is narrower: arm the agency with findings it can put in front of a client and act on together. The reports exist so agency and client can work from the same document toward the same goal — growing the brand’s visibility with AI assistants — with every finding pointing at a fix, and every fix pointing at the number that will show it worked.

In practice: the question pack is drafted for you by AI — or drag in your own list from a template — with the competitors, question types, and weighting under your control, so an audit can be an investigation, not a preset. The citation audit reads the sources channel by channel, at the depth you choose, down to full YouTube transcripts. Every pulse is versioned, and a second run reports its movement against the last one at the question-type level. And the client report — 200 to 250 pages, white-labeled to the agency, section-selectable — closes on a recommended-actions page where every action names an owner and the number that will move next pulse — the same discipline this series has practiced throughout: no claim without a number, no action without a re-measurement behind it.

The complete audit for every brand in this series is published below — every question, every answer, every cited page read. Click any brand to open its full report.

Each of these reports is white-labeled — generated under the agency’s own name and presentable straight to the client, and you control which sections every report includes. The agency name on these ten is a stand-in.

A personal note, to close the series

I’ve been building websites since 1995 — hundreds of them, hand-coded, optimized, fussed over — and I watched the whole SEO industry grow up around that work: the early tools, the guesswork, the slow arrival of real measurement. That experience is why I added an AI visibility tool to StyleForge: agencies deserve to see what’s actually happening under the hood of these answers, and to have a way to act on it with their clients — so no brand finds out secondhand how the assistants talk about them.

Just as the analytics tools of early SEO evolved and got better year over year, I think that’s exactly what’s happening with GEO right now. This series is my contribution to it: measurements on the table, the full record published.

If you’d like to see your own brand — or your clients — through the same lens, you’re warmly invited to try the AI Visibility Pulse, and whether you do or not, I genuinely welcome your feedback.

Glenn · Founder, styleforge.io · September 2026

Appendix: how it was run

Every check in the study, written out in full.

The checkHow it was run
Citations vs. mentionsComputed on the 1,500 answers from the three engines that disclose searching. “Cited” means the brand’s own site appears in the answer’s sources; subdomains count, third-party pages don’t. The 0-of-406 is unconditional; the mention comparison carries the question-type confound stated in the body.
Consensus attributionA page’s engine count is the number of distinct engines citing it anywhere in the study: 425 pages at three-plus, 35 at four. Species rates use the 1,828 read pages.
Gap inventory“Never names the brand” is the reading layer’s judgment on the full page text — 557 of 1,828 read pages. Per-brand rates use each brand’s own read set.
The repeat experimentEvery question re-asked with identical wording, same engines, same US locale, within a twenty-four-hour period (individual gaps ranged from roughly 4 to 16 hours). 500 paired answers per engine. Page-recurrence is computed per question at exact-URL and website level; stripping tracking parameters changed no figure by more than 0.1 points, so the raw numbers are published.
What the repeat can’t sayOne day of spacing measures short-run stability only. Decay over weeks is a separate measurement, currently unpublished because it hasn’t finished being run.
Internal cross-checkAll totals recomputed from raw rows before drafting; every rate in this part traces to the same records behind Parts One through Three.

This study was run on StyleForge

Every measurement in this series ran on StyleForge’s AI Visibility Pulse — and the Pulse is one tool of many. StyleForge is a creative operating system for agencies and brand teams: intelligence, design, production, publishing, and measurement in one connected system that helps small agencies work like teams five times their size.

Explore the AI Visibility Pulse