Which AI visibility tools does AI itself recommend?

We pointed our own frozen, pre-registered instrument at our own market — 6 real brands, 8 real buyer questions, two engines — and published the answer, including our zero.

Semrush tops both engines — named in 36.8% of ChatGPT answers and 58.8% of Claude answers. Whether any tool gets named at all depends on the question on ChatGPT (3 of the 8) while Claude names tools on all 8. And we appear in 0 of 320 answers, measured under the same pre-registered rules we sell.

2 August 2026Results registered before data was seen

A marketing director choosing an AI-visibility tool can now ask the AI first — and the vendor cannot see that answer. We build the instrument that measures exactly this, so we pointed it at our own market: same rules, same controls, results published either way, with ourselves on the same grid as the competitors we sell against.

The measurement returned one incumbent topping both engines, a question shape that decides whether any tool gets named on ChatGPT, and our own name in 0 of 320 answers.

One settled shortlist, however you ask

The model a buyer holds: AI carries one settled "best AI visibility tools" list, and that list surfaces whenever anyone asks anything nearby. Under that model, checking one prompt on one engine tells a vendor where it stands. The data breaks this assumption twice — once by which engine you ask, and once by whether the question literally asks for tools.

Eight real questions, our own frozen rules

Locked before looking

The measurement's inputs — the questions, the roster, the planted controls — were fixed and hash-pinned before any answer existed.

Questions mined
8 real questions, quoted verbatim
Roster pinned
6 real brands, hash-locked
Controls planted
2 invented brands
Grid dispatched
17 repeats × 2 engines
Matches read back
every name-match read in full
Every step was fixed before the first answer was seen; changing the roster after seeing results is barred by the study's own filed rules.

We did not write the questions. The 8 questions came verbatim from Reddit threads and search-bar entries real people typed — none authored or paraphrased by us or by a model. Each question was asked 17 times per engine on two engines, ChatGPT and Claude, through their APIs: 136 answers per engine, 272 in the main phase on 2026-07-30, and a small validation pass on 2026-07-27 brings the study to 320 answers.

The roster — 6 real brands, ourselves plus 5 competitors (Profound, Peec, Otterly, Semrush, Ahrefs) — was pre-registered and hash-pinned before the first answer was seen, so changing it after seeing results is barred by the study's own rules. Two invented brands with no real product behind them were planted as controls; they prove the matching does not hallucinate names. A brand counts as "named" when the answer text names it, a word-boundary match against the registered name and aliases; appearances that turn up only in a retrieved page are tracked separately and not counted as named.

What we registered in advance: who is measured, the exact questions, and mechanical validity checks — that the controls stay silent, and that the specialist brands' name and domain tables actually work. No headline number was bet in advance. The rates below are what the grid produced, not bets we won.

Semrush first on both engines — and our zero

The shortlist, both engines

Share of each engine's 136 main-phase answers naming each brand, sorted by the Claude rate. The bottom row is our own measured zero; the two invented control brands scored 0 appearances and are not drawn.

Claude
ChatGPT
Semrush
58.8%
36.8%
Otterly
54.4%
19.1%
Profound
49.3%
16.9%
Ahrefs
43.4%
35.3%
Peec
39.7%
14.0%
Us
0.0%
0.0%
Named rate = share of that engine's 136 main-phase answers naming the brand. 8 real buyer questions, 17 repeats per engine, 2026-07-30, via the engines' APIs. The two invented control brands scored 0 appearances of any kind and are not drawn; our own row is the measured zero — 0 of 320 answers across the whole study.

When people ask AI which AI-visibility tools to use, ChatGPT gives a shortlist only when the question literally asks for tools. On both engines that shortlist is led by Semrush, and we — the market's newest entrant — appear in 0 of 320 answers.

Semrush tops both engines: 36.8% of ChatGPT's 136 answers and 58.8% of Claude's 136 answers name it, first on both. Ahrefs is second on ChatGPT at 35.3%; its 43.4% on Claude falls behind two specialists. On ChatGPT the two SEO incumbents lead every specialist. On Claude the specialists close the gap — Otterly 54.4% and Profound 49.3% pass Ahrefs, and Peec reaches 39.7%. Their ChatGPT rates sit far lower — Otterly 19.1%, Profound 16.9%, Peec 14.0%.

We appear in 0.0% of answers on both engines — 0 named appearances, and 0 appearances of any kind, across all 320 answers in the study. The two invented control brands also sit at 0 appearances of any kind across the whole study, so the zero is a measured floor the controls sit on, not a glitch in our own count.

The tools built to measure AI visibility are out-recommended, on ChatGPT, by the two SEO suites that predate the category.

ChatGPT names tools only when you say "tools"

Questions that carried any tool's name

Of the 8 real buyer questions, ChatGPT named tools on 3 — the two best-tools questions and the comparison question — while Claude named tools on all 8.

8
Claude
3
ChatGPT
of 8 real buyer questions carried any tool's name

ChatGPT named any tool on 3 of the 8 questions — the two which-tools questions and the comparison question. It named no tool on the other 5: the how-can-I-see-mentions, how-do-I-track, are-these-tools-worth-it, what-is-GEO, and how-do-you-measure questions. On those 3 it is emphatic: Semrush appeared in 17 of 17 answers to the first which-tools question, Ahrefs in 17 of 17 to the second, and on the comparison question Semrush and Ahrefs each appeared in 17 of 17.

Claude named tools on all 8 questions. On the what-is-GEO definitional question the specialists vanish there too — only Semrush, in 5 of 17 answers, and Ahrefs, in 2 of 17, appear. Otterly is 19.1% on ChatGPT and 54.4% on Claude — same tool, same day, same questions, engines apart.

Our reading, scoped to this evidence: ChatGPT treats how-to and definitional questions as advice tasks and answers without vendors, while Claude treats nearly every question as an occasion to name vendors. A pooled "AI visibility score" averages over this and hides it.

One category, two days, eight questions

These are two one-day snapshots — a validation pass on 2026-07-27 and the main phase on 2026-07-30 — one category, English only, 8 questions. Naming rates move with question wording, so a different question set or a different day can move these numbers; nothing here is a stability claim, and whether these rates hold over weeks is a separate, already-designed kind of study we have not run.

Two roads to the same zero

The two live explanations for our zero converge on the one measured fact; which road produced it is not decidable in this study.

Too new to be in training data
not observable in this study
Known, never chosen
not observable in this study
both roads end here — not separable
0 of 320 answers
the observed zero
The study measures answers, not model internals. The zero is exact in this sample; which road produced it is not decidable here, and we do not guess.

Our zero is exact in this sample, 0 of 320 answers, and is still a snapshot. It says nothing about next month, and it does not separate "too new to be in training data" from "known but never chosen" — this study cannot tell those apart, and we do not guess.

"Named" is presence, not endorsement depth. A brand named 17 times with a caveat each time and a brand named once with praise both count once per answer; reporting is per-answer presence only.

One of the 8 questions is comparison-shaped ("closest to traditional SEO tools"). A question that hands the model a frame can pull answers toward tools that fit it; we replaced the mined comparison question that named a competitor outright, the one used names no brand, and we report per-question numbers above so no pooled rate leans on it silently.

"Profound" is also a common English adjective, so its counts could have been inflated by word-matches that are not the company. All 90 matches were read in full by one reader model and 18 of the 90 by a second, independent one, with 0 disagreements: all 90 are the company, so the corrected rate equals the raw rate and no adjustment was needed.

Being findable here is question-shaped

One list assumed, a grid measured

Left: the buyer's model — one settled list any nearby question surfaces. Right: the measured world — names live in engine-by-question cells, and on ChatGPT only the which-tools cells carry names.

The assumed world
best tools?
compare to SEO tools?
how do I track?
what is GEO?
one settled list
every question surfaces the same names
The measured world
ChatGPT
Claude
best tools?
compare to SEO tools?
how do I track?
what is GEO?
filled = names appear in that cell
A sample of the 8 questions is shown. Measured: ChatGPT carries tool names on 3 of the 8 questions; Claude on all 8.

For a tool vendor in this category: on ChatGPT you exist only inside which-tools answers — a buyer asking "how do I track AI mentions?" gets method advice with zero vendor names, so presence there is not bought by being "better known" in general. On Claude every question type carries names.

The concrete action a reader can take tomorrow: before spending anything on "AI visibility", check which engine and which question shapes actually carry brand names in your own category — the answer decided 3 of the 8 questions versus all 8 here, and it will differ by category.

For us, this number is a published baseline. The study re-runs under the same frozen rules, and the next measurement lands against this one.

How we ran this

The validation pass on 2026-07-27 produced 48 answers and the main phase on 2026-07-30 produced 272, on two engines' APIs — OpenAI and Anthropic — under fixed settings. These are API answers, not the consumer apps; consumer answers can differ.

Each of the 8 mined questions was asked 17 times per engine. The roster held 6 real brands plus 2 invented controls, pre-registered and hash-pinned before dispatch. A brand counts as named on a word-boundary match of the registered name and aliases in the answer text; appearances that surface only in a retrieved page are tracked separately and not counted as named.

For the Profound census, all 90 matches were read in full by GLM 5.2 as primary reader, and Grok 4.5 independently re-read 18 of the 90 with 0 disagreements. The registered validity check passed: both controls stayed silent everywhere, and the specialist brands' aliases and domain tables produced real matches — the only reason the main phase ran. Study spend was US$27.20 against a US$65 pre-approved ceiling.

How we ran this

What we measured
8 real buyer questions × 17 repeats per engine across 2 engines — 272 main-phase answers — over a pre-registered, hash-pinned roster of 6 real brands plus 2 invented controls; a 48-answer validation pass ran three days earlier (320 answers total).
When
Validation pass 27 July 2026; main phase 30 July 2026.
Which AI
OpenAI (API), Anthropic (API)

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
Answers sampled per engine (main phase)136136Measured
Repeats per question per engine1717Measured
Real user questions asked88Measured
Main-phase answers, both engines272272Derived
Validation-pass answers, both engines4848Measured
Answers sampled across the whole study320320Derived
Competitor brands measured55Measured
Real brands on the pre-registered roster66Measured
This product — times named, ChatGPT leg00Measured
This product — named rate, ChatGPT leg0.0%0.0000%Derived
This product — retrieved-page-only appearances, ChatGPT leg00Measured
This product — times named, Claude leg00Measured
This product — named rate, Claude leg0.0%0.0000%Derived
This product — retrieved-page-only appearances, Claude leg00Measured
Profound — times named, ChatGPT leg2323Measured
Profound — named rate, ChatGPT leg16.9%16.9118%Derived
Profound — retrieved-page-only appearances, ChatGPT leg66Measured
Profound — times named, Claude leg6767Measured
Profound — named rate, Claude leg49.3%49.2647%Derived
Profound — retrieved-page-only appearances, Claude leg1818Measured
Peec — times named, ChatGPT leg1919Measured
Peec — named rate, ChatGPT leg14.0%13.9706%Derived
Peec — retrieved-page-only appearances, ChatGPT leg22Measured
Peec — times named, Claude leg5454Measured
Peec — named rate, Claude leg39.7%39.7059%Derived
Peec — retrieved-page-only appearances, Claude leg1212Measured
Otterly — times named, ChatGPT leg2626Measured
Otterly — named rate, ChatGPT leg19.1%19.1176%Derived
Otterly — retrieved-page-only appearances, ChatGPT leg33Measured
Otterly — times named, Claude leg7474Measured
Otterly — named rate, Claude leg54.4%54.4118%Derived
Otterly — retrieved-page-only appearances, Claude leg1717Measured
Semrush — times named, ChatGPT leg5050Measured
Semrush — named rate, ChatGPT leg36.8%36.7647%Derived
Semrush — retrieved-page-only appearances, ChatGPT leg00Measured
Semrush — times named, Claude leg8080Measured
Semrush — named rate, Claude leg58.8%58.8235%Derived
Semrush — retrieved-page-only appearances, Claude leg3535Measured
Ahrefs — times named, ChatGPT leg4848Measured
Ahrefs — named rate, ChatGPT leg35.3%35.2941%Derived
Ahrefs — retrieved-page-only appearances, ChatGPT leg22Measured
Ahrefs — times named, Claude leg5959Measured
Ahrefs — named rate, Claude leg43.4%43.3824%Derived
Ahrefs — retrieved-page-only appearances, Claude leg1212Measured
Invented control brands — total appearances, any kind00Measured
Answers checked for this product320320Derived
Competitor named rate — lowest engine-level14.0%13.9706%Derived
Competitor named rate — highest engine-level58.8%58.8235%Derived
Questions where the ChatGPT leg named any tool33Derived
Questions where the ChatGPT leg named no tool55Derived
Questions where the Claude leg named any tool88Derived
Questions where the Claude leg named no specialist tool11Derived
Semrush — times named of 17, best-tools question, ChatGPT leg1717Measured
Semrush — times named of 17, comparison question, ChatGPT leg1717Measured
Ahrefs — times named of 17, comparison question, ChatGPT leg1717Measured
Ahrefs — times named of 17, tracking-tools question, ChatGPT leg1717Measured
Semrush — times named of 17, definitional question, Claude leg55Measured
Ahrefs — times named of 17, definitional question, Claude leg22Measured
Profound name-matches read in full9090Measured
Profound matches that are the company9090Measured
Profound matches that are the adjective00Measured
Matches re-read by the second reader1818Measured
Second-reader disagreements00Measured
Study spend, all phasesUS$27.2027.20Measured
Approved study ceilingUS$6565.00Measured

What this doesn't settle

  • Two one-day snapshots (2026-07-27 and 2026-07-30), one category, English only, 8 questions — no stability claim.
  • Our zero is exact in this sample (0 of 320) and says nothing about next month; "too new to be in training data" versus "known but never chosen" is not separable here.
  • "Named" is per-answer presence, not endorsement depth.
  • One question is comparison-shaped; per-question numbers are reported so no pooled rate leans on it silently.
  • API answers, not the consumer apps; consumer answers can differ.

Whose names come back in your category?

This piece is the same measurement a client gets — a pre-registered roster, planted controls, two engines, published either way.

Join the waitlist