Which AI visibility tools does AI itself recommend?
We pointed our own frozen, pre-registered instrument at our own market — 6 real brands, 8 real buyer questions, two engines — and published the answer, including our zero.
Semrush tops both engines — named in 36.8% of ChatGPT answers and 58.8% of Claude answers. Whether any tool gets named at all depends on the question on ChatGPT (3 of the 8) while Claude names tools on all 8. And we appear in 0 of 320 answers, measured under the same pre-registered rules we sell.
A marketing director choosing an AI-visibility tool can now ask the AI first — and the vendor cannot see that answer. We build the instrument that measures exactly this, so we pointed it at our own market: same rules, same controls, results published either way, with ourselves on the same grid as the competitors we sell against.
The measurement returned one incumbent topping both engines, a question shape that decides whether any tool gets named on ChatGPT, and our own name in 0 of 320 answers.
One settled shortlist, however you ask
The model a buyer holds: AI carries one settled "best AI visibility tools" list, and that list surfaces whenever anyone asks anything nearby. Under that model, checking one prompt on one engine tells a vendor where it stands. The data breaks this assumption twice — once by which engine you ask, and once by whether the question literally asks for tools.
Eight real questions, our own frozen rules
The measurement's inputs — the questions, the roster, the planted controls — were fixed and hash-pinned before any answer existed.
We did not write the questions. The 8 questions came verbatim from Reddit threads and search-bar entries real people typed — none authored or paraphrased by us or by a model. Each question was asked 17 times per engine on two engines, ChatGPT and Claude, through their APIs: 136 answers per engine, 272 in the main phase on 2026-07-30, and a small validation pass on 2026-07-27 brings the study to 320 answers.
The roster — 6 real brands, ourselves plus 5 competitors (Profound, Peec, Otterly, Semrush, Ahrefs) — was pre-registered and hash-pinned before the first answer was seen, so changing it after seeing results is barred by the study's own rules. Two invented brands with no real product behind them were planted as controls; they prove the matching does not hallucinate names. A brand counts as "named" when the answer text names it, a word-boundary match against the registered name and aliases; appearances that turn up only in a retrieved page are tracked separately and not counted as named.
What we registered in advance: who is measured, the exact questions, and mechanical validity checks — that the controls stay silent, and that the specialist brands' name and domain tables actually work. No headline number was bet in advance. The rates below are what the grid produced, not bets we won.
Semrush first on both engines — and our zero
Share of each engine's 136 main-phase answers naming each brand, sorted by the Claude rate. The bottom row is our own measured zero; the two invented control brands scored 0 appearances and are not drawn.
When people ask AI which AI-visibility tools to use, ChatGPT gives a shortlist only when the question literally asks for tools. On both engines that shortlist is led by Semrush, and we — the market's newest entrant — appear in 0 of 320 answers.
Semrush tops both engines: 36.8% of ChatGPT's 136 answers and 58.8% of Claude's 136 answers name it, first on both. Ahrefs is second on ChatGPT at 35.3%; its 43.4% on Claude falls behind two specialists. On ChatGPT the two SEO incumbents lead every specialist. On Claude the specialists close the gap — Otterly 54.4% and Profound 49.3% pass Ahrefs, and Peec reaches 39.7%. Their ChatGPT rates sit far lower — Otterly 19.1%, Profound 16.9%, Peec 14.0%.
We appear in 0.0% of answers on both engines — 0 named appearances, and 0 appearances of any kind, across all 320 answers in the study. The two invented control brands also sit at 0 appearances of any kind across the whole study, so the zero is a measured floor the controls sit on, not a glitch in our own count.
The tools built to measure AI visibility are out-recommended, on ChatGPT, by the two SEO suites that predate the category.
ChatGPT names tools only when you say "tools"
Of the 8 real buyer questions, ChatGPT named tools on 3 — the two best-tools questions and the comparison question — while Claude named tools on all 8.
ChatGPT named any tool on 3 of the 8 questions — the two which-tools questions and the comparison question. It named no tool on the other 5: the how-can-I-see-mentions, how-do-I-track, are-these-tools-worth-it, what-is-GEO, and how-do-you-measure questions. On those 3 it is emphatic: Semrush appeared in 17 of 17 answers to the first which-tools question, Ahrefs in 17 of 17 to the second, and on the comparison question Semrush and Ahrefs each appeared in 17 of 17.
Claude named tools on all 8 questions. On the what-is-GEO definitional question the specialists vanish there too — only Semrush, in 5 of 17 answers, and Ahrefs, in 2 of 17, appear. Otterly is 19.1% on ChatGPT and 54.4% on Claude — same tool, same day, same questions, engines apart.
Our reading, scoped to this evidence: ChatGPT treats how-to and definitional questions as advice tasks and answers without vendors, while Claude treats nearly every question as an occasion to name vendors. A pooled "AI visibility score" averages over this and hides it.
One category, two days, eight questions
These are two one-day snapshots — a validation pass on 2026-07-27 and the main phase on 2026-07-30 — one category, English only, 8 questions. Naming rates move with question wording, so a different question set or a different day can move these numbers; nothing here is a stability claim, and whether these rates hold over weeks is a separate, already-designed kind of study we have not run.
The two live explanations for our zero converge on the one measured fact; which road produced it is not decidable in this study.
Our zero is exact in this sample, 0 of 320 answers, and is still a snapshot. It says nothing about next month, and it does not separate "too new to be in training data" from "known but never chosen" — this study cannot tell those apart, and we do not guess.
"Named" is presence, not endorsement depth. A brand named 17 times with a caveat each time and a brand named once with praise both count once per answer; reporting is per-answer presence only.
One of the 8 questions is comparison-shaped ("closest to traditional SEO tools"). A question that hands the model a frame can pull answers toward tools that fit it; we replaced the mined comparison question that named a competitor outright, the one used names no brand, and we report per-question numbers above so no pooled rate leans on it silently.
"Profound" is also a common English adjective, so its counts could have been inflated by word-matches that are not the company. All 90 matches were read in full by one reader model and 18 of the 90 by a second, independent one, with 0 disagreements: all 90 are the company, so the corrected rate equals the raw rate and no adjustment was needed.
Being findable here is question-shaped
Left: the buyer's model — one settled list any nearby question surfaces. Right: the measured world — names live in engine-by-question cells, and on ChatGPT only the which-tools cells carry names.
For a tool vendor in this category: on ChatGPT you exist only inside which-tools answers — a buyer asking "how do I track AI mentions?" gets method advice with zero vendor names, so presence there is not bought by being "better known" in general. On Claude every question type carries names.
The concrete action a reader can take tomorrow: before spending anything on "AI visibility", check which engine and which question shapes actually carry brand names in your own category — the answer decided 3 of the 8 questions versus all 8 here, and it will differ by category.
For us, this number is a published baseline. The study re-runs under the same frozen rules, and the next measurement lands against this one.
How we ran this
The validation pass on 2026-07-27 produced 48 answers and the main phase on 2026-07-30 produced 272, on two engines' APIs — OpenAI and Anthropic — under fixed settings. These are API answers, not the consumer apps; consumer answers can differ.
Each of the 8 mined questions was asked 17 times per engine. The roster held 6 real brands plus 2 invented controls, pre-registered and hash-pinned before dispatch. A brand counts as named on a word-boundary match of the registered name and aliases in the answer text; appearances that surface only in a retrieved page are tracked separately and not counted as named.
For the Profound census, all 90 matches were read in full by GLM 5.2 as primary reader, and Grok 4.5 independently re-read 18 of the 90 with 0 disagreements. The registered validity check passed: both controls stayed silent everywhere, and the specialist brands' aliases and domain tables produced real matches — the only reason the main phase ran. Study spend was US$27.20 against a US$65 pre-approved ceiling.
How we ran this
- What we measured
- 8 real buyer questions × 17 repeats per engine across 2 engines — 272 main-phase answers — over a pre-registered, hash-pinned roster of 6 real brands plus 2 invented controls; a 48-answer validation pass ran three days earlier (320 answers total).
- When
- Validation pass 27 July 2026; main phase 30 July 2026.
- Which AI
- OpenAI (API), Anthropic (API)
The full numbers
Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.
| Measure | As shown | Exact | Source |
|---|---|---|---|
| Answers sampled per engine (main phase) | 136 | 136 | Measured |
| Repeats per question per engine | 17 | 17 | Measured |
| Real user questions asked | 8 | 8 | Measured |
| Main-phase answers, both engines | 272 | 272 | Derived |
| Validation-pass answers, both engines | 48 | 48 | Measured |
| Answers sampled across the whole study | 320 | 320 | Derived |
| Competitor brands measured | 5 | 5 | Measured |
| Real brands on the pre-registered roster | 6 | 6 | Measured |
| This product — times named, ChatGPT leg | 0 | 0 | Measured |
| This product — named rate, ChatGPT leg | 0.0% | 0.0000% | Derived |
| This product — retrieved-page-only appearances, ChatGPT leg | 0 | 0 | Measured |
| This product — times named, Claude leg | 0 | 0 | Measured |
| This product — named rate, Claude leg | 0.0% | 0.0000% | Derived |
| This product — retrieved-page-only appearances, Claude leg | 0 | 0 | Measured |
| Profound — times named, ChatGPT leg | 23 | 23 | Measured |
| Profound — named rate, ChatGPT leg | 16.9% | 16.9118% | Derived |
| Profound — retrieved-page-only appearances, ChatGPT leg | 6 | 6 | Measured |
| Profound — times named, Claude leg | 67 | 67 | Measured |
| Profound — named rate, Claude leg | 49.3% | 49.2647% | Derived |
| Profound — retrieved-page-only appearances, Claude leg | 18 | 18 | Measured |
| Peec — times named, ChatGPT leg | 19 | 19 | Measured |
| Peec — named rate, ChatGPT leg | 14.0% | 13.9706% | Derived |
| Peec — retrieved-page-only appearances, ChatGPT leg | 2 | 2 | Measured |
| Peec — times named, Claude leg | 54 | 54 | Measured |
| Peec — named rate, Claude leg | 39.7% | 39.7059% | Derived |
| Peec — retrieved-page-only appearances, Claude leg | 12 | 12 | Measured |
| Otterly — times named, ChatGPT leg | 26 | 26 | Measured |
| Otterly — named rate, ChatGPT leg | 19.1% | 19.1176% | Derived |
| Otterly — retrieved-page-only appearances, ChatGPT leg | 3 | 3 | Measured |
| Otterly — times named, Claude leg | 74 | 74 | Measured |
| Otterly — named rate, Claude leg | 54.4% | 54.4118% | Derived |
| Otterly — retrieved-page-only appearances, Claude leg | 17 | 17 | Measured |
| Semrush — times named, ChatGPT leg | 50 | 50 | Measured |
| Semrush — named rate, ChatGPT leg | 36.8% | 36.7647% | Derived |
| Semrush — retrieved-page-only appearances, ChatGPT leg | 0 | 0 | Measured |
| Semrush — times named, Claude leg | 80 | 80 | Measured |
| Semrush — named rate, Claude leg | 58.8% | 58.8235% | Derived |
| Semrush — retrieved-page-only appearances, Claude leg | 35 | 35 | Measured |
| Ahrefs — times named, ChatGPT leg | 48 | 48 | Measured |
| Ahrefs — named rate, ChatGPT leg | 35.3% | 35.2941% | Derived |
| Ahrefs — retrieved-page-only appearances, ChatGPT leg | 2 | 2 | Measured |
| Ahrefs — times named, Claude leg | 59 | 59 | Measured |
| Ahrefs — named rate, Claude leg | 43.4% | 43.3824% | Derived |
| Ahrefs — retrieved-page-only appearances, Claude leg | 12 | 12 | Measured |
| Invented control brands — total appearances, any kind | 0 | 0 | Measured |
| Answers checked for this product | 320 | 320 | Derived |
| Competitor named rate — lowest engine-level | 14.0% | 13.9706% | Derived |
| Competitor named rate — highest engine-level | 58.8% | 58.8235% | Derived |
| Questions where the ChatGPT leg named any tool | 3 | 3 | Derived |
| Questions where the ChatGPT leg named no tool | 5 | 5 | Derived |
| Questions where the Claude leg named any tool | 8 | 8 | Derived |
| Questions where the Claude leg named no specialist tool | 1 | 1 | Derived |
| Semrush — times named of 17, best-tools question, ChatGPT leg | 17 | 17 | Measured |
| Semrush — times named of 17, comparison question, ChatGPT leg | 17 | 17 | Measured |
| Ahrefs — times named of 17, comparison question, ChatGPT leg | 17 | 17 | Measured |
| Ahrefs — times named of 17, tracking-tools question, ChatGPT leg | 17 | 17 | Measured |
| Semrush — times named of 17, definitional question, Claude leg | 5 | 5 | Measured |
| Ahrefs — times named of 17, definitional question, Claude leg | 2 | 2 | Measured |
| Profound name-matches read in full | 90 | 90 | Measured |
| Profound matches that are the company | 90 | 90 | Measured |
| Profound matches that are the adjective | 0 | 0 | Measured |
| Matches re-read by the second reader | 18 | 18 | Measured |
| Second-reader disagreements | 0 | 0 | Measured |
| Study spend, all phases | US$27.20 | 27.20 | Measured |
| Approved study ceiling | US$65 | 65.00 | Measured |
What this doesn't settle
- Two one-day snapshots (2026-07-27 and 2026-07-30), one category, English only, 8 questions — no stability claim.
- Our zero is exact in this sample (0 of 320) and says nothing about next month; "too new to be in training data" versus "known but never chosen" is not separable here.
- "Named" is per-answer presence, not endorsement depth.
- One question is comparison-shaped; per-question numbers are reported so no pooled rate leans on it silently.
- API answers, not the consumer apps; consumer answers can differ.
Whose names come back in your category?
This piece is the same measurement a client gets — a pre-registered roster, planted controls, two engines, published either way.