Which AI visibility tools does AI itself recommend?
We pointed our own frozen, pre-registered instrument at our own market — 6 real brands, 8 real buyer questions, two engines — and published the answer, including our zero.
A marketing director choosing an AI-visibility tool can now ask the AI first — and the vendor cannot see that answer. We build the instrument that measures exactly this, so we pointed it at our own market: same rules, same controls, results published either way, with ourselves on the same grid as the competitors we sell against.
The measurement returned one incumbent topping both engines, a question shape that decides whether any tool gets named on ChatGPT, and our own name in 0 of 320 answers.
One settled shortlist, however you ask
The model a buyer holds: AI carries one settled "best AI visibility tools" list, and that list surfaces whenever anyone asks anything nearby. Under that model, checking one prompt on one engine tells a vendor where it stands. The data breaks this assumption twice — once by which engine you ask, and once by whether the question literally asks for tools.
Eight real questions, our own frozen rules
We did not write the questions. The 8 questions came verbatim from Reddit threads and search-bar entries real people typed — none authored or paraphrased by us or by a model. Each question was asked 17 times per engine on two engines, ChatGPT and Claude, through their APIs: 136 answers per engine, 272 in the main phase on 2026-07-30, and a small validation pass on 2026-07-27 brings the study to 320 answers.
The roster — 6 real brands, ourselves plus 5 competitors (Profound, Peec, Otterly, Semrush, Ahrefs) — was pre-registered and hash-pinned before the first answer was seen, so changing it after seeing results is barred by the study's own rules. Two invented brands with no real product behind them were planted as controls; they prove the matching does not hallucinate names. A brand counts as "named" when the answer text names it, a word-boundary match against the registered name and aliases; appearances that turn up only in a retrieved page are tracked separately and not counted as named.
What we registered in advance: who is measured, the exact questions, and mechanical validity checks — that the controls stay silent, and that the specialist brands' name and domain tables actually work. No headline number was bet in advance. The rates below are what the grid produced, not bets we won.
Semrush first on both engines — and our zero
When people ask AI which AI-visibility tools to use, ChatGPT gives a shortlist only when the question literally asks for tools. On both engines that shortlist is led by Semrush, and we — the market's newest entrant — appear in 0 of 320 answers.
Semrush tops both engines: 36.8% of ChatGPT's 136 answers and 58.8% of Claude's 136 answers name it, first on both. Ahrefs is second on ChatGPT at 35.3%; its 43.4% on Claude falls behind two specialists. On ChatGPT the two SEO incumbents lead every specialist. On Claude the specialists close the gap — Otterly 54.4% and Profound 49.3% pass Ahrefs, and Peec reaches 39.7%. Their ChatGPT rates sit far lower — Otterly 19.1%, Profound 16.9%, Peec 14.0%.
We appear in 0.0% of answers on both engines — 0 named appearances, and 0 appearances of any kind, across all 320 answers in the study. The two invented control brands also sit at 0 appearances of any kind across the whole study, so the zero is a measured floor the controls sit on, not a glitch in our own count.
The tools built to measure AI visibility are out-recommended, on ChatGPT, by the two SEO suites that predate the category.
ChatGPT names tools only when you say "tools"
ChatGPT named any tool on 3 of the 8 questions — the two which-tools questions and the comparison question. It named no tool on the other 5: the how-can-I-see-mentions, how-do-I-track, are-these-tools-worth-it, what-is-GEO, and how-do-you-measure questions. On those 3 it is emphatic: Semrush appeared in 17 of 17 answers to the first which-tools question, Ahrefs in 17 of 17 to the second, and on the comparison question Semrush and Ahrefs each appeared in 17 of 17.
Claude named tools on all 8 questions. On the what-is-GEO definitional question the specialists vanish there too — only Semrush, in 5 of 17 answers, and Ahrefs, in 2 of 17, appear. Otterly is 19.1% on ChatGPT and 54.4% on Claude — same tool, same day, same questions, engines apart.
Otterly's gap between engines is 48 named answers — 74 versus 26 of 136 answers each. Of the 48, 41 sit on the 5 questions where ChatGPT names no tool at all, Otterly or anyone. On the 3 questions where both engines name tools, the split is 33 to 26, a gap of 7. So 85% of the widest brand gap in this study is the question shape, not the engine's view of the brand.
Condition on the 3 naming questions only, and Otterly's rates become 51.0% on ChatGPT and 64.7% on Claude. The denominator is the 3 naming questions' 17 repeats each. That is a real but far smaller difference than the headline 19.1% on ChatGPT versus 54.4% on Claude. Comparing a brand's pooled rate across engines mostly measures the engines; the comparison that conditions on where naming happens at all is the one that matters.
Both engines also differ in how often they fetch the brands' own sites while answering. Claude fetched otterly.ai during 73 of its 136 answers; ChatGPT during 21. For Semrush the counts are 102 versus 41. Naming usually comes with a fetch on both engines. Of Claude's 74 Otterly-named answers, 56 had otterly.ai fetched; of ChatGPT's 26, 18 had it fetched. But naming without any fetch exists on both engines too: 12 and 9 of the Semrush-named answers, 18 and 8 for Otterly. A fetch is not necessary for a name. Our reading: both engines mix memory and live retrieval, and Claude leans far harder on retrieval — a reading of co-occurrence, not a proven cause.
Our reading, scoped to this evidence: ChatGPT treats how-to and definitional questions as advice tasks and answers without vendors, while Claude treats nearly every question as an occasion to name vendors. A pooled "AI visibility score" averages over this and hides it.
One category, two days, eight questions
These are two one-day snapshots — a validation pass on 2026-07-27 and the main phase on 2026-07-30 — one category, English only, 8 questions. Naming rates move with question wording, so a different question set or a different day can move these numbers; nothing here is a stability claim, and whether these rates hold over weeks is a separate, already-designed kind of study we have not run.
Our zero is exact in this sample, 0 of 320 answers, and is still a snapshot. It says nothing about next month, and it does not separate "too new to be in training data" from "known but never chosen" — this study cannot tell those apart, and we do not guess.
"Named" is presence, not endorsement depth. A brand named 17 times with a caveat each time and a brand named once with praise both count once per answer; reporting is per-answer presence only.
One of the 8 questions is comparison-shaped ("closest to traditional SEO tools"). A question that hands the model a frame can pull answers toward tools that fit it; we replaced the mined comparison question that named a competitor outright, the one used names no brand, and we report per-question numbers above so no pooled rate leans on it silently.
"Profound" is also a common English adjective, so its counts could have been inflated by word-matches that are not the company. All 90 matches were read in full by one reader model and 18 of the 90 by a second, independent one, with 0 disagreements: all 90 are the company, so the corrected rate equals the raw rate and no adjustment was needed.
Being findable here is question-shaped
For a tool vendor in this category: on ChatGPT you exist only inside which-tools answers — a buyer asking "how do I track AI mentions?" gets method advice with zero vendor names, so presence there is not bought by being "better known" in general. On Claude every question type carries names.
The concrete action a reader can take tomorrow: before spending anything on "AI visibility", check which engine and which question shapes actually carry brand names in your own category — the answer decided 3 of the 8 questions versus all 8 here, and it will differ by category.
For us, this number is a published baseline. The study re-runs under the same frozen rules, and the next measurement lands against this one.
How we ran this
The validation pass on 2026-07-27 produced 48 answers and the main phase on 2026-07-30 produced 272, on two engines' APIs — OpenAI and Anthropic — under fixed settings. These are API answers, not the consumer apps; consumer answers can differ.
Each of the 8 mined questions was asked 17 times per engine. The roster held 6 real brands plus 2 invented controls, pre-registered and hash-pinned before dispatch. A brand counts as named on a word-boundary match of the registered name and aliases in the answer text; appearances that surface only in a retrieved page are tracked separately and not counted as named.
For the Profound census, all 90 matches were read in full by GLM 5.2 as primary reader, and Grok 4.5 independently re-read 18 of the 90 with 0 disagreements. The registered validity check passed: both controls stayed silent everywhere, and the specialist brands' aliases and domain tables produced real matches — the only reason the main phase ran. Study spend was US$27.20 against a US$65 pre-approved ceiling.
What we measured: 8 real buyer questions × 17 repeats per engine across 2 engines — 272 main-phase answers — over a pre-registered, hash-pinned roster of 6 real brands plus 2 invented controls; a 48-answer validation pass ran three days earlier (320 answers total).
When: Validation pass 27 July 2026; main phase 30 July 2026.