Can AI know a brand and still rarely recommend it?
We went looking for one of our own three failure modes — a brand known to AI but rarely picked by it — across three unrelated shopping categories, and published what we found either way.
Combined across both engines, in these three categories — CRM software, small-business accounting, and pet food and supplies — on this design, no brand out of eighteen was both present in at least half of answers and picked in fewer than four times in ten. The closest any reading came was a pet retailer's answers on one AI engine, which cleared every threshold on that engine alone — 56.2% presence, a 37.0% pick rate, and a safeguard estimate of 50.37% — before the second engine's own reading brought the combined number back above the bar.
Our own AI-visibility framework claims more than one way a brand can fail — but claiming it and finding it in real data are different things. We went looking for one of those failure modes directly: a brand known to AI but rarely picked by it, across three very different shopping categories, on the same measurement setup we use for client work, with rules built to test for the pattern in three specific categories, not to catch it if it exists anywhere. We publish what turned up either way.
A brand counts as showing up when its name (or, for the pet retailer with a common-word name, its full brand form) appears in an answer; it counts as picked when the answer names it as the one it actually recommends, not merely mentions it.
Being known and being picked are supposed to be two different problems
If a brand never gets recommended, the easy assumption is that AI simply does not know it exists. The theory this piece tests says that is not the only way to lose: a brand can be named or considered constantly and still rarely be the one AI actually recommends — known, but not compelling enough to pick. That is the theory under test here. We do not say yet whether it held up.
Three shopping categories, the same two AI engines, the same rule every time
We picked three categories to be three different shapes of buying decision: which CRM software a small business should buy, which accounting software a small business should buy, and where to buy pet food and supplies online. CRM software is a category where AI answers tend to name many brands in one list. Accounting software is a category with one or two dominant defaults. Pet food is an ordinary consumer purchase, not business software at all.
Each category put eight different real buyer questions to two AI engines — OpenAI's and Anthropic's — many times over, for twenty-four distinct questions across the three categories in total. We measured six real brands by name per category: a couple of clear category leaders plus several smaller or more specialized names, not all equally well-known going in. We also planted two invented brand names with no real company behind them as a control.
For every brand we measured two things: how often it showed up in an answer at all (its presence rate), and, of the times it showed up, how often it was the one AI actually recommended (its pick rate). Both rates are combined across both engines' answers unless a sentence says otherwise.
Present in at least half of all answers (50%). Picked fewer than 4 times in 10 (40%). The generous estimate of the true pick rate stays under 55%. This is the rule itself, not a measured brand — every brand's own reading is checked against it in the next section.
The bar for "known but not picked, and not by chance," evaluated on the combined answers from both engines, had three parts: present in at least half of all answers, picked less than 4 times in 10, and — because a rare pattern can look real in a small sample — the estimated ceiling on how high the true pick rate could plausibly be, even generously, had to sit under 55%.
Each category's own roster of six real brands and two invented names was locked in writing before that category's own answers were collected. All three categories used the same roster shape and the same rule. The three categories were not run as one program decided in advance: the third, pet food, was ordered five days after the first two, once they were already finished and published, specifically to check the result outside business software. We applied no correction for that timing gap, because the pet-food category is a deliberate follow-up check, not a rerun of the first two on matching dates.
Every category ran a small first pass before the full count, to confirm the measurement could tell a real recommendation apart from a brand merely being considered, and that the invented control names stayed silent — not to confirm the target pattern itself was present. That first pass added 48 answers per category, 144 across all three. The reported rates below come from the larger full count that followed: 192 answers per category, 576 across the three, on top of the 144 from the first pass — 720 answers in total. Every brand rate quoted anywhere in this piece is built on the 576-answer full count, never the 720-answer grand total.
Zero of eighteen brands cleared the combined-engine rule
Presence rate against pick rate, combined across both engines. Only one reading ever lands inside the dashed zone — and it is Chewy.com's Anthropic-only reading, not a combined one.
Combined across both engines, in these three categories, on this design, none of the eighteen real brands was both present in at least half of answers and picked in fewer than four times in ten. In the combined answers — 192 per category, 576 across the three — every brand present in at least half of answers had a pick rate of 40% or higher, with no exceptions: none of the eighteen real brands met all three parts of the rule, combined.
The strongest instances of the opposite pattern, one per category, make the point concrete. In CRM software, HubSpot showed up in 100% of answers and was the pick in 97.4% of those. In accounting software, QuickBooks showed up in 100% of answers and was the pick in 94.3% of those. In pet food and supplies, Petco showed up in 89.6% of answers and was the pick in 97.7% of those.
The closest a single engine's own reading came was Chewy.com on Anthropic's answers alone: the pet retailer showed up in 56.2% of answers and was the actual pick in only 37.0% of those, clearing every threshold, including the small-sample safeguard, whose estimated ceiling on the true pick rate sat at 50.37%. On OpenAI's own answers for the same retailer, the pick rate was 87.2%. Combined across both engines — the reading the study's rule actually evaluates — Chewy.com's presence was 68.8% and its pick rate was 66.7%, above the 40% bar, so it does not count. The combined number is the one that counts because the study measures each brand's position across both engines' answers together, not whichever single engine happens to look most interesting; one engine's answers alone are half the evidence, not the verdict.
Low pick rates came from thin samples — except one
The study's lowest pick rate came from a brand barely seen at all: Wave Accounting showed up in just 26.0% of Anthropic's answers and was picked in only 20.0% of those 25 instances. A thin sample can cut the other way too — PetFlow showed up in only 7.8% of answers but was picked 100% of the 15 times it did, the opposite extreme from an equally small base.
Two brands were essentially invisible on both engines: MYOB in accounting and Petsense in pet food each showed up in 0% of answers, on both engines, across the whole count.
Wave Accounting (Anthropic only): 26.0% presence, picked in 20.0% of appearances. PetFlow (combined): 7.8% presence, picked in 100% of appearances. Chewy.com (Anthropic only): 56.2% presence, picked in 37.0% of appearances.
In the combined numbers, every brand whose presence rate cleared the 50% bar also had a pick rate at or above 40%, with no exceptions. The one place a different picture shows up is a single-engine number, not a combined one: on Anthropic's answers alone, Chewy.com's pick rate stayed at 37.0% from 54 in-pool instances, alongside a presence rate above the 50% bar. The failure shape the combined data actually shows is a brand not being seen enough, not a brand being seen often and still turned down — our reading of this dataset, not a general law.
Three categories, one bar deliberately set looser than a real diagnosis
Three categories is diligence, not a survey. CRM software and accounting software ran the same day; pet food ran five days later, once the first two were already finished and published, specifically to check the result outside business software. The consequence is that this is not a claim about categories not tested, and there are far more than three kinds of purchases a brand can be searched for. We applied no correction because the third category was a deliberate, separately-run follow-up, not a rerun of the first two.
Each category is one day's answers, two AI engines, one buyer persona. Whether these rates hold over weeks, or shift with the exact wording of the question, is unknown from this design. We applied no correction because testing stability over time or across wordings is a different, already-designed kind of check this piece does not run — not an oversight, a different study.
The near-miss itself sat close to its own safeguard: the estimated ceiling on Chewy.com's true Anthropic-only pick rate was 50.37%, just under the 55% bar the rule requires. A reading this close to its own limit is a genuine clearance on that one engine, not a strong one.
The bar used here — present in at least half of answers, picked in under 4 of 10, combined across both engines, plus the 55% safeguard — is deliberately looser than the bar a real client's own report uses, which requires a higher presence rate before the same pattern would count. A category clearing this study's looser bar would still need to clear that higher one before it changed a real diagnosis. We applied no correction because this piece is explicitly an existence test, not a client diagnosis — the looser bar is by design, not a limitation to fix.
Each category's two invented control names were tested only within that category's own answers, not across all three categories. Collectively, all six category-specific control names appeared in zero answers, in the category each was planted in, across all 720 answers collected — a measured floor for those specific names in their own category, not a claim about every possible name or category. They exist to prove the measurement is not inventing mentions in the category they were planted in.
These are API answers, not the consumer apps. The results apply to the API measurement, not necessarily to what a consumer sees in the chat apps. We applied no correction because the consumer apps were not sampled in this design.
The rare picks mostly came from being rarely seen
In the combined numbers, in these three categories, on this design, the low pick rates came from brands AI barely mentioned in the first place, not from brands that were well-known and turned down — the one single-engine reading that came close (Chewy.com on Anthropic) did not survive being combined with the other engine. This is a reason to check the presence rate first before assuming a poor showing is a positioning problem, not a proof that positioning problems never occur.
Check both a presence rate and a pick rate, on more than one AI engine, before deciding a brand has a positioning problem rather than a visibility one.
For our own measurement, this is one of three failure modes we watch for — a brand known to AI but rarely picked by it — and we did not detect it here, combined, across these three categories. We do not treat it as an established pattern; we would report it if it ever cleared the combined bar.
How we ran this
Three categories: CRM software and small-business accounting software ran the same day; pet food and supplies ran five days later, once the first two were finished, as a follow-up check. Two AI engines each (OpenAI, Anthropic), API answers under fixed settings, not the consumer apps.
Six real brands per category (eighteen total) plus two invented control names per category (six total); each category's own roster was locked in writing before that category's own answers were collected. The same six-real-plus-two-invented shape and the same combined-answer pass/fail rule ran unchanged in all three categories.
We collected 720 answers in total across the three categories: 144 from a small first pass (48 per category) that checked the measurement itself could tell a real recommendation apart from a mere mention, and that the invented controls stayed silent — followed by 576 from the full count (192 per category) that every reported brand rate in this piece is built on. Total spend across all three categories: US$61.20.
How we ran this
- What we measured
- Six real brands per category (eighteen total) plus two invented control names per category (six total); eight buyer questions per category (twenty-four total), put to two AI engines many times over. 720 answers collected across the three categories: 144 from a small first pass, 576 from the full count every reported rate is built on.
- When
- CRM software and small-business accounting software: 30 July 2026. Pet food & supplies: 4 August 2026, a follow-up check.
- Which AI
- OpenAI (API), Anthropic (API)
The full numbers
Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.
| Measure | As shown | Exact | Source |
|---|---|---|---|
| Shopping categories tested | 3 | 3 | Measured |
| AI engines per category | 2 | 2 | Measured |
| Buyer questions per category | 8 | 8 | Measured |
| Buyer questions across all three categories | 24 | 24 | Derived |
| Real brands measured per category | 6 | 6 | Measured |
| Real brands measured across all three categories | 18 | 18 | Derived |
| Invented control brands per category | 2 | 2 | Measured |
| Invented control brands across all three categories | 6 | 6 | Derived |
| Days between the first two categories and the third | 5 | 5 | Measured |
| Full-count answers per category | 192 | 192 | Measured |
| Full-count answers across the three categories, the basis for reported rates | 576 | 576 | Derived |
| First-pass answers per category | 48 | 48 | Measured |
| First-pass answers across the three categories | 144 | 144 | Derived |
| Total answers collected across the study, all stages | 720 | 720 | Derived |
| Invented control brand appearances across the whole study | 0 | 0 | Measured |
| HubSpot (CRM software) -- presence rate, combined | 100.0% | 100.0000% | Derived |
| HubSpot (CRM software) -- pick rate, combined | 97.4% | 97.3958% | Derived |
| QuickBooks (small-business accounting) -- presence rate, combined | 100.0% | 100.0000% | Derived |
| QuickBooks (small-business accounting) -- pick rate, combined | 94.3% | 94.2708% | Derived |
| Petco (pet food & supplies) -- presence rate, combined | 89.6% | 89.5833% | Derived |
| Petco (pet food & supplies) -- pick rate, combined | 97.7% | 97.6744% | Derived |
| Chewy.com -- presence rate, openai | 81.2% | 81.2500% | Derived |
| Chewy.com -- pick rate, openai | 87.2% | 87.1795% | Derived |
| Chewy.com -- presence rate, anthropic | 56.2% | 56.2500% | Derived |
| Chewy.com -- pick rate, anthropic | 37.0% | 37.0370% | Derived |
| Chewy.com -- presence rate, pooled | 68.8% | 68.7500% | Derived |
| Chewy.com -- pick rate, pooled | 66.7% | 66.6667% | Derived |
| Chewy.com -- estimated ceiling on the true pick rate, Anthropic only | 50.37% | 50.372777% | Derived |
| Chewy.com -- instances behind the Anthropic-only pick rate | 54 | 54 | Measured |
| PetFlow -- presence rate, combined | 7.8% | 7.8125% | Derived |
| PetFlow -- pick rate, combined | 100.0% | 100.0000% | Derived |
| PetFlow -- instances behind its pick rate | 15 | 15 | Measured |
| Wave Accounting -- presence rate, Anthropic only | 26.0% | 26.0417% | Derived |
| Wave Accounting -- pick rate, Anthropic only | 20.0% | 20.0000% | Derived |
| Wave Accounting -- instances behind its Anthropic-only pick rate | 25 | 25 | Measured |
| MYOB and Petsense -- presence rate, both engines | 0.0% | 0.0000% | Derived |
| Pre-registered pick-rate bar (the study's rule, not a measurement) | 40% | 40.0000% | Derived |
| Pre-registered presence-rate bar (the study's rule, not a measurement) | 50% | 50.0000% | Derived |
| Pre-registered small-sample safeguard bar (the study's rule, not a measurement) | 55% | 55.0000% | Derived |
| Spend per category | US$20.40 | 20.40 | Cited |
| Total spend across all three categories | US$61.20 | 61.20 | Cited |
What this doesn't settle
- Three categories is diligence, not a survey — the third, pet food, was a deliberate follow-up run five days after the first two, once they were already finished and published, not a claim about categories not tested.
- Each category is one day's answers, two AI engines, one buyer persona; whether these rates hold over weeks or shift with question wording is a different, already-designed check this piece does not run.
- The near-miss sat close to its own safeguard: the estimated ceiling on Chewy.com's true Anthropic-only pick rate was 50.37%, just under the 55% bar the rule requires.
- The bar used here is deliberately looser than the bar a real client's own report uses, which requires a higher presence rate before the same pattern would count — by design, not a gap to fix.
- Each category's two invented control names were tested only within that category's own answers, not across all three.
- These are API answers, not the consumer apps; results apply to the API measurement, not necessarily to what a consumer sees.
Which of your own failure modes have you tested for?
This piece uses the same frozen, pre-registered measurement a client's diagnosis runs on — published either way, including the near-miss.