Can AI know a brand and still rarely recommend it?

We went looking for one of our own three failure modes — a brand known to AI but rarely picked by it — across three unrelated shopping categories, and published what we found either way.

Combined across both engines, in these three categories — CRM software, small-business accounting, and pet food and supplies — on this design, no brand out of eighteen was both present in at least half of answers and picked in fewer than four times in ten. The closest any reading came was a pet retailer's answers on one AI engine, which cleared every threshold on that engine alone — 56.2% presence, a 37.0% pick rate, and a safeguard estimate of 50.37% — before the second engine's own reading brought the combined number back above the bar.

4 August 2026

Our own AI-visibility framework claims more than one way a brand can fail — but claiming it and finding it in real data are different things. We went looking for one of those failure modes directly: a brand known to AI but rarely picked by it, across three very different shopping categories, on the same measurement setup we use for client work, with rules built to test for the pattern in three specific categories, not to catch it if it exists anywhere. We publish what turned up either way.

A brand counts as showing up when its name (or, for the pet retailer with a common-word name, its full brand form) appears in an answer; it counts as picked when the answer names it as the one it actually recommends, not merely mentions it.

Being known and being picked are supposed to be two different problems

If a brand never gets recommended, the easy assumption is that AI simply does not know it exists. The theory this piece tests says that is not the only way to lose: a brand can be named or considered constantly and still rarely be the one AI actually recommends — known, but not compelling enough to pick. That is the theory under test here. We do not say yet whether it held up.

Three shopping categories, the same two AI engines, the same rule every time

We picked three categories to be three different shapes of buying decision: which CRM software a small business should buy, which accounting software a small business should buy, and where to buy pet food and supplies online. CRM software is a category where AI answers tend to name many brands in one list. Accounting software is a category with one or two dominant defaults. Pet food is an ordinary consumer purchase, not business software at all.

Each category put eight different real buyer questions to two AI engines — OpenAI's and Anthropic's — many times over, for twenty-four distinct questions across the three categories in total. We measured six real brands by name per category: a couple of clear category leaders plus several smaller or more specialized names, not all equally well-known going in. We also planted two invented brand names with no real company behind them as a control.

For every brand we measured two things: how often it showed up in an answer at all (its presence rate), and, of the times it showed up, how often it was the one AI actually recommended (its pick rate). Both rates are combined across both engines' answers unless a sentence says otherwise.

Three thresholds, and a brand must clear all three

Present in at least half of all answers (50%). Picked fewer than 4 times in 10 (40%). The generous estimate of the true pick rate stays under 55%. This is the rule itself, not a measured brand — every brand's own reading is checked against it in the next section.

Present in at least half of all answers
0%
50% — counts as known
100%
Picked fewer than 4 times in 10
0%
40% — counts as rarely picked
100%
Even the generous estimate stays under 55%
0%
55% — not just a small-sample fluke
100%
A brand counts as the pattern only when all three hold at once, combined across both engines.

The bar for "known but not picked, and not by chance," evaluated on the combined answers from both engines, had three parts: present in at least half of all answers, picked less than 4 times in 10, and — because a rare pattern can look real in a small sample — the estimated ceiling on how high the true pick rate could plausibly be, even generously, had to sit under 55%.

Each category's own roster of six real brands and two invented names was locked in writing before that category's own answers were collected. All three categories used the same roster shape and the same rule. The three categories were not run as one program decided in advance: the third, pet food, was ordered five days after the first two, once they were already finished and published, specifically to check the result outside business software. We applied no correction for that timing gap, because the pet-food category is a deliberate follow-up check, not a rerun of the first two on matching dates.

Every category ran a small first pass before the full count, to confirm the measurement could tell a real recommendation apart from a brand merely being considered, and that the invented control names stayed silent — not to confirm the target pattern itself was present. That first pass added 48 answers per category, 144 across all three. The reported rates below come from the larger full count that followed: 192 answers per category, 576 across the three, on top of the 144 from the first pass — 720 answers in total. Every brand rate quoted anywhere in this piece is built on the 576-answer full count, never the 720-answer grand total.

Zero of eighteen brands cleared the combined-engine rule

The zone this pattern would occupy — almost empty

Presence rate against pick rate, combined across both engines. Only one reading ever lands inside the dashed zone — and it is Chewy.com's Anthropic-only reading, not a combined one.

00252550507575100100presence ≥ 50%pick rate < 40%known, butrarely pickedChewy.com, Anthropic onlyOpenAI onlypresence rate (%)pick rate (%)
Filled dots: combined-engine reading (16 of 18 real brands with a measurable pick rate). Hollow dots: Chewy.com's two single-engine readings, connected to its own combined dot. MYOB and Petsense showed 0% presence on both engines and are not drawn.

Combined across both engines, in these three categories, on this design, none of the eighteen real brands was both present in at least half of answers and picked in fewer than four times in ten. In the combined answers — 192 per category, 576 across the three — every brand present in at least half of answers had a pick rate of 40% or higher, with no exceptions: none of the eighteen real brands met all three parts of the rule, combined.

The strongest instances of the opposite pattern, one per category, make the point concrete. In CRM software, HubSpot showed up in 100% of answers and was the pick in 97.4% of those. In accounting software, QuickBooks showed up in 100% of answers and was the pick in 94.3% of those. In pet food and supplies, Petco showed up in 89.6% of answers and was the pick in 97.7% of those.

The closest a single engine's own reading came was Chewy.com on Anthropic's answers alone: the pet retailer showed up in 56.2% of answers and was the actual pick in only 37.0% of those, clearing every threshold, including the small-sample safeguard, whose estimated ceiling on the true pick rate sat at 50.37%. On OpenAI's own answers for the same retailer, the pick rate was 87.2%. Combined across both engines — the reading the study's rule actually evaluates — Chewy.com's presence was 68.8% and its pick rate was 66.7%, above the 40% bar, so it does not count. The combined number is the one that counts because the study measures each brand's position across both engines' answers together, not whichever single engine happens to look most interesting; one engine's answers alone are half the evidence, not the verdict.

Low pick rates came from thin samples — except one

The study's lowest pick rate came from a brand barely seen at all: Wave Accounting showed up in just 26.0% of Anthropic's answers and was picked in only 20.0% of those 25 instances. A thin sample can cut the other way too — PetFlow showed up in only 7.8% of answers but was picked 100% of the 15 times it did, the opposite extreme from an equally small base.

Two brands were essentially invisible on both engines: MYOB in accounting and Petsense in pet food each showed up in 0% of answers, on both engines, across the whole count.

Extreme pick rates often rested on very few appearances

Wave Accounting (Anthropic only): 26.0% presence, picked in 20.0% of appearances. PetFlow (combined): 7.8% presence, picked in 100% of appearances. Chewy.com (Anthropic only): 56.2% presence, picked in 37.0% of appearances.

Wave Accounting · Anthropic only
n = 96
5 picked of 25 that mentioned it
Chewy.com · Anthropic only
n = 96
20 picked of 54 that mentioned it
PetFlow · combined, both engines
n = 192
15 picked of 15 that mentioned it
Not present
Present, not picked
Present and picked

In the combined numbers, every brand whose presence rate cleared the 50% bar also had a pick rate at or above 40%, with no exceptions. The one place a different picture shows up is a single-engine number, not a combined one: on Anthropic's answers alone, Chewy.com's pick rate stayed at 37.0% from 54 in-pool instances, alongside a presence rate above the 50% bar. The failure shape the combined data actually shows is a brand not being seen enough, not a brand being seen often and still turned down — our reading of this dataset, not a general law.

Three categories, one bar deliberately set looser than a real diagnosis

Three categories is diligence, not a survey. CRM software and accounting software ran the same day; pet food ran five days later, once the first two were already finished and published, specifically to check the result outside business software. The consequence is that this is not a claim about categories not tested, and there are far more than three kinds of purchases a brand can be searched for. We applied no correction because the third category was a deliberate, separately-run follow-up, not a rerun of the first two.

Each category is one day's answers, two AI engines, one buyer persona. Whether these rates hold over weeks, or shift with the exact wording of the question, is unknown from this design. We applied no correction because testing stability over time or across wordings is a different, already-designed kind of check this piece does not run — not an oversight, a different study.

The near-miss itself sat close to its own safeguard: the estimated ceiling on Chewy.com's true Anthropic-only pick rate was 50.37%, just under the 55% bar the rule requires. A reading this close to its own limit is a genuine clearance on that one engine, not a strong one.

The bar used here — present in at least half of answers, picked in under 4 of 10, combined across both engines, plus the 55% safeguard — is deliberately looser than the bar a real client's own report uses, which requires a higher presence rate before the same pattern would count. A category clearing this study's looser bar would still need to clear that higher one before it changed a real diagnosis. We applied no correction because this piece is explicitly an existence test, not a client diagnosis — the looser bar is by design, not a limitation to fix.

Each category's two invented control names were tested only within that category's own answers, not across all three categories. Collectively, all six category-specific control names appeared in zero answers, in the category each was planted in, across all 720 answers collected — a measured floor for those specific names in their own category, not a claim about every possible name or category. They exist to prove the measurement is not inventing mentions in the category they were planted in.

These are API answers, not the consumer apps. The results apply to the API measurement, not necessarily to what a consumer sees in the chat apps. We applied no correction because the consumer apps were not sampled in this design.

The rare picks mostly came from being rarely seen

In the combined numbers, in these three categories, on this design, the low pick rates came from brands AI barely mentioned in the first place, not from brands that were well-known and turned down — the one single-engine reading that came close (Chewy.com on Anthropic) did not survive being combined with the other engine. This is a reason to check the presence rate first before assuming a poor showing is a positioning problem, not a proof that positioning problems never occur.

Check both a presence rate and a pick rate, on more than one AI engine, before deciding a brand has a positioning problem rather than a visibility one.

For our own measurement, this is one of three failure modes we watch for — a brand known to AI but rarely picked by it — and we did not detect it here, combined, across these three categories. We do not treat it as an established pattern; we would report it if it ever cleared the combined bar.

How we ran this

Three categories: CRM software and small-business accounting software ran the same day; pet food and supplies ran five days later, once the first two were finished, as a follow-up check. Two AI engines each (OpenAI, Anthropic), API answers under fixed settings, not the consumer apps.

Six real brands per category (eighteen total) plus two invented control names per category (six total); each category's own roster was locked in writing before that category's own answers were collected. The same six-real-plus-two-invented shape and the same combined-answer pass/fail rule ran unchanged in all three categories.

We collected 720 answers in total across the three categories: 144 from a small first pass (48 per category) that checked the measurement itself could tell a real recommendation apart from a mere mention, and that the invented controls stayed silent — followed by 576 from the full count (192 per category) that every reported brand rate in this piece is built on. Total spend across all three categories: US$61.20.

How we ran this

What we measured
Six real brands per category (eighteen total) plus two invented control names per category (six total); eight buyer questions per category (twenty-four total), put to two AI engines many times over. 720 answers collected across the three categories: 144 from a small first pass, 576 from the full count every reported rate is built on.
When
CRM software and small-business accounting software: 30 July 2026. Pet food & supplies: 4 August 2026, a follow-up check.
Which AI
OpenAI (API), Anthropic (API)

The full numbers

Every number in this piece, at full precision. The prose rounds for reading; this table doesn't.

MeasureAs shownExactSource
Shopping categories tested33Measured
AI engines per category22Measured
Buyer questions per category88Measured
Buyer questions across all three categories2424Derived
Real brands measured per category66Measured
Real brands measured across all three categories1818Derived
Invented control brands per category22Measured
Invented control brands across all three categories66Derived
Days between the first two categories and the third55Measured
Full-count answers per category192192Measured
Full-count answers across the three categories, the basis for reported rates576576Derived
First-pass answers per category4848Measured
First-pass answers across the three categories144144Derived
Total answers collected across the study, all stages720720Derived
Invented control brand appearances across the whole study00Measured
HubSpot (CRM software) -- presence rate, combined100.0%100.0000%Derived
HubSpot (CRM software) -- pick rate, combined97.4%97.3958%Derived
QuickBooks (small-business accounting) -- presence rate, combined100.0%100.0000%Derived
QuickBooks (small-business accounting) -- pick rate, combined94.3%94.2708%Derived
Petco (pet food & supplies) -- presence rate, combined89.6%89.5833%Derived
Petco (pet food & supplies) -- pick rate, combined97.7%97.6744%Derived
Chewy.com -- presence rate, openai81.2%81.2500%Derived
Chewy.com -- pick rate, openai87.2%87.1795%Derived
Chewy.com -- presence rate, anthropic56.2%56.2500%Derived
Chewy.com -- pick rate, anthropic37.0%37.0370%Derived
Chewy.com -- presence rate, pooled68.8%68.7500%Derived
Chewy.com -- pick rate, pooled66.7%66.6667%Derived
Chewy.com -- estimated ceiling on the true pick rate, Anthropic only50.37%50.372777%Derived
Chewy.com -- instances behind the Anthropic-only pick rate5454Measured
PetFlow -- presence rate, combined7.8%7.8125%Derived
PetFlow -- pick rate, combined100.0%100.0000%Derived
PetFlow -- instances behind its pick rate1515Measured
Wave Accounting -- presence rate, Anthropic only26.0%26.0417%Derived
Wave Accounting -- pick rate, Anthropic only20.0%20.0000%Derived
Wave Accounting -- instances behind its Anthropic-only pick rate2525Measured
MYOB and Petsense -- presence rate, both engines0.0%0.0000%Derived
Pre-registered pick-rate bar (the study's rule, not a measurement)40%40.0000%Derived
Pre-registered presence-rate bar (the study's rule, not a measurement)50%50.0000%Derived
Pre-registered small-sample safeguard bar (the study's rule, not a measurement)55%55.0000%Derived
Spend per categoryUS$20.4020.40Cited
Total spend across all three categoriesUS$61.2061.20Cited

What this doesn't settle

  • Three categories is diligence, not a survey — the third, pet food, was a deliberate follow-up run five days after the first two, once they were already finished and published, not a claim about categories not tested.
  • Each category is one day's answers, two AI engines, one buyer persona; whether these rates hold over weeks or shift with question wording is a different, already-designed check this piece does not run.
  • The near-miss sat close to its own safeguard: the estimated ceiling on Chewy.com's true Anthropic-only pick rate was 50.37%, just under the 55% bar the rule requires.
  • The bar used here is deliberately looser than the bar a real client's own report uses, which requires a higher presence rate before the same pattern would count — by design, not a gap to fix.
  • Each category's two invented control names were tested only within that category's own answers, not across all three.
  • These are API answers, not the consumer apps; results apply to the API measurement, not necessarily to what a consumer sees.

Which of your own failure modes have you tested for?

This piece uses the same frozen, pre-registered measurement a client's diagnosis runs on — published either way, including the near-miss.

Join the waitlist