Which AI engines actually recommend anyone — 34,411 answers classified
We assumed Gemini would be the engine that never commits. It is not. ChatGPT is the least decisive of the seven we measured: 4.2% of its answers name a firm as the pick, against 35.5% for Grok.
Steen Stones · Reviewed 6 Aug 2026
Being mentioned by an AI engine and being chosen by one are different things, and almost every visibility report measures the first. So we went and measured the second: for every answer in our store, does the engine actually name a pick, or does it hand the reader a list and leave them to it?
34,411 answers, seven engines, three B2B service markets. Every answer classified into one of six categories — a clear recommendation, a pick with close runners-up, a list with no pick, firms mentioned only in passing, a warning, or no firms at all. The result was not what we expected, and it overturned four working theories of our own along the way.
Avouch answer-framing study, 6 August 2026 — 34,411 stored AI answers across three B2B service markets and seven engines, each classified for whether it picks a firm or merely lists them
The engines split into two camps
Share of answers that name a specific firm as the choice, including those that pick one with close runners-up. Grok (1,469 answers) and Brave (1,133) sit on thinner counts than the rest.
Grok, Perplexity and Brave behave like an opinionated colleague: asked who to use, they name somebody. ChatGPT, Gemini and Google's AI Overviews behave like a directory — they will happily produce ten firms and decline to separate them. Claude sits in between. The gap between the top and bottom of that chart is more than eightfold, on the same kinds of question.
The surprise is which engine is at the bottom. Going in, we expected Gemini to be the outlier that never commits, on the strength of our own market sweeps where it named firms hundreds of times and recommended almost none. It is near the bottom. But ChatGPT — the engine most of your buyers are actually using — is lower still.
Why this changes what a visibility number is worth
A citation count tells you where you appear. It does not tell you whether appearing there changes anything, and these figures show the answer depends heavily on which engine you appeared in. Being the most-mentioned firm inside an engine that never chooses is a different asset from being the pick inside one that does. Same visibility, different value.
- If your buyers use Perplexity or Grok, the pick is genuinely up for grabs and being chosen is worth chasing.
- If your buyers use ChatGPT or Gemini, the realistic goal is making the shortlist, because a single named winner is rare.
- Either way, a report that counts mentions without separating the two is measuring the wrong thing.
The caveat that matters most
These are natural buyer questions — the sort a real person types. They are not us instructing the model to choose. That distinction turns out to be the whole ballgame: when a prompt explicitly says name the firm you would recommend, the same engines pick far more often. Our own market sweeps use exactly that instruction, and they show ChatGPT recommending in most answers rather than 4.2% of them.
Both measurements are honest, but they answer different questions. One is how an engine behaves when it is told to pick. The other is how it behaves when a buyer simply asks. Any study that does not say which one it ran is not telling you enough to judge it, so: this one is the second.
How it was measured
- 34,411 stored answers across seven engines and three B2B service markets, classified one by one rather than sampled.
- Six categories, so a heading that reads Recommended Providers above an unranked list scores as a list, not a recommendation.
- A single judge model with a fixed prompt version, validated on a 500-answer pilot with no parse failures before the full run.
- Every verdict is stored with its judge model, prompt version and batch reference, so any figure above can be re-derived rather than taken on trust.
- Grok and Brave carry the smallest counts (1,469 and 1,133), so those two percentages are the least stable on the chart.
We are not neutral about this: we sell measurement, so a study showing measurement is harder than it looks is convenient for us. The counter to that is the working. Every number here comes from a stored, re-runnable classification rather than a claim, and the limitations are on the page rather than in a footnote.
AI can't recommend what it doesn't know.
Start free. No commitment.