
How to check your company’s AI search visibility
Start with a clean conversation, repeat the test and measure what the answer actually says.
Start with a fresh conversation that has no saved information about your business. Ask an unbranded buyer question, then record what each AI tool says. Repeat the exercise using the same settings. This gives you a direction to track, not a complete picture of every buyer’s experience.
First, stop the test from knowing who you are
The easy mistake is to ask an assistant you use for work which companies it recommends. If you have already discussed your business, the answer may reflect that history. Opening another chat is not enough if saved memory still carries information across conversations.
Start a fresh conversation outside your work project or existing thread.
Disable saved memory, previous-chat personalisation and custom instructions that identify your company. Check the tool’s current settings.
Do not upload company files or name your business in the question. Test branded questions separately.
Record the account plan, location, tool, model where shown, search setting and date. Keep these conditions the same when repeating the test.
A separate free account can be a practical starting point. We would still check its memory and settings, and label the results with the plan used. Free is a price, not a guarantee of a clean test.
For repeatable checks, we prefer an API, a way for software to send a request directly to a model. Send each baseline question without earlier messages or a linked conversation. OpenAI allows callers to supply conversation history, so using its API alone does not remove context.
Record the instructions and search settings supplied with the request. Keep API tests separate from app tests: you are measuring that setup, not claiming to reproduce every person’s ChatGPT session.
Which questions should we ask?
Choose questions a buyer might ask before knowing your name. Include the problem they need to solve, the kind of product they want and what matters when choosing. Save that question set before looking at results.
For example, a supplier-discovery question might be “Which project management tools suit a small UK software team?” A separate brand check asks “What are the drawbacks of our product?” Keep discovery and brand checks in different groups. Naming yourself changes what you are testing.
How many times should we repeat each question?
Our suggested starting exercise is five fresh runs per question, per engine. Keep every answer, including the disappointing ones. Five runs can expose variation, but they are a pilot, not enough to estimate your visibility reliably.
Rand Fishkin and Patrick O’Donnell’s study collected 2,961 responses across ChatGPT, Claude and Google’s AI search. Its prompts were run 60–100 times each, and the lists, order and number of recommendations varied. Five is our practical starting point, not the study’s sample size.
Fishkin found that how often a company appeared across many runs could still be useful, even when the order changed. Track frequency over time and expand the sample before drawing conclusions from small movements.
Reliable presence requires repeated testing across multiple sessions per prompt and across multiple tracking periods.
A clean baseline leaves out a buyer’s existing history. Run a separate set of tests with realistic needs, constraints and follow-up questions. For example, ask which shortlisted option suits a smaller budget. Label these scenarios clearly, rather than mixing them into the clean baseline.
What should we record from each answer?
Decide the scoring rules before collecting answers. Being named, linked and recommended are different outcomes. Aleyda Solís makes the same distinction in her measurement guidance.
Swipe across to read all columns.
| Reference | Record | What to write down |
|---|---|---|
| Aleyda Solís | Mention | Does the answer name your company? Count it once per answer, even if the name appears repeatedly. |
| Aleyda Solís | Recommendation | Does the answer suggest choosing your product for the stated need? Keep a neutral reference separate. |
| OpenAI sources guidance | Citation | Which pages are linked as evidence? Record the full URL and distinguish your own pages from third-party pages about you. |
| Aleyda Solís | Sentiment | Is the description positive, neutral, mixed or negative? Save the exact wording and the reason for that label. |
| Bing citation reporting | Share of voice | State the formula and competitor list. A simple option is your company’s answer-level mentions divided by all tracked companies’ answer-level mentions. Do not call this market share. |
| Ahrefs tracking guidance | Test conditions | Save the engine, model where shown, account plan, date, location, prompt, memory settings, search status and complete answer. |
Also report your mention rate: the share of tested answers that name you. Its denominator is all eligible answers in your test. Share of voice uses mentions across a defined group of companies instead. Changing that group changes the result, so keep it fixed when comparing periods.
Read negative and mixed answers closely. A company can gain mentions while being described less favourably. Record the complaint or limitation, not just a sentiment label.
A citation is not a complete reading history
The links shown in an answer may not include every page consulted. OpenAI’s web-search documentation distinguishes inline citations from its fuller sources list. So “G2 was not cited” does not establish “the system never looked at G2”.
Keep answers without live search in the record too. Mark search as yes, no or unknown, using the tool’s recorded activity where available. No visible citation is not enough evidence to choose “no”.
Separate the model’s learned knowledge from saved personal memory. The former is what it learned during training; the latter is information retained about a user. Removing personal memory does not erase learned knowledge. Our recommendation is to report answers with and without confirmed web search separately.
Save an unbranded buyer question and the test settings
Clean baseline
- Remove personal history and company context
- Repeat in fresh conversations or independent API requests
- Score mentions, recommendations, citations and sentiment
Buyer scenarios
- Add a clearly described buyer need or constraint
- Record follow-up questions and full answers
- Report these results separately from the baseline
Compare the same questions and conditions over time. Keep each engine separate; neither branch represents every buyer.
Sources: Ahrefs: context and memory; SparkToro: repeated tests
Can we see the searches behind an answer?
Look for query fan-out: the additional searches used to explore parts of the buyer’s question. Google says AI Overviews and AI Mode may use this approach.
For OpenAI API tests, save the search-call details when available. They usually include the search queries, but not always. This is evidence from that API run, not a complete record of how the ChatGPT app searches.
For Google AI search, distinguish the links you can see from searches you are inferring. Do not ask an assistant to invent its hidden search history and then treat the answer as a log. Label proposed search queries as hypotheses until you have a recorded trace.
What can Google and Bing reports add?
Google includes AI Overviews and AI Mode in overall Search Console Performance reporting under Web. Do not label all of that traffic as AI traffic.
Bing Citation Share measures a site’s citations against all citations for the same grounding query, the search used to find supporting material. It is not traffic share or a content-quality score.
What should we do with the result?
Compare each engine with its own earlier results. Look for patterns that survive repeated checks: missing recommendations, wrong descriptions, dated sources or recurring objections. Use those findings to choose the next piece of work.
Treat the report as a direction of travel. It does not capture every buyer’s wording, memory, location or follow-up conversation. Keep the original answers so you can tell a real change from a change in how you ran the test.