| Topic | Testing AI recommendations |
|---|---|
| Best for | Local businesses and service providers |
| Time needed | About an hour for a first pass |
| Related tool | AI Visibility Tracker |
| Reading time | 5 minutes |
Most owners test this once. They type their own business name into an assistant, see a flattering or unflattering answer, and draw a conclusion. That tells you very little, because real customers rarely type your name. They describe a need and expect a recommendation. A useful test starts from the customer's question.
Step 1: write questions customers actually ask
Build a list of fifteen to twenty questions in four groups. First, discovery questions with no brand in them, such as who to see or which provider to choose in a particular city. Second, comparison questions, such as how two named providers differ. Third, trust questions about credentials, reviews and what to ask in a consultation. Fourth, cost and process questions.
Discovery questions matter most, because they are where a new customer first meets a brand. The prompt builder generates a starting set from a service and a location.
Step 2: test more than one assistant
Assistants differ in training data, in whether they browse the web, and in how they present sources. Test at least ChatGPT, Gemini and one other, such as Claude or Perplexity. Where the product offers a browsing or search mode, note which mode you used, because answers can change between them.
Step 3: run each question more than once
Answers vary. Run each question three times, ideally in fresh conversations so earlier context does not colour the result. Record for each run whether you were named, in what position, and which competitors appeared. A single run is an anecdote. Nine runs across three assistants start to look like data.
Step 4: record the sources
When an assistant lists links or names where its information came from, write them down. Patterns appear quickly. You may find that directories, review sites or a competitor's blog supply most of the citations. Those are the places that shape the answer, and they tell you where your own presence needs work.
Step 5: read the gaps
Sort your results into three buckets. Questions where you appear reliably are strengths to protect. Questions where you appear sometimes are the best opportunities, since a modest improvement can tip them. Questions where you never appear need a closer look at whether you have any page that answers them.
Also check accuracy. If an assistant states the wrong location, services or credentials, fixing the underlying source is more valuable than chasing a mention.
A simple recording template
You do not need special tools to keep the record. A spreadsheet with a row per run is enough. Use these columns: the question, the assistant used, the date, whether your brand was named, your position among the named brands, the competitors named, any links or sources shown, and a note on anything inaccurate. Add a final column for the question group, so you can later total your results by discovery, comparison, trust and cost.
Keep the wording of each question exactly the same from month to month. Small changes in phrasing can change which brands appear, so rewording a question makes comparisons unreliable.
Examples of useful question patterns
Patterns, not scripts, are what to copy. Discovery questions follow the shape of who is a good provider of a service in a place. Qualifier questions add a need, such as budget, a specialism or a time constraint. Comparison questions ask how two named options differ. Process questions ask what happens at a first appointment or what to prepare. Risk questions ask what to watch out for or how to tell a good provider from a poor one.
Mix all five patterns. A set made only of discovery questions will miss problems with how you are described when someone already knows your name, and a set made only of branded questions will flatter you.
Pitfalls that distort the result
- Logged-in personalisation. Assistants can use earlier conversations and settings. Test in a clean session where you can.
- Leading questions. Asking if your brand is the best provider will usually produce a polite yes. Ask neutral questions.
- Too few runs. One answer tells you almost nothing about how often you appear.
- Location mismatch. Assistants sometimes infer location from settings. State the city in the question so the test matches the customer.
- Confusing training with search. Note whether the assistant searched the web, because the cause of a gap differs.
What to do with a weak result
A weak first result is normal and useful, because it gives you a baseline. Resist the urge to change everything. Pick the two or three questions with the highest commercial value where you are absent, and ask what a customer would need to see to choose you. Usually the answer is a page that does not exist, a listing that is incomplete, or an identity detail that is inconsistent. Fix those, wait a few weeks, and rerun the same questions.
Doing it without the spreadsheet
The manual method is a good way to learn, but it gets tedious. The AI Visibility Tracker runs your question set through an assistant, repeats it, counts mentions and positions against the competitors you list, shows which sites were linked, and keeps a history so you can compare month to month. It uses the model without live web search, so treat it as a consistent benchmark rather than a mirror of any one consumer app.
Frequently asked questions
Why does ChatGPT give different answers each time?
Generated answers involve randomness, and the system may draw on different information between runs. That is why repeating each question and looking at the share of answers that include you is more reliable than one result.
Should I ask using my business name?
Include a few, because they show what the assistant says about you. But most of your testing should use questions with no brand name, since that is how new customers ask.
Does a mention mean customers will contact me?
Not necessarily. A mention is a necessary first step. Position, the wording around it and whether a source is linked all affect whether someone acts.
How many questions are enough?
Fifteen to twenty well-chosen questions are enough to reveal patterns. Add more once you know which topics matter most.