This is the whole method, including the parts that make it inconvenient. You can do it yourself for free. Most people checking once are fooling themselves.
We are going to give away the method, because the method is not the hard part. Doing it carefully and repeatedly is.
Not keywords. Questions, in the words a person would use when they do not yet know your company exists.
The last one is the most revealing and the one people skip.
A warning about phrasing. Some questions reliably produce no shortlist at all. In our testing, a purely definitional phrasing — comparing two concepts rather than asking for a provider — returned a definitional answer with no companies named, on both engines, on both runs. If you only test questions like that, you will conclude nobody is mentioned, which is true and meaningless. Test questions with buying intent.
ChatGPT, Perplexity, and a Google search where the AI Overview appears. Five questions, two runs, three engines — thirty runs.
Use a fresh session each time. This is the step people get wrong. These tools carry context: in our own testing a nominally separate ChatGPT conversation opened by referring to something asked in an earlier session, which means it was not a clean run.
Signed out, in a private window, is cleaner still. If you must stay signed in, record that you did and treat the results as "what this account is shown" rather than what a stranger sees.
On Google especially, keep these apart — conflating them is the most common error:
Save the full transcript of every run as a file. Not a summary written afterwards — the actual text. When you re-measure in a month you will want to compare like with like, and memory is not evidence.
Two numbers come out:
That second number is the one people underrate. When we ran this on our own category, the most frequently recommended agency appeared in six of thirty runs. The next four appeared in five each. Around sixty other names appeared once or twice.
Read that shape carefully. A market where the leader holds roughly a fifth of answer-share, with a long tail of names appearing once, is a market with no incumbent. Had one brand appeared in twenty-eight of thirty runs, that would be a very different, much worse, finding.
State the limits plainly, because they determine how much the result is worth:
A finding with its limits stated is usable. A clean-looking number with hidden caveats will mislead you later, when you re-measure and cannot explain why it moved.
Zero out of thirty is common and is not a verdict on your company. It usually means an answer engine has little to work with: a thin site, no structured data, and nothing written about you by anyone else.
It is also a clean baseline. Fix what you can, wait a month, run the identical thirty and compare. If the number has not moved, your changes were not the bottleneck — which is a genuinely useful thing to learn, and most people never set themselves up to learn it.
A 30-run audit across ChatGPT, Perplexity and Google AI Overviews. Every transcript saved. 48-hour turnaround.
Book a free audit