Skip to main content
AI Search

How to Test Local Business Visibility in AI Answers

Two sources, one method: the first-party Search Console report and a dated prompt test you can re-run. Plus what nobody can promise.

14 min read

To test whether a local business shows up in AI answers you need two things, and they are not interchangeable: the first-party report Google now provides, and a prompt test you run yourself on a fixed schedule with the date recorded every time. The report tells you whether your pages appeared. The prompt test tells you what an actual answer looked like, and who appeared instead of you. Neither one supports a promise about future placement, and this guide ends by saying so plainly.

What does AI search visibility actually mean?

It means one narrow thing: when someone asks an assistant a question your client could answer, does your client's business get named, and does the assistant cite a page you control. That is it. It is not a score, and it is not a ranking position — there is no public ranking to hold a position in.

This matters for how you talk to clients, because the question they ask is usually "why doesn't ChatGPT recommend us", and the honest first move is to convert that into something testable: which questions, asked how, answered when.

Where does the first-party data come from?

In June 2026 Google introduced Generative AI performance reports in Search Console. They show how often URLs from a site appeared in generative AI features in Search and Discover, and break that out by page, by country, by device (for Search results), and by date with hourly through monthly granularity. Google described the reports as rolling out to a subset of websites while it gathers feedback, so the first thing to check on any client property is simply whether the report is there yet.

Two notes for the audit row:

  • Generative AI data is also included in the overall performance report; the new view is a separate, dedicated cut of it.
  • If the report is not available for a property, the correct entry is "not available", not an estimate. An estimate written into a client document becomes a number they repeat.

One boundary belongs here rather than in a footnote: this report covers Google's surfaces. It says nothing about whether an assistant from another provider mentioned the business. Whether any other provider exposes something equivalent to site owners is a question for that provider's own documentation — do not assume one exists, and do not assume one does not. That is also the rule for any claim about a non-Google assistant: it needs that provider's own documentation behind it, not Google's.

How do you build a repeatable prompt test?

Fix the prompts, fix the fields, fix the schedule. Everything that varies between runs should be the answer, not the method.

Fix five kinds of prompt. For a local business, these cover the ways a real person asks:

  1. Recommendation — "who should I call for {service} in {city}"
  2. Comparison — "{competitor} vs {client}" or "best {service} in {city}"
  3. Problem — the symptom a customer types before they know what service they need
  4. Brand — "what is {client}" and "is {client} any good"
  5. Local qualifier — the same service question with a neighborhood or landmark attached

Fix three fields per prompt. For every run, record: was the client mentioned, which other businesses were named instead, and which sources the assistant cited.

Fix the run conditions, and record the ones you can see. Skip this and the test becomes an anecdote: two rows filed under the same assistant name can come from different products, different modes and different accounts, and nothing in the record would show it. At minimum, start every prompt in a fresh session and write down what the interface actually shows you — the product or model name, the mode (whether web search or browsing is switched on), whether you are signed in, whether personalization or memory is switched on, and the locale or location the assistant appears to be using. Signed-in and personalized are two separate facts rather than one: an account can be signed in with memory off.

Anything you cannot see or cannot fix, record as unknown — and be precise about what that buys you. Writing unknown keeps you honest about which variables you did not control. It does not turn two runs into a like-for-like pair. Two rows that both say unknown support a descriptive comparison at best; they do not license reading a change between them as a trend, and they certainly do not license attributing that change to something you published.

Fix the date. Every row carries the date it was run. Without it the record cannot be compared to anything, which means it cannot show change — and showing change is the only reason to keep the record.

What do you do with the answers you get?

Read the failures, not the successes. The prompts where an assistant answers confidently without your client are the most useful output of the whole exercise — but treat a non-mention as a signal to investigate, not as a finding. On its own it does not establish that the client's page was inadequate. The same result can come from retrieval that never reached the site, from weak third-party signals, from location or personalization, from the assistant's own selection among several adequate sources, or from run-to-run variation.

So before you categorize anything, do three cheap checks: look at what the client actually publishes on that topic, look at which sources the assistant did cite, and confirm the run conditions were the ones you intended. Skipping those three is how an agency ends up billing a content rewrite for a retrieval problem.

Then group what survives those checks:

  • Content — no page. The client has nothing that addresses this at all. This is the only category you can act on directly and quickly.
  • Content — a page that reads like a brochure. Something exists, but nothing on it can be lifted out and stand on its own as an answer.
  • Authority — answered by third parties. Directories, review sites and local press are answering for the whole category. This does not have a fast fix.
  • Retrieval or unknown. The client has a page that answers the question well, the run conditions were what you intended, and the assistant still did not use it — or the result moved between runs. Record it, re-test it next cycle, and do not bill a rewrite against it. This bucket existing is what keeps the other three honest.

The brand prompts deserve separate attention. If an assistant cannot say what the business is, where it operates and what it sells, that is an entity problem rather than a content one. Google's own documentation on establishing business details is the useful starting point here: it names the mechanisms a business actually controls — claiming the Business Profile, verifying in Search Console, updating the knowledge panel as a verified representative, adding structured data — and notes that Google's algorithms also assemble information such as a site's name and contact details from what is publicly available across the web. In other words, part of the entity record comes from sources the client does not own, which is why reconciling the parts they do own comes first. The local SEO audit checklist covers how to reconcile that record.

How often should you re-run the test?

Often enough that the record shows movement, rarely enough that the team keeps doing it. Monthly is a defensible cadence for most local clients, aligned to whatever reporting rhythm you already have. What matters more than frequency is that the prompt set does not drift: changing the prompts between runs produces two datasets that cannot be compared, which is the same as having no dataset.

What can you promise a client about AI answers?

Nothing about placement. Not a mention, not a citation, not a position.

This is not caution for its own sake. Assistants change their retrieval and their answers without notice or changelog, and the same prompt can return a different answer on the same day. Anyone selling a guaranteed mention is selling something they cannot control.

What you can commit to is the work and the record: a fixed prompt set, run on a schedule, with dated results, plus the specific pages published in response to the gaps it found. That is a service with a deliverable. "We will get you into AI answers" is not.

What does a finished test record look like?

One sheet, one row per prompt per run. Every row carries two things at once: what you observed in that answer, and the conditions you ran it under.

Column What goes in it
Date The date the prompt was run. Not the reporting period — the day.
Assistant Which assistant answered. One column value, never merged.
Prompt The prompt, verbatim, including the city or neighborhood.
Mentioned Yes or no. Not "sort of" — if the business was not named, it is a no.
Others named Every business the answer named instead, in the order it named them.
Sources cited The URLs the assistant showed, if it showed any.
Product / model Whatever the interface names, verbatim. unknown if it names nothing.
Mode Whether web search or browsing was on. unknown if you cannot tell.
Session Fresh session, yes or no.
Signed in Yes or no.
Personalization / memory On, off, or unknown. A separate fact from being signed in.
Locale / location What the assistant appears to be using, not what you assume.

The top group identifies the run and records what came back. The bottom group records the conditions it came back under — without those you have a pile of dated anecdotes filed under one assistant name.

Within the first group, Others named and Sources cited are where the value accumulates. After a few runs, Others named becomes a list of who the assistants treat as the default answer in that category and city, and Sources cited becomes a list of the pages and directories that category is being answered from. Both are far more actionable than the yes/no column everyone looks at first.

Keep the sheet with the client's other reporting, not in a tool that expires. The record's whole purpose is to still be readable a year from now.

What about third-party AI visibility tools?

Use them if they fit your workflow, and read their numbers as their model of the thing rather than the thing. Google is direct about this in its own guidance: be wary of third-party tools that promise ranking success or claim to use internal Google metrics, because no third-party tool has access to Google's internal ranking or AI systems. Google's advice is to evaluate any tool's suggestions against official guidance.

Practically, that means a vendor score can be a useful trend line inside one tool over time, and it cannot be quoted to a client as a fact about how an assistant works.

Frequently Asked Questions

Why does the same prompt give different answers?

Because assistants sample and re-retrieve, their underlying indexes and models change without public notice, and the conditions you ran under may not have been identical. This is why every row carries both a date and its run conditions: one run is an anecdote, a series of dated runs under recorded conditions is evidence of a trend, and a series under unrecorded conditions is neither.

Does a non-mention mean the client's content is bad?

Not on its own. A non-mention is a signal to investigate. Retrieval that never reached the site, weak third-party signals, location or personalization, the assistant's own selection among adequate sources, and plain run-to-run variation all produce the same visible result. Check what the client publishes on the topic, which sources the assistant did cite, and whether the run conditions were the ones you intended — then categorize.

Is AI search visibility a different job from SEO?

Mostly no. Google states that its generative AI features are rooted in its core Search ranking and quality systems and that SEO best practices remain relevant. The part that is genuinely different is measurement, which is why it gets its own report and its own test record rather than its own strategy.

How many prompts should the test include?

Enough to cover all five kinds, and few enough that the team actually re-runs them. A set that covers the five kinds for the client's two or three main services is a working starting point; a set nobody re-runs is worth nothing regardless of size.

Should I test on more than one assistant?

Test the ones your client's customers plausibly use, and keep each assistant's results as their own rows in the same sheet — the Assistant column is what separates them, which is why the sheet stays one row per prompt per run. What you must not do is merge them into a single blended figure: an average across assistants describes no real user's experience.

Can I guarantee an AI assistant will recommend my client?

No. Nobody can, including anyone who says otherwise. What can be guaranteed is that the test was run, on these prompts, on this date, and that these pages were published in response to what it found.

References

  1. Google Search Central Blog. "Introducing Search Generative AI performance reports in Search Console." developers.google.com
  2. Google Search Central. "Google's Guide to Optimizing for Generative AI Features on Google Search." developers.google.com
  3. Google Search Central. "AI Features and Your Website." developers.google.com
  4. Google Search Central. "Add Business Details to Google." developers.google.com
  5. Google Search Central Blog. "Top ways to ensure your content performs well in Google's AI experiences on Search." developers.google.com
SHARE THIS ARTICLE:
STAY UPDATED

Subscribe to Our Newsletter

Get the latest insights on AI Agent capabilities and stay ahead in your business.