The short answer
Measure AI search visibility with separate records for brand mentions, linked source pages, platform-reported exposure and attributable website visits. Repeat a defined question set under recorded conditions, then assess relevant inquiries separately. Do not combine these signals into one success score: each answers a different question, and none alone proves that AI visibility generated business.
Your brand appears, but your website is not linked
Check: Is this an accurate mention, and which source supports it?
Next: Record the mention separately from citations; investigate incorrect business details before creating more content.
A visibility score improves without more inquiries
Check: Did the question set, platform coverage or scoring formula change?
Next: Compare the unchanged sample, then review attributable visits and inquiry quality separately.
Your page is cited, but few visits are recorded
Check: Does the answer satisfy the question without a visit, or is referral information missing?
Next: Keep the citation evidence; do not classify every unmeasured interaction as lost traffic.
A report claims success across every AI platform
Check: Does it provide distinct evidence for each named platform?
Next: Accept only conclusions supported by platform-specific records.
Define useful visibility for your business
Start with the business question: are potential buyers encountering an accurate description of your business when researching a problem you solve? A brand mention means your name appears. A citation means a page is identified as a source, usually with a link. Neither automatically means a recommendation. An answer can mention your company while advising a buyer to choose something else.
I recommend reporting relevance and accuracy alongside presence. A maintenance provider appearing in a buying comparison matters differently from appearing in a historical explanation. Record whether the answer describes the right service, market and limitations. This gives AI search optimization a clearer brief than “increase our score.” If a proposal mixes several labels, the comparison of GEO, AEO and SEO explains the scope distinctions before you evaluate its reporting.
Build a question set you can test repeatedly
Choose questions from sales conversations, customer emails and actual purchase criteria—not only questions likely to elicit a mention of your brand. Keep discovery questions separate from questions that already name you. A hypothetical commercial refrigeration maintenance business might test: “What should a restaurant check before choosing an emergency refrigeration repair service?” and “When is preventive maintenance a better choice than calling for repairs as needed?” These address different buying decisions; neither establishes how often customers search for the topic.
Keep a stable comparison set and a separate exploratory set. For each run, record the exact question, platform, date, language, location, model if disclosed, and context such as login state and previous conversation. Use a fresh conversation for standardized tests; record follow-ups separately. Save the answer and visible source destinations so someone else can inspect the classification. Investigating related questions, sometimes called fan-out research, does not reveal a model’s private internal queries.
- Define a mention rule, including accepted brand spellings and namesakes to exclude.
- Count a sampled answer once for citation presence, even if it links to several of your pages.
- Keep failed runs separate from completed answers; record whether a Google AI answer appeared.
- Label the output “citation frequency in this sample,” not market-wide share of voice or a ranking.
Keep Google data separate
In 2026, Search Console’s Generative AI performance report provides impression data for AI Overviews and AI Mode. It does not report clicks, leads, full prompts, or activity in ChatGPT or the Gemini app, and it cannot establish citation frequency on other platforms. If the report is unavailable in your account, total Web impressions cannot substitute for an AI-only impression count.
Before publishing a dashboard, recheck which features the actual account offers. Save the selected reporting period, filters and export date. My recommendation is to compare equivalent periods with consistent filters, then inspect the pages associated with the change. Annotate site releases and reporting changes rather than treating every movement as a content result. A Google-specific observation belongs in the Google section of the report, not under a heading claiming improvement everywhere.
Measure attributable visits, then check inquiry quality
GA4 measures attributable website visits within its measurement limits, not every AI interaction. Referral information may be unavailable, consent choices affect collection, and someone may return later through another channel. Review session source and landing page together. Keep your rules for classifying AI referrals documented, including when those rules change. Do not relabel direct traffic as AI traffic because it increased during the same period.
For business evaluation, check what those recorded visitors did: reached a relevant service page, submitted a valid inquiry or completed a purchase. A form event is not automatically a qualified lead; confirm that events work and distinguish genuine inquiries from spam or duplicates. An optional “How did you hear about us?” field can add context, but customer recollection is supporting evidence, not precise attribution. Report observed conversions separately from unverified influence.
Use a report with supporting evidence
Use the template below for one reporting period at a time, keeping platform records separate. Attach the underlying evidence before adding a conclusion. A dash means unfilled, not zero. For sample-based percentages, include both the numerator and denominator and explain which completed observations were eligible.
Worked example—hypothetical business: the refrigeration provider appears more often in the repeated buying-question sample, but one answer incorrectly describes its service area. The useful decision is to identify the supporting source and clarify public service-area information, not announce broader customer acquisition. If the question set also expanded, compare the original questions separately. More appearances after adding easier branded questions would not demonstrate improvement in the original sample.
| Record | Evidence to attach | Current / comparable prior | Decision it supports |
|---|---|---|---|
| Google exposure | Report export, dates and filters | — / — | Which pages need investigation? |
| Sampled mentions and citations | Question-set version, run context and saved answers | — / — | Is presence relevant and accurate? |
| Attributable visits | GA4 source rules and landing-page records | — / — | What do recorded visitors encounter? |
| Qualified inquiries | Validated events and inquiry review | — / — | Is there observed commercial value? |
Use findings to decide which content to improve
Inspect the question and cited page before commissioning another article. Is the answer missing a meaningful selection criterion, using outdated information or citing a page that explains the issue better? My recommendation is to fix a relevant existing page when it already serves the buyer’s need. Create a separate resource when the unanswered task is genuinely different. The refresh-versus-new-content guide helps make that distinction.
For the hypothetical maintenance provider, an unclear emergency repair service page may need exclusions, service hours and the information needed to request help—not a general article about refrigeration. Add these findings to your content plan, with supporting evidence, a person responsible and a change log. On my own Supplement Explained project, I manually review AI-assisted drafts. For pages selected for improvement, review factual accuracy and whether the revised content answers the buyer’s question before publishing.
Agree on reporting standards before hiring help
Ask a prospective provider for a sample report, its counting rules and an explanation of missing data. A useful tool should preserve questions, context, source URLs and history so you can check what changed. Google does not require special AI schema or an llms.txt file for inclusion in its AI search features. Any structured data you use must accurately describe the visible page content.
Agree on the next action each report should help you choose: correct inaccurate information, improve a specific page, fix tracking or continue monitoring. Record when each change was published; timing alone does not show that an edit caused a visibility change. Use the guide to how long SEO takes to plan review periods. Before hiring help, agree on the platforms, question set, reporting schedule and implementation responsibilities. No provider can guarantee rankings or citations.
Common questions
Do I need a paid AI visibility tool?
Not necessarily. A spreadsheet and saved answers can support a small, carefully defined sample. Consider a tool when repeat collection and evidence storage become difficult. Test its exports and classifications before relying on its score.
How often should I repeat the questions?
Choose a cadence you can maintain consistently. I recommend a scheduled review plus checks after material changes, rather than reacting to individual answers. Preserve the original question set when exploring new topics.
What counts as meaningful improvement?
A repeatable change in relevant, accurate appearances under comparable conditions is stronger evidence than one favorable answer. Commercial improvement requires separate evidence about visits, inquiries or purchases; visibility alone does not establish it.