How to Measure AI Visibility: A Practical GEO Measurement Method

By Dfeelings · Updated

Short answer

To measure AI visibility, test a fixed, documented set of prompts in each AI assistant you care about, split by service, intent, country and language, and log every answer the same way: whether your brand is mentioned, whether your site is linked, which domains are cited, how prominent you are and whether the description is accurate. Repeat the same prompts on a schedule, and pair the results with Search Console and GA4 data.

Key facts

  • AI answers vary between runs, so one prompt tested once is an anecdote, not a measurement.
  • Google counts AI Overviews and AI Mode appearances within the Web search type of the Search Console Performance report.
  • GA4's default channel group includes an AI Assistant channel for referrals from assistants such as ChatGPT, Gemini and Copilot.
  • OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
  • Measurement shows whether visibility is changing; no one can guarantee that an AI system will mention or cite a business.

What does measuring AI visibility actually mean?

Measuring AI visibility means recording, in a repeatable way, how often and how well AI assistants mention, link to and describe a business when people ask questions it should be relevant to. AI visibility measurement is not a single ranking position. It is a set of rates calculated over a fixed prompt set, tracked per engine, country and language, and compared over time against a dated baseline.

Classic SEO has a stable unit of measurement: a query, a results page and a position. AI assistants do not work that way. The same question can produce a different answer, different sources and different brands from one run to the next, and the answer may be a paragraph, a list or a table. So the unit of measurement becomes the prompt-and-answer pair, and the useful numbers are proportions across many of them.

A good measurement programme answers four questions: are we mentioned, are we linked, are we recommended, and are we described correctly? Each needs its own metric, because a brand can be mentioned often but never linked, or linked as a source while a competitor is the one recommended.

How do you build a prompt set for AI visibility testing?

A prompt set for AI visibility testing is a fixed, written list of questions real buyers would ask, grouped by service, intent, country and language. Each prompt gets an ID and exact wording that never changes between test rounds. A useful set mixes informational, comparison and recommendation prompts, includes brand and non-brand questions, and is written natively in each language rather than translated.

Start from evidence of what people ask: Search Console queries, sales and support questions, and the wording customers use in calls and messages. Then turn those into natural prompts. Group them by service line, by intent (learning, comparing, choosing a provider, checking a specific brand) and by market. For a business working in Saudi Arabia and Jordan, that means separate Arabic and English prompts for each country, phrased the way people there actually write.

Keep the set small enough to test properly. Fifty well-chosen prompts tested consistently are more useful than five hundred tested once. Freeze the wording, version the list, and when you need new prompts, add them as a new version rather than editing old ones, so trends stay comparable.

  • Informational: what a service is and how it works.
  • Comparison: one approach or provider type against another.
  • Recommendation: which providers to consider in a city or country.
  • Brand check: what the assistant says about your business by name.

Which AI engines should be included in AI visibility measurement?

AI visibility measurement should include the assistants your buyers actually use, typically ChatGPT, Google AI Overviews and AI Mode, Gemini, Perplexity and Microsoft Copilot. Each engine is logged separately because they retrieve and cite sources differently. Record the product, mode or model name as displayed, and whether search or browsing was used, because the same engine can answer differently depending on these settings.

Before testing an engine, check that your site can be reached by it. OpenAI's crawler documentation says OAI-SearchBot is used to surface websites in ChatGPT's search features and that sites opted out of it will not be shown in ChatGPT search answers, although they may still appear as navigational links. A robots.txt rule blocking it would make ChatGPT measurements meaningless. Google, for its part, says there are no special requirements to appear in AI Overviews or AI Mode beyond normal Search eligibility.

Control what you can: test from the target country where possible, use a clean session or a consistent logged-out state, and record anything that might influence the answer, such as location, language settings and account personalisation. You cannot remove all variability, but you can stop it changing silently between rounds.

What should be recorded for each AI answer?

Each AI answer should be recorded with the prompt ID, engine, mode, country, language and date tested, plus whether the brand was mentioned, whether a linked citation to the brand's site appeared, every domain cited, the brand's prominence in the answer, whether it was recommended, and whether the description was accurate. Saving the full answer text matters more than a screenshot, because text can be reviewed and re-scored later.

Prominence needs a simple, consistent rule, for example: 1 if the brand is named first or in the opening sentence, 2 if it appears in the main list, 3 if it is mentioned only in passing or in a footnote. Accuracy needs one too: does the description match the business's actual services, locations and facts, or does it contain errors such as a wrong city, a service the business does not offer, or outdated details?

The list of cited domains is often the most actionable field. If assistants repeatedly cite the same directories, publications or competitor guides for your topic, those are the sources shaping the answer. That tells you where accurate information about your business needs to exist, not only on your own site.

Which metrics best describe AI visibility?

The core AI visibility metrics are brand mention rate, linked citation rate, recommendation rate, prompt share of voice and competitor share of voice, supported by an accuracy rate. Each is a proportion calculated over the same prompt set and reported per engine, country and language. Together they show whether a brand is present, sourced, preferred and described correctly, which a single ranking number cannot show.

Mention rate and citation rate answer different questions. A high mention rate with a low citation rate suggests the assistant knows the brand but relies on other sources to describe it. A citation without a recommendation can mean your page is used as background information while a competitor is the one suggested. Share of voice puts your numbers in competitive context: the same mention rate means something very different if the main competitor is mentioned far less often, or far more often.

Report each metric with its denominator and segment. A statement in the form 'mentioned in N of 40 Arabic recommendation prompts in ChatGPT, Saudi Arabia, tested on a stated date' is evidence. 'AI visibility is up' with no denominator, engine or date is not.

Why is a single AI answer screenshot not evidence of visibility?

A single AI answer screenshot is not evidence of visibility because AI assistants can give different answers to the same prompt across runs, users, locations and days. One screenshot shows that something happened once, under unknown conditions. Evidence requires the same prompts tested repeatedly under recorded conditions, with results expressed as rates and compared against a dated baseline taken before the work began.

Variation comes from several places: the model generating the text, the retrieval step choosing different pages, personalisation, location, and product updates. That is why the method runs each prompt more than once per round, for example three runs, and records the share of runs in which the brand appeared. A brand that appears in one of three runs is less established than one that appears in three of three.

Be wary of anyone, including an agency, who presents a hand-picked screenshot as proof of results. Ask for the prompt list, the dates, the number of runs and the rate before and after. Results that cannot be reproduced from a documented method are marketing, not measurement.

How often should AI visibility be re-tested?

AI visibility should be re-tested on a fixed schedule, commonly monthly for a full prompt set, with the same prompts, engines, countries and recording rules every round. Additional checks make sense after major site changes or noticeable product updates from an AI platform. A consistent cadence matters more than a frequent one, because comparability between rounds is what turns individual answers into a trend.

The first round is the baseline, and it should be dated and stored before any optimisation begins. Without a baseline, there is no way to tell whether later results reflect the work or were already true. Each later round is compared to the baseline and to the previous round, segment by segment.

Annotate the timeline. Note when content was published, when structured data changed, when profiles were updated, and when an engine visibly changed its interface. Correlation is not proof of cause, but an annotated timeline makes honest interpretation possible.

Which Search Console and GA4 data supports AI visibility measurement?

Search Console and GA4 supply the traffic side of AI visibility measurement. Search Console's Performance report gives clicks, impressions, click-through rate and average position, and Google says AI Overviews and AI Mode appearances are counted within its Web search type. GA4 reports session source and, in its default channel group, an AI Assistant channel for visits referred by assistants such as ChatGPT, Gemini and Copilot.

In Search Console, the Performance report can be grouped by query, page, country, device, search appearance and date. Compare Arabic and English pages, and Saudi and Jordanian traffic, separately. Google notes that results vary by user, so a query listed in the report may not show your site when you search it yourself, which is one more reason to rely on aggregated data rather than spot checks.

In GA4, the Session source dimension records the site that referred the session, so referrals from an assistant's domain, such as chatgpt.com, can be isolated where the referrer is passed. Google's current default channel definitions also include an AI Assistant channel, while traffic from Google's own AI Overviews and AI Mode falls under Organic Search. Treat these numbers as a floor: not every AI-influenced visit carries a referrer, and channel definitions can change.

AI visibility metrics: definitions, formulas and caveats
MetricDefinitionFormulaCaveat
Brand mention rateHow often the brand is named in answersAnswers naming the brand ÷ answers testedSays nothing about accuracy or whether the mention is positive.
Linked citation rateHow often the answer links to the brand's own siteAnswers with a link to the brand's domain ÷ answers testedDepends on whether the engine shows sources in that mode.
Recommendation rateHow often the brand is suggested as an option to chooseAnswers recommending the brand ÷ recommendation-intent answers testedOnly meaningful on recommendation prompts; keep that subset fixed.
Prompt share of voiceThe brand's share of all tracked brand mentionsBrand mentions ÷ mentions of all tracked brandsChanges when the tracked competitor list changes; freeze the list.
Competitor share of voiceEach competitor's share of tracked brand mentionsCompetitor mentions ÷ mentions of all tracked brandsUntracked brands are invisible to the metric; review the list each quarter.
Description accuracy rateHow often mentions describe the business correctlyAccurate mentions ÷ all brand mentionsRequires a written definition of 'accurate' and a consistent reviewer.

AI visibility log record: fields to capture for every answer

Capture the same fields for every prompt run so rounds can be compared.

  1. Prompt ID and prompt-set version
  2. Exact prompt text as entered
  3. Language (Arabic or English) and country tested from
  4. Engine, product mode or model as displayed, and whether search was used
  5. Session state: logged out, logged in, personalisation on or off
  6. Date and time tested, and run number within the round
  7. Brand mentioned (yes or no) and prominence score
  8. Linked citation to the brand's site (yes or no) and the URL cited
  9. All cited domains, in the order shown
  10. Recommended as an option (yes or no) and competitors named
  11. Description accurate (yes or no) with notes on any error
  12. Full answer text saved, with screenshot as a supplement only

Frequently asked questions

Can AI visibility be measured with a single tool?

No single tool covers AI visibility completely. Prompt testing shows what assistants say, Search Console shows Google search performance including AI features, and GA4 shows referred visits and what visitors do next. Third-party trackers can automate prompt testing, but their method, prompt list and locations should be documented and checked like any manual process.

How many prompts are enough for a baseline?

A useful baseline usually needs enough prompts to cover each service, intent, country and language combination that matters, with several runs per prompt. For a small business that may be a few dozen prompts; for a multi-service, bilingual business, more. Consistency and coverage matter more than raw volume.

Does Search Console show AI Overview data separately?

Google's documentation says sites appearing in AI Overviews and AI Mode are counted in overall Search traffic in Search Console, within the Web search type of the Performance report. Plan to analyse that traffic as part of web search performance, and use prompt testing for answer-level detail.

Why do I see my brand in ChatGPT but my colleague does not?

AI assistants can produce different answers for different users, sessions, locations and times, and personalisation or memory settings may influence results. That is why one person's screenshot is not a measurement. A documented prompt set tested under recorded conditions shows how often the brand actually appears.

Can measurement prove that GEO work caused an improvement?

Measurement can show that visibility changed after specific work, especially with a dated baseline and an annotated timeline, but it cannot prove sole cause, because AI platforms and competitors change too. Honest reporting presents the before-and-after rates, what changed, and the other factors that may have contributed.

How Dfeelings measures AI visibility

Measurement is stage 07 Measure in the Dfeelings GEO Framework, and the baseline is taken during 01 Discover before optimisation starts. Dfeelings uses a fixed, documented prompt set tested per engine, per country (Saudi Arabia and Jordan) and per language (Arabic and English), recording brand mention, linked citation, cited domains, prominence and date tested, then re-tests on a schedule with the same prompts. It pairs that log with Google Search Console and GA4 data, and it does not promise that any AI system will mention or cite a client.

Official sources