OBVS / Our Thoughts / AI search

What AI Visibility Tools Actually Measure

Evidence reviewed September 9, 2026

AI visibility tools can capture real answers. Understanding what those answers tell us about an audience takes a little more care.

A dashboard says your AI visibility improved. The line is green. A competitor has slipped behind you. Somewhere underneath, an answer mentioned your business.

That last part is something we can investigate.

Open the saved response. Read the question. Check whether the business was recommended, cited as a source, or included in a warning about what to avoid. Those are very different appearances, although each might begin life as a detected brand mention.

The harder question is what that appearance tells us about the people considering a purchase.

Tools such as Semrush’s AI Visibility Toolkit can make this investigation easier. The problem begins when a measurement of collected answers gets interpreted as a measurement of everyone using AI search. A real observation can sit underneath an uncertain estimate. Both deserve labels.

The database is large. The audience is still a question.

Semrush says its database contains more than 317 million prompts and responses. Its documentation describes AI-search clickstream inputs and Google keyword data, with duplicates removed and wording simplified into topics. It says intent is preserved. These are vendor claims, rather than an independently audited census of AI usage. Semrush’s data methodology

Scale matters. A broad collection can reveal questions, competing brands and cited pages that a person checking a few prompts would miss.

But a large sample can still systematically miss part of the population. To judge representativeness, we need to know who contributes data, which interactions are observable, and how the sample is adjusted. Geographic availability alone does not establish balanced coverage across demographics, languages, platforms or industries.

The public material reviewed does not provide enough information to independently establish those balances, or the probability that a particular customer’s question enters the database. Nor does it establish adequate coverage for every smaller or regional brand. That is an unresolved measurement question, not evidence that such brands are necessarily excluded.

Buying intent needs similar care. Semrush’s Prompt Research distinguishes commercial and transactional questions from informational and other intents. That is useful organization, but a classifier’s label does not establish how frequently actual buyers ask a question. Prompt Research documentation

For a high-consideration purchase, “best equipment” and “which equipment can we service locally without stopping production?” may lead to different suppliers. Tidying language can help organize research. Removing a consequential condition can change the research question.

Even faithful paraphrases create a measurement choice. Count every variation equally and a heavily paraphrased topic can dominate. Collapse them all and you may hide sensitivity to wording. A defensible report should disclose both the topic grouping and the questions underneath it.

What the numbers mean

There are four levels of evidence here.

An observation establishes that a saved answer contained something under recorded conditions. Repeated observations can establish a pattern within a defined test panel. A modeled estimate extrapolates or interprets those observations. A proprietary index adds a calculation an outsider may be unable to reproduce.

These levels can coexist in a useful product. They should not be interchangeable in a sales presentation.

Semrush describes its visibility score as combining topic coverage and mention consistency. Topic difficulty considers competitor strength and available opportunities; topic volume combines interaction data with machine learning. The published ingredients do not supply a complete reproducible calculation. Metric methodology

A buyer would still need the weights, normalization, topic boundaries and comparison population to reconstruct the score. The description also refers to prominence without separately explaining a position component in that two-factor account. This is incomplete disclosure, rather than proof of a mathematical contradiction.

The audience metric deserves particular restraint. Semrush defines monthly audience around the estimated query audience of topics where a brand appears. That is not an observed count of distinct people who saw the brand. Its Overview report separately identifies mentions, responses citing the domain, and individual cited pages. Those underlying records are the more directly checkable evidence. Visibility Overview definitions

Share of voice is also more complicated than its familiar name suggests. Semrush illustrates a simple share of mentions, then says Brand Performance incorporates mention position. Its Enterprise calculation also includes topic volume for ChatGPT. The simple equation therefore does not fully reproduce the product calculations; these are distinct measures, not automatically comparable percentages. Semrush’s share-of-voice explanation

Competitor selection matters too. A comparison against three chosen rivals answers a different question from a comparison against every brand in the sample. Semrush’s competitor report offers useful prompt and source gaps, but an absence is only commercially interesting when the question is relevant to the business. Competitor Research documentation

Sentiment adds interpretation. Semrush’s Perception report uses categorized sentiment on non-branded queries. A sentence such as “expensive, but worthwhile when reliability matters” resists a tidy positive-or-negative label. Keep the sentence available and inspect consequential classifications. Brand Performance documentation

There are documentation inconsistencies worth resolving. Semrush’s data page describes non-API collection for its prompt database, while its broader FAQ describes both API and interface collection. The latter also calls the data directional. Different product scopes could explain the difference, but the mapping is unclear. A customer should be told which collection method produced the report being purchased. Data page, Semrush FAQ

Other vendors demonstrate why the category needs individual scrutiny. Peec publishes a straightforward visibility calculation: responses mentioning the brand divided by total responses. That is reproducible with the underlying records. It still describes Peec’s tested answers; extending it to population exposure requires evidence about the sample. Peec’s visibility definition

The question does not stay still after you submit it

OpenAI says ChatGPT search may rewrite a question into several targeted searches, then issue further queries after reviewing results. Location and saved memory can affect that rewriting. Earlier conversation also changes what a follow-up question means. A clean session asking “Which would you choose?” cannot reproduce a conversation that already established budget and requirements. ChatGPT search documentation

Google describes query fan-out: related searches across subtopics and sources used to construct a response. It also says AI Mode and AI Overviews can use different models and techniques, producing different responses and supporting links. Google Search Central

The measurement target therefore includes the wording, timing, interface, platform version, location, account context and retrieval behavior. Repeating the visible question does not guarantee repeating everything that happened behind it.

Research supports taking that variability seriously. Ronald Sielinski’s 2026 preprint found that apparent citation differences could fall within measurement noise. Its scope was limited: three product topics, generated queries and a short observation period, with API-based collection. It supports repeated sampling and uncertainty reporting; it does not establish a universal error rate for consumer AI search. Study and limitations

Nor does a citation reveal exactly how a page influenced the answer. A 2026 preprint by Zhang Kai and colleagues separates source selection from apparent content absorption. Similarity and factual overlap can help investigate that distinction, but do not prove causal influence. Earlier research by Liu, Zhang and Liang evaluated whether citations actually supported answer claims—a separate issue again. Its 2023 results should not be presented as current platform accuracy. Citation selection research, citation-verifiability research

A smaller group of agents, with a better notebook

OpenAI’s September 2026 Navier–Stokes announcement describes roughly 10,000 concurrent agents in the group associated with its proposed resolution. Groups explored variants and approaches, shared findings, and received consolidated insights. That is OpenAI’s account of the project, not an independent endorsement here of its mathematical claim. OpenAI’s project explanation

The useful analogy is coordination. A much smaller group could divide the work of collecting answers, validating evidence and comparing results. Monitoring brands does not require anything approaching that scale, and mathematical problem-solving does not prove that a marketing sample represents its audience.

Parts of this system already exist. OtterlyAI documents daily prompt monitoring with response-level detail. Peec also runs selected prompts daily. Semrush’s Prompt Research connects discovered questions to daily tracking. Continuous collection is already a product feature. OtterlyAI, Peec, Semrush

What I would want is a disclosed panel grounded in customer questions, with controlled paraphrases reviewed for unchanged intent. Keep a stable core for comparisons and a separately labeled exploratory panel for new questions. Do not quietly replace difficult prompts with ones that make the graph look healthier.

A scheduler would repeat tests across permitted platforms and locations. Collectors would preserve complete answers and conditions. Separate validation agents would confirm each detected mention and citation against that saved evidence, with human review for ambiguous recommendations or factual errors. Straightforward counting should use deterministic rules where possible.

The report would show which results repeat, which depend on wording, and which differ by geography. Failed requests would remain visible rather than becoming apparent brand absences. Every reported change would link back to the answers behind it.

Confidence ranges would help, but they need honest boundaries. Repeated answers from the same question family are related observations, not independent customers. Analysis should account for grouping by intent and time. A narrow interval around a poorly chosen panel can still give a confidently misleading view of the market.

The difficult part is operating the measurement

AI answers are probabilistic, but randomness is only part of the problem. Developer APIs and consumer interfaces may differ in models, instructions, search behavior and personalization. Treat them as separate measurement channels until a matched comparison justifies combining them.

Automated access also has constraints. OpenAI’s consumer terms restrict automated extraction and bypassing limits. A monitoring design needs permitted access, appropriate agreements and respect for platform controls; an API is not permission to automate the consumer interface. OpenAI Terms of Use

Location testing is imperfect. Naming a city in a prompt is different from querying from that city. Clean sessions cannot reproduce every personal history. Model and retrieval changes may arrive without sufficient detail to explain a sudden shift. Log what is known and mark what is unknown.

Costs multiply with every added condition. An illustrative panel of 50 questions, three wordings, four platforms, two locations and three repeats requires 3,600 responses per cycle. That is arithmetic for a proposed design, not a recommended minimum. Collection, validation and storage all need a budget.

The business evidence should sit alongside the monitoring results. Google now documents a dedicated generative-AI performance report covering AI Overviews and AI Mode impressions, with page, country, device and date views. Its help page says worldwide rollout completed on August 31, 2026. These are platform-reported impressions under Google’s counting rules, not a panel’s predicted exposure. Search Console report documentation

Connect that evidence with identifiable analytics referrals, qualified leads and conversions. The current dedicated report does not document query or click metrics, so it cannot supply a complete prompt-to-purchase trail. Nor should a rise in monitored mentions be declared the cause of a rise in sales. Attribution gaps and other marketing activity remain.

I would use these tools to find incorrect descriptions, recurring shortlists, relevant sources and unanswered customer questions. Those findings can change what a business publishes or corrects. A combined index might help summarize them, provided its formula, weights, panel, sample size, cadence and limitations remain visible.

Before acting on a visibility score, open the answers. A useful report should make that the beginning of the investigation.

A practical design

Build a record you can check.

This is a conceptual monitoring system, not a deployed OBVS product or a validated audience model.

Define the panel
Keep a dated set of customer questions. Separate branded checks from unbranded buying questions. Review paraphrases for unchanged intent.
Repeat the tests
Schedule permitted platform and location tests. Record the exact wording, date, interface, model when known, session conditions and collection failures.
Preserve the evidence
Save complete answers, citations, page URLs, competing brands and source order. Separate validators must point to the saved mention or link.
Report the variation
Show answer-level frequencies, wording sensitivity, geographic differences, source overlap and uncertainty. Keep the denominator and testing period visible.
Check business value
Compare the observations with available AI impressions, identifiable referrals, qualified leads and conversions. Keep attribution limits explicit.
What should count as a brand mention?

Count an answer once when it names the correct business, using a documented rule for aliases and ambiguous names. Keep a recommendation separate from a neutral mention or criticism.

What should count as a citation?

Count an answer once when it contains a citation to an owned domain. Retain individual page URLs and distinguish inline citations from other links in a source panel. A citation does not by itself establish causal influence.

When does a change deserve attention?

Look for a sustained, commercially relevant change under comparable conditions. Check failed runs, panel edits and platform changes before calling it an improvement or decline. Serious factual errors deserve review even when infrequent.

Sources & editorial approach

Primary documentation and original research checked September 9, 2026. Vendor capabilities and database figures are attributed claims; proprietary scoring accuracy was not independently audited. Research findings retain their study-specific limitations. Examples and the monitoring design are illustrative. This article is not an affiliate comparison.