Agentic SEO Tools
All posts
By Agentic SEO Tools Teamshare of voicemeasurementbrand mentionsGEOmethodology

AI Search Share of Voice: Measure Brand Mentions in Generative AI Answers for Visibility

How to define and calculate share of voice in AI answers: the unit, the formula, the sampling traps, and how to turn the number into work an agent can actually do.

Share of voice is the monitor number in an AI search program. It tells you how much of the conversation your brand owns compared with the brands you compete against. It does not tell you why, and it does not fix anything. We still think it is worth measuring carefully, because a sloppy share of voice figure sends every later step of the loop (diagnose, act, verify) after the wrong problem.

This guide covers the method first. The tooling comes at the end.

Decide what you are counting

Three different events get called "share of voice" in AI search, and vendors rarely say which one they mean.

A mention is the brand name appearing in the answer text. A citation is a link or source card pointing at your domain. A recommendation is the answer actually telling the reader to pick you, often with a ranked position. These move independently. A brand can be mentioned in every answer and cited in none, because the engine learned about it from a review site. Another brand can be cited constantly as a source of definitions and never recommended.

Pick one unit as your headline and report the other two beside it. For most commercial programs we would lead with mention share and track citation share as the diagnostic. If you sell through comparisons, position matters enough to weight it.

The basic calculation

Mention share across a fixed prompt set works like this:

your brand mentions ÷ total mentions of every brand in your competitor set

Run it over the same prompts, on the same engine, in the same period. A hypothetical example to show the arithmetic: you track 40 prompts on ChatGPT, and across one week's runs your brand is mentioned 30 times while you and four competitors are mentioned 150 times in total. Your mention share is 20%. Those numbers are made up to illustrate the formula, not a benchmark.

Two related figures are worth computing from the same data. Presence rate is the share of answers that mention you at all, which is easier for executives to read. Citation share is your cited URLs divided by all cited URLs across the set, and it points at where engines get their information.

What skews the number

Most bad share of voice reports fail in one of these ways.

The most common is blending engines. ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews pull from different sources and answer differently, so one blended percentage can hide a strong ChatGPT position behind a weak Perplexity one. Report per engine first, then roll up if you must.

Next comes a biased prompt set. If half your prompts include your brand name, you will look dominant. Keep branded prompts in a separate bucket. The share of voice that matters comes from unbranded, category-level questions a buyer would ask before they know you exist.

Sampling variance catches people too. The same prompt can return a different answer an hour later, which makes one run per prompt an anecdote. Repeat runs on a steady cadence and read the trend rather than a single week's figure.

Then there is who is asking. Answers change with location and with the context of the person asking, and a share of voice measured from one country says little about another market. If you sell regionally, measure regionally.

Last, a raw tally treats every mention as good. A mention that calls you "the expensive option with weak support" counts the same as a recommendation. Pair share of voice with sentiment, or you will celebrate a number that is quietly hurting you.

A measurement protocol you can defend

Write the protocol down before you look at results, so nobody tunes it to flatter the brand.

  1. Fix the competitor set. Four to six named rivals is enough for most categories.
  2. Build the prompt set from real buyer questions, tagged by topic and intent, branded prompts separated.
  3. Choose engines based on where your buyers are, and report each one separately.
  4. Set locations and personas if they change who you compete with.
  5. Run on a fixed cadence and keep the history. Changing the prompt set resets your baseline, so log every change.
  6. Report mention share as the headline, with citation share, average position, and sentiment underneath.

From the number to the work

A share of voice drop is a symptom. The loop only closes when the number triggers a diagnosis and an action.

Diagnose by looking at citations on the prompts where you lost ground. If the answers that left you now cite a competitor's comparison page, a Reddit thread, or a YouTube review, that is the source you need to answer or appear on. Act by publishing the page that fills the gap, or by fixing the existing page the engine misreads. Verify by watching the same prompts over the next runs, and by checking that the engine's crawler actually fetched the new page before you call the experiment a failure.

That is a lot of manual joining if your share of voice lives in one tool, your citations in a spreadsheet, and your crawl data nowhere.

Measuring it in Promptwatch

Promptwatch is where we would run this protocol, because each step above maps to a feature in the same workspace. Share of voice and competitive benchmarking give you the headline per engine. Prompt tracking supports topics, tags, personas, and country, state, or city targeting, so the protocol's segments exist as settings, not spreadsheet tabs. Search volumes and difficulty scores help you weight the prompt set toward questions people actually ask. Prompt trends show how each prompt moved between checks and what changed.

For diagnosis, citation analytics break sources down by page, domain, Reddit, YouTube, and offsite mentions, with citation trends over time. For the act step, content gap analysis and Content Agents can plan and draft the missing page and publish it to Webflow or Framer after it clears a review inbox (WordPress is listed as coming soon). Unified Actions turns the findings into a to-do list, and Agent Chat lets you ask the data plain questions, such as which prompts a named competitor gained this month. Agent Analytics crawler logs, on Professional and above, cover verification.

Monitoring runs against the real product interfaces of ChatGPT, Gemini, Claude, Perplexity, Grok, Llama, DeepSeek, Mistral, and Copilot, plus Google AI Overviews and AI Mode. Essential costs $95/mo for 50 prompts. Professional, at $245/mo, gives you 150 prompts and the crawler logs. You can start with the free Explore tier (10 prompts, ChatGPT only) to test your prompt set, then move to a paid plan once you need multiple engines. Pricing and trials are on promptwatch.com.

If you want to compare tools before committing to one, our ranked directory scores each platform on how much of the monitor, diagnose, act, verify loop it closes.