As people ask AI assistants instead of searching, the new question is whether the AI mentions and recommends you. Here is a practical way to measure AI search visibility with a prompt set, a scorecard, and honest limits.
AI Search
Not long ago, the question a founder asked me was simple: where do we rank on Google for our money keywords? These days the question that keeps people up at night is quieter and stranger. When a prospective buyer opens an AI assistant and types “who is the best studio to build my D2C storefront in India”, does the answer include us, and does it say something true? At Identiti, we build for the web and we automate the boring parts of running a business, so this shift lands right in our lap. The hard truth is that there is no ranking report for AI answers, the responses change from one day to the next, and attribution is fuzzy at best. That does not make the problem go away. It makes measurement a discipline you have to build yourself, imperfectly and on purpose. This is the practical approach we use, warts and all.
The question has quietly changed
For twenty years, visibility meant a position on a page of blue links. You could pull a rank tracker, see you were number four for a phrase, and reason about clicks from there. That world still exists, but a growing slice of buying research now happens inside a conversation with an assistant. Someone describes their problem in plain language, the assistant summarises options, and often names a handful of providers. Your prospect may never see a results page at all. The visibility that matters is whether you are one of the names that comes up, and whether the sentence around your name helps or hurts you.
This is not a small tweak to the old game. It is a different surface with different physics. A search result is a list you can inspect. An assistant answer is a synthesis you have to interrogate, because it is generated fresh each time and it blends many sources into a single confident paragraph. I have written before about how the fundamentals differ in the split between answer engines and classic search, and the measurement problem is the sharp end of that split. You cannot manage what you cannot see, and right now most brands genuinely cannot see whether the machines that increasingly mediate buying decisions know they exist.
So the job is to make the invisible visible. Not perfectly. Not with a tidy dashboard handed to you by a platform. You build a rough instrument, you point it at the assistants your buyers actually use, and you read the needle. The rest of this piece is how we build that instrument and what we do with the readings.
Why this is genuinely hard to measure
Let me be honest about the obstacles before I sell you the solution, because pretending they are not there is how people end up with numbers they should not trust. The first obstacle is non-determinism. Ask the same assistant the same question twice and you can get two different answers, with different providers named and different framing. There is no single true response to record. There is only a distribution of responses, which means one query tells you almost nothing and you need repetition to see a pattern.
The second obstacle is the absence of a ranking report. Nobody hands you a scoreboard that says you appear in thirty percent of relevant answers. You have to generate that figure yourself by asking, recording, and counting. The third obstacle is attribution. When a prospect who was influenced by an assistant later lands on your site, they frequently arrive with no clean referral trail. They may have read a recommendation in the assistant, then typed your name into a browser, then clicked an organic link. The credit for that visit gets smeared across channels, and the assistant that actually did the persuading is often invisible in your analytics.
The fourth obstacle is that answers are opaque about sourcing. Sometimes an assistant links to pages it drew from, and sometimes it just asserts. When it does cite, the citation may or may not reflect what truly shaped the recommendation. Put these four together and it is tempting to throw up your hands and say measurement is impossible. It is not impossible. It is imprecise. Those are very different things, and treating an imprecise signal as if it were no signal is a mistake that costs you the whole opportunity.
Start with a prompt set your buyers would actually type
Everything begins with a set of representative prompts. This is the single most important input, and most people get it wrong by writing prompts that flatter them rather than prompts their buyers would use. The goal is to build a fixed list of questions, phrased the way a real prospect phrases things, that would naturally surface a provider like you. You then ask these same questions on a regular schedule so your readings are comparable over time.
Build the list in tiers. First, the category questions, where nobody has named you yet: “who builds high converting Shopify stores for Indian brands”, “which studios do design and automation together”, “best agency to fix a website that gets traffic but no sales”. These are the ones that decide whether you get discovered at all. Second, the comparison questions, where the buyer is weighing named options against each other. Third, the branded questions, where your name is already in the prompt: “is Identiti Design any good”, “what does Identiti Design do”, “who are the founders”. Branded prompts test accuracy rather than discovery, and accuracy is its own battle because assistants will confidently state wrong things about you.
Keep the set deliberately small and stable, somewhere in the range of twenty to forty prompts, because you will run it repeatedly and you want to compare like with like. Write each prompt the way a slightly frustrated buyer would speak it, not the way a marketer would. If you are unsure what those phrasings are, the best source is your own inbox and sales calls, because the language your prospects use is the same language they type into an assistant. This is also where a clear, well structured public identity pays off, and I have argued the case for building a brand the machines can actually parse because a fuzzy brand produces fuzzy answers no matter how good your work is.
What to record every time you ask
Once you have the prompt set, you need a consistent record of what came back. Do not just eyeball it and form a vibe. Capture the same fields every time, for every prompt, so you can turn a pile of conversations into something countable. I keep it to a handful of fields because a scorecard nobody fills in is worthless, and the temptation to over engineer this is strong.
Record, at minimum: whether you were mentioned at all, a simple yes or no. Then your position or prominence, meaning were you named first, buried in a long list, or offered as the lead recommendation. Then sentiment, meaning was the surrounding language positive, neutral, or negative. Then accuracy, meaning was what the assistant said about you true, and if not, what it got wrong. Then citation, meaning did the answer link to or reference one of your pages, and if so which one. Those five fields, gathered across many runs, give you almost everything you need.
From these raw fields you derive the metrics that matter. Share of voice is how often you appear across the relevant prompt set, and how that compares to the competitors who keep showing up beside you. Sentiment and accuracy together tell you whether being mentioned is helping or quietly hurting, because a confident wrong statement about your pricing or your services can do real damage. Citation frequency tells you which of your pages the assistants trust enough to quote, which is gold, because those are the pages to strengthen and the template for the pages to build next. Keep the recording boring and repeatable. The insight comes from the pattern across many boring records, not from any single dramatic answer.
Reading the analytics side
Prompt logging tells you what the assistants say. Your analytics tell you what happens next, and you need both because neither is complete alone. Start by watching for AI referral traffic. Some assistant answers include links, and when a prospect clicks through, that visit can show up in your analytics with a referrer that identifies the assistant. Set up your analytics so these sources are grouped and visible rather than lumped into a generic bucket. It will not capture everyone, because plenty of assistant influenced visits arrive with no referral trail, but the slice you can see is a real, directional signal that is worth trending over time.
The more revealing signal is often branded search lift. When assistants start recommending you more, a portion of the people who hear your name do not click a link at all. They go and look you up directly, by typing your name into a search engine or straight into the address bar. So a rising tide of branded searches and direct visits, especially when it is not explained by a campaign or a press hit, is one of the better proxies you have for growing assistant visibility. It is indirect, and you have to rule out other causes, but combined with your prompt logs it turns a hunch into a defensible read.
Tie these back to outcomes wherever you can. If your assistant visibility and branded interest are climbing while your enquiries and their quality are flat, something downstream is broken and you should look at your site, not your visibility. This is where measurement earns its keep, because it stops you from optimising a metric that does not convert. Traffic that lands and bounces is not a win, and the design and clarity that turn a curious visitor into an enquiry are their own discipline, which is why I treat conversion focused web design as inseparable from any visibility effort. Getting mentioned is the start of the sentence, not the end of it.
Building a simple visibility scorecard
Now pull it together into one artefact you actually look at: a visibility scorecard. The point of a scorecard is not precision, it is repeatability and honesty. You want a small, stable set of numbers you refresh on a schedule so you can see movement, not a sprawling dashboard that impresses in a meeting and gets ignored the rest of the month. Simplicity is the feature, not a compromise.
I structure ours around five components, each scored simply. Share of voice: across the prompt set, in what fraction of relevant answers did we appear. Prominence: when we appeared, were we lead, listed, or an afterthought, averaged into a rough score. Sentiment and accuracy: were the mentions positive and correct, flagging any recurring falsehoods for urgent fixing. Citation: how often, and which pages. Downstream signal: the trend in AI referral visits and branded search, marked simply as up, flat, or down. Roll those into a single visibility score if you like a headline number, but keep the components visible underneath, because the components are what tell you where to act.
Sample on a cadence that matches how fast your world moves and how much patience you have. For most small and mid sized businesses, a monthly full run of the prompt set is a sensible rhythm, with a lighter weekly spot check on your most important five or ten prompts to catch anything alarming early. Run each prompt several times per session, not once, because of the non-determinism I keep harping on. One answer is an anecdote. Five answers to the same prompt start to look like a measurement. Log the date every time, because the whole value of a scorecard is the comparison between this month and last.
Let me make this concrete with an illustrative example, using made up but realistic entries so you can see the shape of the thing. Imagine a monthly run of thirty prompts across an assistant or two, each prompt asked five times, and the results tallied into a small table. This is not real Identiti data, it is a template you can copy.
Share of voice: appeared in 12 of 30 relevant prompt sets, so a share of voice of roughly two in five, up from about one in four last month. Prominence: of those twelve appearances, lead recommendation in three, listed among others in seven, mentioned in passing in two, giving a prominence read of medium and improving. Sentiment and accuracy: eleven of twelve mentions positive or neutral, but two answers wrongly claimed we only do logo design, which is a recurring factual error worth fixing at the source. Citation: our services page was linked in four answers and a case study in two, while the rest cited nothing, telling us those two pages are doing the heavy lifting and deserve reinforcement. Downstream signal: AI referral visits trending up modestly, branded searches up, direct enquiries mentioning “found you through an assistant” appearing for the first time.
Written as a compact card, one month might read like this. Visibility score: 6 out of 10, up from 4. Share of voice: 40 percent, rising. Prominence: medium. Sentiment: positive. Accuracy: one recurring error to fix, the “logo only” misconception. Top cited pages: services, flagship case study. Downstream: referral and branded interest both up. Action for next month: correct the services misconception across our own pages and profiles, earn two more third party mentions in places assistants trust, and strengthen the case study that already gets quoted. That is the entire discipline in one card. Nothing exotic, nothing that needs a platform subscription, just a fixed prompt set, honest logging, and the willingness to look every month.
What to do with what you find
A scorecard that does not change your behaviour is a hobby, so let me be concrete about the moves each finding should trigger. When you find wrong facts, fix them at the source first. Assistants synthesise from your public footprint, so correct your own site, your profiles, your listings, and anywhere else that describes you, making the true version clear, consistent, and easy to parse. You will not get an instant correction, because these systems update on their own schedule, but you are changing the raw material they draw from, and that is the only lever you actually control.
When your share of voice is thin, the work is to earn more mentions in the places assistants appear to trust. That means being present and well described across the credible corners of your category, being quoted and referenced by others, and generally increasing the number of independent signals that point to you as a real answer to a real question. This is unglamorous, cumulative work, and it overlaps heavily with the broader playbook I laid out in the tactics that actually move answer visibility. There is no shortcut that survives contact with reality here. You become the answer by being, verifiably and repeatedly, a good answer.
When certain pages keep getting cited, treat that as a map and strengthen exactly those pages. Make them clearer, more complete, more quotable, and more obviously authoritative, because they are already the parts of you the machines reach for. Then build more pages in the same mould. And when you notice which questions buyers are actually asking assistants, feed that back into how you talk to prospects directly, because those questions are a free, honest research stream about what your market wants to know. The same instinct sits behind asking buyers for their own words and preferences and using them, because the language your market volunteers is the language that wins both the assistant answer and the sale.
The honest limitations, and doing it anyway
I want to close where I started, on honesty, because this field is full of people selling certainty they do not have. Your measurement will be imperfect. Non-determinism means your share of voice figure has noise in it, so treat it as a trend and not a decimal you defend to the death. There is no perfect attribution, so your downstream numbers are proxies you triangulate, not proof you can present as a clean chain of cause and effect. Any tool or method that promises exact, deterministic AI visibility scoring is overselling, because the underlying systems do not work that way, and a number that pretends otherwise is worse than a rough number you understand.
Sampling has limits too. You are testing a fixed prompt set on a schedule, which means you see a slice, not the whole. That is fine, as long as you remember it is a slice and you resist the urge to over interpret a single dramatic answer. Keep the prompt set stable so your comparisons hold, refresh it occasionally as your market’s language shifts, and always run each prompt several times before you believe what it tells you. The discipline is in the repetition and the honesty about error bars, not in any single reading.
And yet you do it anyway, because the alternative is flying blind while an increasingly large share of your buyers form their first impression of you inside a conversation you never witness. You cannot improve what you never measure, even imperfectly. A rough instrument pointed at the right question beats a perfect instrument pointed at the old one. Build the prompt set, log the answers, watch the analytics, keep the scorecard, and act on what it tells you. The brands that start measuring this now, while it is awkward and imprecise, will understand their own visibility long before their competitors even realise the question has changed.
A short checklist to start this week
If you want to begin without ceremony, here is the whole thing in prose. Write down twenty to forty prompts your buyers would actually type, in their words not yours, spanning category questions, comparison questions, and branded questions about you specifically. Pick the one or two assistants your market genuinely uses and commit to a monthly full run with a lighter weekly check on your most important handful of prompts. Ask each prompt several times per session and record five things every time: were you mentioned, how prominently, with what sentiment, with what accuracy, and with what citation if any.
Turn those records into a small scorecard with share of voice, prominence, sentiment and accuracy, citation, and a downstream signal drawn from AI referral traffic and branded search lift. Give it a headline visibility score if you like, but keep the components visible because they tell you where to act. Then do the follow through: fix wrong facts at the source, earn more mentions in trusted places, and strengthen the pages that already get quoted while building more like them. Keep the prompt set stable so your months are comparable, accept that the numbers are rough, and treat every reading as a trend rather than a verdict. Do that consistently for a few months and you will know something almost none of your competitors know, which is whether the machines that now sit between you and your buyers are recommending you, and whether they are getting you right.