We built our own Rufus tracker because the numbers we were being shown did not survive contact with the data. Here are the measurement decisions that changed our reporting, why each one moves the number down rather than up, and the questions to ask any vendor selling you AI visibility.
Teardown: what it actually takes to measure Amazon Rufus honestly
By Stephen Honight, Founder of Lmo7
Most AI visibility numbers are wrong in the same direction. They flatter.
That is not usually dishonesty. It is what happens when you build a dashboard before you have settled what a measurement is. Every unresolved definition has a comfortable default and an uncomfortable one, and the comfortable one ships. Enough of those stack up and you get a chart that goes up while nothing in the business moves.
We ran into this when we started tracking Amazon Rufus for clients. The tooling on the market either did not cover Rufus at all or covered it in a way we could not interrogate. So we built our own tracker, and then spent considerably longer than expected arguing with ourselves about what the data actually meant. This is a teardown of that argument, not of the build. The engineering is ours. The measurement thinking should be everybody’s, because right now the category is full of numbers nobody can reconcile.
Every decision below moved our reported numbers down. That is the tell. If a measurement change makes your client look better, check it twice.
If you are short on time, the vendor questions near the end are the usable half of this piece. The short version of the argument: verify which surface you are actually reading before you call a market measurable, never let a refusal count as a brand absence, keep missing data out of your zeros, treat a single read as an anecdote rather than a trend, and check whose denominator sits under every share number. The rest of this post is the reasoning behind each of those.
Rufus on desktop and Rufus in the app are two different products
The first thing that broke was geography.
We wanted to extend tracking into two new European markets. Rufus was live in both. We could see it working. Queries were written, loaded and ready to collect. Then we checked the surface our tracker actually reads, which is the desktop website, and Rufus was not there. It existed in the mobile app and nowhere else on those domains.
Two obvious options. Ship anyway and report thin data, or turn the market on and quietly let the numbers stand for something they do not describe. We took a third: stage the queries, leave the market paused, and tell the client the market is not measurable yet.
The general point matters more than the specific one. AI shopping assistants roll out per surface, per device, per market and per account. “Rufus is live in France” is not a fact about your reporting until you know which Rufus, on which surface, and whether the thing generating your data can see it. If a vendor tells you they track an assistant in a market, ask them which surface they read and how they verified it. If the answer is a press release rather than a screenshot, treat the market as unmeasured.
A refusal is not an absence
This is the change that cost us the most and mattered the most.
Ask an AI assistant a shopping question and you do not always get an answer. Sometimes it declines. Sometimes it replies with a clarifying question instead of a recommendation. Sometimes the read completes and returns nothing usable. Our early build counted all of those as completed reads, which meant they landed in the denominator of every presence and share calculation.
Think about what that does. Your brand was not mentioned, because nothing was mentioned, because the assistant did not answer. The maths reads that as your brand being absent from a shopping answer. It is not. It is the absence of a shopping answer.
We rebuilt the scoring so that only answered reads count in the numerator and the denominator both. A non-answer is now a capture outcome, tracked separately as a data quality signal, and excluded from anything describing brand performance.
Two consequences worth being blunt about. Presence rates moved, because the denominator shrank. And because the score recomputes from the underlying data rather than being stored, the whole history moved with it. Any number we published before that change will not reconcile with the dashboard today. We say that on the methodology page rather than hoping nobody checks an old deck against a new one.
My sense is that this is the single most common unexamined error in AI visibility reporting. Refusal rates on shopping questions are not trivial, they vary by category, and folding them into your denominator quietly penalises brands in the categories where assistants are most cautious. Health, supplements and anything with a regulatory edge get hit hardest, which is exactly where clients are most anxious about the numbers.
Missing is not zero, and neither is unranked
A dashboard has to render something in every cell. The lazy answer is zero, and zero is a claim.
We hold three states apart. A market we do not track reads “not tracked”. A brand that a third-party provider ranks outside its reported top ten reads “outside top ten”. A brand that was genuinely measured and genuinely did not appear reads 0%. Those look similar on a screen and mean completely different things, and collapsing them is how a client ends up believing a competitor has vanished when the truth is you stopped looking.
The same discipline applies to sample size. Every percentage in our reporting carries its numerator and denominator. “44%” is not a number. “44% (4 of 9)” is a number, and it tells the reader immediately that they should not build a quarter’s content plan on it. We surface the pair everywhere, including in chart tooltips, because a percentage without its fraction is the easiest place in a dashboard to hide a thin week.
One read is not a measurement
Ask an AI assistant the same shopping question twice in an hour and you can get two different answers, different brands, different order. That is not a bug in the assistant. It is the nature of the thing.
Which means a single query run is an anecdote. If your tracker asks each question once a week and plots the result, you are plotting noise and calling it a trend, and you will spend client meetings explaining movements that never happened.
We run repeats. The same question set runs multiple times through a collection window, and the reporting aggregates up rather than showing you individual reads as if they were facts. Variability is not something to smooth away, it is a finding in its own right: a brand that appears in nine reads out of ten is in a genuinely different position from one that appears in five, even if a single lucky read shows both present.
We also refuse to smooth the lines. No curve fitting, no interpolation across gaps. If collection failed for three days, the line breaks and the chart says so. A continuous line through missing data is a lie told in a visual language that most people do not read critically.
The day boundary is a real decision
This one sounds like pedantry and is not.
Our collection runs overnight and routinely crosses midnight. When we bucketed data by the calendar date the run started, a single collection loop split across two dates, and the tail of one night merged with the head of the next. Days that looked like sharp movements were just the boundary falling in an awkward place.
We moved the boundary so a collection day is labelled by the morning it finishes rather than the evening it starts. The splits went to zero. Date labels on our historical reporting shifted by a day, which we had to say out loud rather than let people notice.
The lesson generalises to anything that samples an AI surface on a schedule. Define your period boundary once, in one place, and make every chart and every calculation read it from there. If two parts of your system can disagree about what Tuesday means, they eventually will.
Presence is not prominence, and branded queries are a mirror
Two ways a visibility number quietly inflates.
Being mentioned is not the same as being recommended. Appearing ninth in a list is not appearing first, and an assistant answer has a strong positional bias in what the shopper actually acts on. A flat presence rate treats those identically. We weight by position, so a brand named first counts for meaningfully more than a brand named seventh, and the weighting is published rather than proprietary. A visibility score you cannot decompose is a vanity metric with a decimal point.
The bigger one is branded queries. If the question names your brand, you will appear in the answer. That tells you almost nothing about whether an AI system would surface you to someone who has never heard of you, which is the entire commercial question. Leaving branded queries in your competitive metrics gives you a share of voice number that is largely a reflection of how many of your own queries you wrote.
We exclude them from every competitive calculation and label the basis on the chart. They stay visible separately, because branded queries answer a different and still useful question: does the assistant describe you accurately when someone asks about you directly? That is a brand safety measure, not a discovery measure, and mixing the two produces a number that answers neither.
The definitions have to live in one place
A smaller point that has saved us more embarrassment than any other.
Every tunable definition in our system, the scoring weights, the day boundary, the vocabulary for what counts as a non-answer, sits in one configuration that both the calculation and the interface read from. The tooltip explaining a metric is generated from the same source as the metric. They cannot drift.
This sounds like housekeeping. It is actually a client trust mechanism. The failure mode it prevents is the one where someone changes a weight in the code, the tooltip keeps describing the old formula, and eight weeks later a client’s analyst works out that the explanation and the number disagree. You do not recover the room after that.
What to ask the vendor
If you are buying AI visibility measurement rather than building it, the questions worth asking are not about coverage. Everyone claims coverage. Ask about denominators.
What happens when the assistant refuses to answer? Does that read count against my brand?
How many times do you ask each question, over what period, and do I see the variance or just the average?
What does a blank cell mean? Is it zero, unmeasured, or outside your reporting cut?
When you report my share of voice, share of what exactly? If a provider returns a top ten and renormalises share across those ten, you are being shown a share of ten brands, not a share of your category. We have seen that distortion read roughly double the true figure. It is not deception, it is just an undisclosed denominator, and it is the single easiest thing to get wrong when ingesting third-party AI visibility data.
Do branded queries count in my competitive metrics?
If you changed a metric definition six months ago, does my historical reporting reconcile?
A vendor who answers those cleanly is worth paying. A vendor who reaches for coverage stats instead has not done the thinking.
What to do next
Three things, in order.
Establish what you can actually measure before you commission a content plan. Pick your priority market and category, confirm which AI surfaces are genuinely readable there, and set a baseline with repeats rather than a single snapshot. This is cheap and it stops you optimising against noise. If you have an analyst who can run it, our DaaS tier gives you direct data access with enough upskilling to use it properly.
Fix the retrievable layer while you build the measurable one. Product detail page structure, entity-level schema, clear specific claims, comparison content and seeded questions all improve how an AI system reads you, and they move within weeks. This is the quick-wins track and it is where most brands should start. Our Challenger tier is built for this, running Amazon and the agentic channels hands-on while the measurement matures underneath.
Be honest with yourself about the long game. Domain authority remains the persistent bottleneck for AI citation. Content clarity gets you read. Third-party mentions, reviews, expert sources, retail and publisher presence get you cited. Nobody should sell you a short-term content fix as the whole answer, and anyone who does is either not measuring or not telling you what the measurement says. For multi-brand portfolios where the constraint is internal alignment rather than execution capacity, our Enterprise workshops are usually the faster route in.
The uncomfortable version of all this is that honest AI visibility numbers are lower than the ones you are currently being shown. My recommendation would be to take the lower number. It is the only one you can build on.
Stephen Honight is the founder of Lmo7, an AI-native agentic commerce agency helping consumer brands win in AI-powered discovery across Amazon, ChatGPT, Google AI Overviews and Gemini.