A Peec AI study of 37,804 AI responses found that most ways a shopper phrases a prompt are nearly identical in meaning, so chasing every keyword variation is wasted effort. The real risk sits in unbranded middle-of-funnel queries. Here is how to build an AI visibility tracking programme that holds up.
Intent Beats Keywords: What a 37,804-Response Study Says About Tracking AI Visibility
By Stephen Honight, Founder of Lmo7
One of the first questions a brand asks when it starts tracking AI search is some version of “how many prompts do we need to track”. The instinct is to capture every way a shopper might phrase a question, then watch all of them. It feels thorough. It is also where most tracking programmes quietly fall over, because the list gets enormous, the cost climbs and the signal gets harder to read, not easier.
A study published by Peec AI in June 2026 is useful here because it puts numbers against that instinct. The headline finding is counter-intuitive in a helpful way. You do not need to track every phrasing. You need to track intent, funnel stage and prompt format. That is a much smaller, much sharper job, and it happens to be how I think a credible measurement programme should be built anyway.
Let me walk through what the study found, then layer on what I would actually do with it.
What Peec AI measured
Peec AI is a visibility tracking tool, so this is their research and the figures below are theirs, not Lmo7’s. Worth saying that plainly up front. It is also a sponsored piece, so treat the numbers as directional evidence rather than fixed laws.
The scale is what makes it worth reading. They analysed 1,754 prompts and 37,804 AI responses across five sectors and 18 sub-verticals, run through ChatGPT, Gemini, Perplexity, Google AI Mode and Google AI Overviews. They ran two parallel studies. One looked at 288 real human-written prompts to measure how varied actual phrasing is. The other took 54 base prompts and expanded each into dozens of tiny variations to measure what small wording changes do to which brands get mentioned. Every prompt was run several times to account for the fact that AI engines do not give the same answer twice.
The clever part is how they compared prompts. Rather than eyeballing the words, they turned each prompt into a semantic embedding and measured the distance between them with cosine similarity, scored from 0 to 1. A score near 1 means two prompts mean almost the same thing even if the words differ. A score near 0 means they have drifted apart in meaning. This lets you ask a precise question: when does changing the words actually change which brands surface?
The findings worth acting on
There are five things in here that matter commercially.
Human prompts only look different on the surface. Around 88% to 92% of human prompt pairs sat above 0.50 cosine similarity, and about 95% sat above 0.40. Fewer than one in ten showed real semantic drift. People phrase the same need in many ways, but mathematically most of those phrasings are close together. The long list of “different” prompts you were about to track is mostly the same prompt wearing different clothes.
Wording only moves brand mentions once you cross a threshold. The average probability of any given brand being mentioned was 4.9%. Visibility stayed broadly stable as long as prompts sat above roughly 0.50 to 0.60 similarity. It was only when prompts dropped into the lowest bin, around 0.35 to 0.39, that visibility fell by about 2.40 percentage points, which is close to a 50% relative drop. The big losses live in the left tail, where very few real shoppers actually type. So tracking a thousand near-identical phrasings tells you almost nothing, because they all sit in the stable zone.
Then there is style, which changes what surfaces as much as meaning does. This is the bit most teams miss. Comparison, table, list and ranking prompts surfaced more brands than open-ended questions, by around 20% on average. Concise keyword-style prompts like “best CRM small business 2026” beat conversational persona-engineered prompts like “you are an IT consultant, recommend a tool” by up to 25%, because persona prompts tend to broaden into educational answers that mention fewer brands. Constraints behaved differently by engine. Adding budget or feature requirements cut the brand count in ChatGPT and Perplexity but increased it in Gemini and Google AI Overviews. Prompt length and filler words had effectively no impact at all.
The middle of the funnel is where wording decides winners. Top-of-funnel prompts like “what is a CRM” are stable because they are educational. Bottom-of-funnel prompts are also stable, but Peec AI calls this false stability, because it is anchored on a brand name you already typed. The volatile zone is the unbranded commercial middle, prompts like “best CRMs for a small remote team”. These shifted which brands appeared even at fairly high similarity, up around the 0.60 to 0.65 range. Their suggested tracking split is 25% top-of-funnel, 50% middle, 25% bottom.
The engines themselves do not behave the same way. Gemini’s sensitivity to wording faded fastest. Google AI Overviews held the most persistent middle-of-funnel sensitivity. ChatGPT only started losing brands once prompts dropped below the 0.60 to 0.64 band. The practical takeaway is that you have to report each engine separately before you blend anything, or you average away the differences that actually matter.
There is one honest caveat in their data that I would not skip over. High similarity does not always mean matching intent. “Car rental Charleston” and “Car rental Charlestown” score 95% similar but serve completely different shoppers. If a core qualifier changes, a location, a product variant, a demographic, a brand name, treat it as a new intent even if the cosine score looks high. The maths is a guide, not a judge.
What this confirms about how to track
The reason I wanted to write about this study is that it gives independent backing to something we keep telling clients, usually while they are looking at a tracking quote and wondering why it is not ten times bigger.
The mechanism, before the recommendation. AI engines cluster human phrasings tightly in meaning, so brand visibility is stable across most of the ways a real shopper will ask. The commercial risk is not spread evenly across a thousand prompts. It concentrates in a specific place: unbranded middle-of-funnel queries, where the answer genuinely shifts depending on phrasing, format and engine. That is the ground you actually fight on.
So a tracking programme that tries to be comprehensive by volume is solving the wrong problem. A tracking programme that is precise about intent, funnel stage and format is solving the right one, and it costs less to run.
Here is the shape I would build it in.
Segment by funnel stage and weight toward the middle. The 25/50/25 split is a sensible starting point. The middle is where you can win or lose a recommendation, so that is where most of your tracked prompts should live. Track a few top-of-funnel and bottom-of-funnel prompts to confirm stability, not to chase movement.
Tag every prompt by format. A ranking prompt and an open question are not the same test, even on the same topic. If you mix them in one number you will see noise and call it a trend. Tagging by format also tells you something useful for content work, because the formats that surface more brands are the formats worth being ready for.
Report each engine on its own line. Blending ChatGPT, Gemini and AI Overviews into a single “AI visibility” score feels tidy and hides exactly the differences a brand needs to act on. We see this constantly. The variance clients flag between tools is real, and a lot of it is this. Different engines, reported as one.
Watch the qualifier blind spot. When you build your prompt set, be deliberate about location, pack size, format and use case. Two prompts that look almost identical can serve different buyers, and if you collapse them you will miss a gap that is costing you sales.
Track the stable zone lightly and the volatile zone closely. You do not need a hundred variations of a query that barely moves. You need good coverage of the middle-of-funnel queries where the recommendation is genuinely up for grabs.
Where the proof sits in our own work
This is not theory for us. The Share-of-Model work we ran for Haleon across Voltarol, Sensodyne and Centrum was built on exactly this logic, tracking mention rate and average position per brand and per engine across ChatGPT and Gemini rather than chasing a giant keyword list. For Voltarol that produced a clean, defensible read: a 100% mention rate and a number one average position in its category. That number is only meaningful because it is anchored to a specific category intent and reported per engine, not averaged into mush.
With Trip Drinks the same discipline showed movement that mattered commercially. Over a 60-day engagement their average position across the major models went from 7th to 3rd and AI referral traffic rose by 33%. We could see that movement because we were tracking position on consistent, intent-led prompts, not watching a thousand phrasings drift around in the stable zone.
My sense is that the brands getting the most out of AI visibility tracking are not the ones tracking the most prompts. They are the ones tracking the right prompts and reading them per engine.
A note on honesty
Two things to keep straight. The Peec AI figures are theirs and they come from a sponsored study, so I would treat the exact percentages as directional. Models change, regulated categories like healthcare can behave differently and the precise numbers will move. What is likely to hold is the mechanics: tight clustering of human phrasings, stability across the bulk of real queries and concentrated risk in the unbranded middle.
The other thing, which is our standing view rather than theirs. Tracking tells you where you stand. It does not move you up on its own. The lever that decides whether an engine recommends you is still a mix of content clarity and domain authority, and authority is the slow one. So treat measurement as the thing that points your effort in the right direction, not the thing that does the work. A sharp tracking programme makes your content and authority spend land in the right place. It does not replace it.
So what should you do next
If you are setting this up, here is where I would start, in order.
Start with the data layer. If you have an internal team that can read it, our DaaS tier gives you direct access to the source data across Amazon, Meta, Google and Shopify plus the upskilling to interpret it, from £250 plus VAT a month. This is the low-friction way to get an intent-led, per-engine view running without a full retainer.
Build the tracking properly if you carry a portfolio. For multi-brand businesses where several stakeholders need to align before anyone acts, our Enterprise Share-of-Model tracking sets up the funnel-weighted, format-tagged, per-engine programme described above across your brands, with a baseline and quarterly updates. This is the Haleon shape.
Close the loop if you want the fixes done. If the point is not just to measure but to move, our Challenger tier runs the track, diagnose, fix, re-track loop on Amazon and AI search as a monthly rhythm. Measurement feeds the content and ads work directly, so you are acting on the volatile middle-of-funnel queries rather than admiring them.
If you take one thing from the Peec AI study, make it this. The list of prompts you were about to track is mostly one prompt. Find the handful that genuinely move, watch them per engine and put your effort there.
Stephen Honight is the founder of Lmo7, an AI-native agency helping consumer brands win in AI-powered discovery and agentic commerce. Lmo7 works with brands including Trip Drinks, Veloforte, Brown-Forman, Haleon, Pelotan and Symprove across Amazon and AI search. If you want to know where your brand stands in AI-mediated discovery, that is where we start.