We ran 70,764 queries through Amazon Rufus over 30 days across the UK, Germany and Spain. The lead brand changed constantly, answers named about two brands, and review depth read like an entry requirement. First-party data on why AI shelf space behaves like a nightly re-draw, not a ranking.
We Ran 70,764 Queries Through Rufus in a Month: AI Shelf Space Is a Nightly Re-Draw, Not a Ranking
By Stephen Honight, Founder of Lmo7
Over the last 30 days we ran 70,764 queries through Amazon Rufus. The same prompts, every night, across amazon.co.uk, amazon.de and amazon.es. Not a spot-check, not a screenshot, a tracker running 1,082 distinct prompts over 392 nightly runs from 25 July to 23 August 2026.
The headline finding is simple. The brand that leads a Rufus answer tonight is often not the brand that led it last night. AI shelf space behaves less like a ranking and more like a nightly re-draw, and most brands are still measuring it like a ranking.
If you only take five things from this piece, take these. The lead brand rotates constantly, so a one-off screenshot of your AI visibility measures a coin toss. An answer names about two brands, and a third of the time a brand named in the text has no buyable product card. Rufus refuses almost nothing except age-restricted products and head-to-head brand verdicts. Review depth looks like an entry requirement, not a tiebreaker. And Rufus generated 127,561 of its own follow-up questions in a month, so a static keyword list cannot keep up. The rest of this piece is the evidence and what to do about it.
Every figure in this piece comes from our own Lexym Rufus Visibility Tracker over that window. No vendor study, no extrapolated panel, no borrowed headline. That matters because the last big Rufus dataset to circulate was the Publicis “Decoding Rufus” study, which we read closely on this blog a fortnight ago, and the biggest question it left open was variability. What does Rufus visibility do when nobody is optimising anything? This is our answer, or at least the first month of it.
The lead brand is far less stable than people assume
Of the 885 prompts we ran ten or more times in a single market, an average of 3.7 different brands took the first-named slot at some point. Only 13% of prompts had one brand lead consistently across the month. Nearly half had four or more different brands lead at some point.
Sit with that for a second, because it breaks the mental model most teams bring to this. In classic Amazon search, rank moves, but it moves for reasons you can mostly trace: a price change, a stock-out, an ad push, a review spike. In Rufus, the same prompt on consecutive nights can produce a different opening brand with nothing visible changing on the shelf.
The mechanism is worth understanding before you change anything. Rufus is a generative system. It composes an answer each time rather than reading one off a fixed index, so the output is a draw from a distribution, not a lookup. Your brand does not hold a position in that world. It holds a probability. Some brands hold a high one and lead most nights. Most brands in our set are taking turns.
The practical consequence: if you check your Rufus visibility once, screenshot it and put it in a deck, you have measured a coin toss. A good result means little and a bad result means little. The only reads that mean anything are how often you lead, how often you appear at all and whether those rates are trending up or down over weeks.
An answer names about two brands
Across 69,101 completed answers, Rufus named an average of 2.2 brands per answer, with a median of two, alongside an average of 5.6 product cards. We counted brands against our matched set of 250 tracked brand names and aliases, and an independent measure based on the product cards themselves agreed at 2.10, which is why I am comfortable saying “about two”.
It is a short list. Page one of classic Amazon search gives a category maybe 40 visible slots once you count ads. A Rufus answer gives it about two named brands and a handful of cards. Being fourth or fifth in your category does not get you a soft landing further down the page. It gets you left out of the sentence entirely.
There is a second split inside this finding that I think matters more than the headline. In 34% of answers where a brand was named in the text, that brand had no matching product card in the carousel. Being talked about and being buyable are two different jobs. The text mention comes from what Rufus knows and retrieves about brands. The card comes from a product being eligible and surfaced for that query. A brand can win the conversation and still miss the transaction, and a third of the time in our window, someone did.
If you are tracking Rufus at all, track those two things separately. A mention rate without a card rate flatters you.
Rufus almost always answers, but not everything
Of 69,368 completed query runs, only 267 came back as something other than an answer. That is 0.4%: 91 refusals, 161 clarifying questions and 15 empty responses.
The refusals cluster in two places. Age-restricted categories, where the whisky prompts in our set did most of the refusing, and direct head-to-head brand comparisons. Rufus will happily discuss attributes, use cases and trade-offs. Ask it flat out whether brand X is better than brand Y and it tends to decline the referee role.
The clarifying questions cluster somewhere else entirely: broad discovery prompts and gifting prompts. Ask it something vague and it asks you a question back rather than guessing.
The commercial read is that comparison content on your listings has a different job than most teams assume. Rufus is not going to declare you the winner of a head-to-head, so stuffing “better than [competitor]” claims into content is aiming at an answer Rufus will not give. What it will do is answer attribute questions, and the brand whose listing content actually covers those attributes is the brand whose language gets used.
Review depth reads like an entry requirement, not a tiebreaker
Across 387,366 product cards surfaced in the window, the average rating was 4.50. Only 3.6% of cards were rated below 4.0. Just over half, 50.8%, carried 1,000 reviews or more.
That is not a picture of an assistant weighing up plucky newcomers. It is a picture of a shortlist assembled from products that have already cleared a bar. The Publicis study said the same thing from the optimisation side: behavioural signal filters products before content gets read. Our tracker sees the output of that filter at scale, and the output is overwhelmingly well-reviewed, highly rated products.
Worth noting alongside it: across all 392 runs in this window, on these three markets, we found zero sponsored markers in the carousels. One extractor version, one month, three markets, so I hold that lightly. But as of this window, the Rufus carousel we observed was an organic surface. You could not buy your way in. You had to qualify.
For a challenger brand the sequencing implication is honest and a bit uncomfortable. Review depth and rating are not tiebreakers you polish at the end. They look like the entry ticket, and the content work pays off fully once you are in the room. That is not a reason to delay content. The two run together, because when the eligibility bar is cleared, the content decides what Rufus actually says about you. It is a reason to run reviews, retail basics and content as one programme rather than treating content as the whole plan.
Rufus writes its own question set
Rufus offered an average of 3.95 follow-up questions per answer. Over the month that came to 127,561 distinct follow-up questions.
The practical point is short. A content plan built on a static list of 50 target questions is planning for a surface that invented 127,561 new ones last month. You cannot chase that long tail question by question. What you can do is make sure your listing content covers your category’s attribute space, the use cases, the fit language and the honest limitations, so that whichever branch of the question tree a shopper walks down, the retrievable substance is there.
What this does not tell you
Honesty about the limits, because we ask for it from everyone else’s data. This is one month, three markets, one tracker. Volume is heavily UK-weighted: 63,784 completed queries on amazon.co.uk against roughly 3,000 each in Germany and Spain, so the German and Spanish reads are signals rather than settled numbers. The zero-sponsored observation is scoped to this window and one extractor version. Brands-per-answer counts our matched set of 250 tracked brands, cross-checked against the independent card measure. And none of this includes client account data. Every figure is aggregate tracker output.
This is also a European read. Rufus behaviour on amazon.com may differ, and we have not measured it yet, so I would treat these numbers as directional for the US rather than transferable. Extending the tracker’s market coverage is on the roadmap.
One month also cannot tell you whether the volatility is seasonal, model-version driven or permanent. That is rather the point. The only way anyone finds out is by measuring continuously, which is what we will keep doing, and we will publish what changes.
What to do with this
If AI shelf space is a nightly re-draw, the measurement job changes. A rank on a Tuesday is noise. The signals worth managing are your lead rate and appearance rate over weeks, your mention rate and card rate tracked separately, and whether your review base clears the bar the surfaced set actually shows. We have seen what happens commercially when a brand works this loop properly: over a 60-day engagement Trip Drinks moved from 7th to 3rd average position across the major LLMs and grew AI referral traffic 33%, measured as a trend, not a screenshot.
My recommendation would be to sequence it like this.
1. Get a baseline that respects the volatility. One screenshot is a coin toss, so the first job is repeated measurement over at least a few weeks: lead rate, appearance rate, mention-vs-card split, and where your category’s review bar actually sits. Our DaaS tier exists for exactly this, direct data access with upskilling from £250 + VAT per month per marketplace, so your team reads the distribution rather than the anecdote. Owner: your ecommerce lead, with Lmo7 setting up the tracking.
2. Run eligibility and content as one programme. The surfaced set says reviews and ratings gate entry, and the follow-up data says broad attribute coverage decides what gets said once you are in. That is retail fundamentals and listing content moving together. Our Challenger tier does this hands-on: Amazon Retail Ops at £750 per month for the eligibility side, Agentic Tracking & Content Optimisation at £1,000 per month for the track-diagnose-fix loop. Owner: Lmo7, reporting monthly against the trend lines from step 1.
3. If you run a portfolio, standardise the measurement before you scale the optimisation. Multi-brand teams have the extra problem of five brands measured five ways. Our Enterprise engagements, from £5,000, set a common measurement standard across brands and markets and put a roadmap behind it that internal teams can execute.
The uncomfortable version of the takeaway is that most Rufus measurement being shared in decks right now is a photograph of a moving object. The useful version is that trend tracking is not expensive, the data is collectable and the brands that measure properly will be the ones that notice, weeks before anyone else, when the re-draw starts falling their way.
Stephen Honight is the founder of Lmo7, an AI-native agency helping consumer brands win in AI-powered discovery and agentic commerce. Lmo7 works with brands including Trip Drinks, Veloforte, Brown-Forman, Haleon, Pelotan and Symprove across Amazon, AI search visibility and agentic enablement.