Publicis Just Published the Best Data We Have on Rufus. Read the Appendix Before You Believe the Headline

Analytics & Measurement | 10 | Published:

By , Founder of The Lmo7 Agency

The appendix approach works because of how COSMO infers purchase intent from product content and shopper behaviour rather than keyword presence.

Publicis Commerce, Profitero+ and Mars United Commerce ran a real test-and-learn on Amazon Rufus across 8 brands and 45 products. The headline says content optimisation lifted agent visibility by 11 percentage points. The statistical appendix tells a quieter story, and the most useful finding in the whole report is the one that undercuts the sell.

Publicis Just Published the Best Data We Have on Rufus. Read the Appendix Before You Believe the Headline

By Stephen Honight, Founder of Lmo7

Most of what gets written about Amazon Rufus is opinion. Publicis Commerce has done something better than that. Working with Profitero+ and Mars United Commerce, they ran an actual test-and-learn across 8 brands and 45 products, changed the content, measured what happened, and published the method. The report is called Decoding Rufus and it came out in March 2026.

I want to give it proper credit before I get into the numbers, because it does something rare. It publishes a finding that makes its own commercial pitch harder to sell.

Then I want to talk about page 18, which is the statistical appendix, because the appendix says something quieter than the front cover.

What they actually did

Worth saying why this matters before the method. Amazon reports Rufus attracted more than 300 million users in 2025, and that those users were 60% more likely than mainstream searchers to buy the product it recommended. Those are Amazon’s own numbers, cited in the report, so treat them as a vendor’s framing rather than an independent read. Even discounted, it is a surface worth measuring.

The team put a definition around a metric called SOAR, short for Share of Agent Recommendations:

SOAR = days recommended, divided by total observable days for that prompt set

So for a given ASIN and a given shopper prompt, you watch over a window and count the proportion of days Rufus put that product in its answer. That is a sensible way to measure a system that gives you a different answer every time you ask.

They built the prompt sets by asking Rufus directly what prompts a shopper would use for a category, expanding those into real shopper questions, checking the suggested questions on each participating product page, then getting brand-partner sign-off. They picked ASINs with mixed performance levels, avoided products distorted by campaigns or promotions, and controlled for in-stock rate and promotional activity including Prime Day in the sales analysis.

That is a proper method. It is more rigour than most of the AI visibility market shows its buyers.

The finding that deserves the attention

The most useful thing in the report is not the headline. It is the eligibility gate.

Products with the highest SOAR had an average Best Seller rank 2.5 times stronger than the products in the next few tiers of visibility. The table underneath makes the drop-off obvious. The most visible product by SOAR sat at an average Best Seller rank of 4,220. The hundredth most visible sat at 34,190.

The review picture is starker. The top-ranked product by SOAR averaged 20,890 reviews at 4.6 stars. Even at the hundredth position, average reviews were still 7,395. Star ratings clustered at 4.5 and above across both high and low SOAR, which the report reads as a floor rather than a lever. Getting from 4.5 to 4.7 probably does not move you. Being under 4.5 probably keeps you out.

The report also cites Amazon’s own Rufus engineering team on the mechanism. Rufus is trained and activated on products with sufficient behavioural signal, things like search-buy and co-buy activity. Low-signal products get filtered out early to keep noise down. Which means the model has already decided whether your product is in the running before it reads a single one of your bullets.

Publicis put this in a two-stage frame. Eligibility, then optimisation. Their own line for it is blunt and correct: you can’t tweak your way into gaining agent visibility from a standing start.

Read carefully though, because the finding is not that content does not matter. It is that content cannot do its job on a product the model has already filtered out. Get the product eligible and the content is what wins the recommendation. The model still reads your titles, bullets and attributes to decide who it names, so publish into a product with no behavioural signal and you are writing for a page the model never reaches. That is a strange thing for an agency group to spell out, and they did it anyway. Good on them.

Now the appendix

The headline result is that content-optimised products saw their SOAR rise by 4 percentage points on average, creating an 11 percentage-point delta against the control group.

The chart underneath gives you all four numbers. The test group went from 39% to 43%. The control group went from 48% to 41%.

So the test group rose 4 points. The control group fell 7. Most of that 11-point delta is the control group decaying, not the test group climbing.

Page 18 is where it gets more interesting. The statistical table, described as active pairs only, reports a mean delta of 0.001 for the test group. That is a tenth of a percentage point. The median is 0.009. The control group mean is minus 0.072, which does reconcile with the 7-point fall in the chart. The t-statistic is 2.185 and the p-value is 0.030, on 79 active ASIN and prompt pairs in the test group against 85 in the control.

Two things follow from that. The test group’s own movement, on the report’s own primary analysis, is close to nothing. And the statistical significance is being carried substantially by the control group falling, not by the intervention working.

I am not saying the study is wrong. The reported p-value of 0.030 clears the usual bar, and a control group that degrades while a test group holds is a real result worth knowing about. My sense is that the honest reading is narrower than the headline: content optimisation on already-eligible products may protect visibility that would otherwise erode, and any lift beyond that is small and hard to separate from noise on this sample size.

There is also a gap between the plus 4 percentage points in the chart and the 0.001 in the table. They are computed on different bases, since the table says active pairs only and the report explains that choice as avoiding dilution from structurally inactive pairs. Fair enough. The two numbers are still never reconciled for the reader, and the chart figure is the one that travels.

The number that is missing

Here is the question I would put to any provider selling AI visibility measurement, and it is the question this report cannot answer.

The control group moved 7 percentage points while nobody touched it.

So how much does SOAR move on its own?

There is no variability baseline anywhere in the document. No measurement of how much a product’s recommendation rate drifts week to week with no intervention at all. The team did build in a sensible stability criterion, requiring at least five recommended days in both the pre and post periods, roughly a 10% recommendation-rate threshold, specifically to filter out noise. That helps. It is not the same as establishing how noisy the underlying metric is.

Without that baseline you cannot tell the difference between a 4-point lift and a 4-point wobble. Given that the untouched control group in this very study moved 7 points, that distinction is the whole ball game.

This is not a criticism unique to Publicis. Almost nobody in this market publishes a variability baseline. It is the single most common gap in AI visibility measurement and it is why we build ours before we report a client number rather than after.

The case study nobody will quote

Four brand case studies sit in the middle of the report. Three of them look spectacular. Monster Energy went from 3% to 45% SOAR on “best energy drinks without carbs” after a description rewrite. Nature’s Bakery went from 9% to 49% on a nut-free lunchbox prompt after an attribute fix, with organic rank up 5 places and unit sales up 20%. Revelyst went from 2% to 29% on “golf rangefinder with GPS”, with unit sales up 62%.

Read those against a mean delta of 0.001 and you can see what they are. They are the tail of the distribution, not the average. Single ASINs, mostly on single target prompts, chosen to illustrate the upside. Monster’s is the broadest of them, with improvements across 8 target prompts in total.

The fourth one is Sellstrom, and it is the one I would actually put in front of a client. They enhanced the bullets and refreshed the backend keywords. SOAR did not increase at all. Organic rank rose 24 places for “roofing knee pads”. They landed on page 1 for a Spanish-language term they were not on before. Unit sales rose 45%.

A brand sold 45% more units and their agent visibility did not budge. The commercial result came from classic Amazon listing optimisation, filed under a report about agentic commerce.

I like that Publicis included it. It is the most useful page in the document. It also quietly tells you where the money is right now for most brands.

What this means if you are not a £100m brand

Look again at who is in this study. Monster Energy. Alcon. Revelyst, which owns Bushnell. Lily’s Kitchen. The top-ranked product by SOAR averaged 20,890 reviews.

If you are a consumer brand doing £5m with 400 reviews on your hero SKU, this study contains no brand that looks like you. Not because the work is poor, but because you sit below the eligibility gate it identifies. The visible tier skews heavily towards category winners, and the measured effect is on products that were already in the game.

That is uncomfortable, and it is more useful than a flat “just write better bullets”. The lesson is not that content is optional. It is that content and eligibility are two halves of one job, and most brands pay for only one.

Get the eligibility inputs into shape so the content has something to work on. Review velocity. Rating floor at 4.5 and above. Buy Box ownership. In-stock reliability. Sales momentum on the SKUs you actually want recommended. None of that is glamorous and all of it is the precondition for the content to land.

Then rewrite the content for Rufus, because that is the work that actually wins the recommendation, and it pays even before agent visibility moves. Look again at Sellstrom. SOAR did not budge, and a bullet and backend rewrite still moved organic rank 24 places and lifted units 45%. That is content doing its job and paying its way while the eligibility signal catches up. For most challenger brands the content rewrite is the highest-return thing on the list, not the thing to skip.

Anyone selling you a Rufus content retainer while your hero SKU sits at 4.3 stars with 200 reviews is doing half the job. So is anyone who fixes the fundamentals and then leaves the listing reading like it was written for 2019. You need both, and the content is the half that gets you named.

There is a longer-run constraint underneath all of this that we have written about before. Domain authority is the persistent ceiling on AI citation off Amazon, and behavioural signal is the equivalent ceiling on Amazon. Both are slower and less fun to fix than content. Both set the limit on what content can do.

On the word SOAR

Small point with commercial weight. SOAR is a good name for a real metric, and Publicis has the distribution to make it the category term. We already measure the same thing in our Rufus tracking, and my recommendation would be that the industry adopts the vocabulary rather than everyone inventing a private label for share of agent recommendations.

Using the standard metric and measuring it properly is a stronger position than inventing your own and asking buyers to trust it.

What to do next

Three practical steps, in order of what will move your number soonest.

Run an eligibility check alongside the content brief, not instead of it. Pull Best Seller rank, review count, star rating, Buy Box ownership and in-stock rate for the SKUs you want recommended. If a hero product is under 4.5 stars or thin on reviews, fix that in parallel so the content you are about to write has a product the model will actually read. This diagnostic maps to our DaaS tier from £250 + VAT per month if you want direct data access and your team runs the read. It is also the first thing we do inside any Amazon engagement, right before the content work.

Rewrite the content for Rufus, and fix the eligibility inputs alongside it. Titles, bullets, A+ content, backend attributes and SPN-mapped copy are the work that gets you recommended and gets you found, and they sit in Content Optimisation for Search Visibility at £750 per month. Review velocity, rating recovery, Buy Box and stock reliability sit in Retail Ops at £750 per month. Run them together, because the Sellstrom result shows content pays through organic rank and sales even in the months where agent visibility stays flat. The full Amazon Stack including ads is £2,500 per month.

Ask your measurement provider for their variability baseline. Not their methodology deck. The actual number: how much does this metric move when nobody touches anything. If they cannot answer, treat every lift they report as directional. For multi-brand portfolios where several teams need to agree what the numbers mean before anyone acts, that conversation is usually a workshop rather than a dashboard, and our Enterprise engagements start at £5,000.

We track Rufus visibility for consumer brands alongside ChatGPT and Gemini, and we report ranges and trends rather than a single rank, for exactly the reason this study demonstrates. Trip Drinks moved from 7th to 3rd average position across the major models over a 60-day engagement, with AI referral traffic up 33%. Haleon’s Voltarol holds a 100% mention rate at number one position in its category on ChatGPT and Gemini. Those are our own measured results over the windows stated, on brands that were already eligible to be found, and they took content work, site signals and citation-source work together. Apply the same test to them that I have just applied to Publicis. That is the point of publishing the window.

If you want to know whether yours is, that is a short piece of work and worth doing before you spend anything else.


Stephen Honight is the founder of Lmo7, an AI-native agency helping consumer brands win in AI-powered discovery and agentic commerce. Lmo7 works with brands including Trip Drinks, Veloforte, Brown-Forman, Haleon, Pelotan and Symprove across Amazon, ChatGPT, Gemini and Amazon Rufus.

Source: “Decoding Rufus: An analysis of the key factors influencing AI-enabled product recommendations”, Publicis Commerce, powered by insights from Profitero+ and Mars United Commerce, March 2026. All figures cited in this piece are theirs. Brands named in their case studies are their study partners, not Lmo7 clients. Rufus user figures are Amazon’s own, cited via the report.

Explore More

AI Search Optimisation Services | LLM Visibility Framework | Free AI Search Audit | News & PR | Alexa Shopping Radar

Related Articles