Client system · Search + AI answers
AI rewrote the product pages. We didn't ship any of it on faith. The new pages had to beat the old ones on real traffic first. One of them earned 23.75% more revenue per visitor before it went live.
An Australian ecommerce operator with two stores
Two Shopify stores
SEO + AI answer engines
50/50 split testing
Human decisions on what ships
The same page, before and after
The page before
The old description: paragraphs written for how Google worked years ago. Nothing an AI assistant could quote.
The product page as archived before the project.
Brand masked.
+23.75%
revenue per visitor, new page vs old, live traffic
+13.99%
average order value in the same test
+8.56%
conversion rate, moderate confidence
9.6%→11.04%
The client runs two Shopify stores with hundreds of products between them. The descriptions had been written for how Google worked years ago, and competitors had moved on to structured, specific pages.
Search itself had moved too. Buyers increasingly ask AI assistants what to buy, and those assistants quote pages that answer questions plainly. A paragraph stuffed with old keywords gives them nothing to quote.
Rewriting the catalogue by hand had been on the list for months and kept slipping. The real question was never whether AI could write the pages. It was whether the AI pages were actually better, and how you would prove that before touching a live store.
Five stages, from raw product to a page that proved itself.
Each product is pulled from the store and compared against the strongest competing pages. The generator works from what already ranks, not from a blank prompt.
One format per store: what the product does, what you will feel, who it is for, how it works, a comparison table, full specifications and a plain question FAQ, plus descriptive alt text for the product images.
A human reviews the draft before anything goes near the store. New pages stage on a hidden copy of the shop, never straight onto the live one.
For the first rollout, new pages went into live 50/50 split tests against the ones they were meant to replace, on real store traffic.
Not a paragraph of keywords. A structured page a buyer can scan and a machine can parse. Pages produced through this workflow carry plain question FAQs, the questions buyers actually ask, answered directly. This is the part AI assistants lift answers from.
For the first rollout, every new page faced the one it was meant to replace in a live 50/50 split test. The rule was agreed before results came in: same or better goes live, worse gets discussed. It held even when it was inconvenient.
The rule produced honest outcomes. A tie was called a tie and went live for findability, not on a conversion claim. Not every rewrite won. One later variant trailed the control from the first weeks and was stopped before rollout. It never shipped. Not every page since has needed its own test. The early ones earned the workflow its trust, and new batches still go through testing.
Same or better rule, agreed before results
Ties disclosed as ties
Losing versions stopped, not shipped
The catalogue has a standard Pages produced through the workflow follow one structure built for both Google and answer engines, instead of a decade of accumulated copy styles. Rollout across the catalogue is still in progress, deliberately, at the pace the evidence allows.
Decisions run on evidence, not enthusiasm The rewritten pet bed page earned its place on live traffic. It produced 23.75% more revenue per visitor than the old page across its full test, and it went live. A second test closed up 23.18% on conversion. A tie went live with the tie disclosed. A loser didn't ship. Two early tests simply needed more data, so they got more data.
The voice lesson was learned cheaply The clinical voice worked on some health store pages and missed the pet store audience entirely. The pet pages were rewritten in a warmer voice instead of being defended.
The client's team runs it The process was handed over with a master prompt and a shared sheet, on the client's own API key. Their team produces new pages now. The workflow's completion reports still land in their Slack after every run.
This build fits operators who recognise their catalogue here.
Ecommerce teams with catalogues too big to rewrite by hand.
Stores whose product pages were written for a previous era of search.
Brands that want visibility in AI answers without trusting AI copy blind.
Operators who would rather own a testing loop than buy a one off copy project.
Teams who want the capability in house, running on their own keys.
We can map the catalogue, build the generation and testing loop, and hand your team a process they can run themselves. Built on your stack, owned by you, with people keeping the final say.
store add to cart rate through the rollout
The rule was agreed before any results came in: same or better than the old page means the new one goes live. Worse means it doesn't. No exceptions for enthusiasm.

Live on the client's store.
Product name masked.
The rewritten pet bed page went up against the old one in a live 50/50 split test on real store traffic. The full test ran four weeks: 1,816 visitors and 120 orders. The new page produced 23.75% more revenue per visitor. Average order value was 13.99% higher. Conversion rate was 8.56% higher. The platform's own read: roughly a 76 to 79% probability that the new page is the better one. Ahead on every metric, but short of the 95% bar for a conclusive result, so we say exactly that. Under the rule agreed in advance, the new page went live.
The live test platform, final read after the test closed. Ahead on all three metrics, short of the significance bar.
Product name and revenue figures masked.
A second rewritten page, a pet training product, ran the same gauntlet over the same four weeks: 1,225 visitors and 96 orders. Its new description closed 23.18% up on conversion and 22.72% up on revenue per visitor. Order value was unchanged. Confidence was similar. Same rule, same outcome: it went live.
The second test, final read. Conversion and revenue per visitor both up, order value flat.
Product name and revenue figures masked.
The client watches the same numbers we do. This is their monthly funnel across product pages. Add to cart climbed from under 10% to over 11%, and conversion from under 5% to 5.4% through the rollout, both ahead of the same season a year earlier. We don't claim the description project caused all of that. Ads, seasonality and pricing move these numbers too. The controlled tests are what decisions were based on.
The storewide funnel, as shared from the client's reporting.
Once the process held up, it was automated and handed over. The client's team runs new products through it now. The workflow posts a completion report in their Slack after every run: descriptions, answer blocks, FAQs and image alt text per product, with the error count visible.
The workflow's completion message, as posted after a real run.
Store and product names masked.
A secondary signal, kept in its place. The client's own reporting breaks out visitors from ChatGPT, Perplexity, Gemini, Copilot, Claude and other assistants as their own channel, separate from Google organic. Volume is still small, and it is reported as exactly that: small, real, and watched from zero. The controlled page tests above are the proof. This is the early instrumentation.
The client's live channel report.
Figures masked. The AI Search row is the point.
Wrong voice output rewritten, not defended
What we learned Don't ask whether AI can write the page. Ask how you'd prove the new page is better before it replaces the old one, and build that proof into the process. The storewide trend moved the right way through the rollout, and the client credits the project for it. We point at the controlled tests when people ask why we believe it.
2026-06-18
Read the case study