Most A/B Tests Waste Your Traffic. Here's What to Test Instead.

Microsoft's experimentation team, after running thousands of controlled tests, reported that only about one third of ideas actually improved the metric they were designed to improve. The other two thirds did nothing or made things worse. That is a team with world-class traffic, tooling, and statisticians.

Now think about your own testing programme. Smaller sample sizes, fewer tests, less rigour, and probably a lot more pressure to declare a winner by Friday. The odds are not in your favour.

The maths kills most tests before they start

Run the numbers on a typical ecommerce site. Average conversion rate sits around 2%. To detect a 10% relative lift, moving 2% to 2.2%, with standard confidence and power, you need roughly 78,000 visitors per variation. That is 156,000 visitors for one test.

Most sites do not have that traffic in a reasonable window. So tests run for three weeks, produce a 6% "lift" with no statistical significance, and someone ships it anyway. Six months later nobody can find the revenue.

The practical fix is uncomfortable but simple. Only test changes big enough to produce a 20% swing or more, which needs closer to 20,000 visitors per variation. Everything smaller than that is unmeasurable at your traffic level, so stop pretending otherwise and just make a judgement call.

You are testing the wrong layer

Button colours, headline tweaks, and hero image swaps are popular because they are easy to build. They are also the least likely things to change a buying decision. Nobody abandoned a cart because the call to action was teal.

The things that actually move revenue sit deeper. Offer structure. Shipping thresholds. Whether the price is explained or just displayed. How many steps stand between intent and payment. Baymard Institute's research found that better checkout design could recover around 35% of lost checkout conversions, which is a far bigger prize than any colour test will ever deliver. If you want a place to start, look at where your checkout page is leaking money before you touch anything above the fold.

There is a second problem with page-level testing. It treats every visitor as the same person. A returning customer who has viewed a product four times and a first-time visitor from a cold ad get the identical experience, and the test result averages both into a number that describes neither.

Test moments, not pixels

The most useful thing you can test is not what a page looks like. It is what happens at a specific point in a specific visitor's session.

A hesitation on the shipping field is a moment. A second visit to the same pricing page in 48 hours is a moment. Scrolling to the bottom of a product page and going back to search is a moment. These behaviours carry intent, and they are where offers and reassurance actually land. Your analytics show you the exit, not the reason, which is why page-level testing keeps producing flat results.

This is the work Pounce does. It watches behaviour in the session, reads what the visitor is signalling, and acts at the point where the decision is still live rather than after the tab closes. You are no longer testing a page against a page. You are testing a response against a moment.

Test three things in that category and you will learn more in a month than a year of headline variants. Try a timed reassurance message for visitors stalling at checkout. Try a returning-visitor offer that first-timers never see. Try holding back a discount entirely and leading with delivery speed instead.

Measure revenue, not conversion rate

Here is the trap that catches good marketers. A variation lifts conversion rate 12% and everyone celebrates. Nobody checks that average order value fell 15% because the test pushed a discount to people who would have paid full price.

Revenue per visitor is the only number that settles the argument. It absorbs conversion rate and order value together, so you cannot win the test and lose the month. If your traffic is fine but your revenue per visitor isn't, no amount of page testing will fix that, because you are optimising the wrong variable.

One more rule. Run every test for a minimum of two full weeks and never stop it early because the line looks good. Early stopping is the most common way a testing programme produces confident, expensive nonsense.

Most testing programmes fail because they test small things on small samples and measure the wrong outcome. Test bigger changes, target specific moments, and judge everything on revenue per visitor. That is how testing starts paying for itself instead of filling slides.

Frequently asked questions

Why do most A/B tests fail to produce a winner?
Most tests fail because the change being tested is too small to move behaviour, or the site does not have enough traffic to detect the difference. Microsoft's experimentation team has reported that only around a third of tested ideas improve the metric they were meant to improve.
How much traffic do I need to run a valid A/B test?
At a 2% conversion rate, detecting a 10% relative lift with standard confidence takes roughly 78,000 visitors per variation. Detecting a 20% lift takes around 20,000 per variation. Below that, your result is noise.
How long should an A/B test run?
Run tests for at least two full weeks to cover weekday and weekend buying patterns, and never stop a test early because it looks like it is winning. Stopping early is the single most common cause of false positives.
What should I test instead of button colours and headlines?
Test offers, pricing presentation, checkout steps, and the timing of what you show people. Changes to what you are selling and when you say it move revenue far more than cosmetic changes to how a page looks.
Should I measure conversion rate or revenue per visitor in a test?
Measure revenue per visitor. A variation can lift conversion rate while lowering average order value, which means you win the test and lose money.
← All posts