VezertVezert
Back to Resources

A/B Testing: When It Pays, What It Costs, and What You Need First

A/B testing explained for buyers: required traffic volume, run time, priority tests by page type, and how SEO split testing differs from standard CRO tests.

Updated November 20, 20269 minLena Tarhonska · Co-founder & CEO at Vezert
A/B Testing: When It Pays, What It Costs, and What You Need First

A/B testing is running two versions of a page at the same time, splitting traffic between them, and keeping the one that produces more of the outcome you care about. The mechanics are simple and widely explained. The part that decides whether it works for you is not the mechanics, it is whether you have enough traffic for the answer to mean anything.

Most articles about A/B testing are written by the companies selling the tools, so they skip the question of whether you should be testing at all. This one starts there, because for a large share of sites the honest answer is not yet.

When A/B Testing Is Worth It and When It Is Not

The threshold is set by arithmetic, not enthusiasm. To detect a realistic improvement, say a fifth more conversions, you need roughly 100 conversions per variant before the result separates from chance. That is 200 conversions per test, and at a 2% rate it means about 10 000 sessions on the page you are testing.

Below that, tests do not fail loudly. They produce a winner that is noise, you ship it, and the effect never appears in the monthly numbers. Two rounds of that and the team concludes testing does not work, when what did not work was testing at that volume.

The arithmetic behind that threshold is worth seeing once rather than taking on trust. Evan Miller's sample size calculator lets you put in your own baseline rate and the smallest lift you would act on, and it will tell you how many visitors per variant that costs. Most people who run it for the first time discover their planned test needed four times the traffic they had, which is a cheaper way to learn it than a month of waiting.

The number that decides everything

About 200 conversions per test, roughly 100 per variant, before a result separates from chance at a realistic effect size. At a 2% conversion rate that is 10 000 sessions on the page being tested. Anyone who sells you a testing programme without asking your session count is selling you coin flips.

What You Need in Place Before the First Test

Three things, and each takes hours rather than weeks. Missing any one produces results you cannot act on, which is more expensive than not testing, because the conclusions get quoted for a year afterwards.

Conversion tracking that fires once from every path, verified by hand. A baseline of at least eight weeks so you know what normal variation looks like. And a written hypothesis with a number in it: not "a shorter form will do better" but "cutting the form from nine fields to four will lift submissions by at least 15%". The number is what stops a flat result from being reinterpreted as a small win.

What to Test First, by Page Type

Nobody has the traffic to test everything, so the order matters more than the list. The table below is what we start with by page type, drawn from audits rather than from theory, along with roughly the traffic each needs to give a clean answer within a month.

The pattern behind it: the first test on any page should be a removal, not an addition. Removals are faster to build, easier to interpret, and they win more often, because most pages carry more than they need rather than less.

Page typeFirst testWhy it usually winsSessions needed
Lead form pageRemove two fieldsEvery field is a decision under doubt~4 000/month
Pricing pageShow a range instead of Contact usRemoves the reason to leave and compare~6 000/month
Service pageMove proof beside the askBelief is missing at the moment of asking~5 000/month
Category pageCut filters to the three usedChoice paralysis, not missing features~8 000/month
CheckoutShow total cost one step earlierLate costs are the top abandonment cause~3 000 orders/month

How Long a Test Takes at Your Traffic

Divide the sessions you need by your weekly sessions on that page, then round up to whole weeks. Round to weeks because traffic behaves differently on weekdays and weekends, and stopping mid-week tilts the result toward whichever days happened to be included.

Two rules keep tests honest once they are running. Do not look at the result before the planned end, because you will stop at a peak and call it a win. And do not run two tests on the same flow at once unless your tool handles the interaction, which most do not by default.

How SEO A/B Testing Differs From the Normal Kind

SEO split testing answers a different question and works differently in practice. Instead of splitting users, you split pages: half a set of similar templates gets the change, half does not, and you compare organic traffic between the groups over several weeks.

The reason for the difference is that a search engine sees one version of a URL, not two, so user-level splitting cannot measure a ranking effect. It also means the usual worry is misplaced. Google's own documentation on website testing states plainly that A/B testing does not carry a ranking penalty when implemented normally, and the cloaking concern applies only to serving different content by user agent.

SEO tests need more patience: four to six weeks minimum, and a set of at least twenty comparable pages before the comparison holds.

What Not to Test, and Why Button Colours Keep Coming Up

The examples that circulate hardest are the ones that photograph well, and button colour is the worst of them. It survives because a famous case study from 2009 showed a large lift, and because it is the easiest thing to change. In our audits it has never been the biggest lever on a page, and it is usually a test spent to learn nothing.

The rule that saves the most time: do not test anything you would ship regardless. If the current version is broken, unreadable, or contradicts the ad that brought the visitor, fix it. A test on a broken page measures which broken version is less bad.

Also skip anything that cannot plausibly move the number by a fifth at your volume. A change worth 2% is real and worth shipping, but you cannot detect it below roughly 40 000 sessions, and pretending otherwise burns a month per test.

How to Read the Result Without Fooling Yourself

A test gives you three possible outcomes and most teams only plan for one. A clear winner is the rare case. A clear loser is genuinely useful, because it tells you the assumption behind the page is wrong and saves you from building more of it. And a flat result, which is the most common, means the change was too small to matter at your volume, not that the idea was wrong.

Write the interpretation rule before you start: what number counts as a win, what counts as flat, and what you will do in each case. Deciding afterwards is how a flat result becomes a small win in the retrospective.

Nielsen Norman Group's work on quantitative usability covers the sample size question in more depth, and it applies to testing as much as to research.

What Goes Wrong Most Often

The failures we see repeat, and none of them are about statistics. They are about process, which is why buying a better tool rarely fixes them.

  • Stopping early on a good day, which turns a coin flip into a decision
  • Testing something so small that even a real effect would be undetectable at your volume
  • Changing two things at once and learning which pair won rather than which change worked
  • Never writing down the expected result, so every outcome gets reported as a success
  • Running a test over a seasonal peak and applying the conclusion to the rest of the year

"The tests that teach us something are the ones where we wrote down what we expected and were wrong." - Vezert conversion lead

What A/B Testing Costs to Run

Tooling is the smallest line. A serviceable testing tool runs €0 to €200 a month at the volumes most B2B sites have, and the free tier of several is enough for one test at a time. The real cost is the build and the analysis.

Budget three to eight hours per test for implementation on a normal page, plus an hour to read the result properly. At agency rates that is €400 to €1 200 per test, and a programme that runs one test a month sits around €1 000 to €2 500 including the tool and the reporting. Below roughly 10 000 sessions a month on your key pages, that money returns more if it goes to traffic instead.

One cost that never appears in the quote is the decision time. A test produces a result that somebody has to act on, and in most organisations that meeting is harder to schedule than the build. Programmes die from unread results far more often than from bad tooling.

Where to Start

If your key page clears 10 000 sessions a month, start with the removal test from the table and write the expected number down first. If it does not, spend the same budget on traffic and use a CRO audit to find the problems that do not need a test to be obvious.

Most of what a first audit finds is not testable anyway. It is broken, and broken things get fixed rather than tested. Our conversion optimization work separates the two before anything gets built.

Related Articles

Explore more articles on similar topics to deepen your understanding

Explore All Articles

Frequently Asked Questions

Find answers to common questions about this topic