A/B testing, also called split testing, is a controlled experiment that splits your traffic randomly between the current page and one changed version, then lets statistics decide which converts better. It replaces the loudest opinion in the room with evidence from real visitors. The method is simple; the discipline is not, and most failed tests fail on discipline rather than tooling. They start without a hypothesis, run without enough traffic, or get stopped the moment the numbers look good. This post walks the whole process, from hypothesis to rollout: which type of test fits which job, the sample size math in plain numbers, why peeking fakes winners and which tools tolerate it, and what to test first on each type of page.
The short version
- Every test starts as a falsifiable hypothesis written from data, not a brainstorm.
- Set the sample size before launch. It decides how long the test runs, not your patience.
- Peeking at a fixed-horizon test and stopping early fakes winners. Know which statistics your tool runs.
- Test where traffic and money meet first: landing page headlines, checkout friction, pricing display.
- Only about 1 in 7 tests produces a clear winner, and that rate is normal. Losers still teach.
- 1 in 7 tests produces a significant winner
- 95% confidence is the industry standard
- 1 week minimum run time, whatever the traffic
- 2x the per-variation sample = total test traffic
What A/B testing is and which kind of test fits the job
The classic A/B test changes one element, splits traffic between the original and the variant, and measures a single primary metric. Use it for headlines, calls to action, images, forms, and pricing display. Two siblings exist for special cases. Split URL testing sends traffic to entirely different pages, which fits redesign validation where an element swap cannot express the change. Multivariate testing combines several changed elements to find interactions, but it needs roughly ten times the traffic, so it belongs on pages with 50,000 or more monthly visitors. When in doubt, run the plain A/B test. It answers one question cleanly, and a series of clean answers beats one muddled experiment every time.
A test is only as strong as its hypothesis
Testing random ideas is how programs stall. A real test starts from a problem the data already demonstrated: a checkout step losing 60% of entrants, a landing page where recordings show visitors hunting for pricing, a form abandoned at the phone-number field. Analytics finds the leak, heatmaps and session recordings suggest why, and the hypothesis turns that why into a falsifiable sentence. The format worth memorizing has three parts: if we change X, metric Y will move by roughly Z, because of a stated reason. "If we change the CTA from Submit to Get My Free Quote, form submissions rise 15%, because visitors will understand what they receive." Written that way, even a losing test teaches you something about the audience, because the reason was either confirmed or contradicted. A test without a hypothesis can only tell you what happened, never why, and the why is what compounds across a testing program.
Sample size honesty happens before launch
Before the test goes live, feed three numbers into any sample size calculator: your baseline conversion rate, the smallest lift you would act on (the minimum detectable effect), and the standard 95% confidence with 80% power. The output is the visitors each variation needs, and it is usually larger than people hope.
| Baseline conversion | Lift you want to detect | Visitors per variation |
|---|---|---|
| 2% | 10% | ~15,700 |
| 2% | 20% | ~3,900 |
| 2% | 50% | ~630 |
| 5% | 10% | ~5,900 |
| 5% | 20% | ~1,500 |
| 5% | 50% | ~240 |
Figures assume 95% confidence and 80% power; double them for total test traffic. Read the table backwards and it explains low-traffic strategy: a site with 3,000 monthly visitors cannot detect a 10% lift in any reasonable time, so it should test bold changes capable of 50% swings, accept longer runs, or spend its energy on fixes that need no test at all, like page speed. Pretending the math away does not make a small sample significant; it makes the result a coin flip with a dashboard.
Peeking is the mistake that fakes winners
The classic testing statistics are fixed-horizon: they assume you pick a sample size, run to it, and read the result exactly once. Checking the dashboard daily and stopping the moment significance appears breaks that assumption. Each look is another chance for random noise to cross the line, so a team that peeks every day at a stated 5% false positive rate is really running something closer to 20% or worse. This is why so many "winners" evaporate after rollout: they were noise that got caught at a lucky moment.
The modern tools changed this, and it matters which kind you run. Optimizely, VWO, GrowthBook, and Statsig now use sequential or Bayesian statistics designed for continuous monitoring, where the numbers on screen stay valid whenever you look. Fixed-horizon calculators and older setups do not. The worst combination in practice is a fixed-horizon test read like a sequential one: a calculator that assumed one look, a dashboard checked daily, and a stop decision made at the first green number. Know which statistics your platform runs, and if the answer is fixed-horizon, set the sample size, wait, and read it once.
What to test first on each type of page
Rank test ideas by traffic times proximity to money. A 10% lift on a page with 50,000 monthly visitors is worth more than a 50% lift on one with 500, and a lift next to the buy button is worth more than one three pages upstream. Within a page, the starting element differs by template.
- Landing pages. Headline first, CTA second. A landing page exists to move one number, and the headline plus button pair carries most of its persuasion. Message match with the ad that sent the click is the highest-leverage version of this test.
- Product pages. Image set and proof placement. Lifestyle versus plain product shots, review position, and shipping or returns messaging near the add-to-cart button decide more purchases than copy edits further down.
- Pricing pages. Tier order, plan naming, which plan is highlighted, and whether annual billing is the default toggle. Small presentation changes here move revenue directly because every visitor is late-funnel.
- Checkout and forms. Field count, guest checkout, error message clarity, and trust signals beside the payment step. Removing friction usually beats adding persuasion this deep in the funnel.
- Content pages. CTA format and placement. HubSpot found anchor-text CTAs inside blog posts beat button CTAs by 121% for lead generation, because the less promotional format read as part of the content.
The tool reality after Google Optimize
Google shut down Optimize and Optimize 360 on September 30, 2023, and never replaced them. GA4 has no native testing feature; it is the measurement layer, not the experiment engine. The working pattern in 2026 is a dedicated platform running the experiment while pushing variant data into GA4 as event parameters or audiences, so you can read test results against revenue, engagement, and everything else GA4 already tracks. Wiring that connection correctly is half the setup work, and it is the half our Google Analytics service exists for. Two more realities shape tool choice. Client-side tools inject changes in the browser and can flicker the original page before the variant paints, which both annoys visitors and contaminates results; server-side and edge testing avoid it at the cost of developer time. And consent banners shrink your measurable sample, since visitors who decline analytics vanish from the results, so calculate sample sizes from consented traffic rather than raw sessions.
Run the test and read it like a skeptic
One primary metric decides the test; choose it before launch and let nothing dethrone it afterward. Guardrail metrics ride along to catch collateral damage: revenue per visitor, bounce rate, page speed, and support contacts make sure a winner is not quietly hurting something the primary metric cannot see. Run at least one full week regardless of traffic so weekday and weekend behavior both count, and avoid reading anything into the first 48 hours, when returning visitors react to novelty rather than merit. When the test completes, read segments before declaring: a flat overall result can hide a mobile win canceled by a desktop loss, and paid traffic often responds differently than organic. Then document everything, winner or loser. The hypothesis, the result, the segment quirks. A written test log is the difference between a program that compounds and a team that retests the same button every eighteen months.
Testing and SEO get along fine
Google explicitly permits A/B testing and runs thousands of its own. Three rules keep you safe: canonical split URL variants to the control, never show Googlebot different content than visitors see, and end experiments once they answer instead of letting a 50/50 split run forever. Every mainstream platform ships these defaults.
Testing is one leg of a bigger process
The wins come when research picks the targets. The full conversion rate optimization cycle we run at Egochi pairs quantitative analysis and user research with the experiments, so every test attacks a friction point visitors actually demonstrated. The engagement problems that surface along the way, like the ones covered in our post on reducing bounce rate, usually become the first test backlog.
Questions people ask about A/B testing
How long should an A/B test run?
Until it reaches the sample size you set before launch, and never less than one full week so day-of-week behavior evens out. Most sites land at 2 to 4 weeks; low-traffic sites at 6 to 8. Duration is a function of traffic and effect size, not patience.
How much traffic do I need for A/B testing?
It depends on your baseline conversion rate and the lift you want to detect. Detecting a 20% lift on a 5% conversion rate needs roughly 1,500 visitors per variation; a 10% lift on a 2% rate needs closer to 16,000 per side. Low-traffic sites should test bigger swings.
What is statistical significance in A/B testing?
The probability your result is not random chance. The 95% standard means only a 5% chance the observed difference happened by luck, assuming you read the result once at the planned end. It does not mean version B is 95% better, only that a real difference likely exists.
Is it ever safe to check A/B test results early?
Only if your platform runs sequential or Bayesian statistics built for continuous monitoring, as Optimizely, VWO, GrowthBook, and Statsig do. With a classic fixed-horizon test, every early look at which you might stop inflates false positives well past the stated 5%.
What replaced Google Optimize?
Google shut down Optimize and Optimize 360 on September 30, 2023, and GA4 has no built-in testing feature. Former users moved to VWO, Optimizely, Convert, AB Tasty, and open-source options like GrowthBook, all of which can send experiment data into GA4 for analysis.
Does A/B testing hurt SEO?
Not when implemented properly. Google explicitly permits testing. Use rel="canonical" to the control URL for split URL tests, never show Googlebot different content than users, and end experiments once they answer. Modern platforms handle these defaults for you.