Digital Marketing

A/B testing and how to run split tests that pay for themselves

A/B Testing cover with the title in gold on black and a chrome humanoid robot holding out an open palm against a yellow panel

A/B testing, also called split testing, is a controlled experiment that shows two versions of a page, email, or ad to randomly split traffic and lets statistics decide which one converts better. It replaces the loudest opinion in the room with evidence from real visitors. The method is simple; the discipline is not, and most failed tests fail on discipline: stopped early, changed too little, or measured wrong. This post covers the process, the math, and the traps.

The short version

  • Split traffic randomly between a control and one changed variant.
  • Set the sample size before you start and do not peek early.
  • Only about 1 in 7 tests produces a clear winner, and that is normal.
  • Test big, meaningful changes on high-traffic pages first.
  • 1 in 7 tests produces a significant winner
  • 95% confidence is the industry standard
  • 10-25% typical lift from a winning test
  • 1 week minimum run time, whatever the traffic

What A/B testing is and when to use which type

The classic A/B test changes one element, splits traffic between the original and the variant, and measures a single primary metric. Use it for headlines, calls to action, images, forms, and pricing display. Two siblings exist for special cases. Split URL testing sends traffic to entirely different pages, which fits redesign validation where an element swap cannot express the change. Multivariate testing combines several changed elements to find interactions, but it needs roughly ten times the traffic, so it belongs on pages with 50,000 or more monthly visitors. If in doubt, run the plain A/B test.

What to test first

Rank test ideas by traffic times potential impact. A 10% lift on a page with 50,000 monthly visitors is worth more than a 50% lift on one with 500. The elements with the strongest track record: headlines, since they are the first thing read and the biggest lever on whether anyone keeps reading; call-to-action text and placement; hero images and video; form length; pricing presentation; and social proof position. On a landing page, the headline plus CTA pair usually holds the most upside, because the whole page exists to move one number.

How to run a test that produces a real answer

  1. Find the problem in the data

    Start from analytics, heatmaps, and session recordings, not brainstorms. The best tests attack a specific measured drop-off, like a checkout step losing 60% of entrants.

  2. Write a falsifiable hypothesis

    "If we change the CTA from Submit to Get My Free Quote, form submissions rise 15%, because visitors will understand what they receive." A hypothesis makes even a losing test teach you something.

  3. Calculate the sample size before launch

    Feed your baseline rate, minimum detectable effect, and 95% confidence into a sample size calculator. This single step prevents the most common failure, which is stopping the moment the early numbers look good.

  4. Build one meaningful change

    Change one thing, and change it enough to matter. Testing two shades of the same red produces noise; testing a benefit-led headline against a feature-led one produces an answer.

  5. Track a primary metric and guardrails

    One conversion goal decides the test. Secondary metrics like revenue per visitor and bounce rate make sure the winner is not quietly damaging something else.

  6. Run to completion without peeking

    Checking results repeatedly and stopping at a good moment inflates false positives well past the stated 5%. Set it, wait, then read it once.

  7. Read segments before declaring

    A flat overall result can hide a mobile win and a desktop loss. Break results down by device, source, and new versus returning before you ship anything.

  8. Implement, document, repeat

    Winners become the new control. Losers become documented learning. Either way the next hypothesis starts from more knowledge than the last.

The sample size math in plain numbers

Baseline conversionLift you want to detectVisitors per variation
2%10%~15,700
2%20%~3,900
2%50%~630
5%10%~5,900
5%20%~1,500
5%50%~240

Figures assume 95% confidence and 80% power; double them for total test traffic. Read the table backwards and it explains low-traffic strategy: small sites cannot detect small lifts, so they should test bold changes capable of 50% swings and accept longer runs.

Tests that made the case studies

  • HubSpot. Anchor-text CTAs inside blog posts beat button CTAs by 121% for lead generation. The less promotional format read as part of the content.
  • Booking.com. Real-inventory urgency lines like "Only 2 rooms left at this price" lifted bookings; the company famously runs about 1,000 concurrent experiments.
  • Obama 2008. "Learn More" beat "Sign Up" by 18.6%, a family photo beat headshots, and the combined winner lifted signups 40.6%, an estimated $60 million in added donations.
  • Humana. Stripping clutter from a promotional banner and sharpening one CTA lifted click-through 433%. Removing elements is also a test.

Testing is one leg of a bigger process

Testing random ideas produces mediocre results; the wins come when research picks the targets. The full conversion rate optimization cycle we run at Egochi pairs quantitative analysis and user research with the experiments, so every test attacks a friction point users actually demonstrated. Engagement problems found along the way, like the ones covered in our post on reducing bounce rate, often become the test backlog.

Questions people ask about A/B testing

How long should an A/B test run?

Until it hits your pre-set sample size at statistical significance, and never less than one full week so day-of-week behavior evens out. In practice that means 2 to 4 weeks for most sites and 6 to 8 for low-traffic ones. Duration is a function of traffic and effect size, not patience.

How much traffic do I need for A/B testing?

It depends on your baseline rate and the lift you want to detect. Detecting a 20% lift on a 5% conversion rate needs roughly 1,500 visitors per variation; a 10% lift on a 2% rate needs closer to 16,000 per side. Low-traffic sites should test bigger swings.

What is statistical significance in A/B testing?

The probability your result is not random chance. The 95% standard means only a 5% chance the observed difference happened by luck. It does not mean version B is 95% better, only that you can be 95% confident a real difference exists.

What happened to Google Optimize?

Google discontinued Optimize and Optimize 360 on September 30, 2023. Former users have moved to VWO, Optimizely, Convert, and AB Tasty, which offer the same core capability with active development.

Does A/B testing hurt SEO?

Not when implemented properly. Google explicitly permits testing. Use rel="canonical" to the control URL for split URL tests, never show Googlebot different content than users, and do not run experiments indefinitely. Modern platforms handle these defaults for you.

Written by , Head of Advertising & Content Marketing at Egochi. Every post on this blog comes from the person who runs that work for clients, not a content mill.

Want this handled for you?

Egochi is a US digital marketing agency working with local businesses through enterprise brands from offices in New York, Miami, Milwaukee, and Madison. Tell us what you are trying to grow and we will send back a plan with real numbers in it.

Get a Free Proposal Call (888) 644-7795

Grade Your Website in About 30 Seconds

Egochi's free audit scores any page for technical SEO, content, and AI search readiness. The report renders on screen, and an analyst reviews every run.