Definition
A/B testing (also called split testing) is a method of comparing two versions of an email to determine which one achieves a higher performance on a specific metric. In email marketing, A/B tests are commonly used to optimise subject lines, sender names, preview text, content layout, calls-to-action, images, and send times.
The standard approach is to send version A (control) and version B (variant) to a small percentage of your list — typically 10-20% each — while holding back the remaining 60-80%. Once the test reaches statistical significance, the winning version is sent to the holdout group.
Formula
Statistical significance for A/B tests is typically calculated using a chi-squared test or a z-test for proportions. The standard threshold is 95% confidence (p < 0.05).
Z = (pA - pB) / √(p(1-p)(1/nA + 1/nB))
| Variable | Description |
|---|---|
| pA | Conversion rate of variation A |
| pB | Conversion rate of variation B |
| p | Pooled conversion rate across both variations |
| nA | Sample size of variation A |
| nB | Sample size of variation B |
Average Benchmark
| Test Element | Typical Uplift Range |
|---|---|
| Subject line | 5% - 25% improvement in open rate |
| CTA button text | 10% - 30% improvement in CTR |
| Sender name | 3% - 15% improvement in open rate |
| Preview text | 2% - 10% improvement in open rate |
| Send time | 3% - 12% improvement in engagement |
A well-structured A/B testing program typically yields 10-20% improvement in key metrics over a 6-month period as learnings compound.
How to Improve A/B Testing Results
- Test one variable at a time: Testing multiple differences between A and B makes it impossible to know which change caused the result. Isolate a single variable per test.
- Use a statistically significant sample: Most email platforms require at least 1,000 opens or clicks per variation to reach significance. Smaller samples produce unreliable results.
- Let the test run to completion: Do not declare a winner early. Set a minimum test duration of 4-24 hours depending on your send volume and engagement rate.
- Document and reuse learnings: Build a testing roadmap based on past results. Subject line tests are cheap and fast; content and layout tests require more volume but produce deeper insights.
- Test consistently not occasionally: Run tests on every campaign when possible. The cumulative effect of many small improvements is greater than occasional large tests.
Example Calculation
If you send subject line A to 2,000 subscribers and get 400 opens (20%), and subject line B to 2,000 subscribers and get 460 opens (23%):
Observed Difference = 23% - 20% = 3 percentage points
Using a statistical significance calculator with these inputs would determine whether the 3-point difference is reliable at 95% confidence. With these sample sizes and rates, the result would likely be statistically significant (p < 0.05), meaning subject line B is the clear winner.
Related Glossary Terms
Abandoned Cart Email
An abandoned cart email is an automated message sent to customers who added items to their online shopping cart but left without completing the purchase. It is one of the highest-converting email types in ecommerce.
AMP for Email
AMP for Email is a Google-developed framework that allows email messages to include interactive elements like forms, carousels, accordions, and live content. It turns static emails into dynamic, interactive experiences directly inside the inbox.
Bounce Rate
Email bounce rate is the percentage of emails that were rejected by the receiving server before reaching the recipient. It is a key indicator of list health and data quality.
CAN-SPAM Act
The CAN-SPAM Act is a US law that sets rules for commercial email. It requires accurate subject lines, a physical address, a clear opt-out mechanism, and prompt processing of unsubscribes. Violations can result in penalties up to $51,744 per email.
Click-Through Rate
Click-through rate (CTR) is the percentage of email recipients who clicked one or more links in your email campaign. It measures how compelling your content and call-to-action are.
Click-to-Convert Rate
Click-to-convert rate measures the percentage of email clicks that result in a desired conversion action such as a purchase, signup, or download. It shows how effective your post-click experience is at turning interest into results.
Frequently Asked Questions
Run the test long enough to reach statistical significance, but not so long that timing effects skew the results. For most email campaigns, 4-6 hours is sufficient for high-volume lists (100,000+), while 12-24 hours is needed for smaller lists. Avoid testing across different days because open behaviour varies by day of the week.
The required sample size depends on the expected effect size and your baseline metric. For subject line tests, you typically need 1,000-3,000 opens per variation. For click tests, you need 500-2,000 clicks per variation. Most email platforms recommend testing with at least 10% of your list per variation.
Yes, this is called multivariate or A/B/n testing. However, each additional variation requires a larger sample size to maintain statistical power. For most email marketers, testing more than 3-4 variations simultaneously is impractical unless you have a very large list. Stick with 2-3 variations for reliable results.
Start with the elements that have the highest impact with the lowest effort: subject lines, sender names, and preview text. These are easy to change and directly affect open rates. Next, test CTA buttons, offers, and content layout. Test send times last, since the improvements are typically smaller and harder to measure reliably.
Most email platforms display a confidence level or significance score with their A/B testing results. A confidence level of 95% or higher (p < 0.05) means there is a 95% chance the difference is real and not due to random variation. If your platform does not calculate this automatically, use an online A/B test significance calculator — they are free and widely available.