Definition
Subject line testing, also known as split testing or A/B testing, is the practice of sending two or more subject line variants to randomly assigned segments of a subscriber list to determine which version yields the highest open rate. The variant that performs best is then sent to the remaining subscribers. Subject line testing removes guesswork from copywriting decisions and provides data-driven evidence for what resonates with a specific audience.
Statistical validity in subject line testing requires adequate sample size. For a typical campaign with an expected open rate of twenty per cent and a minimum detectable effect of ten per cent relative improvement, each variant needs approximately fifteen thousand recipients to achieve ninety-five per cent confidence with eighty per cent statistical power. Smaller sample sizes risk false positives that lead to incorrect conclusions. Test duration is equally important: a minimum of four to six hours is recommended, and tests should run until each variant receives at least one hundred opens to ensure stable results.
Best Practices
Test one variable at a time in subject line experiments. Combining length, personalisation, emoji usage, and urgency cues in a single test makes it impossible to attribute the result to any specific element. Isolate variables across sequential tests: run a length test first, then a personalisation test, then an emoji test.
Run subject line tests for a minimum of four to six hours to capture variations in open behaviour across different times of day and time zones. Avoid ending tests prematurely because early results can be misleading. Use your ESP's built-in testing functionality to automate winner selection and holdout management.
Consider the metric holistically: a higher open rate is only valuable if it does not reduce click-through rate or increase unsubscribe rate. A subject line that tricks subscribers into opening but fails to deliver on its promise damages trust and ultimately reduces long-term engagement.
Use the results to build a subject line playbook for your brand. Document which variables perform best, by what magnitude, and note any segment-level differences. A structured playbook reduces future decision-making time and maintains consistent performance across campaigns and team members.
Related Glossary Terms
Email Campaign Timing
Campaign timing strategy and optimisation including optimal send time analysis, time-based triggers, timing A/B testing methodology, and send time personalisation.
Email Conversion Optimisation
Conversion rate optimisation (CRO) specifically for email traffic focuses on landing page alignment, CTA testing, and friction reduction to maximise email click-through conversion.
Email Experimentation Framework
An email experimentation framework structures multivariate testing, factorial designs, and Bayesian statistical approaches to systematically improve campaign performance.
Email Experimentation
Email experimentation applies the scientific method to email marketing with hypothesis-driven A/B and multivariate testing. Statistical significance at 95% confidence is standard.
Email Preview Text Optimisation
Preview text optimisation goes beyond the basics to address character limits across email clients, emoji use, brand reinforcement, and A/B testing methodology.
Email Read Rate
Measurement of actual email content consumption distinct from open rate, including read time analysis, read-through rate calculation, and comparative analysis.
Frequently Asked Questions
Test two variants for most campaigns. Testing more than three variants reduces the sample size available for each variant and requires a much larger total list size to maintain statistical significance. Two variants provide a clear winner while preserving statistical power.
A minimum of fifteen thousand recipients per variant is recommended for typical open rates. Smaller lists can still test but should expect wider confidence intervals and a higher risk of inconclusive or misleading results.
Run subject line tests for a minimum of four to six hours. Longer durations of twelve to twenty-four hours are preferable because they account for time zone variations and different open behaviour patterns throughout the day.
Yes, but you must accept lower statistical confidence. For lists under five thousand, consider sequential testing: send one subject line to the whole list this week and a different one next week, comparing results while controlling for day-of-week effects.
Start with the variables most likely to affect your open rate: length (short versus long), personalisation (with or without first name), and emotional appeal (curiosity versus urgency versus benefit-driven). Test these one at a time across separate campaigns.