Definition
Email experimentation is the practice of applying the scientific method to email marketing decisions. The process follows a five-step cycle: formulate a hypothesis based on data or observation, predict the expected outcome, design and run an experiment, analyse the results, and make a decision to adopt, reject, or iterate on the change. This structured approach replaces opinion-based decisions with evidence-based optimisation and compounds learning over time through repeated experimentation.
Test variables in email experimentation span subject lines, preheader text, sender names, content layout, imagery, call-to-action placement, personalisation depth, send time, and audience segments. Sample size calculation before the experiment ensures the test has sufficient statistical power to detect the expected effect size. The industry standard confidence level is 95 per cent, meaning there is a 5 per cent probability that the observed result is due to random chance. Experimentation velocity — the number of tests completed per month — is a key metric for programme maturity. An experimentation maturity model progresses from ad hoc testing (no structured process) through basic A/B testing (one variable at a time) to multivariate testing and finally to automated continuous experimentation.
Best Practices
Write a clear hypothesis before every experiment using the format: "If we [change], then [metric] will [increase/decrease] because [reason]." A well-formed hypothesis includes the independent variable, the dependent variable, the predicted direction of change, and the rationale. This structure forces clarity and makes the test results interpretable.
Calculate the required sample size before launching the experiment. Use a sample size calculator with inputs for baseline conversion rate, minimum detectable effect, desired statistical power (80 per cent standard), and confidence level (95 per cent standard). Running an experiment with insufficient sample size wastes time and produces inconclusive results.
Test one variable at a time in A/B tests to isolate the cause of any performance difference. Multivariate testing tests multiple variables simultaneously but requires larger sample sizes and more complex analysis. Start with simple A/B tests and progress to multivariate testing only when you have sufficient traffic and testing experience.
Run experiments for a predetermined duration and do not stop early based on intermediate results. Early stopping biases results and invalidates statistical conclusions. Use a fixed sample size or a fixed duration calculated before the experiment begins. The only exception is stopping for ethical or business damage concerns.
Document every experiment including the hypothesis, design, sample size, duration, results, and decision. A testing log creates an institutional knowledge base that prevents repeating failed experiments and accelerates learning. Review the testing log quarterly to identify patterns across successful and unsuccessful experiments.
Related Glossary Terms
Control Group
Control or holdout group testing withholds a random subscriber segment from a campaign to measure incremental lift in engagement, revenue, and conversion.
Email Frequency Engagement
Email frequency vs engagement analysis finds the optimal send frequency per subscriber. Over-sending increases unsubscribes and dormancy; under-sending diminishes brand recall. Testing methodology identifies the frequency sweet spot.
Email Landing Page
Email landing pages align post-click experience with email messaging using dedicated pages per campaign to optimise conversion rates typically between 2 and 5 per cent.
Email Preview
Email preview tools test rendering across 100+ email clients, providing spam scoring, code analysis, accessibility checks and collaboration features for quality assurance.
Email QA
The process of validating email campaigns before send to catch rendering, content, link, and deliverability issues.
Seed List
Seed list testing uses planted test addresses on subscriber lists to monitor inbox placement rates across major mailbox providers including Gmail, Outlook, and Yahoo.
Frequently Asked Questions
Run the experiment until the predetermined sample size or duration is reached. For a typical email campaign with 100,000 subscribers and a 50/50 split, the experiment runs until 50,000 subscribers per variant receive the email and sufficient conversions accumulate. This usually takes one to seven days depending on send frequency.
Minimum sample size depends on the baseline conversion rate and the minimum effect you want to detect. For a baseline conversion rate of 2 per cent and a minimum detectable effect of 10 per cent relative lift, you need approximately 100,000 to 150,000 total recipients at 80 per cent power and 95 per cent confidence.
Use multivariate testing when you have high traffic (500,000+ recipients per test), you need to test interactions between multiple variables, and you have the analytical capability to interpret complex results. For most programmes, a series of sequential A/B tests is more practical and produces clearer learnings.
The experimentation maturity model has five stages: ad hoc (no formal testing), foundational (basic A/B tests with simple metrics), developing (consistent testing with statistical rigour), advanced (multivariate tests, personalisation experiments, automated testing), and optimising (continuous experimentation, AI-driven test design, automated decision-making).
Inconclusive results are valuable data points. Document the hypothesis, the observed results (even if not statistically significant), and possible reasons for the null result such as insufficient sample size, too small an effect, or an incorrect hypothesis. Use the learning to refine the next experiment rather than repeating the same test.