A/B testing compares two versions of a marketing element to determine whether one produces a stronger result against a defined objective. It can be used to evaluate advertising, landing pages, email, offers, messaging, calls to action, creative, and other marketing experiences when the test is designed carefully enough to support a meaningful comparison.
The AAMA A/B Testing Guide for Marketers provides a practical framework for planning, running, interpreting, and documenting marketing experiments. It is designed for advertising and marketing professionals, agencies, brands, students, academics, analysts, and organizations that want to replace assumptions with stronger evidence.
A/B testing is most useful when the experiment begins with a clear hypothesis, isolates a meaningful variable, uses appropriate measurement, and runs long enough to produce interpretable results. A test should answer a specific question rather than simply generate another dashboard.
Start With the Decision
Before deciding what to test, define the decision the result is intended to support. A/B testing is useful when two or more plausible alternatives exist and the organization needs evidence to choose between them.
Useful testing questions might include:
- Which headline produces more qualified registrations?
- Does a shorter form improve completion?
- Which offer generates more purchases?
- Does a new product image increase conversion?
- Which email subject line produces more clicks?
- Does changing the call to action improve response?
- Which advertisement generates lower customer-acquisition cost?
- Does a revised landing page improve lead quality?
The decision should be meaningful enough to justify the time and traffic required to conduct the experiment.
Define the Objective
Every test should have one primary objective. The objective determines the metric used to evaluate the result.
Possible objectives include:
- Increase conversion rate
- Generate more qualified leads
- Increase purchases
- Reduce cost per acquisition
- Improve click-through rate
- Increase email clicks
- Improve registration completion
- Increase average order value
- Reduce abandonment
- Increase trial starts
Avoid choosing a metric simply because it is easy to observe. The metric should connect directly to the behavior the organization wants to influence.
Develop a Hypothesis
A hypothesis states what change is being tested, what outcome is expected, and why that outcome might occur.
A useful structure is:
If we change [element], then [metric] will improve because [reason].
For example:
“If we replace the generic headline with a benefit-focused headline, landing-page conversion rate will increase because visitors will understand the value of the offer more quickly.”
The hypothesis forces the team to explain the reasoning behind the test rather than simply trying random variations.
Define the Control
The control is the existing or baseline version against which another version is compared. It is commonly referred to as Version A.
The control might be:
- Current webpage
- Existing advertisement
- Current email subject line
- Existing offer
- Current checkout process
- Existing call to action
The control establishes the performance level the new variation needs to improve upon.
Define the Variation
The variation is the alternative version being tested. It is commonly referred to as Version B.
The variation should reflect the hypothesis being evaluated.
For example:
Version A: “Marketing Analytics Platform”
Version B: “See Which Campaigns Are Actually Driving Revenue”
If the hypothesis concerns benefit-focused messaging, the difference between the two versions should primarily reflect that messaging change.
Test One Major Variable at a Time
A strong A/B test isolates the variable being evaluated whenever practical.
If Version A and Version B use different headlines, images, calls to action, page layouts, offers, and colors, it may be impossible to determine which difference caused the performance change.
Potential variables include:
- Headline
- Image
- Offer
- Call to action
- Button language
- Form length
- Price presentation
- Subject line
- Landing-page layout
- Message
- Video thumbnail
- Advertisement format
Testing one major variable at a time produces clearer learning.
Understand Multivariate Testing
Multivariate testing evaluates combinations of several variables at the same time. It can help determine which combination of elements performs best when sufficient traffic and technical capability are available.
For example, a test might compare several headlines and several images in different combinations.
Multivariate testing usually requires substantially more traffic than a simple A/B test because the audience is divided among more combinations. It should not be used simply because it sounds more sophisticated.
Prioritize High-Impact Tests
Not every element deserves the same testing priority. Changing the color of a small icon may be easy to test but unlikely to materially affect business performance.
Higher-value testing opportunities often involve:
- Value proposition
- Offer
- Price presentation
- Audience
- Landing-page structure
- Product positioning
- Call to action
- Form requirements
- Creative concept
- Major campaign message
- Checkout friction
Prioritize tests where a meaningful improvement could affect an important objective.
Consider Test Effort
Testing opportunities should also be evaluated according to the effort required to implement them.
A useful prioritization approach considers:
- Potential impact
- Confidence in the hypothesis
- Implementation effort
- Traffic requirements
- Business importance
- Risk
A simple headline test may be worth running quickly, while a complete checkout redesign may require stronger evidence before development resources are committed.
Choose the Primary Metric
The primary metric determines which version wins the test.
Examples include:
- Conversion rate
- Purchase rate
- Lead completion rate
- Revenue per visitor
- CPA
- CAC
- Email CTR
- Registration rate
- Trial-start rate
Define the metric before the test begins.
Changing the success metric after seeing the results can lead teams to select whichever number makes the preferred version appear successful.
Define Secondary Metrics
Secondary metrics provide additional context but should not replace the primary test objective.
For example, a landing-page test might use:
Primary Metric: Completed registrations
Secondary Metrics: CTR, bounce rate, time on page, form abandonment
Secondary metrics may help explain why a result occurred or identify unintended effects.
Identify Guardrail Metrics
Guardrail metrics help ensure that an improvement in one area does not create a serious problem somewhere else.
For example, a more aggressive offer might increase conversion rate while reducing average order value or increasing cancellations.
Possible guardrail metrics include:
- Refund rate
- Cancellation rate
- Average order value
- Lead quality
- Customer-support volume
- Unsubscribe rate
- Complaint rate
- Profit margin
A winning test should improve the intended outcome without causing unacceptable damage elsewhere.
Define the Audience
Determine who should participate in the test.
Possible audiences include:
- All website visitors
- New visitors
- Returning visitors
- Existing customers
- Mobile users
- Desktop users
- Paid-media traffic
- Organic traffic
- Specific geographic markets
- Particular customer segments
The audience should match the decision the test is intended to support.
A result from existing customers may not necessarily apply to first-time visitors.
Randomize Assignment
Participants should generally be assigned randomly to test versions so the groups are as comparable as possible.
Random assignment helps reduce the likelihood that differences in audience composition explain the result.
Without randomization, Version A might receive more high-intent customers while Version B receives more casual visitors, creating a misleading comparison.
Keep Participants in the Same Version
When possible, a participant should continue seeing the same version throughout the test.
If someone sees Version A on one visit and Version B on another, the experience may become inconsistent and the test may be harder to interpret.
Most testing systems use cookies, identifiers, accounts, or other methods to maintain consistent assignment.
Avoid Overlapping Experiments
Running several experiments on the same audience at the same time can create interactions that make results difficult to interpret.
For example, a landing-page headline test and an offer test running simultaneously may affect one another.
Large organizations may have systems designed to manage overlapping experiments, but smaller teams should generally avoid unnecessary overlap when the same audience and conversion path are involved.
Estimate the Required Sample Size
A test needs enough observations to distinguish a meaningful effect from normal variation.
The appropriate sample size depends on factors such as:
- Current conversion rate
- Minimum improvement worth detecting
- Desired confidence level
- Statistical power
- Traffic volume
- Number of variations
Tests with low conversion rates or small expected improvements usually require larger samples.
Do not assume that a few dozen clicks or conversions are enough to establish a reliable result.
Define the Minimum Detectable Effect
The minimum detectable effect, often abbreviated MDE, is the smallest performance difference the test is designed to detect reliably.
For example, if the current conversion rate is 5%, the organization might decide that only an improvement of at least 10% relative to the baseline would justify changing the page.
Defining the MDE helps determine whether the test has enough traffic to answer a business-relevant question.
Understand Statistical Significance
Statistical significance helps evaluate whether an observed difference would be unlikely under a specified statistical model if no real difference existed.
A statistically significant result does not automatically mean the difference is large or important.
A very small difference can become statistically significant with enough data, while a potentially meaningful difference may fail to reach significance when the sample is too small.
Statistical evidence should be considered alongside practical business value.
Understand Statistical Power
Statistical power refers to the probability that a test will detect an effect of a specified size when that effect actually exists.
Low-powered tests may miss real differences and produce inconclusive results.
Power depends on the sample size, baseline performance, effect size, significance threshold, and test design.
Avoid Peeking & Stopping Early
One of the most common A/B testing mistakes is repeatedly checking the results and stopping the experiment as soon as one version appears to be winning.
Early results can fluctuate substantially because the sample is still small.
If the test was designed around a specific sample size or duration, follow that plan unless there is a legitimate operational or ethical reason to stop.
Stopping whenever a preferred result appears can substantially increase the risk of false conclusions.
Run the Test Through Normal Business Cycles
Tests should generally run long enough to account for meaningful patterns in customer behavior.
Depending on the business, performance may vary by:
- Day of week
- Weekend versus weekday
- Payday
- Time of day
- Season
- Promotional period
- Media schedule
A test that runs only on Monday and Tuesday may not represent behavior during the rest of the week.
Where practical, include complete business cycles in the test duration.
Avoid Running Tests During Unusual Conditions
Major external events can distort test results.
Examples may include:
- Site outages
- Major promotions
- Pricing changes
- Holiday periods
- Product shortages
- Major media launches
- Public controversies
- Technical problems
If unusual conditions affect one part of the test disproportionately, the results may not generalize to normal conditions.
Document significant external events during the test.
Validate the Technical Setup
Before launching, confirm that the experiment is functioning correctly.
Check:
- Version assignment
- Tracking
- Conversion events
- URLs
- Page loading
- Forms
- Mobile behavior
- Analytics
- Revenue tracking
- Audience eligibility
- Version persistence
A technically broken experiment can generate precise-looking numbers that measure the wrong thing.
Conduct an A/A Test When Appropriate
An A/A test compares two identical experiences.
The purpose is not to improve performance but to validate whether the testing and measurement system behaves as expected.
A/A testing can help identify:
- Tracking problems
- Uneven traffic allocation
- Statistical-method issues
- Platform inconsistencies
It is most useful when establishing a new experimentation system or investigating suspicious results.
Launch the Experiment
Once the hypothesis, audience, versions, metrics, sample requirements, and tracking are established, launch the experiment.
Document the start date and avoid unnecessary changes after launch.
If creative, targeting, pricing, tracking, or another important factor must change during the experiment, record the change and consider whether the test should be restarted.
Monitor for Technical Problems
Monitoring should focus primarily on whether the experiment is operating correctly.
Check for:
- Broken pages
- Missing conversions
- Incorrect allocation
- Tracking failures
- Unexpected traffic sources
- Large performance anomalies
- Device-specific problems
Avoid treating normal short-term performance fluctuations as reasons to alter the experiment.
Compare the Primary Outcome
After the planned sample or duration has been reached, evaluate the primary metric.
Ask:
- Which version performed better?
- How large was the difference?
- Is the result statistically credible?
- Is the difference large enough to matter?
- Did guardrail metrics remain acceptable?
- Was the audience representative?
- Were there unusual external conditions?
The answer should be based on the test design established before launch.
Calculate Relative Lift
Relative lift describes the percentage improvement of one version over another.
Formula: Relative Lift = (Variation Rate – Control Rate) ÷ Control Rate × 100
For example, if Version A converts at 5% and Version B converts at 6%, the absolute improvement is one percentage point.
The relative lift is:
(6% – 5%) ÷ 5% × 100 = 20%
Both the absolute and relative difference can be useful, but they describe the change differently.
Consider Absolute Difference
Absolute difference is the direct difference between the two rates.
Using the same example:
Version A = 5%
Version B = 6%
The absolute difference is one percentage point.
Reporting both absolute difference and relative lift can prevent results from sounding more dramatic than they are.
Evaluate Business Impact
A statistically credible result should still be evaluated in business terms.
For example, a 2% relative increase in conversion may be highly valuable on a website with millions of transactions, while the same improvement may have little practical impact on a small campaign.
Consider:
- Incremental revenue
- Incremental conversions
- Implementation cost
- Profit
- Customer quality
- Operational impact
- Long-term effects
The purpose of the test is to improve decisions, not simply produce a statistically interesting result.
Check Secondary Metrics
Secondary metrics can help explain the primary result.
For example, a new landing page may improve conversion because visitors:
- Click the call to action more frequently
- Abandon the form less often
- Understand the offer more quickly
- Navigate less before converting
Secondary metrics should support interpretation rather than become an excuse to declare a losing test successful.
Review Guardrail Metrics
Confirm that the winning version did not create unacceptable side effects.
For example:
- Conversion increased but refunds also increased
- Leads increased but lead quality declined
- Email clicks increased but unsubscribes doubled
- Orders increased but average order value fell sharply
Performance should be evaluated across the larger business context.
Segment Results Carefully
After the main analysis, teams may examine results across relevant segments such as:
- Device
- Geography
- New versus returning visitor
- Customer type
- Traffic source
- Audience segment
Segment analysis can reveal meaningful differences, but excessive slicing creates opportunities to find random patterns.
Treat unexpected subgroup findings as hypotheses for future testing rather than automatically as definitive conclusions.
Beware of Multiple Comparisons
The more metrics, variations, and audience segments examined, the greater the chance of finding an apparently significant difference purely by chance.
If a test examines dozens of outcomes, at least one may appear interesting even when no meaningful underlying effect exists.
Keep the primary hypothesis focused and treat exploratory findings appropriately.
Distinguish Inconclusive From Failed
A test that does not produce a clear winner is not necessarily a failed experiment.
An inconclusive result may indicate:
- The variation did not materially affect behavior
- The sample was too small
- The difference was smaller than expected
- The variable was not important
- Both versions perform similarly
Learning that an element does not matter much can itself be useful because the organization can redirect effort toward more important variables.
Do Not Automatically Implement Every Winner
A winning variation should still be evaluated for brand, legal, operational, customer-experience, and long-term implications.
For example, a more aggressive promotional message might increase short-term conversion while damaging positioning or customer expectations.
Experimentation should inform professional judgment rather than replace it.
Document the Result
Every completed test should produce a short record that includes:
- Test name
- Research question
- Hypothesis
- Control
- Variation
- Audience
- Primary metric
- Secondary metrics
- Guardrail metrics
- Start and end dates
- Sample size
- Result
- Statistical interpretation
- Business impact
- Decision
- Lessons
- Follow-up questions
Documentation prevents organizations from repeating the same experiments and helps new team members understand why earlier decisions were made.
Maintain a Testing Library
A testing library creates institutional knowledge across campaigns and teams.
Useful categories may include:
- Messaging tests
- Creative tests
- Landing-page tests
- Email tests
- Offer tests
- Pricing tests
- Audience tests
- Checkout tests
Over time, the library can reveal patterns that are more valuable than any individual experiment.
Test Messaging
Messaging tests can evaluate:
- Benefit statements
- Value propositions
- Headlines
- Proof points
- Objection responses
- Product descriptions
- Calls to action
Messaging tests should reflect a clear strategic hypothesis rather than simply comparing random wording.
Related AAMA Resource: Positioning & Messaging Framework
Test Advertising Creative
Advertising tests may compare:
- Creative concepts
- Images
- Video
- Headlines
- Formats
- Calls to action
- Offers
- Product demonstrations
Performance should be evaluated according to the campaign objective.
A creative version that produces the highest CTR may not produce the highest conversion rate, revenue, or customer quality.
Test Landing Pages
Landing-page tests can evaluate:
- Headline
- Layout
- Form length
- Calls to action
- Offer presentation
- Social proof
- Product imagery
- Navigation
- Content length
Landing-page tests should preserve consistent traffic quality across versions.
Related AAMA Resource: Landing Page Evaluation Checklist
Test Email
Email A/B tests may compare:
- Subject lines
- Preview text
- Sender name
- Email copy
- Calls to action
- Offers
- Images
- Send times
Be careful when optimizing exclusively for open rate because privacy technology can make open tracking unreliable.
Clicks, conversions, revenue, and other behavioral outcomes may provide stronger evidence depending on the objective.
Test Offers
Offer testing can evaluate:
- Discounts
- Free trials
- Bundles
- Guarantees
- Shipping offers
- Promotional structures
- Bonus products
- Limited-time incentives
Offer tests should account for profitability and customer value rather than focusing only on conversion rate.
Test Forms
Forms can be tested for:
- Number of fields
- Field order
- Required fields
- Page structure
- Labels
- Calls to action
- Multi-step versus single-page flow
Shorter forms may improve completion but may also reduce the amount of information available for qualification.
The best design depends on the value and quality of the resulting conversion.
Test Audiences Carefully
Testing different audience groups is different from a traditional A/B test when participants are not randomly assigned.
For example, comparing one social-media audience with another may reveal performance differences, but the groups themselves may differ in many ways.
These comparisons can still be useful, but the result should not automatically be interpreted as causal evidence about the targeting variable.
Testing & Brand Building
Not every important marketing outcome can be optimized immediately through short-term conversion testing.
Brand awareness, memory, preference, positioning, trust, and long-term customer value may develop over longer periods.
A/B testing should therefore be one component of a broader measurement system rather than the sole method used to make marketing decisions.
Avoid Local Optimization
Local optimization occurs when teams improve one small metric while damaging the larger customer journey or business objective.
For example:
- Increasing CTR with misleading headlines
- Increasing form completion by removing qualification questions
- Increasing purchases with unprofitable discounts
- Increasing email clicks through exaggerated subject lines
A successful experiment should improve the overall system rather than simply maximize one number.
Recommended A/B Testing Process
A practical testing process can follow this sequence:
- Define the Decision
- Define the Objective
- Develop the Hypothesis
- Identify the Control
- Create the Variation
- Define the Primary Metric
- Define Secondary and Guardrail Metrics
- Define the Audience
- Estimate Required Sample Size
- Establish the Test Duration
- Validate Tracking
- Launch the Test
- Monitor Technical Performance
- Complete the Planned Test
- Analyze the Primary Result
- Evaluate Statistical Evidence
- Evaluate Business Impact
- Review Guardrail Metrics
- Document the Result
- Decide Whether to Implement, Retest, or Reject
- Add the Learning to the Testing Library
The process can be adapted to the scale of the organization, but the major decisions should be made before the results are visible.
Pre-Launch A/B Testing Checklist
Before launching an experiment, confirm:
- The decision is clear
- The hypothesis is documented
- One primary objective is defined
- The control is established
- The variation reflects the hypothesis
- Major variables are isolated
- The primary metric is defined
- Secondary metrics are identified
- Guardrail metrics are defined
- The audience is appropriate
- Assignment is randomized where possible
- Sample-size requirements are understood
- Test duration is planned
- Tracking is working
- Conversion events are correct
- Mobile and desktop experiences are tested
- Overlapping experiments have been reviewed
- External campaign changes are documented
- The decision rule is established before launch
This checklist helps prevent the team from redesigning the rules of the experiment after seeing the outcome.
Good Testing Creates Knowledge
The value of A/B testing extends beyond identifying which version wins. A well-designed experiment can teach an organization something about its audience, offer, message, product, or customer journey.
Over time, disciplined experimentation creates a body of evidence that improves marketing judgment. Teams can stop debating certain questions based entirely on preference because they have documented evidence from previous tests.
The strongest experimentation programs therefore focus not only on optimization, but also on learning.
Related AAMA Resources
Continue developing your measurement approach with the Marketing Research Methods Guide, How to Design a Marketing Survey, Marketing Metrics & KPI Reference, Common Marketing Formulas, Campaign Measurement Framework, Positioning & Messaging Framework, Landing Page Evaluation Checklist, and AAMA Calculators & Tools. These resources provide additional guidance for designing research, selecting metrics, evaluating campaigns, and translating evidence into marketing decisions.
The AAMA Resource Library will continue expanding with testing worksheets, measurement templates, calculators, and professional resources designed to support evidence-based advertising and marketing.

