Imagine a retention team tests two messages. Message A explains a product feature. Message B offers 15% off. B produces more purchases, so it becomes the default.

The result looks clear. It may also hide the decisions that matter.

An average winner is a useful fact. Turning it into one permanent default is a separate decision.

Three ways a good test gets stuck

1. It optimizes a proxy

Subject lines are often judged on opens. Messages are judged on clicks. Neither proves an incremental order, booking or retained customer. The easiest metric to move can pull the system away from the business outcome.

2. It compresses different contexts

A first-time buyer with a delivery problem, an active customer comparing plans and a dormant customer may all enter the same “at risk” segment. The average effect across them can be positive while an action harms one group and helps another.

3. It ends when the winner is declared

Customer behaviour, inventory, products and seasonality change. A fixed winner decays. Teams rarely have enough time to rerun every branch, so yesterday’s result becomes a permanent rule.

Move from variants to policies

A variant is one message. A policy is a repeatable rule for choosing among safe actions using the customer context available at that moment.

QuestionVariant testPolicy test
What is being compared?Message A against message BThe current fixed journey against adaptive selection among approved actions
Who receives it?Everyone in the campaign segmentEligible customers, with action selected from their state
Can it choose silence?Usually noYes, no send is a real action
What learns?The team records an average winnerThe system updates which action works in which context
What proves value?Often opens, clicks or attributed revenueIncremental business outcomes against a persistent holdout

Keep the action space small

Adaptive does not mean unconstrained. Begin with a handful of actions a marketer would be comfortable sending manually. For a stalled marketplace user, those might be:

  1. send nothing yet;
  2. show a shortlist based on verified recent searches;
  3. explain how the first conversation works;
  4. offer human help for a high-value or repeated failed attempt.

Each action has eligibility, required facts, a frequency limit and approved content. The system chooses within those boundaries. That makes learning possible and failures debuggable.

Preserve exploration and a holdout

If the system always chooses the action that currently looks best, it stops learning about alternatives. It can also confuse naturally high-converting customers with successful treatment.

Use two different controls:

Exploration learns which treatment works. The holdout measures whether intervening beats the counterfactual.

Make the reward honest

Every objective creates behaviour. If a system optimizes clicks, it will find clickable messages. If it optimizes gross orders, it may overuse discounts. If it optimizes short-term saves, it may delay inevitable churn while increasing frustration.

A useful objective combines:

The marketer’s role gets more important

Automation should remove repetitive allocation, not strategy. The marketer defines the valuable outcome, creates safe actions, decides which customer facts are appropriate, sets the brand and commercial boundaries, and reviews where the policy is uncertain.

The machine gets better at repeated choices. The team gets more time to create better choices.