Consider a consumer subscription brand selling a starter product followed by repeat deliveries. Its current lifecycle is tidy: a welcome email, a usage guide, a reminder before the next charge and a cancellation offer. It is also almost blind to what changed after checkout.
This example is a composite, not a report about a specific company. The numbers are illustrative. The method is the important part.
Define the first valuable outcome
Do not begin with “reduce churn.” It is too broad and too slow. Begin with a moment that is frequent, measurable and economically meaningful.
Primary outcome: increase the percentage of starter customers who complete a second order within 45 days, without reducing gross margin or increasing refunds.
The second order is not the full definition of loyalty. It is an early proof that the customer found enough value to continue.
Build customer state, not a giant profile
The system needs only the data required to make this decision. A useful first state might contain:
- starter product, variant and purchase date;
- expected next order date and any cadence changes;
- product pages viewed since purchase;
- delivery status and support issues;
- previous messages, clicks and unsubscribes;
- skip, swap, pause and cancellation actions;
- consent and contact frequency.
Demographic enrichment is probably less useful than knowing that a customer viewed a different variant twice after delaying their next order.
Separate the obstacles
| Observed state | Possible obstacle | Candidate action | What not to assume |
|---|---|---|---|
| Next order due soon, browsing another variant | Preference has changed | Offer a one-tap switch with the relevant option preselected | That the price is too high |
| Low expected usage, next order due soon | Cadence is too fast | Make delay or lower frequency easier than cancellation | That the product failed |
| Delivery delay and recent support contact | Trust and supply problem | Acknowledge the issue and resolve it before any upsell | That a cheerful reminder is appropriate |
| High engagement, no friction, order due | No clear obstacle | Send the normal reminder or nothing | That personalization must always add content |
| Cancellation page after repeated variant views | Product fit may be recoverable | Offer switch, pause and human help in that order | That the highest discount is the best save |
Create a bounded action library
Do not allow a model to invent an unlimited campaign for every person. Give it a set of safe actions and clear eligibility.
- No send. The customer is likely to continue or has already received enough contact.
- Usage help. A short guide tied to the product and stage they actually reached.
- Change cadence. A direct path to bring the next order date in line with observed usage.
- Switch product. A one-tap alternative based on verified browsing or preference data.
- Service recovery. Acknowledge the known failure and provide an approved resolution.
- Price intervention. Use an approved incentive only where price friction is plausible and margin rules permit it.
The model may choose the action, subject, approved creative, timing and channel. It may not invent a price, claim a preference without evidence or contact someone who has withdrawn consent.
Design the test around decisions
A conventional test might compare subject A with subject B across everyone. The more useful experiment compares policies.
- Control: the existing fixed journey.
- Treatment: the adaptive policy chooses from the approved action library.
- Primary metric: second order within 45 days of the starter purchase.
- Guardrails: gross margin, refunds, unsubscribes, support contacts and discount rate.
- Decision log: eligible state, candidate actions, chosen action, content version, reason and outcome.
Within treatment, exploration can compare safe alternatives. A permanent holdout should remain untouched so the team can estimate incremental value.
What the system should learn
The goal is not “Citrus subject lines win.” The goal is a more conditional understanding:
- customers who repeatedly browse an alternative respond to simple switching;
- customers with slower usage retain better after a cadence change;
- discounts add value only after non-price interventions fail for a narrow state;
- service recovery must happen before normal marketing resumes;
- some highly engaged customers need no retention message.
This becomes a policy that keeps learning as products, preferences and seasonality change. It is more valuable than a longer flow because the decision is continuously revisited.
The first 30 days
Week one: agree the outcome, baseline, eligible population, guardrails and data fields. Audit the existing journey.
Week two: connect a read-only event source and current sender. Build four approved actions and a decision preview.
Week three: launch with approval required and a holdout. Review wrong or low-confidence decisions daily.
Week four: move proven low-risk decisions to constrained autopilot. Report early leading indicators, but do not declare retention lift until the outcome window closes.
That is enough to learn whether the approach deserves a larger build. No new CDP, editor or cross-channel platform is required.
