How to use a holdout to measure coupon controls

A practical holdout plan for coupon-extension controls: choose the cohort, keep assignment stable, pick one outcome, and report uncertainty.

Author:Benson
MeasurementUpdated 1 October 20265 min
Keep it useful

Jump to a section

In short

A holdout is the cleanest way to value a coupon control: randomly leave a stable slice of eligible sessions untouched, expose the rest to the change, and compare the same outcome in both arms. Choose the outcome before launch, keep the assignment sticky, and report uncertainty instead of turning a small difference into a win.

If you need to know what a coupon control changed, leave a comparable slice of eligible traffic alone. That slice is the holdout, or control arm. The rest sees the storefront rule, checkout policy, or widget you are evaluating. Compare the same outcome in both arms, keep assignment stable, and report the range around the estimate.

Why is before-and-after not enough?#

An after-period contains every other change that happened at the same time: campaign mix, payday, weather, product availability, sale creative, and traffic source. It also contains a different set of shoppers. A before-and-after number can be a useful smoke test, but it cannot tell you which change caused the movement.

The Benson guide to extension impact describes the common failure modes in more detail. A holdout gives you a contemporaneous comparison, so both arms experience the same broad calendar and store conditions.

Google describes an A/B test as a randomized experiment where variants are shown to random samples of users at the same time. Its GA4 A/B test guidance also makes clear that GA4 interprets experiments from a third-party testing tool rather than running the assignment itself. The lesson for a merchant is simple: choose the assignment mechanism and event definitions deliberately, then use analytics to read the result.

What belongs in the eligible cohort?#

Define eligibility before launch. If the control is for a storefront pop-up rule, include the extension sessions that could encounter that pop-up on the relevant page. If the control is for a checkout code policy, include the tagged extension sessions that reach the code field. If the test is for a widget, define the traffic that can actually see the widget.

Do not use “all sessions” just because it is convenient. Sessions that cannot see the control dilute the readout and can hide the decision you are trying to make. Do not use installation as a substitute for eligibility either: an installed extension can be quiet, and an observed extension session can come from a browser that was not eligible for the separate installation check. The visibility versus installation guide explains the distinction.

Write down the exclusions as well. Common examples are bots and automation, unsupported browsers, checkout pages that the control cannot reach, and visitors who cannot satisfy the store's consent settings. Keep those rules constant across both arms.

How do you keep the comparison fair?#

Randomize eligible sessions into treatment and control. Assignment should be sticky for the unit you chose: a visitor, a browser, or a session. A shopper who switches arms halfway through the journey can receive both experiences and make the result hard to interpret.

Do not remove the control because it is inconvenient to support. The control is the thing that turns a movement into evidence. If customer support or legal requirements mean some sessions must always receive a code or message, encode that rule in eligibility and record it.

For a Benson holdout, the control arm is left alone and the treatment arm receives the configured rule. The result can report conversion lift, revenue per session lift, and a confidence interval when the underlying counts meet the readout's minimums. That is a description of the measurement contract, not a promise that every store will reach a significant result.

Which outcome should be primary?#

Choose one primary outcome before opening the result. Revenue per session is a direct commercial measure when order data is connected. Conversion rate helps explain whether the change moved completed orders. Discount depth shows how much value the store gave away. Extension share and pop-ups hidden describe reach and behaviour, but they are not revenue outcomes by themselves.

Keep the supporting measures visible, especially when the primary result is surprising. A policy might reduce discount depth while leaving conversion unchanged. A widget might lift code application but reduce revenue per session if the discount is too generous. The measurement plan should let a merchant see that path instead of asking one metric to explain everything.

Do not change the primary metric halfway through because a secondary metric looks better. If you need a new question, start a new readout or label the analysis as exploratory.

How long should a holdout run?#

Run it until each arm has enough eligible sessions and conversions for the readout to become informative, then stop at a pre-agreed decision point. There is no universal traffic number that makes every store's answer trustworthy. A low-volume store may need more calendar time; a seasonal store may need to cover the period in which the control matters.

Check the arms for assignment and instrumentation problems while the test runs. If the control is unexpectedly receiving the treatment, the selector or policy is not scoped as intended. If order events are missing in one arm, fix the join before drawing a conclusion. Keep a log of changes, pauses, and eligibility updates.

How do you read uncertainty?#

A point estimate is only one part of the answer. A confidence interval shows a range of effects that remains plausible under the readout's assumptions. If it crosses zero, the test has not shown a clear lift or loss. That can be a useful result: the control may be safe to keep, or the decision may not justify further complexity.

Avoid “the result is up, therefore the control worked” language when the interval is wide or the sample is small. Say what the test can support: “Revenue per eligible session was higher in treatment, but the interval includes no change.” Then choose whether to run longer, narrow the cohort, or leave the control off.

What should the final decision record contain?#

Save the hypothesis, eligible cohort, assignment unit, launch and stop dates, primary outcome, secondary measures, treatment description, control description, and readout. Include the code allowlist or storefront rule version that was active. Link the relevant playbook so the next operator can repeat the setup.

The record should end with a decision and an owner: keep, change, or stop the control; what evidence supports that choice; and when it will be reviewed again. A holdout is valuable because it leaves a clean trail from a merchant question to a reversible action.

Sources#

Find out what's leaking. 14 days of Own, no card.