Technical

What Kill Criteria Should You Use for Distribution Tests?

Kill criteria for distribution tests: stop rules, sample thresholds and decision dates set before launch so losing formats die early and winners run on.

kill criteriastop rulestest governancedistribution experiments

Kill criteria for distribution tests are pre-committed rules, a minimum sample, a metric threshold, a decision date and a cost ceiling, that determine when a test stops, so decisions rest on evidence instead of mood. Write them before launch. If the rule is negotiable afterward, it is not a rule.

The alternative is deciding in the moment, which biases every team toward the format they personally like. Our content pilot success metrics page defines the numbers that feed these rules.

Why Do Tests Need Pre-Committed Stop Rules?

Because most test variants lose, and stopping on feeling wastes the budget. VWO's compilation of testing data reports that in the travel sector only 40 percent of test variations outperform their control, per VWO's A/B testing statistics. A base rate like that demands a mechanical rule for calling losers.

Peeking makes the arithmetic worse. Statistician Evan Miller showed that repeatedly checking an experiment can push an apparent 5% false-positive rate as high as 26.1%, per his analysis of repeated significance testing. A fixed sample and a fixed decision date are the countermeasure.

Without one, teams keep losing tests alive past the decision date, hoping for a late spike. The spike usually does not come, and the budget that should have funded the next test is gone. The rule exists to protect that budget for the next hypothesis, which is usually worth more than the test you are nursing.

What Should the Stop Rule Actually Contain?

Four parts: a minimum sample, usually enough posts and days to stabilize the median; a performance threshold tied to your baseline; a decision date; and a cost ceiling. All four must hold for the test to continue.

The sample requirement is the one teams skip. Statistical significance depends on sample size, and without enough observations a result can be pure noise. Our statistical significance for social tests guide covers how to size that window.

Written down and agreed before launch, those four numbers remove the argument that always follows a disappointing result. A stop rule negotiated after the data arrives is not a rule at all.

Why Do Teams Struggle to Kill Tests on Time?

Because most teams never defined the moment to stop. VWO's data also shows that 52.8 percent of conversion-rate professionals lack a standardized stopping point for A/B tests, per the same VWO statistics roundup. Without a stopping point, no test ever really ends.

A written decision date fixes this. Put it on the calendar when the test launches, and treat it as binding.

What Are the Legitimate Exceptions?

Extend only when leading metrics are rising and the sample is genuinely thin, or when a platform outage corrupted part of the window. Document the reason and the new date. Never extend because the result is close and the team wants a different answer.

There is a real difference between a slow ramp and a dead test. Our guide to telling a slow ramp from a failure walks through the pattern before you invoke an exception.

How Do You Kill a Test Without Losing the Learning?

Retire the format, keep the data. Record what failed, on which accounts, under which conditions, and archive the assets. A killed test is a priced lesson, and the lessons compound just as reach does. Store the cohort data alongside the test so a future team never repeats a mistake you already paid for. Our when to cancel a content pilot page shows how to exit cleanly and recycle the findings.

Then redeploy the freed budget and accounts to the next hypothesis, so the portfolio keeps moving instead of nursing dead formats.

How Conbersa Makes Kill Decisions Clean

Conbersa runs tests across isolated accounts on real physical smartphones, not emulators or browsers, which keeps each observation independent and the sample honest. Warmup and account isolation remove the confounders that normally blur a stop decision.

At fleet scale you can run a control and several variants in parallel, reach your sample size quickly, and apply kill criteria on schedule. Clean measurement is what makes a rule enforceable. See how the fleet is operated at conbersa.ai.

Neil Ruaro
Founder, Conbersa

We run agentic distribution on a fleet of real phones — and write up what we learn helping founders escape the cold start. Got a topic you want covered? Tell us.

FAQ

Frequently asked questions

Kill criteria are pre-agreed rules that tell you when to stop a test: a minimum sample, a metric threshold to beat, a decision date and a cost ceiling. They are written before launch so results cannot be rationalized after the fact.
Kill it when it fails your leading-indicator threshold after a full sample window, when it breaks the cost ceiling, or when retention never improves despite clean execution. Kill on rules, not on a bad week or an emotional read. The rule should exist before the test starts.
No. Virality is not the goal of a test, repeatable performance is. If median views and retention meet your baseline, the format passed even without a breakout post, and cutting it would remove a working asset from your mix. Passing the baseline is enough to keep it.
The Conbersa Blog

New guides, straight to your inbox.

Tactics on organic distribution and the cold-start problem. What's actually working, no fluff.