TEST CALCULATOR
Test Duration & Sample-Size Calculator
How long to run an A/B or incrementality test, and how much traffic you need before you can trust the result.
New to this? Read how to actually design a test first, then come back and size yours.
Take it further
Designing a real incrementality test?
Go through the design with a former Meta growth lead who has run hundreds of them. 30 minutes, no pitch, no retainer.
How to use this calculator
Enter three numbers: your current conversion rate, the lift you want to be able to detect, and how many people enter the test each day. Choose your control split, confidence and power, and the calculator returns the sample size you need, how long that takes at your volume, and the conversions you should expect to see. Everything updates as you type.
What each input means
- Baseline conversion rate. Your current rate in the control group: purchases per visitor, leads per click, subscribes per install. The lower it is, the more traffic a test needs.
- Lift to detect, the minimum detectable effect. The smallest improvement worth catching, as a relative lift. Setting 10% means detecting a 2.0% rate rising to 2.2%. Smaller effects need far more data, because the required sample size grows with the square of the effect.
- People entering the test per day. Reach or sessions per day across both groups. This is what turns a sample size into a timeline.
- Control or holdout size. A 50/50 split is a standard A/B test and reaches significance fastest. A small holdout of 10 to 20% suits an incrementality test where you do not want to withhold much spend, but the small control group becomes the bottleneck and the test runs longer.
- Confidence. Your tolerance for a false positive, calling a change real when it is not. 95% is standard.
- Power. Your chance of detecting a real effect when it exists. 80% is the common target.
How the calculation works
The tool uses the standard sample-size formula for comparing two proportions, adjusted for an unequal control split so a small holdout correctly shows a longer runway:
Here p₁ is your baseline rate, p₂ is the rate after the lift, c is the control fraction, and zα and zβ come from your chosen confidence and power, for example 1.96 and 0.84 at 95% confidence and 80% power. The result is the total number of people who must enter the test across both groups, and the duration is that number divided by your daily volume.
How long should an A/B or incrementality test run?
At minimum, long enough to reach the sample size above. On top of that, run at least one full purchase cycle so delayed conversions land, and avoid confounding events like a major sale unless that is what you are measuring. Do not stop early or read the result the moment it looks good, because peeking at an underpowered test is the fastest way to a confident wrong answer. If the calculator says many weeks, your effect is small or your volume is low, and the honest options are to test for a bigger lift, add traffic, or accept that a change that small cannot be measured reliably at your scale.
For the full method, from choosing a hypothesis to reading the result, see how to design a test in the measurement guide.
Common questions
Long enough to gather the sample size that can detect the effect you care about, and at least one full purchase cycle so delayed conversions land. Enter your numbers above to get the exact figure. If the answer is many weeks, your effect is small or your volume is low, and you should raise the minimum detectable effect or accept that you cannot measure it reliably.
The smallest improvement worth detecting, expressed as a relative lift. An MDE of 10% means catching a change from, say, a 2.0% conversion rate to 2.2%. Smaller effects need dramatically more data, because the required sample size grows with the square of the inverse of the effect.
Confidence is your tolerance for a false positive, calling a change real when it is not, and 95% is standard. Power is your chance of detecting a real effect when it exists, and 80% is the common target. Higher confidence and power both require more sample, so most tests use 95% and 80% as a sensible default.
There is no magic number. You need enough that random noise cannot masquerade as your effect, and for a small lift on a low conversion rate that can mean thousands per group. The calculator shows the conversions to expect for your inputs. If it is only a few dozen, the test can only detect a very large effect.
A 50/50 split reaches significance fastest and suits a straight A/B test. A small holdout of 10 to 20% suits an incrementality test where withholding half your spend is too costly, but the smaller control group makes the test run longer for the same certainty. The calculator shows the trade-off directly.
An underpowered test does not have enough conversions in each group to separate signal from noise, so it produces confident but unreliable results. The usual causes are too little traffic, a conversion rate that is too low, an effect that is too small, or stopping too early. Size the test before you run it, and do not read results until you reach the required sample.
Related: the paid social measurement guide and the break-even calculator.