How to read this Bayesian sample-size planner
Use this before you launch (or when you are deciding whether a test is even worth running). You assume a baseline conversion rate and a true effect size for B vs A, then ask: under our chance-to-beat decision rule, how many visitors do simulations usually need before a winner is called?
It is a planning companion to the Bayesian A/B results calculator — not a frequentist power calculator, and not a guarantee of the sample size your live test will need. An example with the default assumptions is shown immediately; change the inputs and press Update results when you want a fresh plan.
What to enter
- Baseline conversion rate (A) — your expected control rate (e.g. 10%).
- Effect size — relative % (B is 20% higher than A) or absolute percentage points (B is +2pp). The line under the field shows the other unit so you can sanity-check the assumption.
- Traffic to A — default 50/50; change it if you will run an unequal split.
- Winner threshold — the bar for naming a winner (90%, 95% or 99%). Same language as the results calculator.
- Max expected loss (optional) — if set, a simulation only “decides” when chance-to-beat clears the threshold and expected loss on the leading arm is under this ceiling.
- Max total visitors — the cap for each simulated test. If many simulations are still undecided at the cap, raise this or lower the threshold.
- Weekly visitors (optional) — turns the median into calendar weeks, and always shows the visitors/week needed to hit the median in about four weeks.
Share results copies a link with those inputs (not a frozen median), so anyone opening it gets the same assumptions and a fresh simulation.
How to read the output
Median visitors to a decision is the middle of many simulated tests under your assumed true effect — including runs that never decide and sit at the visitor cap. It is not a promised sample size for your next experiment.
The p10–p90 range is the typical spread of stopping times. A wide range means the answer is sensitive to noise even when the true lift is fixed.
Illustrative counts at the median convert that total into approximate visitors and conversions per arm if rates match your assumption. Handy for briefing, not a forecast of what you will observe.
Sensitivity (half / stated / 1.5× effect) shows how the median moves if the true lift is smaller or larger than you stated. If half-effect needs far more traffic than you can afford, rethink the test or the effect you are designing for.
Correct-call rate (among decided simulations) is how often the planner picked the truly better variation when a decision was made under the assumed lift. It is not the same as frequentist power.
Null false-positive rate is the share of zero-lift simulations that still called a winner by the cap. Each simulation is one test that grows over time; checking chance-to-beat along that path inflates this — treat it as a caution about peeking, not a Type I error guarantee.
The stopping-time chart shows where simulations decided. The last bar is the share still undecided at max visitors.
A worked example
Defaults: 10% baseline, +20% relative lift for B (≈ +2pp), 50/50 traffic, 95% winner threshold, 50,000 max visitors.
You should see a modest median (often in the low thousands in these simulations), a sensitivity strip where half-effect needs more traffic than the stated effect, a high correct-call rate among decided runs, and a material null false-positive rate — because sequential checks make “false wins” more common than a single end-of-test look.
Try absolute mode with +2pp instead of +20% relative: near-equivalent assumptions should produce a similar planning picture. Add weekly visitors if you want the calendar framing.
When to be careful
- Assumed effect, not observed data. Optimistic lifts produce optimistic sample sizes. Prefer an effect you would actually ship for, not a stretch goal.
- Not frequentist power / MDE. If your stakeholder asks for “80% power at 5% significance”, this tool answers a different question. Say so up front.
- Peeking. The null false-positive callout exists because the planner checks along the path. Live peeking without a sequential plan has the same problem.
- Cap pile-up. If undecided share is high, the median understates how often you would still be waiting. Follow the on-screen cap advice.
- Two variations only. Multivariate or bandits need a different design.
When this is the right tool
Use it to plan a Bayesian A/B test you intend to decide with chance-to-beat (and optionally expected loss): set a believable baseline and effect, pick the threshold your team uses, and read median plus honesty metrics together.
When the test is live or finished, paste the real counts into the Bayesian A/B results calculator instead.
If you want a hand sizing a messy programme (multiple metrics, uneven traffic, or a hard calendar constraint), get in touch.
