You set up a Product Page Optimization test, it has been running for weeks, and the confidence number still won't reach 90%. Sometimes the improvement even flips from +12% to −3% and back.
The test isn't broken. It is almost always one of two things: not enough impressions, or a change too small to measure. This guide shows how many impressions a test really needs and how to design one that finishes inside Apple's 90-day limit.
What "Confidence" Means in App Store Connect #
For each treatment, App Store Connect shows impressions, conversion rate, the improvement over the original, and a confidence level. Apple labels a treatment as performing better or worse once it reaches 90% confidence.
In plain words: 90% confidence means the difference you see is unlikely to be random luck. Below that, the variant might be better, or you might be looking at noise. Apple recommends waiting until at least one treatment reaches 90% before you apply it.
The test doesn't stop itself when it gets there. It runs until you stop it, apply a treatment, or hit 90 days.
How Many Impressions a Test Needs #
The number of impressions a test needs depends mostly on how big a lift you're trying to detect. The table below uses a standard two-proportion test at 90% confidence and 80% power. Apple doesn't publish its exact method, so treat these as planning estimates, not promises.
Impressions needed per variant (the original needs the same amount):
| Current conversion rate | +5% lift | +10% lift | +20% lift | +30% lift |
|---|---|---|---|---|
| 2% | 248,000 | 64,000 | 17,000 | 7,700 |
| 3% | 164,000 | 42,000 | 11,000 | 5,100 |
| 5% | 96,000 | 25,000 | 6,400 | 3,000 |
Two things jump out:
- Halving the lift roughly quadruples the impressions. Detecting +5% takes about 15 times as many impressions as detecting +20%.
- A lower conversion rate needs more traffic. Rare events are harder to measure.
What That Means in Days #
Here is the same math as days, for a 3% conversion rate, one variant, and 50% of traffic sent to the test:
| Daily impressions | +5% lift | +10% lift | +20% lift | +30% lift |
|---|---|---|---|---|
| 500 | 656 days ❌ | 168 days ❌ | 44 days | 21 days |
| 2,000 | 164 days ❌ | 42 days | 11 days | 6 days |
| 10,000 | 33 days | 9 days | 3 days | 2 days |
An app with 500 impressions a day can't detect a 10% lift within 90 days, no matter how long it waits. That is the most common reason a test never reaches confidence: it was never going to.
Try your own numbers in the free App Store A/B test calculator. It also shows the smallest lift your traffic can detect in 90 days.
Five Ways to Make a Test Finish #
1. Test a bigger change #
This is the most effective lever by far. Instead of changing one word in a caption, change the first screenshot entirely, or the whole caption style. A bold change that might move conversion by 20% or more is measurable on modest traffic; a tweak worth 3% is not. Our list of 25 screenshot A/B test ideas has plenty of bold ones.
2. Use fewer variants #
Every variant needs the full amount of impressions. With 2,000 impressions a day, a 10% lift and 50% traffic to the test, one variant needs about 42 days. Three variants need about 126 days, which is over the limit. If traffic is tight, test one variant at a time.
3. Send more traffic to the test #
With one variant, a 50/50 split is fastest. With more variants, send more to the test: the fastest split puts the same number of visitors on every version, which is 67% to the test with two variants and 75% with three. In the example above, moving three variants from 50% to 75% cuts the test from 126 days to 84.
The trade-off: if a variant is worse, more people see it while the test runs.
4. Test where your traffic is #
A test includes the localizations you choose, and the impressions from those markets are what count. Testing in one small market means waiting for that market's traffic. Start in your highest-traffic market, or include several markets that you expect to react the same way. See A/B testing localized screenshots.
5. Run it in full weeks, and leave it alone #
Weekday and weekend visitors behave differently, so run tests for at least one or two whole weeks even if the numbers look decided sooner. And don't stop the test the first day it crosses 90%. Checking every day and stopping at the first good-looking number makes false wins far more likely.
When to Give Up on a Test #
If a test has had the impressions your estimate called for and confidence is still low, that is a result: the variant doesn't make a big enough difference to matter. Stop the test, keep the original, and try a bolder idea. That isn't a failed test. You now know that thing doesn't matter much to your users.
Make Testing Cheap Enough to Repeat #
Traffic limits how many tests you can run in a year, so each one should be worth it. Making variants shouldn't be what slows you down. In Screenshot Studio, one design change applies to every language and device size, and the variant uploads straight into your Product Page Optimization test. Also read the Product Page Optimization mistakes that waste a test.