Skip to main content
← Back to articles

Developer resources

Why Your Product Page Optimization Test Never Reaches Confidence

Why App Store Product Page Optimization tests stall below 90% confidence, how many impressions a test really needs, and how to design one that finishes in 90 days.

You set up a Product Page Optimization test, it has been running for weeks, and the confidence number still won't reach 90%. Sometimes the improvement even flips from +12% to −3% and back.

The test isn't broken. It is almost always one of two things: not enough impressions, or a change too small to measure. This guide shows how many impressions a test really needs and how to design one that finishes inside Apple's 90-day limit.


What "Confidence" Means in App Store Connect #

For each treatment, App Store Connect shows impressions, conversion rate, the improvement over the original, and a confidence level. Apple labels a treatment as performing better or worse once it reaches 90% confidence.

In plain words: 90% confidence means the difference you see is unlikely to be random luck. Below that, the variant might be better, or you might be looking at noise. Apple recommends waiting until at least one treatment reaches 90% before you apply it.

The test doesn't stop itself when it gets there. It runs until you stop it, apply a treatment, or hit 90 days.


How Many Impressions a Test Needs #

The number of impressions a test needs depends mostly on how big a lift you're trying to detect. The table below uses a standard two-proportion test at 90% confidence and 80% power. Apple doesn't publish its exact method, so treat these as planning estimates, not promises.

Impressions needed per variant (the original needs the same amount):

Current conversion rate +5% lift +10% lift +20% lift +30% lift
2% 248,000 64,000 17,000 7,700
3% 164,000 42,000 11,000 5,100
5% 96,000 25,000 6,400 3,000

Two things jump out:

  1. Halving the lift roughly quadruples the impressions. Detecting +5% takes about 15 times as many impressions as detecting +20%.
  2. A lower conversion rate needs more traffic. Rare events are harder to measure.

What That Means in Days #

Here is the same math as days, for a 3% conversion rate, one variant, and 50% of traffic sent to the test:

Daily impressions +5% lift +10% lift +20% lift +30% lift
500 656 days ❌ 168 days ❌ 44 days 21 days
2,000 164 days ❌ 42 days 11 days 6 days
10,000 33 days 9 days 3 days 2 days

An app with 500 impressions a day can't detect a 10% lift within 90 days, no matter how long it waits. That is the most common reason a test never reaches confidence: it was never going to.

Try your own numbers in the free App Store A/B test calculator. It also shows the smallest lift your traffic can detect in 90 days.


Five Ways to Make a Test Finish #

1. Test a bigger change #

This is the most effective lever by far. Instead of changing one word in a caption, change the first screenshot entirely, or the whole caption style. A bold change that might move conversion by 20% or more is measurable on modest traffic; a tweak worth 3% is not. Our list of 25 screenshot A/B test ideas has plenty of bold ones.

2. Use fewer variants #

Every variant needs the full amount of impressions. With 2,000 impressions a day, a 10% lift and 50% traffic to the test, one variant needs about 42 days. Three variants need about 126 days, which is over the limit. If traffic is tight, test one variant at a time.

3. Send more traffic to the test #

With one variant, a 50/50 split is fastest. With more variants, send more to the test: the fastest split puts the same number of visitors on every version, which is 67% to the test with two variants and 75% with three. In the example above, moving three variants from 50% to 75% cuts the test from 126 days to 84.

The trade-off: if a variant is worse, more people see it while the test runs.

4. Test where your traffic is #

A test includes the localizations you choose, and the impressions from those markets are what count. Testing in one small market means waiting for that market's traffic. Start in your highest-traffic market, or include several markets that you expect to react the same way. See A/B testing localized screenshots.

5. Run it in full weeks, and leave it alone #

Weekday and weekend visitors behave differently, so run tests for at least one or two whole weeks even if the numbers look decided sooner. And don't stop the test the first day it crosses 90%. Checking every day and stopping at the first good-looking number makes false wins far more likely.


When to Give Up on a Test #

If a test has had the impressions your estimate called for and confidence is still low, that is a result: the variant doesn't make a big enough difference to matter. Stop the test, keep the original, and try a bolder idea. That isn't a failed test. You now know that thing doesn't matter much to your users.


Make Testing Cheap Enough to Repeat #

Traffic limits how many tests you can run in a year, so each one should be worth it. Making variants shouldn't be what slows you down. In Screenshot Studio, one design change applies to every language and device size, and the variant uploads straight into your Product Page Optimization test. Also read the Product Page Optimization mistakes that waste a test.

👉 Download Screenshot Studio →

Frequently Asked Questions

What confidence level does Product Page Optimization use?

Apple marks a treatment as performing better or worse than the original once it reaches 90% confidence. Apple recommends waiting until at least one treatment reaches that level before you apply it.

How long should a Product Page Optimization test run?

Until at least one treatment reaches 90% confidence, and at least one to two full weeks so weekday and weekend visitors are both counted. A test can run for up to 90 days. If your estimate is longer than that, change the test design rather than hoping.

How many impressions does an App Store A/B test need?

It depends mostly on the lift you want to detect. At a 3% conversion rate, detecting a 20% lift takes roughly 11,000 impressions per variant, while a 10% lift takes about 42,000 and a 5% lift about 164,000. Those are planning estimates from a standard two-proportion test.

Why does my test show a big improvement but low confidence?

Early on, a few extra downloads swing the percentage a lot. The improvement number is noisy until each variant has enough impressions. Low confidence means the difference could still be chance, so don't apply the variant yet.

Developer resources

Make bolder variants, faster

Screenshot Studio lets you build a variant that changes one thing boldly, across every language and device size, and upload it straight into your Product Page Optimization test.