Pick out the pages that converted best last month and watch them this month. They will do worse. Pick the worst and watch those, and they will do better — whether or not anybody touched them.
This is worth measuring rather than asserting, because the effect is large enough to be mistaken for a real one, and because the way it is usually explained — the winners got complacent, the fixes worked — is exactly the wrong lesson.
The setup, and the control built into it
A thousand pages, each with a true conversion rate drawn around ten percent. Two periods. Each period sends the same number of visitors to every page and records what converted.
The crucial part of the setup is that every page's true rate is fixed and identical in both periods. Nothing is improved, degraded or optimised. That is the control: any movement between periods cannot be performance, because performance is held constant by construction.
What happens with a small sample
Take the top and bottom tenth by the first period's measurement, then look at the same pages in the second. At two hundred visitors per page per period:
period 1 period 2 true rate
top 10% 15.37% 12.37% 12.36%
bottom 10% 5.21% 7.52% 7.52%
The best pages appear to fall by 3.0 points. The worst appear to gain 2.31 points. Nothing changed.
The column that explains it
Look at the third column against the second. The top decile's second-period result, 12.37%, sits almost exactly on its true rate of 12.36%. The bottom decile's 7.52% lands on 7.52%.
Period two is not a decline. It is a return to the truth. Period one was the anomaly: a group selected partly for being genuinely good and partly for having got lucky, and the luck does not come back for a second showing.
That is the whole mechanism, and it has a name: regression to the mean. Selecting on a measurement means selecting on the true value and on random noise. The truth persists into the next period and the noise does not, so the selected group moves back toward the average by however much noise contributed to putting it at the extreme.
How much depends entirely on the sample
visits per period apparent decline of the top decile
200 3.00pp
1,000 0.84pp
5,000 0.19pp
25,000 0.04pp
The effect is not a property of the pages or their performance. It is a property of the measurement's noise, and it shrinks to nothing as the sample grows. At twenty-five thousand visits a page there is almost no luck left to unwind, so there is almost nothing to regress.
Which gives a practical rule of thumb that requires no statistics: if the ranking would look different on a different week, expect the extremes to move regardless of what you do to them.
The trap this sets
The dangerous version is not the winners declining. It is the losers improving.
Take the worst-performing tenth, apply an intervention — a redesign, a new headline, a fix — and measure again. In this simulation, with no intervention at all, they improve by 2.31 points. Any intervention applied to that group will therefore look like a success. The effect will seem large, and it will be reported as a win.
Worse, the pattern repeats. The following period, the newly-selected worst tenth improves again. A programme of fixing the bottom performers can run for years, generating a consistent record of success, without any of the fixes doing anything.
The same applies to teams, salespeople, regions, campaigns and suppliers. Anywhere the underperformers are selected for attention, the attention will appear to work.
What to do about it
Use a control group. Take the bottom decile, treat half at random, and compare the treated half against the untreated half. Both regress; the difference between them is the effect. This is the entire fix and nothing else is as good.
Compare against the right baseline. If a control group is impossible, do not compare "before" against "after". Compare the new performance against the regression you would expect from noise alone, which requires estimating that noise — harder, weaker, and much better than nothing.
Shrink the estimates before ranking. Pages with few visitors should be pulled toward the overall average before they are ranked at all, in proportion to how little data they have. That is what empirical Bayes methods do, and it stops the extremes being populated by the smallest samples.
Or simply wait for more data. The table above is also a table of how much patience buys.
How this was measured, and what it does not show
A thousand simulated pages per run, four hundred runs per sample size, true rates drawn normally around ten percent with a two-point spread, binomial sampling each period, seeded and reproducible.
True rates were held constant deliberately. That is the point of the simulation, and also its limit. Real pages do change. This cannot tell you whether a particular decline was real; it establishes how much apparent movement requires no cause whatsoever, which is the quantity people omit when they explain a decline.
The distribution matters. True performance here is normally distributed, so extreme measurements are usually ordinary pages having a good week. If performance were heavy-tailed — a few pages far better than the rest — the extremes would more often be real and would regress less.
Ten percent is one cut. The more extreme the selection, the stronger the effect: taking the top one percent regresses harder than the top ten, for the same reason and by more.
This and testing many metrics at once are the same failure wearing different clothes — a result produced by selection rather than by anything that happened.