Hi, we did a blanco test to know if we can trust t...
# experimentation
l
Hi, we did a blanco test to know if we can trust the winner to be real. All versions of the test where exactly the same, however we got a significant winner. Can we improve the settings in some way, to be sure there is a statistical winner? Do we need to run te test longer? ....
r
Hi David, This is normal. With two variations you get two chances for random noise to look like a win, so the odds are roughly 1 in 10, and 96% is only just over the line. Two things worth checking: • The traffic split looks slightly off (10.777 / 10.992 / 10.430). Can you check the Health tab for an SRM? If it flags, that's an assignment or tracking issue rather than chance. • Pages per Session counts sessions, but you randomise users. One user has several sessions and those aren't independent, so results can look more certain than they are. A per-user metric avoids that. Running longer won't fix it, because extending a test after seeing the results just gives noise more chances to cross the line. To get fewer borderline winners, raise the threshold to 99%, turn on multiple-comparison correction, and enable sequential testing. I hope this helps.
l
Thx for the clear explanation, I'm not a data analist, but I want to be sure I understand the numbers in a statistical correct manner. 🙂 I've added the pageviews/user and that metric doesn't show a significant difference. The sessions/user is added as well, and this is indeed sign lower for var2, so pages/sess must be higher for var2 since thte pageviews per user stays about the same. The health looks fine.
In a real test with these outcomes, you could conclude that users returned significantly less after being in var2, and found in 1 session what they where searching for.
r
Your thinking is right, but there's a catch. Those three metrics aren't separate measurements. they're tied together by maths. Page views per user is just pages per session multiplied by sessions per user. So if one goes up and page views per user stays flat, the third one has to go down. You're not seeing two findings, you're seeing the same bit of randomness twice. The health check is fine, no SRM. The uneven numbers I flagged before were sessions, not users, so that was the same randomness. Page views per user didn't move. That's the metric where each user is counted once, so it's the most reliable metric. It shows no difference, which is what you'd want for in a test where both versions are identical. Fewer visits is usually a bad sign. To tell the difference you'd need to look at something like conversions or whether people completed what they came to do. I hope this makes sense.
l
Yeah that's clear. Thx for the feedback