Hi David,
This is normal. With two variations you get two chances for random noise to look like a win, so the odds are roughly 1 in 10, and 96% is only just over the line.
Two things worth checking:
• The traffic split looks slightly off (10.777 / 10.992 / 10.430). Can you check the Health tab for an SRM? If it flags, that's an assignment or tracking issue rather than chance.
• Pages per Session counts sessions, but you randomise users. One user has several sessions and those aren't independent, so results can look more certain than they are. A per-user metric avoids that.
Running longer won't fix it, because extending a test after seeing the results just gives noise more chances to cross the line. To get fewer borderline winners, raise the threshold to 99%, turn on multiple-comparison correction, and enable sequential testing.
I hope this helps.