It's not necessarily happening already.
Example: You have a baseline rate of some metric at 5%. You launch an experiment with the default MDE of a 10% relative uplift. To achieve the standard power of 80% to detect a different between control/intervention, one would need roughly 31 000 users in each group. However, with the default "display results" threshold at 150, you would begin to see results already at 150/0.05 = 3000 users in each group, i.e. one tenth of the necessary population. Consequently, results that are underpowered are being displayed which means that they are susceptible to type 2, type S and type M errors.