Let me see if the team has thoughts on this and also eager to hear if other community members have ideas!
I'll call out a new feature we added recently, too, that might help: feature evaluation diagnostics.
It give you logs on how your flags are evaluating in production with whatever data you want to include. I think it'd be helpful for just this situation, where you want to ensure everything is evaluating as expected:
https://docs.growthbook.io/features/features/diagnostics