Hi! I have a question about how to evaluate an ex...
# experimentation
l
Hi! I have a question about how to evaluate an experiment’s impact on global metrics. Let’s say we have overall retention measured from the user’s first app launch, and an experiment where users are assigned only after completing onboarding. It’s clear that retention within the experiment will be measured relative to exposure time, so the values will not match. In that case, we measure the effect only for users who completed onboarding, which seems correct for this specific experiment. But how can we estimate the overall change in retention? For example, something like this: if the experiment leads to a +10% retention uplift (measured from exposure), and only half of users complete onboarding, then the experiment produces roughly a +5% impact for the entire user base. That is, of course, a very rough calculation. Ideally, this would be available directly on the experiment page.
c
May I ask what you would use this very rough number for? Who'd you present it to and how would you frame the uncertainty of it?
l
The number itself is not approximate in general — only in my example. I think it is very important to estimate the global impact. For example, we have a Retention team, and their goal is, by definition, to improve retention. Their overall KPI is improving global retention. Let me show an example: • Hypothesis 1: add push notifications for new users. This applies to the entire new-user audience. The effect is +2% retention uplift from exposure within the experiment, and the same +2% improvement in the app’s overall metric. • Hypothesis 2: add an onboarding tutorial for users interested in surfing (let’s say they make up 1% of users). We get a +50% effect within the experiment, but the impact on the app’s global metric is only roughly +0.05%. Is hypothesis 2 strong? Yes. Did it have a bigger impact on our business than hypothesis 1? No.
As for uncertainty, I would frame it as an estimate of the experiment’s expected impact on the global metric, based on the share of users who are actually eligible for the experiment. So it should not replace the experiment’s core exposure-based result, but rather complement it with a business-level view of the total impact.
c
Aah, I thought it was for multiple experiments cumulatively, not only a single one. Gotcha.
👍 1
c
Hey Vladislav! If I've understood correctly, you have exposure mid-funnel but you want the effect expressed on your full top-of-funnel sample, since that's how your global retention metric is defined? Quick clarifier first: how much time typically passes between first app launch and completing onboarding? If it's minutes or within a day, the rest of this is mostly a measurement-window nuance. If it's days or weeks, the distinction really matters. Two practical options: 1. Move exposure to first app launch. Every new user enters the experiment at launch, regardless of whether they reach the onboarding step where the treatment shows up. This goes against general guidance (usually we recommend assigning at the point of divergence), but in your case the math largely works out: users who don't complete onboarding contribute zeros in both arms, so they dilute the effect but also reduce variance, and the two roughly cancel. Rough directional guidance on the power cost: for baseline retention rates below ~50%, the loss is marginal; above that it can get substantive. Happy to go deeper on the math if useful. With this setup, you'd just read the aggregate result directly: measured from app launch, on your global sample. That's your global-impact number. 2. Keep assignment at onboarding completion and scale manually. You already described this one: take the measured effect and multiply by (users who complete onboarding) / (top-of-funnel users) over the same window. Safe fallback if option 1 isn't practical to implement. Bonus if you go with option 1: Set an activation metric of "completed onboarding" on the experiment, and then in the results view, add the activation dimension. You'll see results split between "Activated" (effect among onboarding completers) and "Not Activated" (effect among non-completers, should be ~0, which is a nice sanity check that the treatment genuinely only fires post-onboarding). Both slices are measured from first app launch, preserving your global measurement window. Note on a subtle feature difference, in case it comes up: GrowthBook's activation metric has two modes. • As a filter (no dimension added): drops non-activated users and measures metrics from the activation timestamp (onboarding completion). This changes your measurement time window, so probably not what you want here. • As a dimension (what I described above): keeps all users, splits them in the results, and measures everyone from exposure (first app launch). Preserves your measurement window. For your use case, I believe the dimension approach is the right one. What do you think?
l
Hi @crooked-rainbow-90115, Thanks, that’s a really great answer. Usually, less than 10 minutes pass between the first app launch and completing onboarding (although of course there are users who complete onboarding a month later). As for the options: 1. We used to do that before, but eventually came to the conclusion that assigning users right before they reach the screen works better. I didn’t know about using the activation metric as a dimension — that’s very cool! It would be even better if the total were shown somewhere nearby as well. In that case, it would really make sense to switch to triggering the feature request at app launch — everything would be in one place: the total and the numbers for activated users, not activated doesn't make sense here for us 2. That’s exactly what we do now, but manually. GrowthBook has already saved us a huge amount of manual work, and if this part were automated too, it would fully cover all our experimentation needs.
c
Good to hear this can work for you! I'll take your feedback back to the team 🙏
👍 1