Work backward from the minimum detectable effect. You need the baseline rate, the smallest lift worth shipping, the significance level, and the desired power (usually 80 percent). Those give a required sample per arm; divide by daily eligible traffic to get days. Then round up to whole weeks so you cover weekday cycles, and fix the duration before you start.
Why interviewers ask this
This tests whether you can prevent the classic failure of peeking at a test that was never powered to detect anything. The interviewer wants the four inputs named, and wants the conversation about the minimum detectable effect, because that is where the business decision hides. Committing to a duration up front and covering full weekly cycles are the details that show you have actually run tests.
How to structure your answer
- Name the four inputs a power calculation needs.
- Turn the minimum detectable effect into a business question.
- Convert required sample into days using eligible traffic.
- Insist on whole weeks and a duration fixed before launch.
Example answer
I ask four things: what is the current conversion rate, what is the smallest lift that would actually make us ship this, what alpha are we using, and what power do we want. The second one is the real conversation, and it is a business question rather than a statistical one. If they say any lift at all, I show them the math: at a 3 percent baseline, detecting a relative half percent takes millions of users per arm, and we do not have that, so we would be running a test that cannot succeed. Usually they land on something like a 5 percent relative lift, which at our traffic might be eleven days. Then I round up to fourteen so we cover two full weekly cycles, because weekend users behave differently and a ten day test overweights whichever days it happens to include. And I write the stop date down before launch. The habit that kills teams is checking daily and calling it the moment p dips under 0.05, which badly inflates the false positive rate. If they genuinely need early reads, I set up a sequential test properly instead of pretending peeking is free.
Walking into this interview soon? GhostPilot listens to your live call, spots the question the moment it is asked, and puts a structured answer on your screen in real time. Try it on your next mock, or grab a $29 Session Pass, no subscription, for the real thing.
See how it worksFollow-up questions to expect
- How much does peeking daily actually inflate the false positive rate?
- What is a sequential test and what does it cost you?
- How would you handle a metric with a very long tail, like revenue per user?
Related data scientist questions
Your interviewer will ask their own version of this. Paste your actual job description into the free Question Predictor and get the 20 questions that role is most likely to ask, with what each one is really probing.
Predict my questions