Define one primary metric and the minimum effect worth detecting, then compute the required sample size and run duration before launching. Randomize at the user level, not the session level, so a returning user always sees the same variant. Run for whole weeks to cover weekly cycles, check the randomization held, and decide in advance what result triggers a launch. Track guardrail metrics alongside the primary one.
Why interviewers ask this
Experiment design is where analysts most influence product decisions, and interviewers want to see discipline rather than enthusiasm. The strongest signals are pre committing to a primary metric and sample size, randomizing at the right unit, running full weekly cycles, and refusing to peek and stop early when the number looks good.
How to structure your answer
- Pick one primary metric and state the minimum detectable effect.
- Compute sample size and duration before launch.
- Randomize at the user level and verify balance after launch.
- Define guardrail metrics and the launch decision rule up front.
- Commit to the end date rather than stopping on a good day.
Example answer
I would start by pinning down one primary metric, which for checkout is completed purchases per user entering checkout, not clicks, and then ask what size of improvement would actually justify shipping. That minimum detectable effect drives the sample size, and the sample size plus our daily traffic tells us the runtime. If that comes out at eleven weeks, the honest answer is that this test is not worth running and we should decide another way. Randomization goes at the user level with a persistent identifier, since session level assignment means a returning customer sees two different checkouts, which both pollutes the data and annoys them. I run in whole week multiples because weekday and weekend behavior differ a lot. I also fix the guardrails in advance, so refund rate and support contacts, and the launch rule. And I do not stop early on a good day; peeking repeatedly at a running test inflates false positives badly, which I have watched happen.
Walking into this interview soon? GhostPilot listens to your live call, spots the question the moment it is asked, and puts a structured answer on your screen in real time. Try it on your next mock, or grab a $29 Session Pass, no subscription, for the real thing.
See how it worksFollow-up questions to expect
- What is a minimum detectable effect and how do you choose it?
- What is the problem with peeking at results daily?
- How would you test something with too little traffic for a clean test?
Related data analyst questions
Your interviewer will ask their own version of this. Paste your actual job description into the free Question Predictor and get the 20 questions that role is most likely to ask, with what each one is really probing.
Predict my questions