Data Analyst Interview Question

What is Simpson's paradox, and have you seen it in real data?

What the interviewer is probing, how to structure your answer, and a spoken example you can adapt.

Quick answer

Simpson's paradox is when a trend that appears in every subgroup reverses when the groups are combined, because group sizes and base rates differ. A classic case is a treatment that performs better in both mild and severe patient groups yet looks worse overall because it was given disproportionately to severe cases. The fix is to segment on the confounding variable rather than trusting the aggregate.

Why interviewers ask this

It is a direct test of whether you interrogate aggregates. Interviewers use it because business dashboards are aggregates by default, and an analyst who never segments will eventually report a reversed conclusion with total confidence. A real example from your own work is worth far more than the textbook admissions case.

How to structure your answer

  • Define the reversal clearly with the mechanism, not just the name.
  • Give one concrete example with the confounding mix.
  • Explain the practical defense: segment before concluding.
  • Say how you decide which variable to segment on.
  • Note that the segmented view is not automatically the right one either.

Example answer

Spoken example, first person

It is when each subgroup shows one direction and the pooled data shows the opposite, driven by uneven group sizes. I ran into a version of this comparing two acquisition channels. Channel A had a worse overall conversion rate, so the obvious move was to cut its budget. When I split by device, channel A converted better on both mobile and desktop. The reason the aggregate flipped was that channel A sent overwhelmingly mobile traffic, and mobile converts worse for everyone, so the channel was being penalized for its mix rather than its quality. Cutting it would have been the wrong call. Since then my default is to check the aggregate against two or three obvious segments before I present anything, usually device, geography, and new versus returning. I would add that segmenting is not automatically correct either. If you slice far enough you can find a story in noise, so I decide which variables to segment on based on what plausibly causes the outcome, not by hunting for the split that agrees with me.

Walking into this interview soon? GhostPilot listens to your live call, spots the question the moment it is asked, and puts a structured answer on your screen in real time. Try it on your next mock, or grab a $29 Session Pass, no subscription, for the real thing.

See how it works

Follow-up questions to expect

  • How do you decide which variables to segment by?
  • How do you avoid finding false patterns when slicing many ways?
  • When is the aggregate number the right one to report?

Related data analyst questions

Your interviewer will ask their own version of this. Paste your actual job description into the free Question Predictor and get the 20 questions that role is most likely to ask, with what each one is really probing.

Predict my questions

Rehearse the hard questions before they are asked

Practise with a live copilot, then walk in ready. A $29 Session Pass gets you through the interview with no subscription and no lock-in.

Get GhostPilot