Probability - Conditional Probability and Bayesian Inference Interview Problems
Interview Experience
Round 1 ML / Probability
Problem
You are given several probability problems typical of ML Engineer and data science interviews. Solve each clearly, showing your reasoning.
Q1: A test for a disease has 99% sensitivity and 95% specificity. The disease affects 1% of the population. Given a positive test result, what is the probability the patient actually has the disease?
Q2: Two fair dice are rolled. Given that the sum is at least 9, what is the probability that at least one die shows a 6?
Q3: An ML model outputs a score in [0, 1]. You observe that it is well-calibrated: P(y=1 | score=p) = p. You have two predictions: 0.7 and 0.8. What is the probability that both underlying events occur?
Expected Reasoning
Q1: P(disease | positive)
= P(pos | disease) * P(disease) / P(positive)
= (0.99 * 0.01) / (0.99*0.01 + 0.05*0.99)
~= 16.7%
Follow-ups
- How does the base rate (disease prevalence) affect the PPV? What if prevalence drops to 0.1%?
- Explain the difference between frequentist and Bayesian interpretations of these answers.
- In a recommender model, calibration matters for ranking vs. revenue optimization differently. How?
- How do you detect and correct miscalibration in a deployed classification model?
Full Details
Round 1 ML / Probability
Problem
You are given several probability problems typical of ML Engineer and data science interviews. Solve each clearly, showing your reasoning.
Q1: A test for a disease has 99% sensitivity and 95% specificity. The disease affects 1% of the population. Given a positive test result, what is the probability the patient actually has the disease?
Q2: Two fair dice are rolled. Given that the sum is at least 9, what is the probability that at least one die shows a 6?
Q3: An ML model outputs a score in [0, 1]. You observe that it is well-calibrated: P(y=1 | score=p) = p. You have two predictions: 0.7 and 0.8. What is the probability that both underlying events occur?
Expected Reasoning
Q1: P(disease | positive)
= P(pos | disease) * P(disease) / P(positive)
= (0.99 * 0.01) / (0.99*0.01 + 0.05*0.99)
~= 16.7%
Follow-ups
- How does the base rate (disease prevalence) affect the PPV? What if prevalence drops to 0.1%?
- Explain the difference between frequentist and Bayesian interpretations of these answers.
- In a recommender model, calibration matters for ranking vs. revenue optimization differently. How?
- How do you detect and correct miscalibration in a deployed classification model?
About This Question
This is a candidate experience report from a stackadapt interview during the onsite round.