What is the multiple comparisons problem?
When you run many hypothesis tests, the chance of at least one false positive grows quickly. With 20 independent tests at alpha 0.05, the probability of a false positive is about 1 - 0.95^20, roughly 64 percent. Reporting the single significant result is misleading.
Mitigations:
- Bonferroni correction: divide alpha by the number of tests, controlling the family-wise error rate but conservative.
- Holm-Bonferroni: a step-down improvement that is uniformly more powerful.
- Benjamini-Hochberg: controls the false discovery rate, the expected proportion of false positives among rejections, better for large exploratory screens such as genomics.
from statsmodels.stats.multitest import multipletests
reject, p_adj, _, _ = multipletests(p_values, method="fdr_bh")
Pre-register hypotheses and distinguish confirmatory from exploratory analyses. Segment mining in A/B tests is a common trap.