How do you handle missing data in an analysis?
First understand why data is missing, because the mechanism determines the right approach.
- MCAR: missing completely at random, safe to drop rows with a small loss of power.
- MAR: missing at random given observed variables, so imputation using other features is reasonable.
- MNAR: missing not at random, where the missingness itself carries information, such as high earners refusing to report income. Simple imputation then biases results.
Options: drop rows or columns, mean or median imputation, model-based imputation such as MICE or k-NN, or flag missingness with an indicator variable and let the model use it. For time series, forward-fill carefully and never leak future values backward.
df["income"] = df["income"].fillna(df["income"].median())
df["income_missing"] = df["income"].isna().astype(int)
Always report how much data was missing and test sensitivity across methods.