Understanding, diagnosing and handling missing data in Data Analytics and Research. Move from complete-case analysis to Multiple Imputation by Chained Equations (MICE), with practical application on real NHANES data.
A dataset may contain thousands of observations, but participants may still have missing laboratory measurements, income information, questionnaire responses, clinical measurements or other important variables.
How these missing values are handled can affect sample size, statistical power, standard errors, confidence intervals, precision, representativeness, statistical estimates and research conclusions. Handling missing data is therefore not simply about deleting incomplete observations or filling empty cells.
Missing Data Analysis Using Stata is a 4-week applied course that takes participants from the foundations of missing data through complete-case analysis and Multiple Imputation by Chained Equations (MICE), using NHANES data and Stata throughout.
For clarity and easier navigation, the 40 modules have been organized into 7 major learning units on this page. These units are not individual modules, each brings together several related modules taught progressively during the 4-week course. Participants receive training across the full 40-module curriculum.
Begin by understanding what missing data is, why it occurs and why it matters for Data Analytics and Research.
Develop a clear understanding of the assumptions behind missing-data analysis, and move beyond simply saying "my data have missing values" to asking what process may have produced them.
Learn how to investigate who has missing information and whether observed participant characteristics are related to missingness. Finding predictors of missingness does not prove that the data are MAR.
Understand what happens when incomplete observations are excluded, and why apparently simple missing-data solutions can create new statistical problems. Complete-case analysis should be a deliberate analytical decision, not an automatic response to missing data.
Progress from deleting incomplete observations to understanding modern multiple-imputation methods, and Stata's mi framework.
The major practical component of the course. Participants apply missing-data methods using NHANES 2017 to March 2020 Pre-Pandemic data and Stata, connecting statistical theory directly to real applied research practice.
Bring the complete missing-data analysis process together and finish the course with the professional missing-data workflow.
Basic knowledge of datasets and introductory statistics is recommended.
It is a statistical and research problem. Deleting incomplete observations without understanding why they are missing can alter the analytical sample and potentially affect research conclusions. Filling empty cells without accounting for uncertainty can produce misleading results.
Survey weighting and missing-data handling solve different analytical problems. Multiple imputation does not replace NHANES survey weights, strata or Primary Sampling Units (PSUs). Similarly, using survey weights does not automatically solve problems created by missing data. Both issues must be considered when conducting rigorous NHANES analysis.
Do not automatically delete incomplete observations. Do not automatically replace missing values with averages. Do not assume missing data are harmless. Learn to identify missingness, investigate the missing-data process, understand MCAR, MAR and MNAR, conduct complete-case analysis, perform multiple imputation using Stata and report your analytical decisions professionally.