
How can we identify meaningful structure in complex global health data?
My Master’s thesis focuses on developing and evaluating machine learning approaches to detect, quantify, and validate structure in often sparse and highly biased global health datasets. A central challenge is distinguishing meaningful patterns from random variation and determining whether a dataset contains sufficiently robust structure to support methods such as clustering. Using a novel HIV dataset from Zimbabwe as a real-world application, the thesis investigates how different data representations, clustering techniques, and statistical validation methods can be used to characterize latent structure and assess its relevance for answering global health research questions.
Leave a Reply