CUDIA: Bridging the Resolution Gap in Multi-Source Healthcare Data
A probabilistic imputation framework for predictive analysis using variably aggregated, multi-source healthcare data
This paper introduces CUDIA (Clustering Using features with DIfferent levels of Aggregation), a probabilistic framework for predictive analysis using multi-source data with mixed resolutions. It effectively bridges the gap between individual-level observations and aggregated-level summaries (such as state or county averages) to perform clustering and feature imputation.
TL;DR
CUDIA (Clustering Using features with DIfferent levels of Aggregation) is a novel probabilistic framework designed to handle the "Ecological Fallacy" in healthcare data. It allows researchers to perform high-precision predictive modeling even when some features are only available as aggregated summaries (e.g., state averages) due to privacy constraints. By treating aggregation as a generative process governed by the Central Limit Theorem, CUDIA statistically "reclaims" individual-level features.
The "Data Silo" Paradox and the Ecological Fallacy
In the ideal world of a data scientist, every feature for every patient exists in a single, tidy table. In the real world of healthcare, privacy regulations like HIPAA and reporting norms mean that while you might have individual hospital bed counts, your "surgical discharge rates" might only be available as a single average for an entire state.
The easiest—and most dangerous—solution is to assign that state average to every hospital in the state. This leads to the Ecological Fallacy: the assumption that individual members of a group mirror the group's aggregate characteristics. This ignores intra-group variance and leads to weak, biased models.
Methodology: The Generative Logic of CUDIA
The core of CUDIA is a Bayesian Directed Graphical Model. The authors' brilliant insight is to leverage the Central Limit Theorem (CLT). If we assume individuals belong to one of latent clusters, then an aggregated summary (like a mean) is effectively a weighted sum of those cluster-specific distributions.
As the number of individuals in a partition (like a state) grows, the sample mean naturally follows a Normal distribution. CUDIA uses this property to back-calculate the likely characteristics of individuals based on the group mean and the cluster memberships.

Efficient Learning: From Gibbs to Bregman
Learning this model isn't trivial because the posterior distribution is intractable. The authors propose two solutions:
- Approximated Gibbs Sampling: A probabilistic approach that uses Metropolis-Hastings for sampling the mixture coefficients.
- Deterministic Hard Clustering: For massive datasets, they derived a version that maps the problem to Bregman Divergences. This allows CUDIA to handle not just Gaussian data, but also Poisson or Multinomial distributions with the same linear time complexity as K-Means.
Experimental Results: Recovering Lost Information
The authors tested CUDIA on a variety of real-world healthcare sources, including the Dartmouth Health Atlas, CDC, and U.S. Census Bureau.
In one case study, they predicted hospital bed counts using a mix of individual hospital data and state-level discharge rates.
- No Imputation: Low predictive power.
- State-level Imputation (Naive): Only marginal improvement.
- CUDIA Imputation: Showed a significant boost in , capturing non-linear relationships that naive averages completely missed.
In the Old Faithful benchmark, CUDIA significantly reduced Mean Squared Error compared to baseline methods, effectively "un-mixing" the aggregated data.
Critical Insight: Why This Matters
The value of CUDIA isn't just in the accuracy boost; it's in its generality. Most healthcare data is inherently multi-resolution. By providing a mathematically grounded way to integrate state-level "macro" metrics with hospital-level "micro" metrics, CUDIA opens the door for richer epidemiological studies that respect privacy laws without sacrificing the power of the individual signal.
Limitations
A current constraint is that the number of clusters cannot exceed the number of partitions . This means if you only have data from 5 states, you can realistically only model 5 clusters. Future work incorporating more partitions or hierarchical priors could further relax this restriction.
Conclusion
CUDIA proves that aggregated data is not a dead end for individual analysis. By viewing aggregation as a probabilistic sampling task, we can "see" through group averages to the individuals beneath, providing a powerful tool for the next generation of privacy-aware predictive healthcare.
