GEP-CV: Breaking the Unimodal Barrier of Expectation Propagation

Generalizing expectation propagation with mixtures of exponential family distributions and an application to Bayesian logistic regression

2019-01-30
Shiliang Sun, Shaojie He
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Generalized Expectation Propagation (GEP), a novel deterministic approximate inference framework that extends traditional EP by utilizing mixtures of exponential family distributions. To address convergence issues caused by noisy stochastic gradients, the authors further develop GEP-CV, which incorporates control variates for variance reduction, achieving superior performance in Bayesian logistic regression tasks.

TL;DR

Expectation Propagation (EP) has long been a staple of Bayesian inference, but its reliance on single-mode exponential family distributions often leaves it "blind" to the multi-modal reality of complex posteriors. This paper introduces Generalized Expectation Propagation (GEP), which adopts a mixture-model approach to capture multi-modality, and GEP-CV, which utilizes control variates to solve the stability issues inherent in stochastic optimization.

Academic Positioning

This work sits at the intersection of Deterministic Approximate Inference and Stochastic Optimization. It refines the "energy function" perspective of EP (similar to Power EP or BB-) but pushes the boundary of expressivity by moving beyond unimodal approximations.

Problem & Motivation: The "Averaging" Trap

Traditional EP attempts to match the moments of a target distribution. When the true posterior is multi-modal, a single Gaussian (the most common choice) will attempt to cover all modes, often resulting in a high-variance "average" that represents none of the modes accurately.

Furthermore, traditional EP has two major engineering flaws:

  1. Convergence: It is not guaranteed to converge, especially with non-smooth likelihoods.
  2. Memory: It requires storing "site parameters" for every single data point, a nightmare for Big Data.

The authors' insight is to treat the problem as a global optimization of the KL divergence where is a mixture distribution, and use stochastic methods to bypass the memory bottleneck.

Methodology: Mixture Models and Variance Control

1. From Single to Mixture

Instead of a single , GEP uses: This allows the approximation to "shape-shift" and fit multiple peaks in the probability landscape.

2. The GEP-CV Architecture

While stochastic gradients allow the model to scale, they are notoriously "noisy." To fix this, the authors introduce Control Variates (CV). By defining a function that is highly correlated with the gradient but has a known expectation, they can "subtract" the noise:

GEP-CV Algorithm Flow The algorithm iteratively draws samples from the mixture components and applies variance-reduced updates to both the mixing weights and the natural parameters .

Experimental Results: Stability and Performance

The authors tested GEP-CV on both synthetic three-modal data and several UCI real-world datasets for Bayesian Logistic Regression.

Quantitative Edge

GEP-CV consistently outperformed standard EP and Variational Inference (VI). In high-dimensional datasets like Madelon (501 dimensions), GEP-CV achieved significantly better log-likelihoods, proving that its mixture approach captures the complexity that unimodal methods miss.

Performance Comparison on Synthetic Data Figure 1: RMSE comparison tracking iterations. GEP-CV (bottom line) shows the fastest and lowest error convergence.

The Power of Control Variates

The visual evidence of variance reduction is striking. When comparing the "Change of Approximating Mean" over time, GEP-CV exhibits significantly fewer oscillations than standard GEP.

Smoothing Effect of Control Variates Figure 2: Top row (GEP-CV) shows a much smoother convergence of parameters compared to the fluctuating bottom row (Standard GEP).

Critical Analysis & Conclusion

Takeaway

GEP-CV effectively solves the expressivity-stability trade-off. By using mixture models, it handles complex posteriors; by using control variates, it makes the training of those complex models feasible.

Limitations

  • Computation: While memory-efficient, the sampling process for mixtures is computationally intensive compared to closed-form EP updates.
  • Heuristics: The selection of the number of components () remains a hyperparameter. As seen in Table 4, more components aren't always better; was the "sweet spot" for 3-modal data, but larger increased optimization difficulty.

Future Outlook

The integration of Reparameterization Tricks (like those used in VAEs) could further enhance GEP, potentially allowing it to compete with modern Deep Bayesian frameworks in even higher-dimensional latent spaces.

Find Similar Papers

Try Our Examples

  • Which recent papers have expanded on the "Black-box Alpha-divergence" framework to improve the robustness of Expectation Propagation in deep learning models?
  • What are the original theoretical foundations of using control variates in stochastic variational inference (SVI), and how does this paper's Taylor expansion approach differ from the delta method?
  • How can mixture-based Expectation Propagation be scaled to high-dimensional Latent Dirichlet Allocation or Large Language Model fine-tuning tasks where the parameter space is significantly larger?
Contents
GEP-CV: Breaking the Unimodal Barrier of Expectation Propagation
1. TL;DR
1.1. Academic Positioning
2. Problem & Motivation: The "Averaging" Trap
3. Methodology: Mixture Models and Variance Control
3.1. 1. From Single to Mixture
3.2. 2. The GEP-CV Architecture
4. Experimental Results: Stability and Performance
4.1. Quantitative Edge
4.2. The Power of Control Variates
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook