MADM: Eliminating Discretization Bias in Diffusion Models with Metropolis Adjustments
Metropolis-Adjusted Diffusion Models
This paper introduces Metropolis-Adjusted Diffusion Models (MADM), a framework that integrates Metropolis-Hastings (MH) and Barker's accept-reject steps into the Predictor-Corrector diffusion sampling process. By utilizing a score-based line integral identity to estimate intractable density ratios, the authors provide the first exact correction method for diffusion models and a highly efficient Simpson's rule approximation to eliminate discretization bias.
TL;DR
Metropolis-Adjusted Diffusion Models (MADM) bridge the gap between MCMC theory and diffusion sampling. By replacing biased Langevin "corrector" steps with a novel, score-based accept-reject mechanism, MADM eliminates persistent discretization artifacts. It introduces an exact Bernoulli factory-based sampler and a hyper-efficient Simpson’s rule approximation that improves FID across standard benchmarks like ImageNet and FFHQ.
Motivation: The Hidden Bias in "Correction"
Diffusion models typically transform noise into data by simulating a reverse-time SDE. In the popular Predictor-Corrector (PC) framework, a predictor (like an ODE solver) moves the sample between noise levels, and a corrector (usually the Unadjusted Langevin Algorithm, or ULA) refines the sample.
However, there is a fundamental flaw: ULA is itself a biased sampler. Because it discretizes continuous Langevin dynamics, it never truly converges to the target distribution . In classical MCMC, we fix this using the Metropolis-Adjusted Langevin Algorithm (MALA). But MALA requires the ratio of densities , which is inaccessible in diffusion models—we only have the score .
Methodology: Score-Based Accept-Reject
The core insight of MADM is that while we don't know the density, we can calculate the Log-Density Ratio as a line integral of the score:
1. The Exact Approach: Two-Coin Bernoulli Factory
The authors propose the first exact Barker adjustment for diffusion. It uses a "Bernoulli factory"—a method to flip a coin with probability given flips of a coin with probability . By using a Poisson-truncated power series, they can decide to accept or reject a proposal based on the score function without ever approximating the integral.
2. The Practical Approach: Simpson’s 1/3 Rule
For large-scale image generation, the authors introduce a deterministic approximation using Simpson's Rule. By evaluating the score at the midpoint of a Langevin step, they achieve an error bound of . This is essentially "free" performance, requiring only one extra score evaluation per step.
Figure 1: Comparison of ODE Predictor, ULA Correction, and the proposed MADM Correction. Notice how MADM moves samples closer to the true data support.
Experiments and Performance
The researchers tested MADM on synthetic 2D Manifolds and high-resolution image datasets (CIFAR-10, FFHQ, ImageNet-64).
Key Findings:
- Outlier Removal: On synthetic datasets (Spirals, Pinwheels), MADM successfully moved "hanging" outliers back onto the data manifold where ULA failed.
- FID Gains: MADM provided systematic improvements in Fréchet Inception Distance (FID). Notably, on ImageNet, adding MADM to a simple Euler solver made it perform better than the significantly more expensive Heun solver.
Table 1: MADM consistently lowers FID scores across various datasets and ODE solvers.
Critical Analysis & Conclusion
MADM is a mathematically elegant solution to a long-standing "shrugged-off" problem in diffusion sampling.
Pros:
- Compatibility: Works with any pre-trained score-based model (Llama-gen, Stable Diffusion, etc.).
- Theoretical Rigor: Provides the first exact MCMC adjustment for score-only targets.
- Efficiency: The Simpson approximation is scalable for production environments.
Limitations:
- The exact Bernoulli factory method is computationally expensive (random number of score hits).
- It still assumes the learned score is accurate; if the neural network provides a poor score estimate, the "exact" adjustment will correctly sample a "wrong" distribution.
Takeaway: As generative models move toward high-precision applications (like protein design or architecture), the "good enough" discretization of ULA will no longer suffice. MADM provides the toolkit to turn diffusion models into precise statistical samplers.
