DiffusionFF: Redefining Deepfake Detection via Iterative Artifact Localization

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

2025-01-01
Siran Peng, Haoyuan Zhang, Li Gao, Tianshuo Zhang, Bao Li, Zhen Lei
Summary
Problem
Method
Results
Takeaways
Abstract

DiffusionFF is a novel joint framework for face forgery detection and fine-grained artifact localization that utilizes a denoising diffusion model as a decoder. It achieves State-of-the-Art (SOTA) performance across multiple benchmarks, including a 97.24% AUC on the Celeb-DF dataset.

TL;DR

DiffusionFF is a breakthrough framework that treats deepfake detection not just as a classification problem, but as a generative localization task. By leveraging a Denoising Diffusion Model to reconstruct fine-grained artifact maps (DSSIM), it provides human-interpretable evidence of forgery while simultaneously boosting classification accuracy to SOTA levels (e.g., 97.24% AUC on Celeb-DF).

The "Blurry" Problem in Modern Forensics

In the arms race against increasingly realistic Deepfakes, binary "real vs. fake" classification is no longer enough. We need to know where the forgery is to build user trust. However, current localization methods face a dilemma:

  • Mask-based methods provide only coarse "blobs" of manipulated regions.
  • Regression-based DSSIM methods (like LiSiam) theoretically capture pixel-level details, but in practice, they produce blurry, smoothed-out maps that lose subtle manipulation traces.

The authors' core insight is that Direct Regression is ill-suited for the complex, high-frequency signals of forgery artifacts. Instead, they turn to the iterative power of Diffusion Models.

Methodology: Encoder-Decoder Reimagined

DiffusionFF establishes a novel architecture by repurposing existing components:

  1. The Artifact Encoder: A pretrained forgery detector (like ConvNeXt) extracts multi-scale features that represent "forgery-ness" at different semantic levels.
  2. The Artifact Decoder: A diffusion model takes these features as conditions to iteratively transform random noise into a detailed DSSIM map (the pixel-wise difference between a forged image and its latent real counterpart).
  3. Feature Fusion: The high-fidelity DSSIM map isn't just an output; it's fed back into a gating mechanism to help the classifier focus on the most "suspicious" pixels.

DiffusionFF Overall Architecture

Why It Works: The Iterative Advantage

Unlike standard U-Nets that try to guess the artifact map in one "shot," the Diffusion decoder progressively refines the result. This allows the model to capture sharp, fine-grained inconsistencies (like blending edges or texture mismatches) that regression models simply smooth over.

Quantitative and Qualitative Excellence

The results are striking. On the FaceForensics++ (FF++) dataset, the Fréchet Inception Distance (FID)—a measure of map quality—dropped from 241.6 in previous work to just 43.1, indicating a massive leap in the realism and precision of the localization maps.

Visual Comparison of Artifact Localization

SOTA Performance and Generalization

One of the harshest tests for deepfake models is Cross-Dataset Evaluation (training on one dataset and testing on another). DiffusionFF consistently outclassed 16 different SOTA baselines:

  • DFDC: 85.05% AUC (surpassing SBI and KFD).
  • Celeb-DF-v2: 97.24% AUC.
  • Robustness: The model maintains higher accuracy than competitors under real-world degradations like JPEG compression and Gaussian blur.

Cross-Dataset Performance Table

Critical Analysis & Conclusion

Takeaway

DiffusionFF proves that high-quality visual explanations (localization maps) are not just "nice to have"—they are features that, when fused back into the model, directly improve classification performance.

Limitations

  • Inference Speed: Like all diffusion models, the iterative denoising process (even at ) is slower than simple CNN or Transformer inference.
  • Scope: The method relies on finding facial artifacts and may not be directly applicable to full-image GAN generations where no original facial alignment exists.

Future Outlook

The "Inverted ControlNet" strategy used here—freezing the conditioning encoder and training the diffusion decoder—could be a blueprint for other specialized vision tasks like medical imaging anomalies or industrial defect detection.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize diffusion models specifically for pixel-level forgery localization or image forensics rather than image synthesis.
  • What are the original theoretical foundations of using Structural Dissimilarity (DSSIM) maps for image quality assessment and how was it first adapted for deepfake detection?
  • Explore research that applies multi-scale feature conditioning from frozen encoders to optimize diffusion-based decoders in non-generative tasks.
Contents
DiffusionFF: Redefining Deepfake Detection via Iterative Artifact Localization
1. TL;DR
2. The "Blurry" Problem in Modern Forensics
3. Methodology: Encoder-Decoder Reimagined
4. Why It Works: The Iterative Advantage
4.1. Quantitative and Qualitative Excellence
5. SOTA Performance and Generalization
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook