RLSBI: Mastering Explainable Deepfake Detection via RL-Enhanced Self-Blending
Explainable Deepfake Detection with RL Enhanced Self-Blended Images
The paper introduces RLSBI, an explainable deepfake detection framework that utilizes Reinforcement Learning (RL) and a Chain-of-Thought (CoT) data generation pipeline based on Self-Blended Images (SBI). By fine-tuning Multimodal Large Language Models (MLLMs) with automated forgery attribution, it achieves SOTA-level cross-dataset generalization (e.g., 0.963 AUC on Celeb-DF v2) without requiring dedicated external detectors.
Executive Summary
TL;DR: RLSBI is a novel framework that bridges the gap between high-performance deepfake detection and human-readable explanations. By combining Self-Blended Images (SBI) with Reinforcement Learning (specifically GRPO), the authors created a system that not only detects forgeries with SOTA accuracy but also explains why it thinks an image is fake by pointing to specific regional anomalies.
Background: Most deepfake detectors are "black boxes." While recent trends shift toward using Multimodal Large Language Models (MLLMs), these models usually lag behind dedicated detectors because they lack high-quality, annotated training data. RLSBI solves this by automating the data generation and using RL to refine the model's reasoning.
The Core Challenge: Reward Sparsity and Annotation Fog
Deepfake detection is fundamentally a binary classification problem (Real vs. Fake). However, training a Large Language Model using Reinforcement Learning (RL) on binary labels is notoriously difficult—the feedback signal is too "sparse" to guide the model through complex reasoning steps.
Furthermore, creating "Chain-of-Thought" (CoT) data—where a model explains its logic—requires expert human annotation, which is expensive and doesn't scale. Previous methods tried to use GPT-4o to describe differences, but these often suffer from "hallucinations" or imprecise localization.
Methodology: Reinforcing the "Why"
The authors' insight was to turn the deepfake generation process itself into a supervisor.
1. Automated CoT Generation
Instead of asking humans to describe a fake image, the system uses the metadata from the Self-Blended Images (SBI) process. Since the algorithm knows exactly where it blended pixels and what transformations (hue, contrast, scaling) it applied, it can automatically generate precise labels like "There is a slight blurring on the right eye boundary due to aggressive blending."

2. Group Relative Policy Optimization (GRPO)
To optimize the MLLM, the authors adopted the GRPO algorithm (popularized by DeepSeek-R1). To solve the reward sparsity, they introduced a Keyword-Driven Reward Mechanism:
- Accuracy Reward: Did it get the Real/Fake label right?
- Format Reward: Is the output in the correct XML/Markdown structure?
- Localization Reward: Does the text description of the "tampered region" match the actual mask used during synthesis? (Calculated via Jaccard similarity).
3. Adaptive Feedback Loop
The architecture includes a dynamic feedback mechanism. If the model is performing well (high rewards), the synthesis engine generates harder fakes (subtle blending). If the model struggles, it generates easier fakes. This creates a curriculum learning effect.

Performance and SOTA Comparison
RLSBI was tested against industry heavyweights like RECCE and SBI. The results show that it doesn't just match dedicated detectors—it often beats them while providing explanations.
- Celeb-DF v2: Achieved 0.963 AUC, outperforming the original SBI (0.886) and X2-DFD (0.955).
- Ablation Insight: Moving from standard supervised fine-tuning (SFT) to RL with feedback increased the AUC from 0.881 to 0.905 on challenging cross-dataset tasks.
| Method | MLLM | CDF2 (AUC) | DFD (AUC) |
|---|---|---|---|
| SBI (Original) | No | 0.886 | 0.827 |
| X2-DFD | Yes | 0.955 | 0.957 |
| RLSBI (Ours) | Yes | 0.963 | 0.965 |
Critical Analysis & Conclusion
Takeaway: RLSBI proves that MLLMs can be superior deepfake detectors if we stop treating them as simple classifiers and start treating them as reasoning agents. By grounding the model's rewards in the physical reality of how the image was manipulated (the SBI mask), the authors successfully mitigated the "hallucination" problem common in LLM-based vision tasks.
Limitations: The model performs slightly worse on the DFDC dataset. The authors attribute this to the fact that DFDC contains many non-blending artifacts (e.g., full-frame GAN generation) and extreme environmental degradation that the SBI-style augmentation doesn't fully cover.
Future Outlook: The integration of RL with synthetic data generation strategies represents a powerful shift. Future work will likely extend this to video-level reasoning where temporal inconsistencies (flickering) can be synthesized and rewarded in a similar CoT fashion.
