DGRU: Bridging Feature Interaction and Interest Evolution for Smarter Social Advertising
A Hybrid Neural Network Architecture to Predict Online Advertising Click-Through Rate Behaviors in Social Networks
The paper introduces DGRU, a hybrid neural network architecture for Click-Through Rate (CTR) prediction in social networks. It combines DeepFM for high-order feature interaction learning with a Gated Recurrent Unit (GRU) to model user interest evolution based on simplified binary click sequences.
TL;DR
In the high-stakes world of online advertising, understanding not just what a user is, but how their interests evolve is the key to conversion. This paper presents DGRU, a hybrid model that fuses the feature-engineering prowess of DeepFM with the temporal memory of GRU. By creatively using binary sequences (1s and 0s) of past behaviors rather than heavy item IDs, DGRU captures user interest evolution with minimal overfitting and maximum efficiency.
Background: The Sparse Data Trap
Traditional Click-Through Rate (CTR) models face a persistent dilemma. To understand users, they need historical data. However, using large-scale item IDs (e.g., millions of specific products) as input often leads to overfitting because most items have sparse interaction data. Furthermore, many models focus only on what users did click, ignoring the rich information hidden in what they ignored.
The authors position DGRU as a robust solution that balances "Memorization" (feature combinations) with "Generalization" (interest evolution) without the computational overhead of item-heavy sequences.
Methodology: The Best of Both Worlds
The DGRU architecture is a dual-engine system designed to tackle feature sparsity and temporal dynamics simultaneously.
1. The Deep Component (DeepFM)
This module handles the static attributes. It uses a Factorization Machine (FM) to capture 1st and 2nd-order feature interactions and a Multi-Layer Perceptron (MLP) for high-order interactions. This ensures that complex relationships (e.g., "young users on mobile devices at night") are automatically mined.
2. The GRU Component (Temporal Memory)
This is the core innovation. Instead of feeding the model a sequence of item embeddings, DGRU feeds it a simple sequence of 0s (not clicked) and 1s (clicked).
- Why 0s and 1s? It avoids the "dimensionality curse" of item IDs.
- Why GRU? Gated Recurrent Units are adept at managing long-term dependencies and capturing how a user’s interest shifts over time.

The outputs of both components are fused through a sigmoid function to produce the final click probability.
Experiments & Results: Performance Meets Efficiency
The model was stress-tested on three massive real-world datasets: ZhiHu (Q&A social network), Turing Federation (Video), and Panshi (Advertising).
SOTA Comparison
DGRU consistently outperformed established baselines like LR, Wide&Deep, and even the sophisticated xDeepFM.
- Accuracy: AUC increased by up to 2.50% over standard deep models.
- Error Reduction: LogLoss and RMSE saw significant improvements (up to 18.29% LogLoss reduction on the Panshi dataset).

Computational Efficiency
One of the most impressive findings is that despite adding a recurrent unit (GRU), the training time remains low. Because the GRU processes simple binary inputs rather than high-dimensional embeddings, it avoids the bottleneck seen in models like xDeepFM, which can be 5x slower.

Critical Insights: Why Does it Work?
- The Power of Negative Feedback: Most models treat "no-click" as noise. DGRU treats the sequence of non-clicks as a signal of "disinterest," which is just as important for filtering out irrelevant ads.
- Conciseness as Regularization: By stripping away the specific identities of items and focusing on the pattern of behavior (click vs. no-click), the model becomes much more robust against overfitting—a common "death sentence" for deep recommendation models.
Summary & Future Outlook
DGRU proves that you don't always need "big data" in the sense of high-dimensional inputs to get "big results." A clever, lightweight representation of user behavior, combined with a strong hybrid architecture, can outperform much heavier models.
Future Directions: The authors suggest extending this to multi-category feedback (e.g., distinguishing between "liking," "following," and "disliking") and integrating social graph data to further refine the interest modeling.
