AMEN: Capturing the Long-Term Pulse of Social Media Personalization
Modeling the Long-Term Post History for Personalized Hashtag Recommendation
The paper introduces the Adaptive neural MEmory Network (AMEN), a novel memory-augmented architecture for personalized hashtag recommendation. By leveraging an external memory module, the system models long-term user post histories, achieving SOTA performance on Twitter datasets through the joint encoding of textual content and hashtag semantics.
TL;DR
Recommending the perfect hashtag requires more than just understanding a single tweet; it requires knowing the user's history. This paper presents AMEN (Adaptive neural MEmory Network), a specialized architecture that utilizes long-term post history and character-level hashtag semantics. By introducing a unique "Out-of-Memory" (OOM) gate, AMEN knows exactly when to rely on a user's past and when to treat a new tweet as a fresh start, setting a new SOTA on Twitter recommendation benchmarks.
Problem & Motivation: The Short-Sightedness of Current Models
Hashtags serve as the connective tissue of social media, yet recommending them is notoriously difficult. Most current SOTA models suffer from two primary limitations:
- Temporal Myopia: They only "see" the last few posts (fixed-length short-term history). Long-term shifts in interest or niche recurring topics are lost.
- Semantic Blindness: They treat hashtags like
#Christmasand#ChristmasEveas distinct, independent IDs, ignoring the character-level similarity that links them. - The Fresh Start Paradox: Sometimes users tweet about breaking news or trending events that have nothing to do with their history. Standard memory networks try to force-fit history into these scenarios, leading to poor recommendations.
Methodology: The Dual-View Memory Architecture
The core innovation of AMEN is how it manages its "Memory Module." Unlike standard Neural Turing Machines (NTMs), AMEN processes the content and the hashtags of historical posts as two different views.
1. Representation Learning
- Content (CNN): Uses a multi-filter convolutional layer to extract local n-gram features from the tweet text.
- Hashtags (char-RNN): Instead of one-hot encoding, AMEN uses a recurrent network over the characters of the hashtag. This allows the model to capture the relationship between similar-sounding tags.
2. The Adaptive Memory Read/Write
AMEN doesn't just store data; it filters it.
- The Write Module: Uses a gating mechanism () to decide if a historical post is worth remembering. This prevents "noisy" or irrelevant re-tweets from polluting the user's long-term profile.
- The Read Module and OOM Gate: When making a recommendation, the OOM gate () analyzes the current tweet. If the query is drastically different from everything in the memory, the gate shuts, and the system relies solely on the current post's content.
Figure 1: The overall architecture of the Adaptive neural MEmory Network.
Experiments & Results: Proving the Value of History
The researchers tested AMEN against a battery of baselines, ranging from traditional topic models (TPLDA) to sophisticated memory networks (HMemN2N).
Key Findings:
- Long-Term is King: AMEN outperformed short-term memory models by substantial margins (~8% better than HMemN2N on Hits@5).
- The OOM Advantage: The inclusion of the OOM gate accounts for a ~4% jump in Hits@1 accuracy. This proves that knowing when to forget is as important as knowing what to remember.
- Ablation Success: Removing the "Hashtag History" (only storing text) caused a significant drop, proving that what users tagged in the past is a stronger predictor of future tags than the text alone.
Table 1: AMEN vs. SOTA Baselines on Twitter Dataset.
Critical Analysis & Conclusion
AMEN makes a compelling case for "intelligent" memory management in social media. By treating hashtags as semantic sequences rather than static labels, it bridges the gap between NLP and recommendation systems.
Limitations:
- Memory Saturation: As shown in Table 4, increasing memory size () beyond 5 locations actually decreases performance, likely due to overfitting. This suggests the model still struggles with extremely high-dimensional history.
- Computation: Character-level RNNs and CNNs per history item add overhead that might be challenging for real-time inference at the scale of millions of users without heavy optimization.
Future Outlook: The "Dual-View" memory approach and the OOM gate are highly transferable. We could see this architecture adapted for personalized news feeds or even LLM context management, where the system must intelligently decide which parts of a "long context" are actually relevant to the current prompt.
