Beyond Real-Time Interaction: Stabilizing Social Similarity with Evolutionary and Neural Learning
Using Social Information to Compose a Similarity Function Based on Friends Attendance at Events
This paper proposes the Social Similarity Function Generator Method, which leverages stable social variables (age, education, etc.) to learn personalized similarity functions for ranking friends based on event attendance. The study demonstrates that an Artificial Neural Network (ANN) approach significantly outperforms generalist methods, achieving a Spearman Ranking Correlation (SRC) improvement of nearly 69% in specialist settings.
TL;DR
Determining how "similar" or "influential" a friend is often relies on tracking every like, share, or check-in—data that is notoriously volatile. This paper introduces the Social Similarity Function Generator, a method that replaces fleeting interaction data with stable social variables (like age and education) to predict event-based affinity. By using Artificial Neural Networks (ANN) and Genetic Algorithms (GA), the researchers achieved up to a 0.918 correlation with actual real-world social behavior.
Background: The Perishability of "Likes"
In the context of the Social Web, "Similarity" is the engine of recommendation. However, researchers Luiz Mario Lustosa Pascoal and colleagues argue that interaction-based similarity is perishable. A friend you went to concerts with last summer might not be the one you follow to conferences this winter.
The core insight of this paper is that while behavior changes, the underlying social profile (age, gender, education, relationship status) that binds friends together is relatively stable. If we can map these stable variables to a ranking of a person's most "influential" friends at events, we create a similarity function that lasts longer and requires less retraining.
Methodology: Ranking via Optimization
The authors proposed two distinct modeling paths to discover the "hidden" similarity function:
- Universal Function Approximator Methods (UFAM): Using ANN (Multilayer Perceptron) and SVM. These are "black box" models capable of finding complex, non-linear relationships between a user's profile and their friends' importance.
- Populational Evolutionary Methods (PEM): Using Genetic Algorithms (GA), Evolution Strategy (ES), and Particle Swarm Optimization (PSO). These evolve a linear weighted sum (e.g., ), providing more "interpretable" results regarding which variable matters most.

The system extracts a Control List (CL) from Facebook, ranking friends by their frequency at events. The optimization algorithms then try to replicate this ranking using only the four social variables.
The "Specialist" vs. "Generalist" Battle
A key contribution of the paper is the experimental split between two approaches:
- Generalist: Training on data from months 1-6 and testing on months 7-11.
- Specialist: Retraining the model for each specific month to capture the highest possible accuracy.
Performance Comparison
The results (measured by Spearman Ranking Correlation) highlight a massive divide:
| Algorithm | Specialist SRC | Generalist SRC |
|---|---|---|
| ANN (Neural Net) | 0.918 | 0.221 |
| GA (Genetic) | 0.651 | 0.175 |
| SVM | 0.583 | 0.146 |
| Random | 0.005 | 0.005 |

Deep Insights
- Non-linearity Wins: The gap between ANN (0.918) and the linear methods (~0.65) suggests that social affinity is not a simple weighted sum. There are complex interactions—perhaps age similarity matters more for certain education levels—that only deep learning layers can catch.
- The Cost of Generalization: The sharp drop in the Generalist approach shows that even "stable" variables have a shelf life. A similarity function learned in January is significantly less effective by July, though it still outperforms random selection.
- Evolutionary Interpretation: For those needing to explain why a recommendation was made, the PEM (Evolutionary) methods are valuable. They produced coefficients that allow researchers to see if "Age" or "Relationship Status" was the primary driver for a specific user.
Critical Analysis & Future Outlook
While the ANN Specialist approach is the clear winner for accuracy, its "black box" nature and the need for constant retraining (high computational cost) are drawbacks.
The authors acknowledge a significant real-world hurdle: Data Sparsity. Many users do not fill out their Facebook profiles fully, leading to noisy data. Future iterations of this work aim to incorporate more "behavioral" variables like likes and shares, and contextual data like geolocalization, to bridge the gap between the Specialist and Generalist performance.
Takeaway: This research moves the needle from tracking what people do to understanding who they are, providing a more robust foundation for social recommendation systems.
