Relational Classifiers in a Non-relational World: Unlocking the Power of Homophily
Relational Classifiers in a Non-relational World: Using Homophily to Create Relations
This paper introduces a systematic framework for applying Statistical Relational Learning (SRL) to non-relational, single-table datasets by artificially constructing relations based on attribute similarity. Using the weighted-vote Relational Neighbor (wvRN) classifier, the method achieves superior performance over traditional machine learners on the majority of 31 UCI benchmark datasets.
TL;DR
Statistical Relational Learning (SRL) has long been the gold standard for networked data, but most of our data lives in flat, non-relational tables (like UCI benchmarks). This paper presents a breakthrough "relationalization" framework: by treating similarities between attributes as "virtual links" and weighting them through a homophily-based metric (Assortativity), we can use relational classifiers to outperform traditional models like SVMs and Logistic Regression.
Motivation: Why Bridge the Gap?
Standard machine learning treats data points as "islands"—the iid assumption. However, reality is rarely independent. Two patients with similar symptoms or two financial transactions with similar timestamps are inherently "related." The author argues that even if a dataset doesn't have explicit links (like citations or hyperlinks), we can invent them to leverage Collective Inference—a process where the label of one instance helps settle the label of its neighbors.
Methodology: How to Build a Network from a Table
The transformation from a CSV-style table to a Graph involves three steps:
1. Relation Construction
For every attribute in the dataset, a potential edge type is created:
- Categorical Attributes: An edge exists if .
- Numerical Attributes: Edge weights are calculated based on normalized distance:
- Instance Similarity: A global link based on the inverse of Euclidean distance across all attributes.
2. Identifying Signal (Assortativity)
Not all relations are useful. The paper uses the Node-based Assortativity Coefficient () to filter the noise. If a specific attribute link doesn't result in "like-linking-with-like" (low homophily), the relation is pruned.
3. Collective Inference (wvRN + RL)
The model uses the weighted-vote Relational Neighbor (wvRN) algorithm, which estimates a node's class by averaging its neighbors' probabilities. This is refined through Relaxation Labeling (RL)—a simulated annealing process that iteratively settles the entire network into a consistent state.

Experimental Showdown
The author tested this approach against heavyweights: J48 (C4.5), k-NN, Logistic Regression, Naive Bayes, SMO (SVM), and Transductive SVM.
Key Findings:
- Superiority: wvRN won 12 out of 31 datasets, more than any other single classifier.
- The Threshold Effect: Relational inference is a "data hungry" link-builder. When training data is sparse (10%), it struggles because the Assortativity weights cannot be estimated accurately. However, once training data hits 50%, collective inference (wvRN-RL) becomes the dominant force.

Critical Insight: The "Relationalness" of All Data
The most striking takeaway is that relational learning isn't just for social networks—it's a general-purpose tool for any data where "similarity implies shared identity."
Limitations & Future Work
- Parameter Tuning: The current study used "vanilla" settings. Hyperparameter optimization on the Assortativity thresholds could lead to even higher gains.
- Sparsity: The method remains sensitive to the initial "seed" of labeled data. Future research could explore semi-supervised methods to better bootstrap the relation weights.
Conclusion
By converting attributes into relations, we move from viewing data as isolated points to viewing it as a rich, interconnected manifold. This paper provides the mathematical and empirical justification to stop treating UCI datasets as "flat" and start treating them as networks.
