KNN meets LSI: Bridging the Accuracy Gap in Blog Gender Prediction
Gender prediction on a real life blog data set using LSI and KNN
The paper introduces a text classification framework for gender prediction on real-life blog data by combining Latent Semantic Indexing (LSI) with the K-Nearest Neighbor (KNN) algorithm. By leveraging Singular Value Decomposition (SVD) for dimensionality reduction, the authors successfully mitigate KNN's performance degradation in high-dimensional feature spaces, achieving a predictive accuracy of 69%.
TL;DR
Predicting author gender from social media posts is a classic yet challenging text classification task. While K-Nearest Neighbor (KNN) is a popular and intuitive algorithm, it traditionally collapses under the weight of high-dimensional text data. This paper proposes a robust solution by integrating Latent Semantic Indexing (LSI) to distill noisy blog posts into a low-dimensional semantic space. The result? A significant boost in accuracy to 69%, nearly matching the gold-standard performance of Naïve Bayes.
Problem & Motivation: The Curse of High Dimensions
In the world of Information Retrieval, gender prediction is often treated as a supervised learning problem. However, social media data (like blog posts) is notoriously messy. Traditional approaches using a Bag of Words (BOW) create feature spaces with thousands of dimensions—most of which are "noise" or irrelevant terms.
For a "lazy learner" like KNN, this is a nightmare. In high-dimensional spaces, the distance between any two documents becomes nearly uniform, making it impossible for KNN to find meaningful "neighbors." The authors’ core insight was that we don't need every word to identify gender; we need the latent concepts hidden within the text.
Methodology: The LSI + KNN Framework
The researchers developed a modular architecture to transform raw blog text into a refined coordinate system.
1. The Preprocessing Pipeline
The text undergoes a rigorous four-step cleaning process:
- Tokenization: Breaking down sentences into words.
- Stop Word Removal: Eliminating non-informative terms (e.g., "the", "is").
- Stemming: Using the Lancaster Stemmer to map "stemming" and "stemmed" to a single root "stem."
- TF-IDF Weighting: Scoring words based on their relative importance across the corpus.
2. Dimensionality Reduction via LSI
This is the "secret sauce" of the paper. Instead of using the raw TF-IDF matrix, the authors apply Singular Value Decomposition (SVD). By decomposing the matrix into and keeping only the top k singular values, they project the documents into a dense, low-rank space where semantic relationships are preserved.
Fig 1: The modular architecture of the proposed gender prediction system.
3. Classification
Once the validation documents are projected into the reduced -dimensional space, KNN uses Cosine Similarity to find the top 7 nearest neighbors. The predicted gender is determined by a majority vote among these neighbors.
Experiments & Results: Finding the "Sweet Spot"
The authors discovered that the performance of the model is highly sensitive to the Rank-k approximation. If is too small, critical information is lost; if is too large, the model plateaus or degrades due to noise.
- The Peak: The model achieved its highest accuracy of 69% at .
- Closing the Gap: While Naïve Bayes still holds the lead at 71%, the KNN+LSI approach is far more competitive than the traditional KNN+BOW (which lagged by 8%).
Fig 2: The impact of Rank-k on prediction accuracy, showing the optimal performance at k=34.
| Method | KNN Accuracy | Naïve Bayes Accuracy | Gap |
|---|---|---|---|
| Facebook (BOW) | 65% | 73% | 8% |
| Blog Posts (LSI) | 69% | 71% | 2% |
Critical Analysis & Conclusion
The Takeaway
The primary contribution of this work isn't just the implementation of KNN, but the focused analysis on Rank-k optimization. It proves that by carefully selecting the dimensionality of the latent space, we can empower simple, non-parametric algorithms to perform at near-SOTA levels.
Limitations & Future Work
The current study relies on an "enumeration" method to find the best . As the authors note, the next academic frontier is developing a mathematical heuristic to determine the optimal automatically without exhaustive testing. Additionally, while 69% is a strong result for a 1200-sample dataset, modern transformer-based embeddings (like BERT) would likely push this boundary further, albeit with much higher computational overhead.
In conclusion, the KNN+LSI method remains a highly efficient and interpretable choice for large-scale text classification where computational resources and semantic clarity are paramount.
