KNN meets LSI: Bridging the Accuracy Gap in Blog Gender Prediction

Gender prediction on a real life blog data set using LSI and KNN

2017-01-01
Jianle Chen, Tianqi Xiao, Jie Sheng, Ankur Teredesai
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a text classification framework for gender prediction on real-life blog data by combining Latent Semantic Indexing (LSI) with the K-Nearest Neighbor (KNN) algorithm. By leveraging Singular Value Decomposition (SVD) for dimensionality reduction, the authors successfully mitigate KNN's performance degradation in high-dimensional feature spaces, achieving a predictive accuracy of 69%.

TL;DR

Predicting author gender from social media posts is a classic yet challenging text classification task. While K-Nearest Neighbor (KNN) is a popular and intuitive algorithm, it traditionally collapses under the weight of high-dimensional text data. This paper proposes a robust solution by integrating Latent Semantic Indexing (LSI) to distill noisy blog posts into a low-dimensional semantic space. The result? A significant boost in accuracy to 69%, nearly matching the gold-standard performance of Naïve Bayes.

Problem & Motivation: The Curse of High Dimensions

In the world of Information Retrieval, gender prediction is often treated as a supervised learning problem. However, social media data (like blog posts) is notoriously messy. Traditional approaches using a Bag of Words (BOW) create feature spaces with thousands of dimensions—most of which are "noise" or irrelevant terms.

For a "lazy learner" like KNN, this is a nightmare. In high-dimensional spaces, the distance between any two documents becomes nearly uniform, making it impossible for KNN to find meaningful "neighbors." The authors’ core insight was that we don't need every word to identify gender; we need the latent concepts hidden within the text.

Methodology: The LSI + KNN Framework

The researchers developed a modular architecture to transform raw blog text into a refined coordinate system.

1. The Preprocessing Pipeline

The text undergoes a rigorous four-step cleaning process:

  • Tokenization: Breaking down sentences into words.
  • Stop Word Removal: Eliminating non-informative terms (e.g., "the", "is").
  • Stemming: Using the Lancaster Stemmer to map "stemming" and "stemmed" to a single root "stem."
  • TF-IDF Weighting: Scoring words based on their relative importance across the corpus.

2. Dimensionality Reduction via LSI

This is the "secret sauce" of the paper. Instead of using the raw TF-IDF matrix, the authors apply Singular Value Decomposition (SVD). By decomposing the matrix into and keeping only the top k singular values, they project the documents into a dense, low-rank space where semantic relationships are preserved.

Implementation Architecture Fig 1: The modular architecture of the proposed gender prediction system.

3. Classification

Once the validation documents are projected into the reduced -dimensional space, KNN uses Cosine Similarity to find the top 7 nearest neighbors. The predicted gender is determined by a majority vote among these neighbors.

Experiments & Results: Finding the "Sweet Spot"

The authors discovered that the performance of the model is highly sensitive to the Rank-k approximation. If is too small, critical information is lost; if is too large, the model plateaus or degrades due to noise.

  • The Peak: The model achieved its highest accuracy of 69% at .
  • Closing the Gap: While Naïve Bayes still holds the lead at 71%, the KNN+LSI approach is far more competitive than the traditional KNN+BOW (which lagged by 8%).

Accuracy vs. Rank-k Fig 2: The impact of Rank-k on prediction accuracy, showing the optimal performance at k=34.

MethodKNN AccuracyNaïve Bayes AccuracyGap
Facebook (BOW)65%73%8%
Blog Posts (LSI)69%71%2%

Critical Analysis & Conclusion

The Takeaway

The primary contribution of this work isn't just the implementation of KNN, but the focused analysis on Rank-k optimization. It proves that by carefully selecting the dimensionality of the latent space, we can empower simple, non-parametric algorithms to perform at near-SOTA levels.

Limitations & Future Work

The current study relies on an "enumeration" method to find the best . As the authors note, the next academic frontier is developing a mathematical heuristic to determine the optimal automatically without exhaustive testing. Additionally, while 69% is a strong result for a 1200-sample dataset, modern transformer-based embeddings (like BERT) would likely push this boundary further, albeit with much higher computational overhead.

In conclusion, the KNN+LSI method remains a highly efficient and interpretable choice for large-scale text classification where computational resources and semantic clarity are paramount.

Find Similar Papers

Try Our Examples

  • Search for recent papers that compare Latent Semantic Analysis (LSA) and Word2Vec embeddings for improving the accuracy of KNN-based text classifiers.
  • Which seminal paper first introduced Latent Semantic Indexing for information retrieval, and how has its application in authorship attribution evolved since then?
  • Explore research that applies Singular Value Decomposition (SVD) for dimensionality reduction in multi-modal social media classification tasks involving both text and images.
Contents
KNN meets LSI: Bridging the Accuracy Gap in Blog Gender Prediction
1. TL;DR
2. Problem & Motivation: The Curse of High Dimensions
3. Methodology: The LSI + KNN Framework
3.1. 1. The Preprocessing Pipeline
3.2. 2. Dimensionality Reduction via LSI
3.3. 3. Classification
4. Experiments & Results: Finding the "Sweet Spot"
5. Critical Analysis & Conclusion
5.1. The Takeaway
5.2. Limitations & Future Work