Decoding the Oval Office: Can Machines Predict Political Parties from Speeches Alone?

Distinguishing Political Parties From Inaugural Speeches Using Machine Learning

2020-08-01
SeungJin Lim, Hyoil Han
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the classification of US presidential inaugural speeches by political party (Democratic vs. Republican) using supervised machine learning. By utilizing lexical features like N-grams and stopword removal across historical data from 1789 to the present, the authors achieved 100% classification accuracy using Support Vector Machines (SVM).

Executive Summary

TL;DR: Researchers have developed a machine learning framework capable of identifying a US President's political party with 100% accuracy based solely on the text of their inaugural address. By analyzing 59 speeches from 1789 to the present, the study demonstrates that despite the "unifying" tone of inaugurals, distinct lexical signatures remain.

Background Positioning: This work bridges the gap between political science and computational linguistics. It moves beyond simple word clouds to prove that high-dimensional vector spaces (SVMs) can perfectly separate Republican and Democratic rhetoric without needing any semantic "understanding" or metadata.

Problem & Motivation: The Hidden Language of Power

Inaugural speeches are unique; they are designed to heal campaign wounds and speak to the entire nation. Because of this, one might assume the language is too homogenized to be easily categorized by a machine.

Prior work in text mining often focuses on sentiment or topic modeling. However, the authors of this study sparked a specific question: Are there underlying lexical "fingerprints" that transcend time and individual personality, tying a president to their party's platform? The challenge is filtering the noise—the common patriotic vocabulary like "nation" and "people"—to find the subtle discriminators.

Methodology: From Words to Vectors

The researchers treated each speech as a "Bag-of-Words" but refined the input through several specific technical stages:

  1. Preprocessing: Lemmatization to reduce words to their base forms (e.g., "running" to "run") and strict stopword removal to eliminate non-partisan words like "the" or "and."
  2. Feature Engineering: The use of N-grams (sequences of 1 to 5 words). Interestingly, the study found that Unigrams (single words) often provided the cleanest signal.
  3. Weighting: Applying a log-smoothing factor to term frequencies and normalizing vectors using the L2 norm (Euclidean distance) to ensure speech length wouldn't bias the results.

Model Architecture Table

The study compared three primary "standard" classifiers against the more robust SVM:

Classifier Categories

Experiments & Results: The SVM Breakthrough

The results were polarized. While Logistic Regression and Naive Bayes struggled—achieving accuracies barely better than a coin flip in some configurations (62-66%)—the Support Vector Machine (SVM) proved to be the "silver bullet."

SOTA Comparison

As shown in the table below, once the SVM was applied with Linear, Polynomial, or RBF kernels, the classification error dropped to zero.

SVM Performance Results

Why did SVM win? While Logistic Regression looks for a linear probability, SVM looks for the "Maximum Margin Hyperplane." In a high-dimensional space where unigrams act as dimensions, the SVM found a clear, albeit complex, boundary that perfectly separated the two parties' vocabularies.

Key Findings from N-grams

The authors noted that increasing complexity (using 4-grams or 5-grams) actually degraded performance for non-SVM models. This suggests that political identity in formal speeches is found in word choice (vocabulary) rather than specific phrases (which change too much over 200 years).

Critical Analysis & Conclusion

Takeaway

The study successfully provides a framework for "Political Narrative Analysis." It proves that the "flavor" of Democratic and Republican rhetoric is mathematically distinct enough to be captured by classical supervised learning.

Limitations

  • Data Size: With only 59 speeches, the model is highly susceptible to the "curse of dimensionality," though the 10-fold cross-validation helps mitigate this.
  • Kernel Sensitivity: The Sigmoid kernel performed poorly (52.17%), indicating that the data is not "linearly inseparable" in a way that the sigmoid function can easily map without heavy tuning.
  • Historical Drift: The Republican party of 1861 (Lincoln) is vastly different from the modern era. The model manages to find a thread that connects them, but it doesn't explain what that thread is (e.g., is it a focus on "duty" vs "progress"?).

Future Outlook

The next frontier for this research is moving from lexical (word-based) to semantic (meaning-based) analysis. Could a model trained on these speeches predict the party of a hypothetical "Independent" candidate based on their inaugural? This framework provides the baseline for such high-stakes political forecasting.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Transformer-based models (like BERT or RoBERTa) to classify US presidential speeches and compare their results to traditional SVM approaches.
  • Which study first established the use of Support Vector Machines for stylistic authorship attribution, and how does this paper's feature selection differ from that foundational work?
  • Explore research that applies lexical classification methods to modern social media posts by politicians to see if party identity is more or less distinct than in formal inaugural addresses.
Contents
Decoding the Oval Office: Can Machines Predict Political Parties from Speeches Alone?
1. Executive Summary
2. Problem & Motivation: The Hidden Language of Power
3. Methodology: From Words to Vectors
3.1. Model Architecture Table
4. Experiments & Results: The SVM Breakthrough
4.1. SOTA Comparison
4.2. Key Findings from N-grams
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook