Mining the Cabinet: How Automated Social Network Analysis Explains Presidential Popularity
Automatic Mapping of Social Networks of Actors from Text Corpora: Time Series Analysis
This paper introduces the WORDij approach for the automatic identification of social networks by mining large-scale text corpora. Applied to news stories covering the Clinton and Bush administrations, the method extracts co-occurrence networks among cabinet members to analyze the relationship between network centrality and presidential job approval ratings.
TL;DR
By processing over 200,000 news stories, researchers used the WORDij 3.0 toolkit to map the shifting social networks of the Clinton and Bush cabinets. The study reveals a fascinating mechanical link: how "central" a president appears in the news compared to their cabinet members directly impacts their Gallup job approval ratings.
Executive Summary
In the era of the 24-hour news cycle, a leader's public image is not just about what is said, but who they are associated with. This paper demonstrates a robust methodology for Automatic Mapping of Social Networks from text. By synchronizing news-mined networks with Gallup polls, the authors prove that media-portrayed social structures have high predictive validity for political outcomes.
The Problem: Beyond the "Bag of Words"
Standard text mining often treats documents as a "bag of words," losing the structural context of how people interact. Furthermore, static network analysis fails to capture the evolution of influence over time.
The authors argue that the media has a "negative information orientation." If a President is the sole central figure (High Centrality), they become a "lightning rod" for criticism. If the cabinet is more central (Structural Dispersion), the negativity is distributed, often leading to higher presidential approval.
Methodology: Proximity-Based Network Extraction
The researchers used a sophisticated pipeline to ensure the networks reflected actual social distance:
- Corpus: 114,511 stories for Clinton; 89,810 for Bush.
- Proximity Window: Instead of whole documents, a 100-word window was used to define co-occurrence, better reflecting human readability and journalistic intent.
- Eigenvector Centrality: Unlike simple "betweenness," this measure accounts for the quality of connections—being connected to other important people increases your score.
- Time-Slicing: Data was segmented into 22-30 day intervals to perfectly match the frequency of Gallup polling data.
Figure 1 & 2: The aggregate social structure of the Clinton (left) and Bush (right) administrations as portrayed by the New York Times and Washington Post.
Key Results: The "Lightning Rod" Effect
The findings confirm that the topology of an administration in the press predicts its popularity:
- The Clinton Pattern: Support for the "lightning rod" theory was strong. When Clinton was highly central, his approval dropped. When his cabinet members took the spotlight (Average Cabinet Centrality increased), his approval rose by 11%.
- The Bush Deviation: For G.W. Bush, the effect was "synchronous" rather than lagged, likely due to a faster news cycle. Interestingly, for Bush, his own centrality slightly helped his approval (+5%), suggesting different media strategies or public perceptions of leadership.
- Sentiment Analysis: Both administrations operated in a "negative" environment (positivity ratios below the "Losada Line" of 1.0), but Bush's coverage contained higher extremes of both positive and negative emotion.
Figure 6: Cross-correlation showing how Bush's approval reacted to network centrality over multiple time lags.
Critical Insight & Conclusion
The significance of this work lies in its Cross-System Predictive Validity. It proves that the "media-constructed reality" of social networks isn't just an abstract graph—it is a tangible predictor of public opinion.
Takeaway for Tech & Pol-Sci:
- Structural Dispersion matters: For organizational leaders, distributing media visibility among a team can buffer the leader from negative press.
- Methodology: Proximity-based co-occurrence remains a powerful, computationally efficient alternative to complex NLP models for large-batch longitudinal studies.
Limitations
The study is limited to two administrations, and the "optimal window size" (100 words) was determined by intuition rather than rigorous benchmarking. Future work could benefit from Ontological Categories (automatically identifying new actors) rather than relying on a-priori lists of names.
