Who Am I? 2CLSTM: Decoding Personality through Textual Structure
Who Am I? Personality Detection Based on Deep Learning for Texts
This paper introduces 2CLSTM, a hybrid deep learning model combining bidirectional LSTMs and CNNs for personality detection from social media texts. By targeting the "Big Five" personality traits, the model specifically extracts structural features of language to achieve competitive results on both long-form and short-form datasets.
TL;DR
Researchers from Southeast University have developed a novel deep learning architecture, 2CLSTM, which bridges the gap between reading (RNN) and composing (CNN) text to identify personality traits. By introducing the concept of Latent Sentence Groups (LSGs), the model identifies how the logical structure of our writing—not just the words we choose—reveals whether we are introverted, neurotic, or open to new experiences.
The Missing Piece: Why Words Aren't Enough
Predicting personality via the Big Five Model (Openness, Conscientiousness, Extroversion, Agreeableness, and Neuroticism) has long relied on "what" people say. Tools like LIWC (Linguistic Inquiry and Word Count) count the frequency of positive words or pronouns.
However, the authors argue that personality is also hidden in "how" we organize our thoughts. Current models often miss the structural flow of a document. For example, a conscientious person might structure their arguments more logically than someone high in neuroticism. Capturing this "latent structure" is the primary challenge addressed by this work.
Methodology: The 2CLSTM Architecture
The 2CLSTM model operates on a "Dual-Network" logic:
- Bi-directional LSTMs (The Readers): These mimic the human reading process, scanning text forward and backward to understand the context surrounding every word.
- CNN with LSG (The Architects): The output of the LSTMs is fed into a CNN. Here, the authors introduce the Latent Sentence Group (LSG). Instead of looking at fixed word windows, the CNN uses varied kernels (1, 2, and 3-grams) to identify clusters of sentences that are logically or semantically linked.

Fig. 1: The 2CLSTM pipeline, moving from Word Embedding to Contextual Encoding (LSTMs) and finally Structural Feature Extraction (CNN).
The LSG Insight
The "Latent" in LSG refers to relationships between sentences that exist within the high-dimensional vector space. These sentences don't have to be physically adjacent; the CNN identifies these groups as a "synthesis of sentence vectors closely connected in specific coordinates."
Experiments and Results
The study is notable for testing on two very different data sources:
- Stream-of-consciousness essays: Long-form, introspective texts (Average 648 words).
- YouTube Transcripts: Short-form, conversational texts (Average 526 words).
Performance Comparison
The 2CLSTM model consistently landed in the top-performing tier across nearly all personality traits. Interestingly, the model performed slightly better on the YouTube dataset, suggesting that our "outer-perception" (how others see us) might be easier for AI to decode than our "autognosis" (how we see ourselves in essays).

Table 1: Macro Precision comparison across models. 2CLSTM (bottom row) shows dominant performance, particularly in Neuroticism and Extroversion.
Critical Insight & Future Directions
The core achievement of this paper is proving that textual structure is a valid biometric of personality. While a person can consciously choose specific words to appear a certain way, their underlying cognitive structure—mirrored in how they group sentences—is harder to mask.
Limitations: The model currently treats each of the Big Five traits as a separate classification task. Human personality is a spectrum where these traits interact; future models might benefit from Multi-task Learning (MTL) to capture the correlations between, for instance, Extroversion and Agreeableness.
Future Work: The team plans to expand this structural analysis to multi-modal data, including social media photos and voice patterns, moving toward a truly holistic digital "soul" profile.
Conclusion
The 2CLSTM model represents a shift from "bag-of-words" analysis to "architecture-of-thought" analysis. By combining the strengths of RNNs and CNNs, it provides a more nuanced lens through which AI can understand the complex human psyche.
