TACCTI: Scaling the Identification of Cultural Capital in STEM Education
Using Text Analytics on Reflective Journaling to Identify Cultural Capitals for STEM Students
The paper introduces TACCTI (Text analysis and mAchine learning for Cultural Capital Theme Identification), a computational framework designed to automatically detect "Cultural Capital Themes" (CCT) in reflective journals written by STEM students from underrepresented backgrounds. Utilizing models ranging from Logistic Regression to BERT, the system identifies themes like 'Attainment' and 'First Generation' status, with the fine-tuned BERT model achieving a state-of-the-art F1-score of 0.92 for the Attainment category.
TL;DR
Researchers at San Francisco State University have developed TACCTI (Text analysis and mAchine learning for Cultural Capital Theme Identification). This framework automates the detection of "Cultural Capitals"—the unique strengths and aspirations of historically underrepresented (HU) students—in reflective journals. By shifting from essay-level to sentence-level analysis and leveraging BERT, the system achieves human-level accuracy in identifying career-oriented goals, paving the way for scalable, personalized student support.
Problem & Motivation: The Bottleneck of Qualitative Analysis
In the "Alma" project, students reflect on deep questions like "Why am I here?" to reinforce their sense of belonging in challenging STEM fields. While these journals are gold mines for understanding student motivation (e.g., being a first-generation student or having specific career "Attainment" goals), analyzing them is an administrative nightmare.
The Alma team initially used 11 researchers to manually code essays. This manual process is:
- Labor Intensive: Requires multiple passes to ensure inter-rater reliability.
- Non-Scalable: Impossible to implement for thousands of students across an entire university system.
- Sparse: Cultural capital mentions are often "needles in a haystack," buried within long reflective passages.
Methodology: From Essays to Sentences
The core insight of TACCTI is granularity. The authors discovered that evaluating an entire essay at once dilutes the signal. As shown in the paper's preliminary tests, switching to sentence-level classification drastically improved F1-scores because it allowed models to focus specifically on the "active" parts of the text.
The TACCTI Architecture
The framework processes text through an Input Data Pre-processing (IDP) module that handles the transition from human-coded essay excerpts back to structured sentence-level labels.

Advanced Feature Engineering
For their baseline models (Logistic Regression and Random Forest), the authors didn't just use words; they engineered "Expert Knowledge" features:
- STEM Similarity: Using WordNet to calculate the semantic distance between student words and core STEM fields.
- Named Entity (NE) Features: Identifying mentions of specific professions or job titles.
- Word Embeddings: 100-dimensional Word2Vec vectors trained specifically on student essays to capture the "local" vocabulary of the CS/Bio/Physics majors.
Experimental Results: The BERT Dominance
The researchers compared four major approaches for the "Attainment" theme:
- Logistic Regression (LR): The interpretable baseline.
- Random Forest (RF): Handled noisy features well.
- Bi-LSTM: Captured some sequence info.
- BERT: The Transformer-based heavy hitter.
| Model | Precision | Recall | F1-Score |
|---|---|---|---|
| Logistic Regression | 0.79 | 0.88 | 0.82 |
| Random Forest | 0.78 | 0.87 | 0.81 |
| LSTM | 0.85 | 0.84 | 0.84 |
| BERT | 0.90 | 0.94 | 0.92 |
Deep Insight: BERT’s success (0.92 F1) stems from its bidirectional self-attention. It can distinguish between "I want to go to medical school" (Attainment) and "My aunt went to medical school" (not the student's own goal), a nuance where simpler models often fail.
The "Small Data" Caveat
Interestingly, for the First Generation theme, where only a tiny amount of training data was available, Logistic Regression (0.91 F1) outperformed BERT (0.83 F1). This serves as a vital reminder for practitioners: Transformers are data-hungry. When you have fewer than a few hundred examples, classical ML often remains king.
Critical Analysis & Future Outlook
The TACCTI framework represents a significant step toward "Asset-Based" pedagogy. Instead of looking for what students lack, it uses AI to find what they bring.
Limitations:
- Negation Handling: Even with advanced features, non-BERT models struggled with negation (e.g., "I am unsure if I want to continue medical school").
- Theme Expansion: The current study only tackles 2 of the 11 identified Cultural Capital Themes.
Future Work: Expansion to the remaining 9 themes (like "Community Consciousness") will likely require Data Augmentation or Synthetic Data generation to overcome the scarcity of labeled examples found in this study.
Takeaway
TACCTI proves that we can use the latest NLP breakthroughs not just for chatbots, but to help educators see the "hidden" strengths of their students, potentially increasing the retention of diverse talent in STEM.
