NLP for Dementia: Moving Beyond the "Detection Trap" to Clinical Impact
A Systematic Review of NLP for Dementia -- Tasks, Datasets and Opportunities
This paper provides a comprehensive systematic review of over 240 studies applying Natural Language Processing (NLP) to dementia research, evaluating tasks like detection, biomarker extraction, and caregiver support. It bridges the gap between medical and technological communities while identifying how Transformer-based models and LLMs are achieving over 90% accuracy in clinical benchmarks.
TL;DR
Dementia is a global health crisis, and language is its most sensitive sensor. This systematic review of 240+ papers reveals that while AI can detect Alzheimer's with >90% accuracy in lab settings, we are failing in real-world application. The field must pivot from simple classification to personalized digital twins, caregiver support, and "degraded" models that simulate the aging brain.
Background Positioning
In the landscape of AI-for-Healthcare, this work serves as an authoritative map. Rather than proposing a single new model, it critiques the entire trajectory of the field—from 1990s SVMs to today’s LLMs—and identifies why, despite impressive SOTA numbers, your doctor isn't using an NLP tool for diagnosis yet.
The Problem: Homogeneous Data and the "Black-Box" Barrier
The core friction in NLP for dementia is the Trust Gap.
- Data Staleness: Most SOTA results rely on the Pitt Corpus (1994). The language of a 67-year-old in 1994 differs significantly from a "digital native" senior today.
- Lack of Rigor: Incredibly, 86% of speech-based detection papers do not report statistical significance. To a clinician, a model reaching 92% accuracy without a p-value is just a correlation, not a diagnostic tool.
Methodology: The Four Pillars of Dementia-NLP
The authors reframe the research into distinct task families, moving away from a one-size-fits-all approach:

- Dementia Detection: The "Saturated" zone. High accuracy, but low clinical utility as it often focuses on advanced Alzheimer's where diagnosis is already obvious.
- Biomarker Extraction: Analyzing "Empty Speech," word-retrieval latencies, and pronoun overuse.
- Caregiver & Patient Support: The "Emerging" zone. Using LLMs to provide emotional scaffolding for the 11 million informal caregivers worldwide.
The "Future Frontier": Artificially Degraded Models
One of the most provocative "How" insights in this paper is the use of Language Retrogenesis. By deliberately degrading specific layers of an LLM or fine-tuning them on impaired datasets, researchers can create a "Digital Twin" of the patient's cognitive decline.
> Physical Intuition: Think of it as "stress-testing" a bridge by simulating rust on specific joints. If we simulate "protein deposits" on a neural network's weights, does the resulting output match the linguistic "slurring" seen in clinical patients?
SOTA Performance vs. Reality
While the migration from SVMs to Transformers has significantly reduced the error rate, the "Applicative State-of-Mind" is missing.

The review notes that ASR (Automatic Speech Recognition) is the hidden bottleneck. If the ASR misses a "silent pause" (a key biomarker), the most powerful GPT-4 based classifier will fail because the input data was "cleaned" of its most vital diagnostic signals.
Critical Insight & Conclusion
The "Takeaway" is clear: We need to stop chasing decimal points on 30-year-old datasets. The future of NLP for dementia lies in:
- MCI-Focus: Detecting the subtle shifts 15 years before the first clinical symptom.
- Synthetic Augmentation: Creating diverse datasets that include non-native speakers and various dialects.
- Ethical Vigilance: Addressing the "Precision-Recall" trade-off—where a false positive for an incurable disease can be psychologically catastrophic.
This review is a call to action for NLP researchers to look past the leaderboard and toward the living room, where AI might finally offer a lifeline to millions of families.
