Beyond the Words: Decoding Age and Gender through the Rhythm of Typing and Mouse Movements
Predicting Age and Gender by Keystroke Dynamics and Mouse Patterns
This paper presents a user modeling study that predicts age and gender by analyzing "unintentional" digital traces—specifically keystroke dynamics and mouse movement patterns. Using data from 1,519 subjects across diverse real-world and experimental contexts, the study achieves high-accuracy classification with F-scores exceeding 0.90 for certain models.
TL;DR
Most user interfaces are designed to respond to what we do—our intentional commands. However, researchers are increasingly looking at how we do it. This paper explores the "unintentional" traces we leave behind through keystroke dynamics and mouse patterns. By analyzing data from over 1,500 subjects, the study demonstrates that we can predict a user's age and gender with surprising accuracy (F-scores > 0.9), even when the user isn't writing long paragraphs.
Context: The Limitations of Content Analysis
In the realm of user modeling, age and gender detection are typically handled by analyzing text content (stylometry). While effective, this approach has a major bottleneck: it requires a substantial amount of text. If a user is just clicking through a questionnaire or writing a short tweet, text mining fails.
The author, Avar Pentel, shifts the focus from linguistic content to behavioral dynamics. The intuition is rooted in history—telegraph operators in WWII could identify each other by their unique keying rhythms. In a modern context, if a certain group (like teenagers) is highly familiar with specific character patterns, that familiarity will manifest as a distinct, faster typing rhythm compared to other groups.
Methodology: Capturing the Digital Pulse
The study utilized a diverse dataset collected between 2011 and 2017 from six distinct sources, ranging from school intranets to psychological experiments. This diversity provides ecological validity, meaning the results are more likely to hold up in the real world than a controlled lab study.
Key Feature Groups
- Keystroke Dynamics: Measured via "Hold Time" (duration a key is pressed) and "Seek Time" (the delay between release and the next press).
- N-Graph Latencies: The total time taken to type a sequence of n characters.
- Mouse Patterns: Defined by distance (ratio of path to straight line), angles, and velocity standard deviations.

Figure 1: Visualization of hold times, seek times, and spatial mouse movements used as input features.
Experiential Differences: The 16-29 "Sweet Spot"
One of the most compelling findings from the statistical analysis was the distribution of typing speeds across age groups. The 16-29 age group emerged as significantly faster than all others.
| Age Group | Comparison to 10-15 | Comparison to 30-39 |
|---|---|---|
| 16-19 | Much Faster (***) | Faster (***) |
| 20-29 | Much Faster (***) | Faster (**) |
| 50+ | Slower (*) | Similar |
Table 1: T-test results showing significant differences between age group means. Key: *** denotes p < 0.0001.
The study suggests this is a combination of motor skill development and "generational experience." Younger children (10-15) are still gaining proficiency, while users over 30 may not have the same level of native "digital fluency" or physical typing speed as young adults.
Results: Can Machines Guess Your Profile?
By feeding these features into machine learning algorithms (Logistic Regression, SVM, Random Forest, etc.), the study achieved remarkable results, particularly with mouse dynamics.

Table 2: F-scores for different models. Note the kNN and Random Forest performance on mouse data.
While the mouse-based results (F-scores up to 0.95) are exceptionally high, the author cautiously notes that in some cases (like Nearest Neighbor), the model might be identifying individual users rather than just the general class—a nuance that requires further study into the boundary between classification and identification.
Deep Insight: The Link Between Text and Rhythm
The most profound takeaway is the overlap between Frequency and Fluency. The author found that n-graphs (character sequences) that are frequent in Estonian text-mining studies are also the ones that show the most distinct timing patterns in this study. Essentially, if your brain "knows" a word well, your fingers "know" how to type it rhythmically.
Conclusion & Future Directions
This research proves that "how we type" is as much a part of our identity as "what we type." For applications like cyber-forensics (identifying pedophiles or preventing underage access) and online marketing, these unintentional patterns offer a non-intrusive way to verify user demographics.
Limitations: The data is specific to the Estonian language and keyboard layouts. Future work will need to explore how these models generalize across different languages and more diverse device types (like the shift from physical keyboards to mobile swipe-typing).
