Moving Beyond Surprisal: How Hidden State Trajectories Predict the "Weight" of Words

Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal

2026-06-01
Elan Barenholtz
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces "Trajectory Extrapolation Error," a novel metric that measures deviations in the hidden-state dynamics of Language Models (LLMs) to predict human reading times. By fitting linear trajectories to sequential hidden states, the authors demonstrate that representational continuity predicts processing costs independently of the standard "Surprisal" metric, achieving a more nuanced understanding of incremental language comprehension.

TL;DR

For decades, psycholinguistics has been dominated by Surprisal Theory: the idea that the harder a word is to predict, the longer it takes to read. However, a new study reveals that this is only half the story. By analyzing the "flight path" of internal representations in models like GPT-2, researchers found that Trajectory Extrapolation Error—a measure of how much a word forces a "sharp turn" in our mental representation—predicts human reading effort entirely independently of surprisal.

The Missing Dimension of Language Processing

Standard Information Theory treats language as a sequence of discrete events. If you see the sequence "The horse raced past the barn...", surprisal tells you how much the final word "fell" shocks the system.

But humans don't just calculate probabilities; we build incremental interpretations. Our mental state has momentum. Imagine driving a car: surprisal is like a sudden obstacle on the road, while trajectory error is like a sharp, unexpected curve. Even if the curve is "expected" on a map, taking it at high speed requires more physical effort than driving straight. The authors argue that current LLMs capture this "momentum" in their hidden states, but we have been ignoring it by only looking at the final probability output.

Methodology: Measuring the "Turn"

The researchers introduced a simple but elegant metric: Trajectory Extrapolation Error.

  1. State Tracking: For every word in a sentence, they extract the hidden state vector from a Transformer (e.g., Layer 6 of GPT-2).
  2. Linear Fitting: They look at the last 3 words and fit a linear "path" through those three high-dimensional points.
  3. Extrapolation: They project where the path should go next.
  4. Error Calculation: They measure the Euclidean distance between that projection and where the actual hidden state for the new word landed.

Model Architecture and Extrapolation Schematic

This distance represents the reorientation cost. If the new word continues the "theme" or "structure" established by the previous words, the error is low. If it forces a sudden shift (like the disambiguating verb in a garden-path sentence), the error is high.

Key Findings: The Great Dissociation

The most striking result is that Surprisal and Extrapolation Error are nearly orthogonal ().

This means a word can be:

  • High Surprisal / Low Extrapolation Error: A rare word that perfectly fits the current topic (e.g., using a specific technical term like "synapse" in a biology text). Reading time increases, but the "mental path" stays straight.
  • Low Surprisal / High Extrapolation Error: A common word that changes the direction of the sentence (e.g., "and" or "but" starting a new clause). Even though the word is common, the system must shift its representational gears.

Heatmap showing independent effects on RT

Robustness Across Scales and Architectures

The effect wasn't just a quirk of GPT-2. The researchers tested:

  • Model Scale: Replicated in GPT-2 Small, Medium, and Large.
  • Architecture: Replicated in Pythia, which uses Rotary Position Embeddings (RoPE) instead of GPT-2’s absolute embeddings.
  • Controlled vs. Naturalistic: The effect held both in laboratory "Garden Path" sentences and in the "Natural Stories" corpus.

Deep Insight: Local vs. Global Continuity

Through a Direction Preservation Analysis, the authors discovered that this trajectory is strictly local. At intermediate layers (Layer 6), the "direction" of the path dies out after just one or two words. This matches the human "planning horizon"—we tend to plan our speech just a few words at a time.

The final layers of the model, interestingly, show much longer-range directional persistence. However, it is the intermediate, locally-focused layers that best predict human reading behavior, suggesting that human cognition is constrained by similar short-term working memory and recency biases.

Critical Analysis & Conclusion

This work offers a compelling solution to the Surprisal Scaling Paradox. As LLMs get bigger and "smarter," they become too good at long-range prediction, whereas humans remain local, dynamical processors.

Takeaways for the Field:

  • Psycholinguistics: We must stop treating "processing cost" as a single-variable problem. Probability is not the only metric of difficulty.
  • AI Development: If we want to generate text that is "easier" for humans to consume, we should optimize for trajectory coherence, not just perplexity.
  • Limitations: The study relies on self-paced reading. Future work using eye-tracking or EEG would be necessary to see if trajectory errors map to specific neural signatures of re-analysis (like the P600 ERP component).

Ultimately, this paper reminds us that language is not just a sequence of tokens to be predicted, but a dynamical system in motion. Understanding the "arc" of that motion is key to understanding the human mind.

Find Similar Papers

Try Our Examples

  • Find recent papers investigating the "surprisal scaling paradox" in large language models and alternative metrics proposed to better align LLM internal states with human brain activity or reading times.
  • Which seminal papers in dynamical systems theory (e.g., by Spivey or Tabor) established the theoretical basis for treating language comprehension as a continuous trajectory in a representational space?
  • Search for studies that evaluate if preserving local trajectory continuity in machine-generated text improves human readability or reduces cognitive load during L2 language learning.
Contents
Moving Beyond Surprisal: How Hidden State Trajectories Predict the "Weight" of Words
1. TL;DR
2. The Missing Dimension of Language Processing
3. Methodology: Measuring the "Turn"
4. Key Findings: The Great Dissociation
4.1. Robustness Across Scales and Architectures
5. Deep Insight: Local vs. Global Continuity
6. Critical Analysis & Conclusion