When Humans Forget: Designing Bias-Aware Systems for Social Stream Processing
Modeling human annotation errors to design bias-aware systems for social stream processing
This paper introduces a bias-aware active learning framework designed for social media stream processing. It proposes a novel "Error-mitigating Sampling" method that models human forgetting behavior using the Ebbinghaus Curve to optimize the annotation schedule, effectively improving classification AUC in environments prone to human errors.
TL;DR
High-quality machine learning requires high-quality human data. But humans aren't perfect—we forget, get tired, and make mistakes. This paper introduces a Bias-Aware Hybrid System that models human forgetting behavior (based on the Ebbinghaus Curve) to decide which data points to show annotators. By optimizing the annotation schedule, the system boosts classification accuracy in social media crisis monitoring by mitigating human error before it happens.
Problem & Motivation: The "Burnout" of the Human Oracle
In the middle of a disaster like Hurricane Harvey, social media streams are a firehose of information. Hybrid Stream Processing Systems (HSPS) rely on humans to label "infrastructure damage" or "rescue efforts" to train AI models in real-time.
The industry's blind spot has long been the assumption that human labels are "Ground Truth." In reality:
- Burnout: High cognitive load leads to deteriorating quality.
- Slips & Mistakes: Forgetting a category exists simply because you haven't seen it in the last 100 tweets.
- The Schedule Gap: Conventional active learning picks the most "uncertain" data for humans but ignores whether the human is in a fit state (cognitively) to label that specific instance correctly.
Methodology: Modeling the Human Mind
The researchers' core insight is that human error isn't random; it's predictable.
1. The Forgetting Curve
They utilize the Ebbinghaus Curve, modeled as a sigmoid function, to calculate a forgetting_score. If a class (e.g., "Utility Damage") hasn't appeared in the stream for a while, the probability of a human "slipping" increases.

2. Error-Mitigating Sampling
Standard Active Learning (Uncertainty Sampling) focuses on the model's confusion. The proposed Algorithm 3 (Error-mitigating Sampling) adds a second filter:
- Predictive Uncertainty: Identify samples where the model is 30-70% confident.
- Bias & Forget Score: Calculate if showing this specific class right now will induce a performance error or if the human has likely forgotten the class representation.
- Discard & Update: Discard instances that are likely to result in erroneous human labels, thereby protecting the model from learning from "poisoned" data.
Experiments & Results
The team tested their approach against two hurricanes (Harvey and Irma) using Twitter data. They simulated three types of oracles: "Slow Forgetting," "Fast Forgetting," and "No Forgetting."
Key Findings:
- Superior Robustness: Even with "Fast Forgetting" (highly error-prone humans), the error-mitigating algorithm maintained a stable and higher AUC compared to random sampling.
- Learning Curve: While simple uncertainty sampling works initially, its performance fluctuates wildly as human errors accumulate. The proposed method acts as a stabilizer.
Fig: AUC scores across datasets. The error-mitigating sampling (Proposed) shows more consistent growth and higher peaks than standard baselines.
Critical Analysis & Conclusion
Takeaway
This paper shifts the focus of Active Learning from "What does the model need to know?" to "What is the human capable of teaching right now?" By treating the human annotator as a dynamic, error-prone component of the system rather than a static oracle, we can design much more resilient AI.
Limitations
- Text-Only: The current study focuses on Twitter text; cognitive load might behave differently for image or multi-modal tasks.
- Parameter Sensitivity: The sigmoid function requires specific parameters () which might vary significantly between different crowdsourcing pools.
Future Work
The next frontier is extending this to Human-AI Collaboration where the system provides "hints" (e.g., showing a reference image of the class) to "refresh" the human's memory when the forget_score gets too high.
