Eye Tracker in the Wild: Bridging the Gap Between What We Say and What We Do

Eye Tracker in the Wild: Studying the delta between what is said and measured in a crowdsourcing experiment

2015-10-13
Pierre Lebreton, Isabelle Hupont, Toni Mäki, Evangelos Skodras, Matthias Hirth, Matthias Hirth
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a non-intrusive, webcam-based eye-tracking framework for crowdsourcing environments using HTML5 and Web-RTC. It introduces an adaptive regression model to map eye positions to screen coordinates and demonstrates a significant discrepancy between users' self-reported decision criteria and their actual physiological gaze behavior.

TL;DR

Researchers have developed a webcam-based eye-tracking framework that works directly in the browser without specialized hardware. By using mouse clicks as "hidden" calibration points, the system reveals a startling truth in crowdsourcing: users often spend their time looking at data (like meta-scores) that they claim didn't influence their decisions.

Background: The Crowdsourcing Blind Spot

In the world of UX research and crowdsourcing, we traditionally rely on Self-Reported Metrics (surveys/questionnaires). However, humans are notoriously bad at describing their own subconscious processes. While laboratory eye tracking (using infrared sensors) can catch these nuances, it is expensive and doesn't scale. This paper moves eye tracking from the "Lab" to the "Wild" using nothing but a standard RGB webcam and a browser.

The Problem: The Fragility of Webcam Tracking

Webcam-based gaze estimation typically fails due to two reasons:

  1. Head Movement: Users in the wild don't sit perfectly still.
  2. Calibration Fatigue: Asking a user to "follow the dot" every 2 minutes ruins the user experience and leads to unnatural behavior.

Methodology: The "Adaptive" Implicit Calibration

The genius of this framework lies in its implicit recalibration. The authors assume the Inductive Bias that when a user clicks an element, they are looking at it.

1. The Architecture

The system splits processing between the user's browser (collecting video and click events via HTML5/Web-RTC) and a server-side backend that processes frames using a Viola-Jones face detector and specialized eye-localization algorithms.

System Architecture and Calibration

2. Adaptive Fitting

Lighting is often uneven in home environments. One eye might be in shadow while the other is clear. The researchers proposed an Adaptive Fitting model. By checking the error at each click point, the system dynamically chooses between:

  • Left-eye regression
  • Right-eye regression
  • Average of both This ensures that even if a user tilts their head or has a lamp on one side, the gaze estimation remains robust.

Results: Metrics vs. Reality

The researchers conducted a "Movie Selection" experiment. Users were asked to pick a movie and then rank why they chose it (Cast, Title, Description, etc.).

The "Delta"

The study found a significant "Delta" (gap) between reported and measured behavior:

  • The Claim: Participants reported that the Movie Description was the most important factor.
  • The Reality: Gaze data (Fixation Maps) showed they spent significantly more time staring at Meta-Scores (Ratings) and Posters.

Experimental Layout with Areas of Interest

Figure: The movie selection interface. Note that the "Next" buttons are placed in corners to maximize the calibration range across the screen area.

Fixation Distribution

As shown in the fixation heatmaps, the gaze distribution changed as users became "familiar" with the task—a learning effect that traditional metrics like "Time to Completion" fail to explain in detail.

Gaze Distribution per Category

Critical Insight & Conclusion

This work provides a powerful tool for Quality Control in crowdsourcing. By monitoring gaze, platforms can identify "random clickers" or "cheaters" who don't even look at the content they are rating.

Takeaway: The study proves that "what is said" is a filtered version of reality. For high-stakes UI design or psychological profiling, physiological measurements like gaze are no longer a luxury of the lab—they are a viable tool for the open web.

Limitations: The vertical accuracy of webcams (usually placed at the top of the monitor) remains a challenge due to the steep viewing angle. However, for horizontal categorization (like comparing columns), the accuracy is more than sufficient for industrial use.

Find Similar Papers

Try Our Examples

  • Which recent papers have improved upon the "click-to-calibrate" method for webcam eye tracking in large-scale remote user studies?
  • What is the theoretical origin of the "eye-mouse coordination" assumption, and how has its accuracy been validated against infrared eye trackers?
  • How are current E-commerce or streaming platforms utilizing non-intrusive gaze estimation to optimize UI layouts compared to traditional A/B testing?
Contents
Eye Tracker in the Wild: Bridging the Gap Between What We Say and What We Do
1. TL;DR
2. Background: The Crowdsourcing Blind Spot
3. The Problem: The Fragility of Webcam Tracking
4. Methodology: The "Adaptive" Implicit Calibration
4.1. 1. The Architecture
4.2. 2. Adaptive Fitting
5. Results: Metrics vs. Reality
5.1. The "Delta"
5.2. Fixation Distribution
6. Critical Insight & Conclusion