EyeCrowdata: Solving the Replication Crisis in Eye-Tracking Research

EyeCrowdata: Towards a Web-based Crowdsourcing Platform for Web-related Eye-Tracking Data

2020-05-27
Naziha Shekh Khalil, Ecem Dogruer, Abdulmohimen K. O. Elosta, Sukru Eraslan, Yeliz Yesilada, Simon Harper
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces EyeCrowdata, a web-based crowdsourcing platform designed specifically for sharing and exploring web-related eye-tracking datasets. It establishes a standardized framework for data features including equipment, participants, and raw gaze coordinates to facilitate the replication and validation of HCI studies.

TL;DR

Eye-tracking is a cornerstone of User Experience (UX) research, yet its data often remains locked in private silos, hindering the replication of findings. EyeCrowdata is a new web-based crowdsourcing platform that standardizes metadata and raw gaze data storage, allowing researchers to upload, search, and download datasets across different studies to enable large-scale meta-analysis and validation.

The "Data Silo" Problem in HCI

In Human-Computer Interaction (HCI), eye-tracking provides deep insights into how users navigate web pages. However, the field faces a significant bottleneck:

  • High Resource Costs: Recruiting participants and managing hardware is time-consuming.
  • Generalist Repositories: While tools like Zenodo exist, they don't understand what a "sampling rate" or "AOI" (Area of Interest) is, making them useless for finding specific types of eye-tracking data.
  • Replicability Crisis: Without raw data coordinates, other researchers cannot re-validate published results or apply new algorithms to old datasets.

Methodology: Building a Domain-Specific Schema

The authors didn't just build a database; they derived its structure from a Structured Literature Review of 45 seminal papers from CHI, WWW, and ETRA. They identified that to make a dataset truly "reusable," a platform must track seven specific groups of features.

Table 2: Data Fields and Features

Key Feature Categories included:

  • Equipment: Eye tracker model, accuracy, and screen resolution.
  • Participants: Demographics and vision abilities (e.g., corrected vision).
  • Raw Data: Ensuring X, Y coordinates and pupil size are stored, not just high-level summaries.

Prototype Architecture & Workflow

The EyeCrowdata prototype uses a LAMP stack (Linux, Apache, MySQL, PHP) to manage a relational database. The workflow is designed to be researcher-friendly:

  1. Upload: Users provide basic study info (DOI, abstract) and upload raw data files.
  2. Search: An "Advanced Search" interface allows users to filter by specific conditions, such as "Participants aged 20-30 using a Tobii tracker."
  3. Download: Researchers can export entire datasets or filtered subsets for their own analysis.

EyeCrowdata Workflow Figure 1

Experimental Insights: What’s Missing in Current Research?

The systematic review revealed a worrying trend: while 33 out of 45 papers discussed computed metrics like fixation duration, only 19 provided enough information about raw gaze points. This gap is exactly what EyeCrowdata aims to bridge. By forcing a standardized entry for "Material Type" (e.g., real web pages vs. screenshots) and "Task Type," the platform ensures that future researchers can compare "apples to apples."

Advanced Search Prototype Figure 3

Critical Analysis & Future Outlook

While EyeCrowdata is a major step forward, the authors acknowledge several hurdles:

  • Scalability: Storing massive raw gaze files requires significant storage capacity, potentially leading to a shift toward "processed data" storage in the future.
  • Privacy: Eye-tracking data can be uniquely identifying. Future iterations will need robust anonymization protocols.
  • Archival Complexity: Since web pages change, simply storing a URL isn't enough; the system may eventually need to store full DOM snapshots or screenshots to maintain the context of the gaze data.

Takeaway: EyeCrowdata represents a move toward "Open Eye-Tracking," providing the infrastructure necessary for the HCI community to share resources and accelerate the development of more intuitive web interfaces.

Find Similar Papers

Try Our Examples

  • Search for recent papers or platforms that focus on the specialized archival and sharing of physiological data in Human-Computer Interaction beyond eye-tracking.
  • Which original publications established the "Scanpath Theory" mentioned in this paper, and how has web-based crowdsourcing evolved to support or challenge this theory?
  • Investigation of privacy-preserving methods for sharing raw eye-tracking data to mitigate the ethical risks of participant re-identification noted in this study.
Contents
EyeCrowdata: Solving the Replication Crisis in Eye-Tracking Research
1. TL;DR
2. The "Data Silo" Problem in HCI
3. Methodology: Building a Domain-Specific Schema
4. Prototype Architecture & Workflow
5. Experimental Insights: What’s Missing in Current Research?
6. Critical Analysis & Future Outlook