The "Almost" Knowing: Why Raw Social Media Data is Never Enough

A Critical Reflection on Social Media Research Using an Autoethnographic Approach

2016-01-01
Jordan Eschler
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a critical autoethnographic study of social media research, specifically within the context of illness narratives. By triangulating public artifacts, private data, and lived experience of Stage IV cancer, the author demonstrates that "big data" and "small data" methods often result in only partial knowledge ("almost" knowing) if users are not directly engaged.

TL;DR

In an era where "big data" dominates, researchers often assume that social media artifacts provide a transparent window into human experience. Jordan Eschler’s autoethnographic study of her own Stage IV cancer journey proves otherwise. By comparing her public posts, private emails, and medical records, she reveals that social media data is a curated performance—one that masks more than it reveals without the context of lived experience.

Background: Beyond the Binary of Big and Small Data

Social media research typically falls into two camps:

  1. Big Data: Computational methods mapping networks and sentiment at scale.
  2. Small Data: Qualitative content analysis or digital ethnography focusing on "close readings."

Eschler argues that even "small data" research falls short because it often treats digital artifacts (posts, photos) like objects in a museum—divorced from the "embodied" context of the person who created them. To bridge this gap, she adopts Autoethnography, turning the researcher into the informant to see where the data and the reality diverge.

Methodology: The Chronological Confrontation

To analyze the interplay between life and data, the author reconstructed her survival timeline between January and October 2014. She merged two distinct streams:

  • The "Objective" Medical Stream: Hospital bills, discharge papers, and treatment schedules.
  • The "Subjective" Digital Stream: Downloads of Facebook, Twitter, and Gmail archives.

By overlaying these, she could track when she was active, when she was silent, and—most importantly—what she chose not to share.

Table of Artifact Types and Research Use

Key Insights: What the "Data" Missed

The results highlight the inherent "skew" in social media datasets:

1. The Invisibility of Crisis

The most traumatic moments (diagnosis) were entirely absent from public feeds. News was shared via email or phone calls. A researcher scraping her public profile during that month would see statuses about laundry and movies, concluding that nothing significant was happening.

2. Mania vs. Engagement

A massive spike in Twitter activity was observed during treatment. A data scientist might interpret this as "increased social support seeking." The reality? It was a side effect of high-dose Prednisone (steroids), causing insomnia and mania. The posts were often incoherent and driven by drug-induced energy, not a social strategy.

Twitter Usage Trends Before, During, and After Treatment

3. Curated Resilience

On Facebook, the author performed the role of the "fighter." She shared "last chemo" posters and celebration photos but self-censored "selfies" taken in hospital rooms where she felt she looked too sick. This creates a "positivity bias" in the data that reflects societal pressures on cancer patients rather than their full emotional reality.

Celebrating the last chemotherapy session

Discussion: Borrowing from Museum Studies and Indigenous Knowledge

The paper makes a profound connection to Indigenous Knowledge systems, which have long criticized Western researchers for "extracting" artifacts from cultures. Just as a physical artifact loses its meaning when removed from its tribe, a tweet loses its meaning when removed from the user’s immediate physical and emotional state.

The author suggests that social media researchers should act more like curators and collaborators rather than extractors. This means:

  • Direct Engagement: Eliciting user reflections through interviews.
  • Contextual Integrity: Understanding that platform affordances (like Facebook lists vs. Twitter public feeds) dictate what is "truth" versus "performance."

Critical Analysis & Conclusion

This work serves as a necessary reality check for the field of Social Computing. While autoethnography isn't generalizable, it exposes the "partial knowledge" that plagues automated and even manual content analysis.

Takeaway: We cannot interpret digital remains without talking to the humans who left them behind. If we ignore the user, we aren't studying people; we are studying ghosts of their self-presentation.

Future Work: This research calls for new ethical frameworks that prioritize "ethics-as-process," moving beyond check-box IRB approvals to designs that respect a participant's right to curate their own narrative post-hoc.

Find Similar Papers

Try Our Examples

  • Search for recent papers in HCI or CSCW that utilize autoethnography to critique the ethics of big data collection in health contexts.
  • Which foundational works in Indigenous Knowledge systems first discussed the "extraction" of knowledge from cultural contexts, and how are these concepts being applied to digital data ownership today?
  • How have researchers applied the "walk-through" method to help vulnerable populations reflect on their own social media timelines for qualitative research?
Contents
The "Almost" Knowing: Why Raw Social Media Data is Never Enough
1. TL;DR
2. Background: Beyond the Binary of Big and Small Data
3. Methodology: The Chronological Confrontation
4. Key Insights: What the "Data" Missed
4.1. 1. The Invisibility of Crisis
4.2. 2. Mania vs. Engagement
4.3. 3. Curated Resilience
5. Discussion: Borrowing from Museum Studies and Indigenous Knowledge
6. Critical Analysis & Conclusion