Automatic Alt-Text: How Facebook Uses AI to Make the Visual Web Audible for Blind Users

Automatic Alt-text: Computer-generated Image Descriptions for Blind Users on a Social Network Service

2017-02-14
Shaomei Wu, Jeffrey Wieland, Omid Farivar, Julie Schiller, Julie Schiller
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents Automatic Alt-Text (AAT), a large-scale system deployed by Facebook that utilizes computer vision to generate descriptive alt-text for blind users. It translates images into "Image may contain..." sentences identifying faces, objects, and themes, significantly improving social network accessibility for the visually impaired.

TL;DR

With over 2 billion photos shared daily, the visual-centric nature of social media creates a "digital divide" for blind users. This paper introduces Automatic Alt-Text (AAT), a pioneering system that uses real-time computer vision to generate "Image may contain..." descriptions. Deployed to thousands of Facebook users, the system proved that AI can significantly reduce the mental and social burden of photo interpretation, even if it doesn't immediately change quantitative social "Like" behaviors.

Context & Motivation: The Latency of Human Help

Before AAT, blind users had three bad choices for understanding photos:

  1. Crowd-sourced apps (e.g., VizWiz): Accurate, but take minutes to return results—far too slow for a fast-moving News Feed.
  2. Social Capital: Asking friends to describe photos, which creates a feeling of being a "burden" or "social debt."
  3. Ignoring Content: Most blind users simply skipped photos, missing out on 50%+ of the social conversation.

The authors’ Research Intuition was simple: A real-time, "good enough" automated description is better than a perfect but delayed human one.

Methodology: Precision over Poetry

Instead of generating poetic, free-form captions that risk "hallucinations" (a term we use today, though the authors focused on "accuracy failures"), AAT uses a structured, tag-based approach.

1. The 97-Concept Vocabulary

The team selected 97 concepts (e.g., "smiling," "selfie," "nature," "pizza") based on two criteria:

  • Prominence: What appears most in Facebook photos.
  • Clarity: Avoiding fuzzy adjectives (red, happy) or sensitive identity markers (gender, age) that the AI might misinterpret.

2. The Information Hierarchy

Users don't want a random list of objects. Through lab studies, the researchers found that People are the most important element. Thus, the system follows a strict order:

  1. People Count & Expression (e.g., "3 people, smiling")
  2. Objects (e.g., "car," "tree")
  3. Settings/Themes (e.g., "indoor," "close-up")

AAT News Feed Experience Figure 1: How AAT appears (visually) to a screen reader user.

Experiments: Perception vs. Reality

The researchers conducted a massive field study with 9,000 users.

  • Perceived Ease of Use: There was a dramatic shift in how users felt. In the control group, almost no one found photos easy to tell apart. In the AAT group, nearly 1/3rd felt it was easy.
  • The "Like" Disconnect: Interestingly, while users said they were more likely to "Like" a photo because of AAT, the actual server logs didn't show an increase in clicking the Like button. This suggests that while AAT makes the user feel more "included," the decision to socially interact is still dominated by who posted the photo and the existing comments.

Ease of Interpretation Results Figure 2: Statistical comparison show AAT (test group) makes photos 2x easier to interpret.

Critical Analysis: The Ethics of Identity

The most requested feature from blind participants was Face Recognition (e.g., "Your friend John is smiling"). However, the authors made a strategic choice to omit this initially.

The trade-off: While describing a person's gender or identity is highly useful, the risk of a "social miscue"—such as misgendering someone or misidentifying a stranger—carries a high social cost. This highlights the unique Inductive Bias required for assistive AI: Accuracy is not just a metric; it’s a prerequisite for social safety.

Summary & Future Outlook

AAT represents a shift from "human-in-the-loop" to "AI-at-scale" for accessibility.

  • Key Contribution: Demonstrating that high-confidence, low-complexity tags are more useful than high-complexity, low-confidence captions.
  • Limitations: The system remains "vague" for many users who want to know why a photo is significant (e.g., "What kind of guide dog?" or "What does the text on the sign say?").

As we move into the era of Multimodal LLMs (like GPT-4o or Gemini), the foundations laid by AAT—ordering information by social priority and managing algorithmic uncertainty—remain the blueprint for truly inclusive technology.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend automatic alt-text by incorporating facial recognition or personalized identity recognition for blind users.
  • Which 2014-2016 studies first established deep learning benchmarks for image captioning specifically aimed at assistive technology?
  • Explore how large language models (LLMs) are currently being used to replace the fixed-tag "Image may contain" format with natural, conversational image descriptions for the visually impaired.
Contents
Automatic Alt-Text: How Facebook Uses AI to Make the Visual Web Audible for Blind Users
1. TL;DR
2. Context & Motivation: The Latency of Human Help
3. Methodology: Precision over Poetry
3.1. 1. The 97-Concept Vocabulary
3.2. 2. The Information Hierarchy
4. Experiments: Perception vs. Reality
5. Critical Analysis: The Ethics of Identity
6. Summary & Future Outlook