Professional Identity and Information Use: On Becoming a Machine Learning Developer

Professional Identity and Information Use: On Becoming a Machine Learning Developer

2019-01-01
Christine T. Wolf
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the intersection of Information Behavior (IB) and professional identity formation among Machine Learning (ML) developers. Through qualitative interviews, it identifies information use as an organizing principle that facilitates occupational entry and sustains the ML "community of practice."

TL;DR

Becoming a Machine Learning (ML) developer isn't just about mastering calculus or Python; it’s an identity-shifting process driven by how one consumes information. This research reveals that while "famous breakthroughs" spark interest, the real transition to a professional involves mastering the "invisible" data work that scholarly papers often ignore. By framing ML as a "technical craft" rather than just a math discipline, we can lower barriers to entry and foster a more inclusive community.

Background: Beyond Code and Math

In the high-stakes world of AI, we often focus on the what—the algorithms and the hardware. However, this study by Christine T. Wolf (IBM Research) shifts the focus to the who. It maps the journey of ML developers through the lens of Information Behavior (IB), arguing that information isn't just a functional tool for coding but a social resource used to build a "professional self."

The "Origin Story": Information as an Identity Catalyst

How does one decide to become an ML developer? The paper highlights that for many, the spark comes from public breakthroughs reported in the popular press.

  • The "Wow" Moment: Challenges like ImageNet or AlphaGo serve as "switching points." Developers reported that seeing these results transformed their perception of ML from a niche academic topic to a viable professional orientation.
  • Elastic Identity: The study finds that "ML Developer" is a flexible identity. For some, it’s about bringing innovation to biology; for others, it's a "data-first" mindset applied to finance.

Participant Metadata and Project Context Note: The study interviewed a diverse range of 11 developers with experience ranging from 2 to 11 years across various domains.

The Gap: "ML on the Books" vs. "ML in Practice"

One of the paper’s most profound insights is the disconnect between formal information (scholarly papers) and daily work.

The "Invisible" Work of Data

Participants noted a massive gap between theory and practice. While academic papers focus on novel algorithms, they often treat datasets as "invisible" or "standardized."

  • The Hardest Part: Cleaning data, tuning parameters, and "entity resolution" occupy the bulk of a developer's time but are rarely the subject of the high-level papers that drew them to the field.
  • Hybrid Knowledge: Success in ML requires weaving together formal information (theory) with informal expertise (experience-based know-how).

Communities of Practice: Social Information Support

If papers don't teach you how to handle "messy data," who does? The study identifies a multi-scalar community of practice:

  1. Micro: Immediate lab colleagues and office-mates.
  2. Meso: Internal company networks and "old-timers" who hold institutional memory.
  3. Macro: Global platforms like StackOverflow and Facebook Groups.

These networks do more than just debug code; they validate the developer’s identity. When a developer asks a question on StackOverflow, they aren't just seeking a function—they are participating in the collective problem-solving that defines the profession.

Critical Insight: Reframing ML for Diversity

The "mathematician" stereotype acts as a gatekeeper. Many potential developers are scared off because they don't see themselves as "theory people."

The author argues that by making the Information Behavior of ML visible—the searching, the trial-and-error, the collaboration—we can reframe ML as a technical craft. This shift is crucial for diversity and inclusion (STEM policy). If we highlight that "good ML" is about being a reflective, curious information-seeker rather than just a math prodigy, we open the door to a much wider array of talent.

Conclusion: Information as the Glue

Information use acts as the "centrifugal force" that binds the ML community together. For practitioners and managers, the takeaway is clear:

  • Value the Data Work: Stop treating data cleaning as a "preliminary" task; it is the core of the professional identity.
  • Foster the Community: Intentional spaces for informal information sharing are just as important as formal training sessions.

Ultimately, becoming an ML developer is a process of learning to navigate a complex information ecosystem—a journey from "who am I?" to "who am I going to be?"

Find Similar Papers

Try Our Examples

  • Search for recent studies on "invisible labor" or "invisible work" in machine learning operations (MLOps) and data engineering.
  • What are the foundational papers on "Communities of Practice" by Jean Lave and Etienne Wenger, and how have they been applied to software engineering identities?
  • Find literature examining the impact of popular science media coverage (e.g., AlphaGo, ChatGPT) on career choice and STEM enrollment trends.
Contents
Professional Identity and Information Use: On Becoming a Machine Learning Developer
1. TL;DR
2. Background: Beyond Code and Math
3. The "Origin Story": Information as an Identity Catalyst
4. The Gap: "ML on the Books" vs. "ML in Practice"
4.1. The "Invisible" Work of Data
5. Communities of Practice: Social Information Support
6. Critical Insight: Reframing ML for Diversity
7. Conclusion: Information as the Glue