Deciphering Social Dynamics: Speaker Role Recognition via SNA and Duration Modeling
Speakers Role Recognition in Multiparty Audio Recordings Using Social Network Analysis and Duration Distribution Modeling
This paper introduces a framework for Speaker Role Recognition in multiparty audio recordings, utilizing Social Network Analysis (SNA) and Duration Distribution Modeling (DDM). Tested on 19 hours of radio news bulletins, the system achieves approximately 85% accuracy in labeling speaker roles by combining relational interaction patterns with intervention timing.
TL;DR
This research moves beyond simple speaker identification to Role Recognition. By treating an audio recording as a social network, the system identifies the "Anchorman," "Guest," or "Interviewer" with over 85% accuracy. It achieves this by combining the topology of interaction (Social Network Analysis) with the statistics of talk-time (Duration Distribution), providing a blueprint for automated meeting summarization and structural audio indexing.
Problem & Motivation: The "Who" vs. The "Role"
In multimedia retrieval, we've gotten good at Speaker Diarization (segmenting "who spoke when"). However, knowing that Speaker A talked for 2 minutes tells us nothing about the function of that intervention.
In structured environments like radio bulletins or board meetings, participants follow an implicit script. An Anchorman isn't just a voice; they are a "hub" in a social network. Prior works often failed because they relied too heavily on lexical cues (keywords) or specific voice IDs. This paper asks: Can we identify a person's role purely by looking at whom they talk to and for how long?
Methodology: Relational Data & Temporal Statistics
The author proposes a pipeline that transforms raw audio into a social graph.
1. The Social Network Analysis (SNA) Path
Instead of analyzing the audio signal for each speaker, the author builds a Sociomatrix.
- Centrality: The Anchorman is defined by their "Closeness Centrality." They interact with almost everyone, making them the most reachable node in the network.
- Interaction Fraction: Roles like "Secondary Anchorman" or "Guest" are identified by the percentage of their interactions that involve the primary Anchorman.
2. The Duration Distribution (DDM) Path
Roles often have "temporal signatures." An Anchorman typically accounts for ~40% of the time, while a "Meteo" (weather) speaker has a very short, specific intervention. The author uses Gaussian distributions to model these likelihoods.
3. Cleaning the Noise with Poisson Processes
Raw speaker clustering is noisy, often creating "spurious turns" (brief interruptions, noise). The author uses a Poisson Stochastic Process (PSP) to model the probability of speaker changes. If a segment's duration is too short to be a valid turn under the PSP model, it is merged, drastically improving the Social Network's clarity.
Figure 1: The proposed system flow, from audio segmentation to SNA/DDM integration.
Experiments & Results
The experiments involved 96 radio bulletins (19 hours total) with an average of 11 speakers per recording.
- SNA Excellence: SNA was incredibly effective at identifying the Anchorman (AM) and Meteo (MT) roles, where interaction patterns are highly distinct.
- The Power of Combination: SNA alone struggles with noise in automatic segmentation. However, when combined with DDM, the system becomes robust. The accuracy jumped from 69.6% (SNA only) to 85.1% (Combined).
- Impact of PSP Smoothing: Filtering the segments using the Poisson model increased role recognition accuracy by ~15% for the SNA approach.
Table II: Accuracy () and Purity () results. Note the significant jump in performance after applying PSP filtering (af) compared to before (bf).
Figure 5: A-posteriori probability distributions showing how different roles are separated by speaking time (fraction of bulletin).
Critical Analysis & Conclusion
Takeaway
The genius of this work lies in its content-agnostic nature. It doesn't need to "understand" the language or recognize the specific person; it only needs to observe the rhythm and flow of the conversation. This makes it highly portable across different languages and news formats.
Limitations
- Short Interventions: Roles like "Secondary Anchorman" and "Interview Participant" (IP) suffered low accuracy because their interventions were often filtered out by the PSP smoothing as "noise."
- Statistically Structured Environments: The system relies on a stable format. It would likely struggle in highly spontaneous, multi-party environments (like a loud dinner party) where roles are fluid.
Future Outlook
As we move toward a world of automated meeting minutes (Zoom/Teams), this SNA-based approach could be the key to identifying the "Decision Maker" or "Mediator" without needing expensive, privacy-invasive transcriptions.
