Learning the Truth: Can Followers Overcome Influencer Manipulation in Dynamic Social Networks?
Learning the Truth by Weakly Connected Agents in Social Networks Using Multi-Armed Bandit OLUSOLA TOLULOPE ODEYOMI , (Graduate Student Member, IEEE)
This paper introduces an online diffusion learning framework for social networks to help "weakly connected agents" (followers) learn an arbitrarily time-varying true state using the Multi-Armed Bandit (MAB) technique. Specifically, it proposes a non-stochastic MAB algorithm that enables misinformed followers to converge to the most stable true state despite the dominating influence of influential personalities.
TL;DR
In the age of social media, "influential personalities" often dictate the narrative, sometimes leading their followers toward false information. This paper proposes a new Multi-Armed Bandit (MAB) algorithm that allows "weakly connected" followers to independently learn an arbitrarily time-varying truth. By treating social learning as an online optimization problem, the study demonstrates that even dominated agents can reach a 100% belief in the truth, despite the inherent lag caused by their network position.
Strategic Position
This work sits at the intersection of Graph Theory and Online Reinforcement Learning. It addresses a critical void: most prior social learning research (SOTA) assumes the truth is static. By introducing dynamic parameters, the author shifts the focus from "reaching consensus" to "minimizing regret" in a fluctuating world.
The Problem: The "Echo Chamber" of Static Truth
In traditional graph-theoretic social learning, agents are categorized into:
- Strongly Connected Subnetworks: Influencers who communicate bi-directionally and conscious of their own opinions.
- Weakly Connected Subnetworks: Followers who primarily receive information and are "dominated" by the leaders' signals.
The Pain Point: Previous models failed because they assumed the "True State" () never changed. In reality, the "truth" (like stock prices or news updates) is a moving target. In these old models, followers often ended up with a belief probability , meaning they never fully "learned" the truth—they just became permanent echoes of their influencers.
Methodology: High-Stakes MAB for Social Graphs
The author re-imagines the social network as a non-stochastic game. Each agent must choose a "state" (a belief) and incurs a loss if that state isn't the truth.
The Core Mechanism: Online Diffusion Learning
The algorithm follows a 5-step loop for every time :
- Step 1-2 (Exploration vs. Domination): Agents calculate an intermediate probability that balances their past beliefs with a "domination number" (), representing the influence of the leaders.
- Step 3-4 (The Bandit Action): An agent picks a state, observes the loss (0 if they found the truth, 1 otherwise), and estimates the losses for states they didn't pick.
- Step 5 (Exponential Update): Beliefs are updated using an exponential function: This ensures that states with higher losses (the "lies") decay exponentially, while the most stable state (the "truth") survives.
Figure 1: A network topology showing strongly connected subnetworks (A, B) dominating a weakly connected subnetwork (C).
Experimental Insights: The Cost of Being a Follower
The experiments conducted in MATLAB simulate 5 potential states where the true state changes arbitrarily.
- The Convergence Gap: Strongly connected agents find the truth faster (average ). Weakly connected agents are 66% slower () because they rely on dominated signals.
- The Learning Rate Lever: By increasing (learning rate), the convergence speed of followers improved by 50%.
- Regret Analysis: The "Regret" (the difference between the agent's performance and an all-knowing oracle) grows at a rate of . While this is higher than the seen in strongly connected networks, it is mathematically "sublinear," meaning the agents eventually do learn.
Figure 2: Probability of belief for agents 6, 7, and 8. Notice all beliefs drop to 0 except for state , which reaches 1.0 (True State).
Critical Insight & Conclusion
The true value of this paper lies in its pessimistic assumption. By using a non-stochastic Multi-Armed Bandit, the author assumes an "adversary" is trying to maximize the agent's regret. If the algorithm can find the truth in this "worst-case" scenario, it is highly robust for real-world social media dynamics.
Takeaway for Future Research: The "Domination Number" is a powerful abstraction. Future work could explore what happens when followers start "unfollowing"—i.e., when the network topology itself is dynamic, allowing weakly connected agents to break the domination of influencers.
Limitations
The current model assumes followers have no "self-loop" (they don't trust their own original observations at all). In reality, humans have some level of independent verification. Adding "self-conscious" weights to followers could potentially bridge the 66% performance gap identified in this study.
