BSMRL: Weaponizing Reinforcement Learning for Bribery Selfish Mining in Ethereum
BSMRL: Bribery Selfish Mining with Reinforcement Learning
This paper introduces BSMRL (Bribery Selfish Mining with Reinforcement Learning), an AI-optimized attack strategy targeting the Ethereum blockchain. By combining selfish mining with bribery mechanisms and utilizing a Markov Decision Process (MDP) solved via reinforcement learning, the authors demonstrate that attackers can achieve higher rewards compared to traditional methods even with significantly lower hashing power.
TL;DR
Researchers have developed BSMRL, a hybrid attack strategy that combines Selfish Mining and Bribery Attacks optimized through Reinforcement Learning. The strategy exploits Ethereum's unique reward structure (Uncle blocks), lowering the profitability threshold for attackers from the typical 16% down to a startling 6.7% of network hash power.
Background: Why Ethereum is Different
In the world of Proof-of-Work (PoW) blockchains, "Selfish Mining" has long been a theoretical and practical threat. However, most research has been centered on Bitcoin. Ethereum introduces a more complex incentive layer through Uncle Blocks—blocks that are mined nearly simultaneously with the main chain block but aren't included in it. Unlike Bitcoin, Ethereum rewards these blocks to reduce the disadvantage of network latency.
The authors of BSMRL realized that these "kind" rewards actually create a more fertile ground for attackers. By using Reinforcement Learning (RL), they sought to answer: Can an attacker use bribery and AI to turn Ethereum's uncle rewards against itself?
Methodology: The BSMRL Architecture
The researchers modeled the attack as a Markov Decision Process (MDP). In this environment, an "Agent" (the attacker) observes the state of the blockchain and chooses actions to maximize relative revenue.
1. State Space & Actions
The system tracks the lead the attacker has over the public chain, whether a fork exists, and whether uncle blocks are available for referencing. The agent can choose:
- Adopt: Accept the public chain (reset).
- Wait: Continue mining secretly.
- Override: Publish the secret chain to invalidate the public chain.
- Match: Publish a block concurrent with an honest block to trigger a race.
2. The Bribery Loop
The "Bribery" component allows the attacker to temporarily "rent" hash power from rational miners. By offering a fee (), the attacker convinces a portion of the network to work on their private branch during a fork, significantly increasing the probability that the attacker's branch wins the race.
Figure 1: The MDP-based strategy logic where RL determines the optimal action based on chain length and bribery status.
Key Results: Lowering the Bar for Attacks
The findings are a wake-up call for blockchain security.
- Threshold Collapse: In standard Selfish Mining (SM1), an attacker usually needs roughly 16.5% of the network's power to be profitable. With BSMRL, this drops to 6.7%.
- Revenue Dominance: At 25% hashing power, BSMRL provides substantially higher relative returns than both honest mining and standard selfish mining.
- The Ethereum Vulnerability: The study confirms that Ethereum's reward for uncle blocks acts as a "buffer" for attackers, effectively subsidizing their failed attempts and making the network more vulnerable than Bitcoin.
Figure 2: Mining Revenue vs. Hashing Power. Note the green line (BSMRL) crossing the honest mining threshold much earlier than others.
Critical Insights & Limitations
The Diminishing Returns of Bribery
Interestingly, the study found that increasing the bribery success rate () doesn't yield linear returns. As increases, the cost of the bribe eventually eats into the profits. There is a "sweet spot" for attackers—usually between 0.7 and 0.8—beyond which the bribery attack becomes less efficient.
Academic Conclusion
The BSMRL paper demonstrates that "Machine Learning is accessory to a tyrant’s crimes." By automating the strategy search, the authors proved that the security assumptions of consensus protocols are often more fragile than they appear when subjected to algorithmic optimization.
Limitations
- Single Attacker Model: The current research assumes one intelligent attacker. In a real-world scenario, multiple selfish miners might compete, leading to a "War of Attrition" that RL models are only beginning to explore.
- PoW Focus: While Ethereum has transitioned to Proof-of-Stake (PoS), the fundamental logic of "Strategic Reference" and "Bribery" (in the form of MEV - Maximum Extractable Value) remains highly relevant to modern chain security.
Future Outlook
The next frontier is Multi-Agent Reinforcement Learning (MARL). As mining pools become more sophisticated, we can expect to see "algorithmic arms races" where different AI agents compete to manipulate block propagation and rewards. Proactive detection systems must now be trained against these RL-driven adversaries.
