Intelligent Crowdsourcing: Solving Task Neglect through DBSCAN Clustering and Proportional Sharing
Analysis of Optimal Pricing Model of Crowdsourcing Platform Based on Cluster and Proportional Sharing
This paper presents an optimized pricing model for crowdsourcing platforms by combining an improved proportional sharing mechanism with the DBSCAN clustering algorithm. The method focuses on task packaging for geographically adjacent missions to enhance fulfillment efficiency, achieving a significant reduction in task failure rates on real-world datasets.
TL;DR
Crowdsourcing platforms often struggle with "unclaimed tasks," where low-value missions are ignored by workers. This paper introduces a dual-strategy approach: geospatial task packaging using DBSCAN and an improved proportional sharing mechanism for pricing. By bundling nearby tasks and factoring in worker credit and real transportation costs, the authors reduced task failure rates from 35% down to 6.3%.
Background & Motivation: The "Low-Value Task" Trap
In mobile crowdsourcing (e.g., information collection, inspection), the irrational pricing of tasks is a fatal flaw. Existing platforms often publish tasks individually. When many tasks are clustered in one area but offered at low prices, workers tend to ignore them in favor of higher-value targets elsewhere.
The author identifies four critical pricing factors:
- Task Location: The primary driver of cost.
- Member Distribution: Where the potential workforce is located.
- Creditworthiness: Determining task priority and reliable completion.
- Task Quota: The maximum workload a single member can handle.
Analysis of a real-world "Take Pictures to Make Money" dataset reveals a staggering 36% failure rate, primarily due to task fragmentation and a lack of incentive for geographically dense, low-price missions.
Methodology: Clustering Meets Mechanism Design
1. The Basic Pricing Model (Proportional Sharing)
The model assumes a static environment where member costs () are calculated based on real-world transportation data (fuel, insurance, maintenance) and a "reputation premium." The pricing mechanism uses a Greedy Algorithm logic:
- Members are ranked by Creditworthiness first, then Cost.
- The platform allocates tasks to ensure maximum budget utilization while meeting the member's "individual rationality" (ensuring they actually profit).
2. Task Packaging via DBSCAN
To prevent members from scrambling for single high-value tasks while leaving others behind, the author uses DBSCAN (Density-Based Spatial Clustering of Applications with Noise) to group adjacent tasks.

- Core Points: Areas with high task density.
- Packaging Strategy: Small tasks are merged into a "large task" package. The cost to complete this package is the sum of costs for individual sub-tasks, making the overall incentive more attractive to efficient workers.
Experiments and Results
The study utilized a dataset of over 2,000 tasks and 1,800 members in Guangdong Province, China. By calling map APIs and calculating weighted average transportation costs (), the author simulated the new pricing logic.

Key Performance Indicators:
- Original Failure Rate: ~35%
- Optimized Failure Rate: 6.3%
- Practical Insight: The model identifies "dead angles"—tasks that even with optimized pricing remain unviable—allowing platforms to proactively adjust or cancel them.
Critical Insight & Conclusion
The genius of this work lies in the spatial-economic synergy. While most pricing models treat tasks as independent variables, this paper recognizes that spatial density is a form of value. By clustering tasks, the platform reduces the "cognitive load" and "travel overhead" for the worker, effectively increasing the net hourly wage without necessarily increasing the total payout.
Limitations & Future Work
The current model is static, assuming members don't move. In future work, incorporating real-time mobility data and Stochastic Dynamic Programming would allow the model to function in high-volatility urban environments. Furthermore, investigating the impact of heterogeneous car types on the transportation cost formula () could yield even higher pricing precision.
