Benchmarking Spark MLlib: Why Naïve Bayes Dominates SVM in Large-Scale Banking Analytics
Evaluation of classification algorithms for banking customer’s behavior under Apache Spark Data Processing System
This paper presents a comparative performance evaluation of two core classification algorithms—Naïve Bayes (NB) and Support Vector Machine (SVM)—within the Apache Spark MLlib framework. Using a massive financial dataset from Santander Bank (13M+ records), the study demonstrates that Naïve Bayes significantly outperforms SVM in predictive accuracy and computational efficiency for multi-class banking customer behavior tasks.
TL;DR
Predicting customer behavior in the banking sector involves processing millions of transactions and personal data points. This study evaluates Apache Spark MLlib's capacity to handle 13 million records from Santander Bank. The verdict? Naïve Bayes (NB) outperforms Support Vector Machines (SVM) across all metrics (Precision, Recall, F-measure) while offering superior computational speed for multi-product prediction tasks.
Problem & Motivation
The financial sector is plagued by the "Big Data" problem—the sheer volume and velocity of data mean that traditional, single-machine algorithms are no longer viable. While Apache Spark offers a distributed solution (MLlib), the selection of the right algorithm remains a critical challenge.
The authors identified a gap: although many frameworks exist, empirical comparisons on specific real-world tasks (like predicting the next banking product a customer will buy) are scarce. The challenge lies in:
- High Cardinality: Features like residence and income ranges require sophisticated encoding.
- Class Complexity: Predicting one of many possible products turns a simple binary task into a complex multi-class problem.
Methodology - The Spark Pipeline
The authors utilized a 13-million-record training set. To make this data digestible for MLlib, they implemented a four-stage preprocessing workflow:
- Categorical Encoding: Converting string variables (e.g., country codes like 'ES', 'FR') into numerical IDs via lookup tables.
- Numeric Binning: Continuous variables like income ('renta') were discretized into 30 bins to reduce noise and handle outliers.
- Data Cleaning: Stripping whitespace and handling null values with default assignments.
- Class Mapping: A unique approach where binary product usage was converted into a single decimal class ID (e.g., usage pattern
00...111becomes ID7).
Architecture & Algorithms
- Naïve Bayes (NB): A probabilistic classifier assuming feature independence. It excels in high-dimensional spaces by calculating conditional probabilities.
- SVM: A binary classifier that seeks the maximum margin hyperplane. In this study, it struggled due to its inherent binary nature, requiring separate models for every financial product.

Experiments & Results
The experiments were conducted on a standalone Spark environment (Core i7, 24GB RAM). The performance disparity between the two approaches was stark.
Performance Metrics Comparison
| Metric | Naïve Bayes (NB) | Support Vector Machine (SVM) |
|---|---|---|
| Precision | 4% | 0.18% |
| Recall | 49% | 10% |
| F-Measure | 7.3% | 0.3% |
The results (summarized in the table below) indicate that Naïve Bayes is significantly more effective at capturing the patterns of customer behavior than SVM in this specific high-volume context.

Why did NB win?
- Multi-class Capability: NB naturally handles many classes, whereas SVM had to be adapted for each product, leading to "time-consuming processes" and poor convergence on sparse behavior.
- Independence Assumption: Despite being "naïve," the independence assumption often provides a strong inductive bias in banking datasets where many features are loosely correlated.
Critical Analysis & Conclusion
Takeaway
The paper confirms that for Spark-based Big Data tasks involving multi-product prediction, probabilistic multi-class models are significantly more efficient than binary hyper-plane models.
Limitations
- Baseline Choice: The precision and recall for both models remain relatively low (4% precision for NB). This suggests that while NB is better than SVM, further feature engineering or more complex models (like Random Forests or Gradient Boosted Trees) might be necessary for production-level accuracy.
- Hardware Constraint: The study used a "standalone" Spark environment rather than a true multi-node cluster, which might mask some scalability bottle-necks of the NB algorithm compared to SVM.
Future Outlook
This work sets the stage for investigating more advanced MLlib utilities, such as automated hyperparameter tuning and the use of the Pipeline API for more streamlined feature engineering.
