Ensemble Learning for Coking Coal Freight: Overcoming the Small Data Sparsity Challenge
Coking Coal Railway Transportation Forecasting Using Ensembles of ElasticNet, LightGBM, and Facebook Prophet
This paper develops machine learning models to forecast coking coal railway transportation for Russian Railways using a hybrid ensemble approach. By combining ElasticNet, LightGBM, and Facebook Prophet, the study achieves high-accuracy short-term forecasts (MAPE as low as 6%) despite the constraints of small, sparse datasets.
TL;DR
Predicting industrial freight volumes is notoriously difficult due to the "small data" trap—limited historical records coupled with highly irregular demand. This study demonstrates that an ensemble of LightGBM and ElasticNet can achieve high precision (6-10% error) in forecasting Russian coking coal transportation, proving that sophisticated regularization and boosting often outperform complex neural networks and specialized time-series tools like Facebook Prophet on sparse datasets.
The "Zero-Value" Bottleneck in Logistics Forecasting
Most modern AI research assumes an abundance of data. However, in the world of heavy industry and logistics, we often deal with "sparse time series." For Russian Railways, forecasting coking coal involves:
- Tiny Datasets: Only about 114 monthly records.
- High Sparsity: Many months show zero transit or import, making it impossible for standard models to learn a continuous trend.
- External Volatility: Dynamics are driven by global GDP, USD/RUB exchange rates, and international coal prices rather than just historical internal trends.
The authors argue that when features are numerous and data points are few, the variance of standard OLS (Ordinary Least Squares) regressions tends toward infinity, leading to total model collapse.
Methodology: The Power of Three
The researchers tested three distinct philosophies to find the optimal balance:
- ElasticNet: A hybrid of Ridge (L2) and Lasso (L1) regression. It excels here because it can perform variable selection (ignoring irrelevant features) while maintaining stability.
- LightGBM: A gradient boosting framework that grows trees leaf-wise. Its main advantage is speed and a natural ability to handle missing values and non-linear relationships without heavy preprocessing.
- Facebook Prophet: A decomposable time-series model. Interestingly, while popular for "at scale" forecasting, it struggled here as it couldn't fully leverage the external economic regressors as effectively as the ensemble.
Model Architecture & Hyperparameters
The authors didn't just dump data into a "black box." They meticulously tuned the models, specifically using the Quantile Loss function for LightGBM to handle the skewed distribution of the coal data.

Experimental Battleground: Export vs. Domestic
The study split the data into a training set (2010–2016) and a test set (2017) to prevent "looking into the future."
1. Export Transportation
Exports are highly sensitive to global markets. The best result came from a weighted ensemble:
- 60% LightGBM (focusing on prices and production indicators).
- 40% ElasticNet (focusing on GDP and macroeconomics).
- Result: MAPE of 10%.
2. Domestic Transportation
Internal demand was more stable. Surprisingly, a pure LightGBM model outperformed all ensembles.
- Result: MAPE of 6.12%.
In the figure above, the red line tracks the blue (actual) values with remarkable accuracy during the test phase (green background).
Critical Analysis: Why Prophet Failed
One of the paper's most salient insights is the failure of Facebook Prophet to improve the ensemble. While Prophet is excellent for handling seasonality and holidays, the coking coal market is driven more by exogenous economic shocks (like exchange rate fluctuations) than by annual cycles. This serves as a warning to practitioners: "fancier" time-series tools are not a silver bullet for every domain.
Conclusion & Future Outlook
This work provides a blueprint for industrial forecasting where data is scarce. By combining the linear stability of ElasticNet with the non-linear flexibility of LightGBM, the authors turned a "noisy" 114-line spreadsheet into a high-precision strategic tool.
Takeaway for the industry: If your data is sparse and column-heavy, skip the deep learning. Focus on regularized ensembles and boosting. Future work could potentially integrate "Zero-Inflated" models to tackle the import/transit segments that were excluded from this study due to their extreme number of zero values.
