Ensemble Learning for Coking Coal Freight: Overcoming the Small Data Sparsity Challenge

Coking Coal Railway Transportation Forecasting Using Ensembles of ElasticNet, LightGBM, and Facebook Prophet

2020-01-01
Vladimir Soloviev, Nikita Titov, Elena Smirnova
Summary
Problem
Method
Results
Takeaways
Abstract

This paper develops machine learning models to forecast coking coal railway transportation for Russian Railways using a hybrid ensemble approach. By combining ElasticNet, LightGBM, and Facebook Prophet, the study achieves high-accuracy short-term forecasts (MAPE as low as 6%) despite the constraints of small, sparse datasets.

TL;DR

Predicting industrial freight volumes is notoriously difficult due to the "small data" trap—limited historical records coupled with highly irregular demand. This study demonstrates that an ensemble of LightGBM and ElasticNet can achieve high precision (6-10% error) in forecasting Russian coking coal transportation, proving that sophisticated regularization and boosting often outperform complex neural networks and specialized time-series tools like Facebook Prophet on sparse datasets.

The "Zero-Value" Bottleneck in Logistics Forecasting

Most modern AI research assumes an abundance of data. However, in the world of heavy industry and logistics, we often deal with "sparse time series." For Russian Railways, forecasting coking coal involves:

  • Tiny Datasets: Only about 114 monthly records.
  • High Sparsity: Many months show zero transit or import, making it impossible for standard models to learn a continuous trend.
  • External Volatility: Dynamics are driven by global GDP, USD/RUB exchange rates, and international coal prices rather than just historical internal trends.

The authors argue that when features are numerous and data points are few, the variance of standard OLS (Ordinary Least Squares) regressions tends toward infinity, leading to total model collapse.

Methodology: The Power of Three

The researchers tested three distinct philosophies to find the optimal balance:

  1. ElasticNet: A hybrid of Ridge (L2) and Lasso (L1) regression. It excels here because it can perform variable selection (ignoring irrelevant features) while maintaining stability.
  2. LightGBM: A gradient boosting framework that grows trees leaf-wise. Its main advantage is speed and a natural ability to handle missing values and non-linear relationships without heavy preprocessing.
  3. Facebook Prophet: A decomposable time-series model. Interestingly, while popular for "at scale" forecasting, it struggled here as it couldn't fully leverage the external economic regressors as effectively as the ensemble.

Model Architecture & Hyperparameters

The authors didn't just dump data into a "black box." They meticulously tuned the models, specifically using the Quantile Loss function for LightGBM to handle the skewed distribution of the coal data.

Hyperparameter Tuning and Model Results

Experimental Battleground: Export vs. Domestic

The study split the data into a training set (2010–2016) and a test set (2017) to prevent "looking into the future."

1. Export Transportation

Exports are highly sensitive to global markets. The best result came from a weighted ensemble:

  • 60% LightGBM (focusing on prices and production indicators).
  • 40% ElasticNet (focusing on GDP and macroeconomics).
  • Result: MAPE of 10%.

2. Domestic Transportation

Internal demand was more stable. Surprisingly, a pure LightGBM model outperformed all ensembles.

  • Result: MAPE of 6.12%.

Forecasting Visualization In the figure above, the red line tracks the blue (actual) values with remarkable accuracy during the test phase (green background).

Critical Analysis: Why Prophet Failed

One of the paper's most salient insights is the failure of Facebook Prophet to improve the ensemble. While Prophet is excellent for handling seasonality and holidays, the coking coal market is driven more by exogenous economic shocks (like exchange rate fluctuations) than by annual cycles. This serves as a warning to practitioners: "fancier" time-series tools are not a silver bullet for every domain.

Conclusion & Future Outlook

This work provides a blueprint for industrial forecasting where data is scarce. By combining the linear stability of ElasticNet with the non-linear flexibility of LightGBM, the authors turned a "noisy" 114-line spreadsheet into a high-precision strategic tool.

Takeaway for the industry: If your data is sparse and column-heavy, skip the deep learning. Focus on regularized ensembles and boosting. Future work could potentially integrate "Zero-Inflated" models to tackle the import/transit segments that were excluded from this study due to their extreme number of zero values.

Find Similar Papers

Try Our Examples

  • Find recent comparative studies on LightGBM vs. Transformer-based models for small-scale industrial time-series forecasting.
  • Which paper first established the theoretical framework for "regularization during variable selection" that inspired the ElasticNet approach used here?
  • Explore research that applies zero-inflated modeling or Tweedie loss functions to improve freight forecasting in datasets with high zero-value frequency.
Contents
Ensemble Learning for Coking Coal Freight: Overcoming the Small Data Sparsity Challenge
1. TL;DR
2. The "Zero-Value" Bottleneck in Logistics Forecasting
3. Methodology: The Power of Three
3.1. Model Architecture & Hyperparameters
4. Experimental Battleground: Export vs. Domestic
4.1. 1. Export Transportation
4.2. 2. Domestic Transportation
5. Critical Analysis: Why Prophet Failed
6. Conclusion & Future Outlook