Forecasting the Unseen: Leveraging Regional Data for COVID-19 Scenarios in Vietnam

Modeling Transmission Rate of COVID-19 in Regional Countries to Forecast Newly Infected Cases in a Nation by the Deep Learning Method

2021-01-01
Le Duy Dong, Vu Thanh Nguyen, Dinh Tuan Le, Mai Viet Tiep, Vu Thanh Hien, Phu Phuoc Huy, Hieu Phan Trung
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a cross-country transfer learning approach using Deep Learning to forecast COVID-19 cases in Vietnam during the Delta strain outbreak. By training LSTM, GRU, and hybrid models on data from regional neighbors (Philippines, Malaysia, Thailand) with earlier outbreaks, the authors successfully simulated various epidemic scenarios for a target nation with initially insufficient data.

TL;DR

When a new COVID-19 variant hits a nation, data is often too sparse for accurate deep learning. This research shifts the focus from local data to regional transmission dynamics, training models on the Delta strain's behavior in the Philippines, Malaysia, and Thailand to project scenarios for Vietnam. The study concludes that Gated Recurrent Units (GRU) are superior to standard LSTM for these high-volatility tasks.

Background & Motivation: The Data Scarcity Traps

During the 4th wave of the pandemic in 2021, the Delta variant introduced a transmission velocity that rendered previous historical models obsolete. For countries like Vietnam, which initially kept cases low, the sudden surge meant there wasn't enough "learning material" for neural networks to understand the curve's trajectory.

The authors' core insight was simple yet powerful: Transmission rates are often geographically and culturally linked. By modeling the "spreading speed" in neighboring South East Asian nations that were further along the infection curve, they could create a representative behavioral set for Vietnam.

Methodology: RNNs and the Hybrid Approach

The study explores a hierarchy of Recurrent Neural Network (RNN) architectures:

  1. LSTM (Long-Short Term Memory): The traditional choice for time-series, using three gates to manage memory.
  2. GRU (Gated Recurrent Unit): A streamlined version of LSTM that merges the cell state and hidden state, using fewer parameters.
  3. Hybrid Models (HLG/HGL): Novel stacks combining GRU layers with LSTM layers to attempt to capture both long-term dependencies and faster training dynamics.

Model Architecture

The models utilize a 15-step look-back window to predict a 15-step look-ahead period. The data was normalized using MinMaxScaler to prevent activation function saturation.

Model Architecture and Hybrid Layers Fig 1: Structure of the Hybrid GRU-LSTM (HGL) and LSTM-GRU (HLG) models.

Experiments and Comparative Performance

The experimental procedure involved validating the "neighbor-trained" models against actual Vietnamese data. A critical finding was the consistent ranking of architectural performance.

The Superiority of GRU

The researchers found that the GRU model achieved the lowest Mean Squared Error (MSE) across all training subsets. This is likely due to the GRU's simpler architecture, which generalises better on the relatively short and highly fluctuating sequences characteristic of pandemic data.

ModelPhilippines Data (MSE)Malaysia Data (MSE)Thailand Data (MSE)
GRU0.001610.001560.00156
HLG0.001670.001820.00177
LSTM0.001810.002170.00186

Training Loss Curves Fig 2: Loss and Val_Loss curves showing stable convergence, indicating the models generalized the regional data well.

Forecasting Scenarios for Vietnam

By applying these models, the researchers generated three distinct scenarios:

  • The "Philippines" Scenario: Predicted a downward trend after a peak in early August.
  • The "Malaysia/Thailand" Scenarios: Predicted a continued upward surge, with cases potentially hitting 20,000 daily—a warning that proved vital for policy makers.

Scenarios Comparison Fig 3: Forecasting scenario for Vietnam based on Malaysia’s transmission rates.

Critical Insight & Conclusion

This paper validates a "proxy-training" strategy for public health crises. The most significant takeaway is architectural: while LSTM is often the "default" for sequences, the GRU and its hybrid derivatives proved more accurate and efficient for epidemiological modeling.

Future Work: The authors suggest that this cross-regional modeling can be applied at a more granular level (e.g., using data from a major city like Ho Chi Minh City to forecast for secondary provinces), provided that local authorities maintain transparent and standardized data reporting similar to the Johns Hopkins University Lab.

Find Similar Papers

Try Our Examples

  • Search for recent papers using meta-learning or transfer learning techniques to forecast infectious disease outbreaks in data-sparse regions.
  • Which original studies compared the efficiency of GRU versus LSTM in short-term epidemiological time-series forecasting, and do they align with this paper's findings?
  • Explore how hybrid RNN models (like GRU-LSTM) have been applied to multi-variant viral transmission modeling in South East Asia.
Contents
Forecasting the Unseen: Leveraging Regional Data for COVID-19 Scenarios in Vietnam
1. TL;DR
2. Background & Motivation: The Data Scarcity Traps
3. Methodology: RNNs and the Hybrid Approach
3.1. Model Architecture
4. Experiments and Comparative Performance
4.1. The Superiority of GRU
5. Forecasting Scenarios for Vietnam
6. Critical Insight & Conclusion