SEFS-DL: Optimizing Urban Mobility via Economic-Inspired Feature Selection
An Effective Search Economics Based Feature Selection Algorithm for Passenger Flow Prediction
The paper introduces SEFS-DL, a feature selection algorithm based on the Search Economics (SE) metaheuristic, specifically optimized for Passenger Flow Prediction in public transport systems. It integrates a Deep Neural Network (DNN) with SE to identify the most relevant features from large datasets (e.g., Taipei MRT flow combined with weather data) to improve prediction accuracy while reducing computational overhead.
TL;DR
Predicting passenger flow in massive metro systems like the Taipei MRT requires processing huge amounts of data (weather, time, location). SEFS-DL (Search Economics for Feature Selection on Deep Learning) introduces a unique "market-investment" approach to feature selection. By treating different feature subsets as "investment regions," it outperforms standard algorithms like PSO and Gray Wolf by reducing prediction error (MAPE) to 9.905% and significantly thinning out irrelevant data.
Background: Why Feature Selection Matters for MRT
Modern smart cities rely on accurate passenger flow prediction to manage crowding and subway schedules. However, deep learning models are only as good as the data they consume. Inputting every possible variable—from historical flow at 109 stations to global radiation levels—often leads to "noise" that confuses the model. The challenge is NP-hard: finding the optimal subset of features among combinations is mathematically impossible via exhaustive search.
The Core Innovation: Search Economics (SE)
Unlike traditional metaheuristics that treat all candidate solutions equally, SEFS-DL adopts a space-aware approach inspired by economic investment:
- Solution Space Partitioning: The "market" (all possible feature combinations) is divided into regions. Each region has a different probability () of selecting features, ensuring some regions look at "dense" feature sets while others look at "sparse" ones.
- Investment Expectation: The algorithm calculates an "Expected Value" for each region. If a region consistently yields high-accuracy prediction models (high "ROI"), the "investors" (searchers) focus more resources there.
- Region Reduction: As the search matures, the algorithm merges similar regions. This prevents redundant searching of the same space—a common pitfall in algorithms like Particle Swarm Optimization (PSO).
Fig 1: The flow of SEFS-DL, from feature masking to DNN evaluation and feedback.
Methodology: The "Region Reduction" Operator
The paper's most significant technical contribution is the Region Reduction mechanism. Early in the training, the algorithm explores broadly across 8 different regions. Every evaluations, it halves the number of regions, doubling the samples in the remaining ones. This mimics the consolidation of a market, allowing the final search stages to focus intensely on the most promising "economic" zones.
Fig 2: Visualization of how solution space regions merge over time to refine the search.
Experimental Results: Beating the SOTA
The authors tested SEFS-DL against three heavyweights: BPSO, BGWO, and BGOA. Using 3 years of Taipei MRT data (2017-2019), they modeled the flow of the Taipei Main Station.
- Accuracy: SEFS-DL achieved a MAPE of 9.905, while the nearest competitor (BGOA) was at 11.14.
- Efficiency: It reduced the total feature count by 52%, ensuring the final DNN model was leaner and faster during inference.
- Convergence: As shown in the graph below, while other algorithms stagnated (local optima), SEFS-DL continued to find better solutions in the late-game phase thanks to its dynamic region management.
Fig 3: Objective value convergence over 1280 evaluations. Note the continued descent of SEFS-DL (black line).
Critical Insight & Conclusion
The success of SEFS-DL highlights a vital lesson for AI researchers: Global search isn't always optimal. By enforcing regional diversity and implementing a structured "market consolidation" strategy, we can handle the high-dimensional complexity of urban time-series data more effectively than with random or purely greedy searches.
Takeaway: If your model is struggling with high-dimensional noise, look beyond standard "feature importance" rankings. Structured metaheuristic searches like Search Economics offer a robust path to discovering non-obvious feature interactions that boost SOTA performance.
