SORCER: Beyond Decision Trees in Climate-Hydrology Forecasting
From climate history to prediction of regional water flows with machine learning
The paper presents a machine learning study on regional water flow prediction using SORCER, a Second-Order Table Compression system. It integrates solar, ocean, and atmospheric conditions to forecast the seasonal water inflows of Lake Okeechobee in Florida, achieving a lower error rate than the standard C4.5 decision tree algorithm.
TL;DR
Predicting local water flows from global climate patterns is a "black box" challenge for resource management. This paper introduces SORCER, a learning system based on Second-Order Table Compression. By transforming raw data into compact, set-based logical rules, SORCER outperforms the classic C4.5 decision tree algorithm, particularly in predicting critical "dry" periods for Florida's Lake Okeechobee.
Context & Positioning
In the landscape of 2000-era AI, the dominant forces for classification were Decision Trees (like C4.5) and emerging Neural Networks. This work identifies a niche for Symbolic Induction—specifically Second-Order Relational frameworks. It positions itself as a "Downscaling" solution, bridging the gap between coarse global climate models and the fine-grained needs of regional water management.
The Problem: The Disconnect Between Global and Local
Hydrological management relies on knowing how much water will flow into reservoirs six months in advance. However:
- Global Models (ENSO, PDO) are too coarse for local grids.
- Statistical Methods struggle with the non-linear "noise" of climate data.
- Extreme Events (Droughts/Floods) are often smoothed out by standard algorithms, yet they are the most vital for policy intervention.
Methodology: The Power of Table Compression
The core innovation is the use of Second-Order Relations. In a standard "flat" database (First-Order), an attribute has one value. In a Second-Order table, an attribute can represent a set of values.
1. Logic Synthesis as Learning
The authors view learning as the task of compressing a massive table of observations into the smallest possible set of consistent rules. This follows Occam’s Razor: the simplest model that fits the data is likely the best.
2. Compression Algorithms
The system uses a distance metric to determine which rows can be merged without introducing inconsistency:
The distance function used in SORCER to favor merging rows with large components that share values.
The algorithm (specifically S2) prioritizes "locally joinable" pairs (equivalence-preserving) before jumping to "consistently joinable" pairs (generalization), ensuring the resulting rules stay grounded in the training data.
Experiments & Results
The researchers tested the models on nearly 90 years of historical data (1912-2000), including indices for solar activity (sunspots, geomagnetic Aa-index) and ocean temperatures.
Performance Mastery over Extremes
While the overall accuracy improvement over C4.5 was modest (~2%), the class-specific performance told a different story:
- Dry Predictions: SORCER had an error rate of only 25%, whereas C4.5 failed 50.2% of the time.
- Very Wet Predictions: SORCER again led with 39.2% error vs. C4.5's 53.3%.
Table showing SORCER's significant lead in predicting extreme "dry" and "very wet" events.
Visualizing the Forecast
The system's ability to track the trend of Lake Okeechobee's inflows is visually demonstrated in the comparison between actual and predicted flows:
Fig. 4: A decade of forecasting shows that while not perfect, the SORCER model captures the cyclical nature of water inflows effectively.
Deep Insights & Conclusion
Why does it work?
SORCER's success lies in its Inductive Bias. By using sets as components, it naturally handles disjunctive concepts better than the axis-aligned splits of a decision tree. It creates "denser" rules that are more robust to the inherent noise in climate indices.
Limitations
The "Very Wet" events remain difficult to predict. The authors attribute this to a lack of sufficient historical data for these rare, high-magnitude events—a common "imbalanced data" problem in ML. Furthermore, SORCER lacks a pruning mechanism, leading to larger rule sets (255 rules vs. C4.5's 53).
Final Takeaway
This paper proves that specialized symbolic logic techniques like Table Compression can hold their own against—and even outperform—refined general-purpose learners in domains where capturing extreme outliers is more important than average precision.
