Uber's mRMR: Solving the "Redundancy Paradox" in Large-Scale Marketing ML

Maximum Relevance and Minimum Redundancy Feature Selection Methods for a Marketing Machine Learning Platform

2019-10-01
Zhenyu Zhao, Radhika Anand, Mallory Wang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a framework for feature selection in large-scale marketing applications using Minimum Redundancy and Maximum Relevance (mRMR). It introduces two novel extensions: Randomized Dependence Coefficient (RDC) for non-linear redundancy and Random Forest importance for relevance, achieving SOTA balancing between model interpretability and predictive accuracy at Uber.

TL;DR

In industrial machine learning, more data does not always mean better models. Uber's latest research demonstrates that selecting the "m best features" is fundamentally different from selecting the "best m features." By extending the Minimum Redundancy Maximum Relevance (mRMR) framework with non-linear measures and model-based importance, Uber achieves higher accuracy with a fraction of the features, drastically reducing engineering overhead and improving model interpretability.

The Problem: The Curse of "Relevant but Redundant"

In marketing ML—predicting churn, cross-sell, or app sign-ups—we often have thousands of features ranging from static demographics to dynamic app usage logs. Standard practice often involves ranking features by their individual correlation to the label (Relevance).

However, this leads to a critical failure: the top 10 features might all represent slightly different versions of the same signal (e.g., "logins in 7 days," "logins in 14 days," "active days"). This redundancy adds noise, increases the risk of overfitting, and bloats the data pipeline infrastructure without adding predictive value.

Methodology: Refining the mRMR Framework

The authors pivot from simple ranking to an optimization problem: Maximize Relevance while Minimizing Redundancy.

1. The mRMR Core Formula

The standard mRMR is defined as:

2. Uber’s Extensions: RFCQ and RFRQ

The researchers identified that standard linear correlation (Pearson) misses non-linear dependencies. They introduced two key upgrades:

  • Non-linear Redundancy (RDC): Using the Randomized Dependence Coefficient to catch features that correlate in complex, non-obvious ways.
  • Model-based Relevance: Using Random Forest (RF) importance scores instead of F-statistics to ensure the feature selection is aligned with the actual downstream model's logic.
  • The Quotient Scheme: Instead of subtracting redundancy from relevance (which fails when scales differ), they use a Quotient (FCQ/RFCQ), which makes the trade-off more robust across different data types.

Automated Machine Learning Platform Architecture

Evaluation: Synthetic and Real-World Impact

The team tested 8 variants across synthetic data and three massive Uber datasets (cross-sell and up-sell scenarios).

Key Findings:

  • Efficiency: Top-tier performance (AUC) was often reached with just 10-15 features, even when the original set contained over 1,300.
  • Stability: The FCQ (F-test Correlation Quotient) emerged as the "production hero"—it is nearly as accurate as complex model-based versions but significantly faster to compute ( minute vs. hours for MI-based methods).
  • Overfitting Prevention: In several real-world datasets, the mRMR-selected feature subset actually outperformed the model using "All Features," proving that removing noise is as important as finding signal.

Synthetic Data Performance Curves

From Research to Production: The Uber Implementation

Uber integrated this into their Scala-Spark based AutoML platform. To handle the scale of millions of users, they implemented several optimizations:

  1. Down-sampling: Feature selection is performed on a representative sample before full training.
  2. Concurrency: Replacing iterative loops with Scala's functional map/reduce operations to speed up the computation of the redundancy matrix.
  3. Dynamic Selection: The platform can runtime-evaluate whether a model-based (RFCQ) or model-free (FCQ) approach delivers better results for a specific campaign.

Critical Analysis & Future Outlook

The paper provides a roadmap for "pragmatic ML." While researchers often focus on the most complex non-linear kernels, Uber’s findings suggest that for production, a linear-based quotient (FCQ) offers the best ROI between speed and accuracy.

However, a noted limitation is the computational cost of the RDC and Mutual Information methods on extremely high-dimensional datasets. Future work could explore Streaming Feature Selection or Differential Privacy constraints within the mRMR framework to protect user data while maintaining relevance.

Takeaway: In the era of Big Data, the most valuable models are those that can do more with less. Uber's mRMR framework proves that diversity in features is just as important as individual predictive power.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply mRMR feature selection specifically within automated machine learning (AutoML) pipelines for tabular data.
  • Who first proposed the Randomized Dependence Coefficient (RDC), and how does it compare to Mutual Information for measuring non-linear feature dependencies?
  • Explore research comparing the performance of filter-based feature selection (like mRMR) versus embedded methods in high-dimensional marketing datasets.
Contents
Uber's mRMR: Solving the "Redundancy Paradox" in Large-Scale Marketing ML
1. TL;DR
2. The Problem: The Curse of "Relevant but Redundant"
3. Methodology: Refining the mRMR Framework
3.1. 1. The mRMR Core Formula
3.2. 2. Uber’s Extensions: RFCQ and RFRQ
4. Evaluation: Synthetic and Real-World Impact
4.1. Key Findings:
5. From Research to Production: The Uber Implementation
6. Critical Analysis & Future Outlook