Why AI Wealth Managers Fail: The Gap Between arXiv and Wall Street

A review of machine learning experiments in equity investment decision-making: why most published research findings do not live up to their promise in real life

2021-04-01
Wojtek Buczynski, Fabio Cuzzolin, Barbara J. Sahakian
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a critical literature review of 27 academic experiments on machine learning (ML) in equity investments, contrasting them with the actual performance of AI-driven funds. It identifies a "cognitive dissonance" where academic claims of 90%+ accuracy fail to translate into real-world profitability or industry adoption.

TL;DR

Despite two decades of academic papers claiming over 90% accuracy in stock market forecasting, real-world AI funds are consistently underperforming the S&P 500. This paper exposes the structural flaws in how we measure ML success in finance: specifically, the misuse of "average" errors and the deceptive practice of "cherry-picking" backtest results.

Context: This is a landmark critical review that shifts the focus from "better algorithms" to "better evaluation protocols" in the intersection of Computer Science and Quantitative Finance.

The "Cherry-Picking" Problem

The authors identified a disturbing trend: academic researchers often run dozens—sometimes hundreds—of model configurations in the testing phase and only report the top performer.

  • The Reality: In real-life investing, you only have one "configuration" (your actual fund).
  • The Data: Out of the 15 experiments analyzed closely, the average number of configurations was 70.7. Picking the best one out of 70 is not science; it’s statistical noise that creates an illusion of SOTA performance.

Why MAPE is a "Four-Letter Word" in Finance

The paper provides a brilliant critique of standard ML metrics like Mean Absolute Percentage Error (MAPE) and Root Mean Square Error (RMSE).

In most ML tasks (like Image Classification), an "average" accuracy of 95% is great. In Finance, "average" accuracy is a trap. If your model is 99% accurate but fails catastrophically on the 100th day (a Black Swan event), your fund liquidates. Finance is governed by geometric compounding, not arithmetic averages.

Effect of Outliers on Metrics Figure: The authors' experiment shows how a MAPE of 5% can still result in a model that is no better than a coin flip (47% hit rate).

Methodology: The Architecture of Failure

The paper categorizes the 27 reviewed experiments into:

  1. Market Forecast: Predicting a benchmark index (Most common).
  2. Individual Equity Forecast: Predicting specific stocks.
  3. Bespoke Portfolio Construction: Autonomous weight allocation (Rarest).

Interestingly, while models became more complex (moving from simple Backpropagation to Ensembles and LSTM), the actual "Hit Rate" (directional accuracy) has not improved significantly since 2000.

Chronological Summary of Techniques Table: A snapshot of the evolution of ML techniques in the reviewed literature.

Real-World Performance: The "AI Fund" Dissonance

If academic models are so accurate, why aren't AI funds dominating? The authors analyzed high-profile cases like Sentient Technologies and Aidya, both of which liquidated after disappointing results.

Comparing AI Hedge Funds to the market reveals a stark reality:

  • S&P 500 (2016-2019): ~45% Return.
  • AI Hedge Fund Index (Eurekahedge): ~7% Return.

The "Black Box" nature of these models also creates a Regulatory Wall. Under UK's SMCR and EU's MiFID II, a human manager is legally liable for "due care." Trusting a black-box model that cannot explain why it is selling a stock is not just risky—it may be illegal.

Deep Insights & The Road Ahead

The core contribution of this paper is the call for a "Finance-First" approach to AI.

Key Recommendations for Researchers:

  • Explainability is Non-Negotiable: Regulators and investors will never accept a tool they cannot audit.
  • One Model, One Test: Stop reporting "best-of" results. Report the full forecast time series.
  • Incorporate Alternative Data: Standard market data is too noisy. The future of Alpha lies in satellite imagery, shipping traffic, and ESG signals.
  • The "Man + Machine" Synergy: The authors argue for AI as a "Cold Cognition" tool (processing data) to support "Hot Cognition" human judgment.

Conclusion

AI is not yet ready to replace the portfolio manager. The "phenomenal" 90% accuracy claimed in papers is often a result of flawed evaluation protocols. Until we solve the transparency problem and the "average-error" fallacy, AI will remain a supplementary tool rather than the master of the market.

Find Similar Papers

Try Our Examples

  • Search for recent papers that propose alternatives to MAPE and RMSE specifically for evaluating sequential financial time-series forecasting.
  • Which researchers pioneered the concept of 'backtest overfitting' in finance, and what are the current SOTA methods for preventing it in ML-based trading?
  • Explore how Explainable AI (XAI) techniques, such as SHAP or LIME, are being integrated into algorithmic trading to meet MiFID II regulatory requirements.
Contents
Why AI Wealth Managers Fail: The Gap Between arXiv and Wall Street
1. TL;DR
2. The "Cherry-Picking" Problem
3. Why MAPE is a "Four-Letter Word" in Finance
4. Methodology: The Architecture of Failure
5. Real-World Performance: The "AI Fund" Dissonance
6. Deep Insights & The Road Ahead
6.1. Key Recommendations for Researchers:
7. Conclusion