Vision-Based Racing: Reaching Champion-Level Performance in Gran Turismo 7
A Champion-Level Vision-Based Reinforcement Learning Agent for Competitive Racing in Gran Turismo 7
The paper introduces a champion-level vision-based Reinforcement Learning (RL) agent for Gran Turismo 7 that achieves superhuman performance in competitive multi-opponent racing. Utilizing an asymmetric actor-critic framework, the agent relies solely on ego-centric camera views and onboard sensors during inference, outperforming the game's built-in AI and matching or exceeding human champion benchmarks.
TL;DR
Researchers have developed a vision-based autonomous racing agent capable of defeating world-class human players in Gran Turismo 7. Unlike previous "superhuman" AI that cheated by using global telemetry (knowing exactly where everyone is on a map), this agent drives like a human: looking through the windshield and feeling the car's momentum. Using an asymmetric actor-critic setup and recurrent memory, it manages to navigate 20-car packs at 340 km/h with professional-grade precision.
Background: The "Global Feature" Crutch
In the world of Reinforcement Learning (RL), Gran Turismo Sophy (2022) was a landmark. However, Sophy relied on "global features"—perfect environmental knowledge. Translating this to a real-world race car is nearly impossible because GPS and LiDAR have latency and noise.
If we want AI that can eventually drive a real race car, it must master Vision-Based RL. This is notoriously difficult because:
- High Dimensionality: Processing raw pixels is computationally expensive.
- Partial Observability: If an opponent is in your blind spot, they "disappear" unless you have a memory of where they were.
Methodology: The Asymmetric Advantage
The core innovation lies in how the agent is trained versus how it drives.
1. Asymmetric Actor-Critic
The authors use QR-SAC (Quantile Regression Soft Actor-Critic) but with a twist. During training, the Critic is "all-knowing"—it sees the global map, every opponent's coordinate, and the exact track geometry. This allows it to give the Actor very accurate feedback on the value of its actions. However, the Actor (the part that actually drives) only sees the 64x64 pixel image and basic car data (IMU/Speed). At race time, the Critic is discarded, leaving a lean, vision-only pilot.
2. Temporal Memory (The RNN)
To solve the "blind spot" problem, the Actor isn't just a static CNN. It incorporates a Gated Recurrent Unit (GRU). This acts as a short-term memory, allowing the car to "remember" an opponent it just passed or a curve it's currently entering, even if the visual cues are temporarily obscured.
Figure 1: The Asymmetric Architecture showing the Actor (vision-limited) and Critic (global-aware).
Experiments & Champion Confrontations
The agent was tested on three iconic tracks: Tokyo Expressway, Spa-Francorchamps, and Le Mans (Sarthe).
Key Results:
- Tokyo: The agent crushed Human Champions. Because Tokyo is tight and walled, the vision-based agent actually outperformed the "global" Sophy baseline because vision allowed it to judge the orientation and width of opponents better than Sophy's point-mass representation.
- Winning Margin: Starting 20th, the agent consistently fought its way to 1st place, maintaining a "Sportsmanship" code by avoiding excessive collisions.
Figure 2: Winning margin distribution. The agent (blue) sits significantly further right (higher margin) than the Human Champion (red).
Deep Insight: What is the AI "Looking" At?
Using Integrated Gradients, the team visualized the agent's attention.
- On Straights: It looks at the horizon and treelines (perspective cues) to keep the car centered.
- In Traffic: It focuses on the lower chassis and shadows of opponent cars to judge the exact distance for an overtake—identical to how professional drivers use visual "anchors."
Figure 3: Heatmaps showing the agent's focus on opponents and the horizon during different race segments.
Critical Analysis & Conclusion
This paper proves that the bottleneck for vision-based RL isn't necessarily the "vision" itself, but how we train the agent to value those pixels. By using a privileged Critic, we can train a "blind" Actor to see.
Limitations: The agent is currently "overfit" to specific car-track combinations. A real champion needs to adapt to changing weather and different car handling characteristics instantly. However, as a proof of concept, this is the first time a vision-only agent has truly earned its place on the top step of a GT7 podium against world-class humans.
Future Work: Expanding this to zero-shot generalization across different vehicle classes and dynamic weather (rain/night) will be the final hurdle before this technology hits the real-world track.
