GRAM: Breaking the Deterministic Ceiling of Recursive Latent Reasoning
Generative Recursive Reasoning
This paper introduces Generative Recursive reAsoning Models (GRAM), a framework that reformulates recursive latent reasoning as a stochastic, generative process. By replacing deterministic state updates with learned stochastic transitions, GRAM achieves state-of-the-art results on structured reasoning tasks like Sudoku-Extreme and ARC-AGI.
TL;DR
Researchers have introduced Generative Recursive reAsoning Models (GRAM), a framework that shifts neural reasoning from a single, deterministic path to a probabilistic exploration of multiple latent trajectories. By treating reasoning as a generative process, GRAM allows models to scale via both depth (more steps) and width (parallel samples), achieving superior performance on ARC-AGI and Sudoku-Extreme while naturally handling problems with multiple valid solutions.
Background: The Problem with One-Way Thinking
In the quest for efficient AI reasoning, Recursive Reasoning Models (RRMs)—like Looped Transformers or TRM—have emerged as a parameter-efficient alternative to massive LLMs. Instead of generating a long "Chain-of-Thought" (CoT) sequence, they refine a persistent hidden state through repeated computation.
However, current RRMs suffer from a "deterministic collapse." Like a hiker following a single preset path, if the first few steps are slightly off, the model gets stuck in a local minimum with no way to backtrack or explore alternatives. This is why tiny recursive networks often struggle with "multi-solution" puzzles or extremely high-constraint tasks where one wrong move invalidates the entire result.
The Core Insight: Stochastic Latent Transitions
GRAM solves this by turning the latent update into a stochastic process. Instead of a fixed function , GRAM samples the next state from a distribution:
The architecture uses a clever Hierarchical Transition mechanism:
- Low-level (Inner Loop): A deterministic refinement that handles fine-grained local computation.
- High-level (Outer Loop): A stochastic update ("Stochastic Guidance") that steers the abstract reasoning trajectory.

By training this system using Amortized Variational Inference, the model learns to propose diverse reasoning paths that are likely to lead to a correct solution.
Beyond Depth: Scaling the "Width"
One of the most exciting aspects of GRAM is Inference-Time Scaling. Traditionally, to make a model "smarter," you either make it bigger (parameters) or run it longer (sequential depth).
GRAM introduces a third axis: Width. Because transitions are stochastic, you can sample reasoning trajectories in parallel.
- The Result: 20 parallel samples at 16 iterations achieve higher accuracy (97.0%) than a deterministic model running for 320 iterations (90.5%).
- Efficiency: Parallel sampling bypasses the sequential latency bottleneck, making "thinking wider" faster than "thinking deeper" on modern GPU hardware.

Experimental Results: Slaying the Chaos
The researchers tested GRAM on several "hard" benchmarks where deterministic models typically fail:
- Sudoku-Extreme: Puzzles with minimal clues. GRAM achieved 97% accuracy, while -mini and DeepSeek-R1 (using standard prompting) struggled to solve these specific constraint-heavy puzzles.
- ARC-AGI: GRAM reached 52% on ARC-1, a significant jump over prior recursive SOTA.
- Multi-Solution Tasks: In N-Queens and Graph Coloring, where multiple valid answers exist, GRAM demonstrated excellent Coverage, finding diverse solutions whereas other models would repeatedly output the same answer.

Deep Insight: Reasoning as Unconditional Generation
Remarkably, GRAM can act as an unconditional generator. If you give it an empty board, it can "reason its way" into creating a perfectly valid, unique Sudoku puzzle from scratch. This bridges the gap between structured reasoning (finding a solution) and creative generation (creating a valid structure), proving that the internal logic of a recursive model can be used to generate data that adheres to complex global constraints.
Critical Analysis & Future Outlook
While GRAM is a breakthrough for compact reasoning models, it still faces challenges:
- Training Efficiency: The sequential nature of deep supervision makes training slower than the massively parallel training of standard Transformers.
- Scalability: While 10M-parameter models perform exceptionally well, it remains to be seen if this stochastic latent approach can be scaled to the 70B+ parameter regime effectively.
The Takeaway: GRAM proves that reasoning shouldn't be a straight line. By embracing uncertainty and stochasticity in the latent space, we can build models that are not only more efficient but also more robust at solving the world's most complex logical puzzles.
Reference: Baek et al., "Generative Recursive Reasoning". Licensed under Creative Commons / MIT for respective benchmarks.
