NF-CoT: Normalizing Flows as the Engine of Latent Reasoning
3
The paper introduces NF-CoT, a latent reasoning framework that models continuous chain-of-thought (CoT) trajectories using Normalizing Flows (TARFlow-style) embedded within an LLM's causal stream. It achieves state-of-the-art results on several code-generation benchmarks (e.g., HumanEval+, MBPP+) by replacing verbose text rationales with compact, stochastic continuous states.
Executive Summary
TL;DR: NF-CoT is a breakthrough framework that brings the efficiency of continuous latent states to the reasoning capabilities of Large Language Models (LLMs). By embedding Normalizing Flows directly into the causal stream of an LLM, the authors allow the model to "think" in a high-bandwidth continuous space before committing to a final text answer. This approach yields a +13.0% Pass@1 improvement on coding tasks while being 2.7x faster than previous diffusion-based latent reasoners.
Positioning: This work moves beyond the "low-bandwidth" bottleneck of text-based Chain-of-Thought (CoT) and the "slow-generation" bottleneck of latent diffusion, establishing a new SOTA for efficient, probabilistic latent reasoning.
Problem & Motivation: The Verbosity Tax
Explicit Chain-of-Thought (CoT) is the current gold standard for LLM reasoning. However, it suffers from the "Verbosity Tax":
- Efficiency: Generating hundreds of tokens of intermediate steps is slow and memory-intensive.
- Representational Bottleneck: Natural language is discrete and serial; it struggles to represent uncertainty or partially formed semantic updates.
- Prior Work Limitations: Methods like Coconut use deterministic hidden states, losing the ability to sample diverse paths. Diffusion-based methods like LaDiR allow for sampling but require expensive iterative denoising and lack a native autoregressive likelihood.
The authors' insight: Normalizing Flows (NF) provide the "missing link"—they offer the exact likelihood and left-to-right sampling of language models but operate in the efficient continuous domain.
Methodology: Thinking in Flows
NF-CoT transforms a standard LLM into a hybrid generator. It uses a shared backbone (e.g., Qwen3-8B) with two distinct heads:
- NF Head: Predicts the parameters () of a Gaussian density for continuous thought tokens.
- LM Head: Predicts the standard logits for the final discrete text answer.
The Unified Architecture
Unlike prior "dual-path" models, NF-CoT uses a Unified Causal Stream. Continuous thoughts () are generated first, followed immediately by answer tokens (). This allows the model to leverage the KV-cache of the thoughts directly for answer generation, making the transition seamless and lightning-fast.

Learning and RL
NF-CoT is trained with a unified likelihood objective. Because the NF provides a tractable likelihood, the model can be further refined using Reinforcement Learning (GRPO). This ensures that the latent reasoning paths are optimized for final answer correctness (e.g., passing unit tests in code).
Experiments & Results: Speed Meets Accuracy
NF-CoT was tested on rigorous Python coding benchmarks (HumanEval+, MBPP+, LiveCodeBench).
- Performance: On MBPP+, NF-CoT (Unified) reached a 72.1% Pass@1, compared to the base model's 53.8%.
- Scaling: Unlike some RL methods that collapse to a single solution, NF-CoT maintains excellent Pass@k diversity, meaning more samples continue to yield better results.
- Efficiency: Compared to the latents-diffusion baseline (LaDiR), NF-CoT reduces the FLOPs per sample from 49.3T to 19.9T.

Deep Insight: Form vs. Function
A fascinating finding in the ablation studies was Latent Perturbation Robustness. The authors found that adding noise to the continuous thoughts changed the form of the code (variable names, implementation style) but rarely broke the correctness. This suggests that the latent space learns a smooth manifold of "strategies" rather than brittle, discrete tokens.
Conclusion & Future Work
NF-CoT demonstrates that we don't need to choose between the speed of latent reasoning and the probabilistic rigor of text-based CoT. By using Normalizing Flows, we can have both.
Future Outlook: The primary limitation is the reliance on a fixed-length VAE-encoded trajectory. Future iterations might explore dynamic-length latent reasoning, where the model decides how many "continuous thoughts" it needs before answering, potentially unlocking even higher levels of cognitive efficiency.
