The Latent Space Revolution: From Verbal Tokens to Machine-Native Thinking

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

2026-01-01
⋆ Xinlei Yu *, Zhangquan Chen *, Yongbo He *, Tianyu Fu *, Cheng Yang *, Chengming Xu *, Yue Ma *, † Xiaobin Hu *, Zhe Cao, Jie Xu, Guibin Zhang, Jiale Tao, Jiayi Zhang, Siyuan Ma, Kaituo Feng, Haojie Huang, Youxing Li, Ronghao Chen, Huacan Wang, Chenglin Wu, Zikun Su, Xiaogang Xu, Kelu Yao, Kun Wang, Chen Gao, Yue Liao, Ruqi Huang, Tao Jin, Zhucun Xue, Cheng Tan †, Jiangning Zhang †, Wenqi Ren, Yanwei Fu, Yong Liu, Yu Wang, Xiangyu Yue †, Yu-Gang Jiang †, Fudan University Tsinghua University Zhejiang University Shanghai Artificial Intelligence Laboratory Renmin University of China The Chinese University of Hong Kong The Hong Kong University of Science and Technology DeepWisdom Nanjing University Shanghai Jiatong University Nanyang Technological University Tencent Hunyuan QuantaAlpha Beijing University of Posts and Telecommunications Zhejiang Lab University of Chinese Academy of Sciences Hong Kong University of Science and Technology (Guangzhou) Sun Yat-sen University https://github.com/YU-deep/Awesome-Latent-Space Shuicheng Yan † * Core Contributors † Core Supervisors ⋆ Organizer National University of Singapore
Summary
Problem
Method
Results
Takeaways
Abstract

This comprehensive survey introduces the "Latent Space" paradigm for language-based models, shifting computation from explicit token generation to a machine-native continuous substrate. It provides a unified taxonomy of Mechanisms (Architecture, Representation, Computation, Optimization) and Abilities (Reasoning, Planning, Perception, Memory, etc.), establishing latent space as a core system paradigm for next-generation AI.

TL;DR

Modern AI is moving beyond "talking to itself" in human tokens. This paper argues that the next leap in intelligence lies in Latent Space—a continuous, machine-native substrate where models reason, plan, and remember without the bottlenecks of natural language. By shifting search and deduction into high-dimensional manifolds, models achieve higher fidelity, better efficiency, and superior multimodal integration.

Background: The Limits of the "Verbal" BottleNeck

For years, we have evaluated AI through its Explicit Space (Verbal Space)—the strings of tokens it generates. However, this "token-centric" view is fundamentally flawed for high-level intelligence. Natural language is:

  • Redundant: Most tokens serve grammar, not logic.
  • Discrete: It forces a "quantization bottleneck" where nuanced internal states are lost.
  • Sequential: It locks the model into slow, step-by-step decoding.

As the survey points out, "Latent Space" is not just a hidden layer; it is a Machine-Native Substrate that allows for Superposition Reasoning—the ability to explore multiple reasoning paths simultaneously in a single continuous vector.

Methodology: How Latent Space Works (The Mechanism)

The paper deconstructs the latent paradigm into four pillars:

1. Architecture: The Engine of Thought

We are seeing a shift from standard Transformers to:

  • Recurrent Backbones: Like Huginn and Ouro, which reuse layers to "ponder" in depth.
  • Auxiliary Models: Using JEPA-style architectures to provide perceptual anchors.

Model Mechanism Overview

2. Computation: From Compressed to Adaptive

Computation in latent space isn't fixed. Adaptive Pondering allows models to spend more "latent cycles" on hard problems and "shortcut" easy ones. Interleaved Computation enables a hybrid flow where a model can switch between silent latent thinking and explicit verbal reporting.

3. Representation: Chain-of-Embedding (CoE)

Instead of a Chain-of-Thought (CoT), we now have Chain-of-Embedding. This preserves fine-grained uncertainty and multi-modal alignment that language simply cannot describe—such as the precise torque for a robotic arm or the 3D spatial layout of a room.

Empowering New Abilities

The survey identifies seven core abilities unlocked by this shift, but two stand out:

  • Perception & Imagination: Latent space allows models to "sketch" internal visual thoughts. Paradigms like Latent Sketchpad enable models to simulate 3D environments internally before acting.
  • Embodied Action: For robotics (VLA models), latent space solves the "correspondence problem." Actions are mapped into a body-agnostic latent space, allowing a model trained on human videos to control a four-legged robot zero-shot.

Abilities Taxonomy

The Critical Analysis: The "Black Box" Problem

As an editor, I must highlight the "Double-Edged Sword" mentioned by the authors. While latent space is more powerful, it is less evaluable.

  • Interpretability: If the model doesn't "speak" its reasoning, how do we audit it?
  • Controllability: Intervening in a continuous manifold is harder than editing a prompt.

The paper calls for a new "Representation Engineering" (RepE) to make these hidden "thoughts" governable.

Final Takeaway: The Systems Paradigm

We are entering an era where Language is the Interface, but Latent is the Workspace. Future SOTA models will likely be judged not just by their parameter count, but by the "depth" and "flexibility" of their latent manifolds. This survey provides the first comprehensive map of this new territory.


For more details on the evolution of latent models, visit the Awesome-Latent-Space repository.

Find Similar Papers

Try Our Examples

  • Search for recent papers that specifically compare the computational complexity of "Reasoning by Superposition" in latent spaces versus traditional autoregressive token generation.
  • Which study was the first to propose the "COCONUT" framework for continuous thought, and how have subsequent works like "Ouro" or "Huginn" improved upon its stability?
  • Find research papers exploring the application of latent action spaces in multi-modal Vision-Language-Action (VLA) models for cross-embodiment robotic transfer.
Contents
The Latent Space Revolution: From Verbal Tokens to Machine-Native Thinking
1. TL;DR
2. Background: The Limits of the "Verbal" BottleNeck
3. Methodology: How Latent Space Works (The Mechanism)
3.1. 1. Architecture: The Engine of Thought
3.2. 2. Computation: From Compressed to Adaptive
3.3. 3. Representation: Chain-of-Embedding (CoE)
4. Empowering New Abilities
5. The Critical Analysis: The "Black Box" Problem
6. Final Takeaway: The Systems Paradigm