Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

Do Foundation Models Know Geometry? Probing Frozen Features for Continuous Physical Measurement
Text-Driven Emotionally Continuous Talking Face Generation
Critical dynamics govern the evolution of political regimes
Stem: Rethinking Causal Information Flow in Sparse Attention
Scalable Training of Mixture-of-Experts Models with Megatron Core
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
\$OneMillion-Bench: How Far are Language Agents from Human Experts?
How Far Can Unsupervised RLVR Scale LLM Training?
Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
Agentic Critical Training
On the Width Scaling of Neural Optimizers Under Matrix Operator Norms I: Row/Column Normalization and Hyperparameter Transfer
Scale Space Diffusion
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
Fish Audio S2 Technical Report
EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery
RecThinker: An Agentic Framework for Tool-Augmented Reasoning in Recommendation
Grow, Don't Overwrite: Fine-tuning Without Forgetting
Introduction to Generalized Symmetries
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA
Reward Prediction with Factorized World States
VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
ReconDrive: Fast Feed-Forward 4D Gaussian Splatting for Autonomous Driving Scene Reconstruction
PlayWorld: Learning Robot World Models from Autonomous Play
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare
ZeroWBC: Learning Natural Visuomotor Humanoid Control Directly from Human Egocentric Video