Blog

告别冗长 PDF,一文读懂最新顶会与核心期刊的创新价值。

VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
Progressive Residual Warmup for Language Model Pretraining
Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation
Data Analogies Enable Efficient Cross-Embodiment Transfer
FlowMotion: Training-Free Flow Guidance for Video Motion Transfer
LATO: 3D Mesh Flow Matching with Structured TOpology Preserving LAtents
MoEMambaMIL: Structure-Aware Selective State Space Modeling for Whole-Slide Image Analysis
PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
Computational Pathology in the Era of Emerging Foundation and Agentic AI -- International Expert Perspectives on Clinical Integration and Translational Readiness
AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
Fly360: Omnidirectional Obstacle Avoidance within Drone View
Unified Learning of Temporal Task Structure and Action Timing for Bimanual Robot Manipulation
XAI for Coding Agent Failures: Transforming Raw Execution Traces into Actionable Insights
NOVA: Next-step Open-Vocabulary Autoregression for 3D Multi-Object Tracking in Autonomous Driving
TADPO: Reinforcement Learning Goes Off-road
CodeScout: Contextual Problem Statement Enhancement for Software Agents
Relational Semantic Reasoning on 3D Scene Graphs for Open World Interactive Object Search
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking
Pinterest Canvas: Large-Scale Image Generation at Pinterest
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
GreenRFM: Toward a resource-efficient radiology foundation model
SCOPE: Scene-Contextualized Incremental Few-Shot 3D Segmentation
Stock Market Prediction Using Node Transformer Architecture Integrated with BERT Sentiment Analysis
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
NEGATE: Constrained Semantic Guidance for Linguistic Negation in Text-to-Video Diffusion