Turning Every Smartphone into a Periscope: Real-Time NLOS Imaging via Motion-Induced Sampling
Imaging Hidden Objects with Consumer LiDAR via Motion Induced Sampling
This paper introduces a multi-frame fusion strategy for Non-Line-of-Sight (NLOS) imaging using smartphone-grade, low-cost consumer LiDAR. By proposing the Motion-Induced Aperture Sampling (MAS) model and a particle filtering framework, the authors achieve real-time 3D tracking, reconstruction, and camera localization by treating indirect light reflections as valid signals.
TL;DR
Researchers from MIT and Dartmouth have unlocked the ability to see around corners using the $100 LiDAR sensors already found in smartphones and vacuum robots. By moving the camera and using a new mathematical model called Motion-Induced Aperture Sampling (MAS), they fuse multiple noisy frames to track and reconstruct hidden objects in real-time.
The "Blind Spot" of Consumer LiDAR
Non-Line-of-Sight (NLOS) imaging sounds like science fiction: using a wall as a "virtual mirror" to see what is hidden behind an obstacle. Until now, this required massive lasers and picosecond-accurate lab equipment.
Consumer LiDARs (like those in an iPhone) are designed for direct depth sensing. When these sensors try to "see" a hidden object, the signal is buried under three major hurdles:
- Low SNR: Eye-safety regulations restrict laser power.
- Low Resolution: Sensors often have only 10x10 or 8x8 pixels.
- Motion Blur: Handheld movement and object movement scramble the delicate timing of light.
The Insight: Motion is an Asset, Not a Bug
The core philosophy of this paper is inspired by Burst Photography and Synthetic Aperture Radar (SAR). Instead of trying to get a perfect image from one static frame, the authors use the motion of the user's hand to scan the wall, creating a large "Synthetic Aperture."
The MAS Model
The authors developed the Motion-Induced Aperture Sampling (MAS) model. By applying a Light-Cone Transform (LCT), they show that the relationship between object shape and the space-time measurements is essentially a 3D convolution. Crucially, a translation in the object's position results in a simple shift of the signal in the transformed space-time.

The figure above illustrates how the MAS model decomposes measurements into object shape (canonical STIR), object motion (translation ∆), and camera pose (sampling function).
Methodology: Particle Filtering for the "Where" and "What"
Because NLOS measurements are incredibly noisy, traditional "backprojection" (the standard algorithm for NLOS) produces garbage. Instead, the team uses a Particle Filter.
- Propagation: 1,000 "particles" (guesses) move according to a motion prior.
- Evaluation: For each guess, the system "renders" an expected LiDAR measurement using the MAS model and compares it to reality.
- Resampling: Guesses that match the actual data thrive; others are discarded.
This probabilistic approach allows the system to maintain "regions of ambiguity"—knowing exactly where it is certain and where the hidden object might be a blur.
Applications and Experiments
The researchers demonstrated three primary capabilities:
- 3D Tracking: Tracking a hidden person or hand moves in real-time.
- Hidden Reconstruction: Building a 3D model of a hidden mannequin by waving a phone.
- NLOS Localization: Helping a robot find its position by looking at a hidden object as a landmark when the visible wall is just a blank, featureless white plane.

Experimental setups showing tracking (blue/black lines) and localization (camera trajectory) using only indirect reflections.
Critical Analysis: Why This Matters
The most impressive part of this work isn't the tracking accuracy (which is ~4.7 cm), but the democratization. They demonstrated successful tracking using an off-the-shelf ST VL53L8CX sensor ($10-20 component).
Limitations:
- The model currently assumes "retroreflective" properties (like a safety vest) for the best results.
- While it works for diffuse (normal) objects, the SNR drops significantly, requiring slower motion or better priors.
- It requires knowing at least two of the three variables: object shape, object motion, or camera pose.
Conclusion
This research marks the transition of NLOS from a "physics lab curiosity" to a "mobile feature." By treating the physical world and sensor motion as parts of a unified computational model, we no longer need $50,000 hardware to see around corners. The "Plug-and-Play" NLOS era has arrived.
For more details, check out the project page at: sidsoma.com/consumer-nlos/
