FloorSense: Evolving Indoor Maps from Geometry to Semantics via CRF

FloorSense: a novel crowdsourcing map construction algorithm based on conditional random field

2019-05-25
Zhuqing Jiang, Jiahao Zhang, Chonghua Liu, Chengkai Huang
Summary
Problem
Method
Results
Takeaways
Abstract

FloorSense is a crowdsourcing-based indoor map construction framework that utilizes Smartphone sensors (IMU, Barometer, Magnetometer) and Conditional Random Fields (CRF) to generate semantic indoor maps. Beyond typical layout reconstruction, it achieves high-accuracy semantic labeling of functional areas like clothing stores and restrooms in complex mall environments.

TL;DR

Researchers have moved beyond simple "skeleton" maps of buildings. FloorSense introduces a novel crowdsourcing framework that doesn't just draw the walls of a mall—it identifies if a room is a clothing store, a restroom, or a cashier. By using Conditional Random Fields (CRF) and smartphone sensor data (IMU/Barometer), it transforms raw walking traces into a rich semantic map with an average accuracy of 2.1 meters.

Problem & Motivation: The "Blind" Grammar Map

Most modern indoor mapping solutions (like those from Google or MazeMap) provide a "grammar map"—a layout showing where walls and corridors are. However, they lack Semantic Intelligence. They don't know the function of the space.

Existing crowdsourcing methods (like Zee or CrowdInside) rely on Hidden Markov Models (HMM) to predict user movement. The problem? HMM assumes that your current activity is only dependent on your current state, ignoring the context of your entire journey. Furthermore, traditional SLAM (Simultaneous Localization and Mapping) requires expensive robots or specialized hardware, making large-scale mall mapping a logistical nightmare.

Methodology: The Three Pillars of FloorSense

The system operates through a sophisticated top-down pipeline that converts kinetic energy (walking) into spatial data.

1. High-Precision PDR (Pedestrian Dead Reckoning)

To track users accurately, FloorSense utilizes a dynamic thresholding formula for step detection: This adaptation allows the system to maintain less than 1% error in distance estimation regardless of the user's walking pace.

2. Geometry Reconstruction (The Grammar Map)

By collecting thousands of "traces" (user paths), the system uses an Alpha-shape algorithm. Imagine a circle rolling along the outer edges of a cloud of points; the lines it traces form the boundary of the reachable space. Model Architecture Figure: The process of breaking traces into segments and clustering them to define corridors vs. rooms.

3. CRF Semantic Inference (The Core "Brain")

This is where the magic happens. Unlike HMM, the Conditional Random Field (CRF) model allows the system to look at the entire sequence of actions. It uses a 6-dimensional vector for each compartment, including:

  • Stop-and-Still (SS) frequency: People stand still longer at cosmetic counters than in restrooms.
  • Turning (TN) frequency: Restrooms involve specific, tight turning patterns.

Inference Logic Figure: The procedural integration of grammar maps into a graph-based CRF for semantic labeling.

Experiments & Results: Real-World Mall Testing

The team tested FloorSense in a 10,000 m² shopping mall.

Positioning Accuracy

Compared to state-of-the-art (SOTA) systems like Zee or PiLoc, FloorSense holds its own without needing pre-existing floor plans or expensive Wi-Fi infrastructure.

  • Reported Accuracy: 2.1 meters (Average).
  • Trace Density: The system reaches peak performance once roughly 300 user traces are collected.

Accuracy Metrics Figure: The correlation between track density and prediction accuracy.

Semantic Precision

FunctionPrecision
Restroom80.00%
Clothing Store68.03%
Cashier64.52%
Cosmetic Counter61.64%

The high precision for restrooms is a direct result of "continuous turning" patterns being highly distinct in IMU data. Meanwhile, stores are slightly harder to distinguish because "standing and browsing" behavior looks similar across different retail types.

Critical Insight & Conclusion

FloorSense successfully proves that human behavior is a secondary "signal" for mapping. By treating a person's movement as a signature of the room's purpose, we can build maps that are not just blueprints, but functional databases.

Limitations: The system still struggles with "slight turns" that smartphones fail to capture in a pocket, and it assumes users walk naturally. Future work could integrate Wi-Fi RTT (Round Trip Time) or Magnetometer fingerprints to further reduce the 2.1m drift.

Takeaway: The future of LBS (Location-Based Services) isn't just knowing where you are, but what you are doing there.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Transformer-based architectures or Graph Neural Networks to replace CRF for semantic indoor map labeling via crowdsourced trajectory data.
  • Which original paper established the use of Alpha-shapes for indoor boundary reconstruction from point clouds, and how does FloorSense modify this for multi-floor scenarios?
  • Explore how the FloorSense framework's reliance on barometric pressure for floor-level detection can be integrated with visual SLAM for more robust indoor navigation in areas with variable atmospheric conditions.
Contents
FloorSense: Evolving Indoor Maps from Geometry to Semantics via CRF
1. TL;DR
2. Problem & Motivation: The "Blind" Grammar Map
3. Methodology: The Three Pillars of FloorSense
3.1. 1. High-Precision PDR (Pedestrian Dead Reckoning)
3.2. 2. Geometry Reconstruction (The Grammar Map)
3.3. 3. CRF Semantic Inference (The Core "Brain")
4. Experiments & Results: Real-World Mall Testing
4.1. Positioning Accuracy
4.2. Semantic Precision
5. Critical Insight & Conclusion