GIVA: Unifying Smart City Silos through Ontology-Based Instance Matching
Ontology-based Instance Matching for Geospatial Urban Data Integration
The paper introduces GIVA, a geospatial data integration framework that employs ontology-based instance matching to unify disparate urban datasets. Specifically, it demonstrates the integration of "Business Licenses" and "Food Inspections" from the City of Chicago to create a unified view of business entities, supporting predictive analytics and situational awareness.
TL;DR
Researchers at the University of Illinois at Chicago have developed GIVA, a framework that solves a major "Smart City" headache: how to tell if two records in different government databases refer to the same physical business when they don't share a unique ID. By combining spatial ontologies with advanced string-matching algorithms, they unified thousands of Chicago's business records with over 96% accuracy.
The Problem: A Tale of Two (or Ten) Datasets
Imagine you are a city health inspector in Chicago. You have one list for Business Licenses and another for Food Inspections. To predict which restaurant might have a health violation, you need to see both. However, "Seven Eleven" in the license database might be "7-11 Inc." in the inspection log. Without a shared primary key (like a Social Security number for businesses), these records remain "information silos."
The challenges are fourfold:
- Data Heterogeneity: Different formats (CSV, OWL, XML) and naming conventions.
- Representational Variations: Typos, abbreviations, and different address formats for the same building.
- One-to-Many Relationships: One business entity might have dozens of annual licenses or inspection reports.
- Spatial Ambiguity: Co-located businesses (multiple shops in one mall) make location-only matching impossible.
Methodology: The GIVA Framework
GIVA (Geospatial data Integration, Visualization, and Analytics) addresses these challenges through a modular, four-phase pipeline.
1. The Strategy: Ontologies and Blocking
To prevent the system from trying to match every single business in Chicago against every other business (which would be computationally impossible), GIVA uses Blocking. It employs a Spatial Ontology (Country -> State -> City -> Zip) to ensure it only compares businesses within the same geographic area. It also uses a Business Ontology to avoid comparing a "Daycare Center" with a "Liquor Store."
2. The Core: Instance Matching Logic
Rather than relying on a single algorithm, GIVA generates a similarity vector. It calculates several scores for Business Names and Addresses, including:
- ISub Similarity: A metric specifically designed for ontology matching that is more robust than standard Levenshtein distance.
- Weighted Jaccard: Measures the overlap of word components.
- String Cleaning: Removing special characters and normalizing cases.

The system uses a high threshold (90% for names, 95% for addresses) to prioritize Precision—ensuring that when the system says two records are a match, they almost certainly are.
Experimental Results: Cleaning Up Chicago's Data
The team tested GIVA on the Chicago ZIP code 60606. In this single area:
- Internal De-duplication: 16,853 business licenses were collapsed into 3,522 unique business entities.
- Cross-Dataset Matching: 2,179 food inspections were processed, and GIVA successfully matched 96.7% of them to the correct business license record.
Table: Examples of matches and blocked pairs based on similarity scores.
Why It Matters: From Data to Action
The output of GIVA is a Unified Business Model served via a RESTful API. This doesn't just clean up spreadsheets; it changes how the city operates.
By integrating this into platforms like OpenGrid, city administrators can see a single "pin" on a map that contains the entire history of a location—licenses, violations, and even Yelp reviews. This visual and analytical unity allows for Predictive Analytics: inspectors can now identify patterns of violations across the city and intervene before a public health crisis occurs.
OpenGrid UI displaying unified data from multiple sources for a single restaurant.
Conclusion and Future Outlook
GIVA proves that semantic web technologies are not just theoretical; they are practical tools for urban management. The authors plan to expand the system to include Google Places and Yelp data, which will bring the "Voice of the Citizen" into official government workflows. While external validation of third-party data remains a challenge (due to the lack of official keys), GIVA’s modular architecture provides a scalable path forward for the next generation of Smart Cities.
