Resolving the Academic Identity Crisis: A BSNMF Approach to Institutional Data Governance

Data Quality Management in Institutional Research Output Data Center

2019-01-01
Xiaohua Shi, Zhuoyuan Xing, Hongtao Lu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a comprehensive data quality management framework for developing institutional research output data centers. By utilizing a hybrid approach of Levenshtein text distance for department matching and Bayesian Symmetric Non-negative Matrix Factorization (BSNMF) for author name disambiguation, the authors achieve a 10.07% accuracy improvement in aggregating multi-affiliation scholar data.

TL;DR

Managing a university's research output is often a data science nightmare characterized by "same name, different person" and "same person, ten different affiliations." This paper presents an end-to-end framework for institutional data quality management that leverages the Levenshtein distance for fuzzy matching and Bayesian Symmetric Non-negative Matrix Factorization (BSNMF) to detect communities within co-author networks. The result? A 10% boost in data accuracy and a cleaner, more authoritative academic record.

Background: The Silica-fication of Research Data

In the era of "Double World-Class" initiatives, universities need unified data centers to support talent evaluation and subject development. However, data from sources like Web of Science or CNKI often lacks uniform standards. A researcher named "Yang Liu" might be listed under a State Key Lab in one paper and a School of Medicine in another. Without a robust Author Name Disambiguation (AND) strategy, the institutional "Big Data" becomes "Bad Data."

Methodology: The Three-Tiered Defense

The authors propose an eight-module pipeline, but the heavy lifting happens in the Matching and Fitting phases.

1. Fuzzy Matching with Levenshtein Distance

To handle the "non-standard abbreviation" problem (e.g., "Sch Agr & Biol" vs. "Coll Agr & Biol Sci"), the system uses an optimized Levenshtein algorithm. Unlike standard edit distance, the authors introduced domain-specific weights—for example, reducing distance penalties for predictable variations like "Affiliated Hospital 1" vs. "Affiliated Hospital 3" to ensure they are correctly distinguished or grouped.

2. Community Detection via BSNMF

The "Fitting" module treats the academic world as a graph. If two "Yang Liu" nodes share a dense network of co-authors, they are likely the same person. The paper employs BSNMF, a matrix learning method that captures the underlying community structure (modularity) of the co-author network.

Conceptual Framework of the Data Processing Pipeline Figure 1: The proposed conceptual framework for quality-controlled data integration.

Experiments: Performance at Scale

The study tested five community detection algorithms on a massive co-author network of 17,908 nodes derived from Shanghai Jiao Tong University's (SJTU) SCI output.

Comparative SOTA Analysis:

Compared to Greedy Community Detection (GCD) and Louvain (BGLL), the BSNMF method achieved a superior Modularity score of 0.7564, indicating it found more cohesive and meaningful academic clusters.

Algorithm Comparison Table Table 2: Comparison of community detection methods. BSNMF achieves the highest Modularity.

Real-world Case: The "Yang Liu" Disambiguation

Through BSNMF, the system successfully identified that "Yang Liu" appearing across five different labels (e.g., "Stem Cell Res Ctr", "Renji Hosp") was actually a single entity in the School of Medicine, while a different "Yang Liu" belonged to the School of Life Science.

Critical Analysis & Future Outlook

The framework's strength lies in its iterative feedback loop (Data Center -> Faculty Claim -> Supervised Learning). By allowing scholars to "claim" or "reject" papers, the system continuously refines its feature set for the "Learning" module.

Limitations:

  • Scalability: While matrix factorization is robust, the computational cost of BSNMF on global-scale networks (millions of nodes) remains a challenge compared to simpler heuristics.
  • Cold Start: For new faculty members without a co-authorship history, the network-based "Fitting" module is less effective.

Conclusion: This research proves that institutional data quality isn't just an IT problem—it's a machine learning problem. By combining text similarity with latent community analysis, universities can finally bridge the gap between fragmented administrative records and a true "Big Data" academic ecosystem.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) for author name disambiguation in Research Information Systems (CRIS).
  • Identify the foundational papers for Bayesian Symmetric Non-negative Matrix Factorization (BSNMF) and examine how this paper adapts the update rules for co-authorship networks.
  • Explore how big data governance frameworks in universities are integrating OAuth 2.0 and row-level access control for multi-departmental data sharing.
Contents
Resolving the Academic Identity Crisis: A BSNMF Approach to Institutional Data Governance
1. TL;DR
2. Background: The Silica-fication of Research Data
3. Methodology: The Three-Tiered Defense
3.1. 1. Fuzzy Matching with Levenshtein Distance
3.2. 2. Community Detection via BSNMF
4. Experiments: Performance at Scale
4.1. Comparative SOTA Analysis:
4.2. Real-world Case: The "Yang Liu" Disambiguation
5. Critical Analysis & Future Outlook