ME-MGNN: Bridging Granularity Gaps in Chinese Healthcare NER
14351_Multiple Embeddings Enhanced Multi-Graph Neural Networks for Chinese Healthcare Named Entity Recognition.
The paper introduces ME-MGNN, a Multiple Embeddings enhanced Multi-Graph Neural Network designed for Chinese Healthcare Named Entity Recognition (NER). It integrates radical, character, and word-level embeddings into an adapted Gated Graph Sequence Neural Network (GGSNN) to achieve a state-of-the-art F1-score of 75.69 on a newly constructed healthcare dataset.
TL;DR
Recognizing medical entities in Chinese is a "double challenge": the language lacks visual delimiters, and the domain is filled with niche, complex terms. The ME-MGNN model solves this by fusing radicals, characters, and words into a Multi-Graph Neural Network. It doesn't just look at strings; it understands the "DNA" of characters (radicals) and maps them against professional medical dictionaries, achieving a 75.69 F1-score while being 6x faster to train than the previous SOTA.
The "Segmentation Trap" in Chinese Healthcare NER
In English, "Schizophrenia" is one word. In Chinese, “思覺失調症” can be incorrectly split into “思覺” (thinking/feeling), “失調” (disorder), and “症” (disease). This error propagation is the nemesis of Chinese NER.
Moreover, the healthcare domain is unique:
- Visual Semantics: The radical “疒” (illness) immediately signals a disease-related character, even if the model hasn't seen the specific character before.
- Dictionary Dependency: Doctors use standardized terminologies (ICD-10) that act as crucial "anchors" for recognition.
Methodology: The ME-MGNN Architecture
The authors propose a three-tier pipeline that treats NER not just as a sequence task, but as a graph-reasoning problem.
1. Multi-Granularity Representation
The model extracts features from three levels:
- Radical Level: Uses CNNs over radical embeddings to capture "hieroglyphic" meaning.
- Character Level: Uses a Conv-BiLSTM to capture local N-grams and global dependencies.
- Word Level: Aligns segmented word embeddings back to characters.

2. Adapted Gated Graph Sequence Neural Network (GGSNN)
Instead of a flat sequence, ME-MGNN builds a Directed Multi-Graph.
- Nodes: Every character is a node. Special "start" and "end" nodes are added for every dictionary match (e.g., a match in the Symptom dictionary).
- Edges: Directed edges connect adjacent characters and link dictionary-matched spans.
- The "Adapted" Part: Traditional GGSNNs treat all edges as equal. This model learns contribution coefficients (), allowing the network to decide how much to trust the character sequence versus the dictionary matches.

Experiments and Results
The authors curated the first-of-its-kind Chinese Healthcare NER dataset with 68,460 entities across 10 types (Body, Symptom, Chemical, etc.).
Performance vs. Efficiency
The model consistently outperformed the competition:
- vs. BERT: ME-MGNN +1.87 F1 points.
- vs. Lattice-LSTM: ME-MGNN +0.47 F1 points.
- Speedup: While Lattice-LSTM took over 6 days to train, ME-MGNN finished in 23 hours.

Ablation Insights
The ablation study (Table IV) reveals that removing word embeddings causes the most significant drop, proving that boundary information is vital. However, removing radicals also hurts performance, confirming their value as an "inductive bias" for medical semantics.
Critical Analysis: Why This Matters
The real brilliance of ME-MGNN is its efficiency. Lattice-LSTM models are notoriously slow because their architecture changes based on the number of potential word matches in a sentence (creating a computational bottleneck). By using a Graph Neural Network with a fixed propagation mechanism, ME-MGNN achieves the same (or better) "Lattice effect" without the quadratic time complexity overhead.
Limitations:
- The "No-Cross" error (missing an entity entirely) accounts for 72% of errors. This usually happens with extreme domain-specific acronyms (e.g., "ICSI") not present in the training set or dictionary.
- The model relies heavily on the quality of external dictionaries.
Future Outlook
This work sets a new baseline for Chinese clinical informatics. By releasing the first annotated healthcare corpus, the authors have provided the community with a "North Star" for future research. Expect to see this "Multi-Graph + Multi-Embedding" approach migrate to other domain-heavy tasks like legal NER or financial entity extraction where specialized jargon and radicals are paramount.
