LLM-Augmented Entity Linking: Bridging the Context Gap with T5 and LlaMA-3
Contextual Augmentation for Entity Linking using Large Language Models
The paper introduces a novel Entity Linking (EL) framework that combines a jointly fine-tuned T5 model for Named Entity Recognition (NER) and Entity Disambiguation (ED) with a LlaMA-3 powered contextual augmentation strategy. By expanding ambiguous mentions into their full Wikipedia forms, the approach achieves state-of-the-art performance on various out-of-domain benchmark datasets.
TL;DR
Entity Linking (EL) is often hamstrung by short, ambiguous texts. This paper proposes a dual-model synergy: a fine-tuned T5 acts as the efficient workhorse for recognition and disambiguation, while LlaMA-3 acts as a "knowledgeable consultant," expanding vague mentions (like "AK") into explicit entities ("Alaska"). This hybrid approach sets new SOTA benchmarks for out-of-domain entity linking.
The Problem: The "Jaguar" Dilemma
In the world of Knowledge Graphs, ambiguity is the enemy. Does "Jaguar" refer to the animal, the luxury car, or the NFL team? Traditional two-stage systems (Candidate Generation + Re-ranking) often fail when:
- Context is sparse: Questions or news headlines provide few clues.
- Long-tail mentions: Rare entities or first-name-only mentions (e.g., "Brad" for Brad Pitt) are easily missed.
- Out-of-domain shifts: Models trained on news fail on sports or technical papers.
The authors identify that the missing link isn't just a better retriever, but richer context.
Methodology: Joint Fine-tuning & LLM Expansion
The architecture simplifies the EL pipeline while supercharging the input.
1. Unified T5 Backbone
Instead of separate models for NER and Disambiguation, the authors fine-tune a T5 model to perform both.
- NER Task: Maps C to
[BEGIN_ENT] mention [END_ENT]. - ED Task: Maps the annotated text to
[BEGIN_ENT] mention [END_ENT] [Wikipedia_Title].
2. Contextual Augmentation (The Secret Sauce)
Before the final linking step, the system uses LlaMA-3 70B to expand mentions. By providing the LLM with the mention and its surrounding context, it prompts the model to generate the "likely Wikipedia title."
Figure 1: The unified framework integrating T5 joint fine-tuning and LlaMA-based expansion.
Experiments & SOTA Results
The authors benchmarked their approach against heavyweights like GENRE and BLINK using the GERBIL framework.
Performance Highlights:
- KORE50 (Hard Dataset): Achieved a Micro F1 of 72.7, a massive jump from the previous 64.5.
- Out-of-Domain Excellence: While performance on in-domain AIDA-test-B was slightly lower (likely due to T5 being trained on AIDA and the LLM adding "too much" context for already resolved entities), the model dominated on diverse datasets like OKE 2015/2016 and Reuters.
| Approach | K50 (F1) | R-500 (F1) | OKE 2016 (F1) |
|---|---|---|---|
| De Cao et al. (GENRE) | 60.7 | 40.3 | 50.0 |
| Zhang et al. (EntQA) | 64.5 | 41.9 | 51.3 |
| T5(NER+ED) + Aug (Ours) | 70.6 | 56.6 | 58.5 |
| Flair & T5(ED) + Aug (Ours) | 72.7 | 56.3 | 61.5 |
Ablation: Does LLM Size Matter?
The authors confirm a clear scaling law: LlaMA-3 70B significantly outperformed its 8B counterpart and the LlaMA-2 series. The larger models were "more stable," producing consistent JSON outputs that didn't hallucinate non-existent Wikipedia anchors.
Critical Insight: Mitigating Hallucinations
A significant challenge in using LLMs for EL is their tendency to invent entities. The authors implemented two clever safeguards:
- Dictionary Matching: Only allow Wikipedia titles that exist in a pre-built URI dictionary.
- Span Constraint: Only expand spans that were explicitly identified by the NER step, preventing the LLM from link-stuffing.
Conclusion
This research demonstrates that we don't necessarily need to replace established architectures (like T5/BART) with massive LLMs for the entire EL task. Instead, using LLMs as a pre-processing expansion layer allows us to leverage their vast internal knowledge to "clarify" text before an efficient, specific model does the heavy lifting of linking.
Takeaway for Practitioners: If your EL system is struggling with short-text ambiguity, don't just retrain—try augmenting the mentions with a zero-shot LLM prompt.
