newOntExt: Scaling Never-Ending Language Learning for the Modern Web
Never-ending ontology extension through machine reading
This paper introduces newOntExt, an enhanced system for the Never-Ending Language Learning (NELL) project designed to automate ontology extension. By combining state-of-the-art Open Information Extraction (OIE) with a novel "divide-and-conquer" computational architecture, it identifies and names new semantic relations between existing categories in a knowledge base with significantly improved feasibility.
TL;DR
The NELL (Never-Ending Language Learning) system has been reading the web since 2010, but extending its ontology (learning new types of relations) remains a bottleneck. This paper presents newOntExt, a revamped architecture that uses advanced Open IE and a hierarchical data indexing strategy to make ontology extension computationally feasible and helps NELL "self-reflect" by identifying errors in its existing knowledge base.
Background: The Never-Ending Learning Challenge
NELL's mission is to move from "learning a task" to "learning to learn forever." While it excels at populating known categories (e.g., identifying that "Neymar" is an "Athlete"), it struggles to discover new predicates (e.g., realizing that a "Drug" can "Treat" a "Disease").
The previous iteration, OntExt, was slow and noisy—the computational overhead of scanning millions of web sentences to find potential new links between billions of instance pairs was simply too high for a 24/7 operating system.
The Bottleneck: Why Ontology Extension is Hard
- Computational Complexity: Comparing every known instance against every sentence in a SVO (Subject-Verb-Object) corpus leads to a comparison space of roughly —an astronomical number for traditional processing.
- Semantic Noise: Most web-extracted triplets are incoherent. Systems like TextRunner often produced "garbage" relations that polluted the KB.
- The Naming Problem: Components like Prophet can predict that a link should exist between two nodes in the KB graph, but they cannot tell you what that link is called (e.g., they see a connection but don't know it means "isMemberOf").
Methodology: The newOntExt Approach
The researchers introduce a multi-pronged strategy to solve these issues.
1. High-Fidelity Extraction
Instead of relying on first-generation OIE, newOntExt adopts ReVerb and R2A2. These systems use syntactic and lexical constraints to ensure that extracted relations are actually informative, doubling the precision of the underlying data source.
2. Hierarchical File Indexing
To solve the search speed problem, the team moved away from sequential scanning. They implemented a three-level directory structure based on noun prefixes:
Extractions/b/ba/ban/banana.txt
This allows the system to jump directly to the relevant facts for any instance, effectively turning a global search into a local file read.
3. Collaborative Naming (The Prophet Pipeline)
This is the most critical conceptual shift. Instead of finding relations in a vacuum, newOntExt works with Prophet, a link predictor. Prophet identifies unnamed relations by mining the KB graph; newOntExt then "reads" the web to find a name for those specific links.
(Note: This diagram illustrates how Prophet identifies a gap in the graph and newOntExt fills it by clustering verb patterns like "can cure" or "is a treatment for" to label the edge.)
Experimental Insights: Self-Reflection
During testing, newOntExt attempted to name categories like sportsleague and sportsteamposition. While it successfully found relations like lodge has crowned, it also uncovered a significant number of "false beliefs" in NELL.
For instance, it found pairs like (water, sport) or (zero, food). When the system tried to find a verb pattern for these pairs, the clustering failed or produced illogical results. The authors argue that this is actually a feature, not a bug: it provides a signal for Auto-Reflection, allowing the system to go back and delete incorrect instances that it previously thought were true.
(Note: The data shows that by using the divide-and-conquer method, the number of required comparisons for the 'animal' category dropped from to , a massive reduction in search space.)
Takeaways & Future Work
The core contribution of newOntExt is not just a faster algorithm, but a more robust self-supervision loop. By attempting to ground its graph-based predictions in real-world text, NELL can verify its internal logic.
Limitations: The system is still highly sensitive to "noise" in the initial seed classification. If the base categories are wrong, the ontology extension will likely fail.
Future Outlook: Integrating this architecture with modern Large Language Models (LLMs) could potentially solve the "semantic ambiguity" problem, using models like GPT-4 to validate the logic of a proposed relation before it is committed to the KB.
