Advancing Bengali NER: Harnessing MEMM and Rich Linguistic Heuristics

A Proposed Model for Bengali Named Entity Recognition Using Maximum Entropy Markov Model Incorporated with Rich Linguistic Feature Set

2020-01-10
Fahmida Alam, Md. Asiful Islam
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a Bengali Named Entity Recognition (NER) model using a Maximum Entropy Markov Model (MEMM) combined with a rich linguistic feature set. The system focuses on classifying named entities into three core categories: Person, Organization, and Location, specifically designed to handle the morphological complexities of the Bengali language.

TL;DR

Named Entity Recognition (NER) is the backbone of Information Extraction, but applying it to the Bengali language presents unique hurdles like the absence of capitalization and complex morphology. This paper introduces a Maximum Entropy Markov Model (MEMM) framework that goes beyond simple sequence tagging by incorporating a "Rich Linguistic Feature Set," specifically designed to identify Person, Organization, and Location entities in Bengali text.

Background & Motivation: Why is Bengali NER Hard?

In English, NER often relies heavily on the "Capitalization" feature—proper nouns are easily flagged by their first uppercase letter. Bengali, however, lacks this distinction. Furthermore, the language features:

  • Free Word Order: Sentences usually follow S-O-V, but entities can appear anywhere.
  • Ambiguity: A word like "Kabita" can mean "Poem" (common noun) or a person’s name (proper noun).
  • Resource Scarcity: Lack of standardized gazetteers and high-quality tagged corpora.

The authors argue that traditional Hidden Markov Models (HMM) are too restrictive because they are generative and cannot easily handle overlapping tokens or external knowledge sources.

Methodology: The Power of MEMM

The core of the proposal is the shift to a Maximum Entropy Markov Model (MEMM). Unlike HMM, MEMM is a discriminative model, allowing it to weigh multiple, non-independent features simultaneously.

1. The Processing Pipeline

The system follows a rigorous four-step pipeline: Data Collection -> Preprocessing (Tokenization, Stemming, Punctuation Removal) -> POS Tagging -> Entity Classification.

2. Feature Engineering

The "Secret Sauce" lies in the features fed into the MEMM:

  • Context Windows: Analyzing surrounding tags to provide clues (e.g., a "Profession" noun appearing before a proper noun).
  • Morphological Suffixes: Identifying "Mohammad" as a person, but "Mohammad-pur" (with the suffix 'pur') as a location.
  • Linguistic Rules: A set of 4 specific rules to group contiguous proper nouns (NNP) and validate them against gazetteer lists for titles (Mr., Director) or organization types (Company, Association).

System Architecture Figure 1: The Proposed NER Model Workflow

Technical Deep Dive: Rules and Logic

The authors implement a logic where the POS Tag Information serves as the primary filter. By focusing on the <NNP> (Proper Noun) tag generated via NLTK, the model then applies linguistic heuristics:

  • Rule 1 (Contiguity): If multiple NNPs appear in a row, they are treated as a single entity (e.g., "Kazi Anis Ahmed").
  • Rule 4 (Postposition Check): If a postposition appears after an entity, the model checks for organization keywords (like "Club" or "Society") to distinguish between a person and a collective body.

MEMM Graphical Representation Figure 2: Graphical representation of the Maximum Entropy Markov Model transitions

Experimental Insights

By testing on news articles, the model demonstrated an ability to resolve complex person-profession associations. For example, in the phrase "Dhaka Bank Director Kazi Anis Ahmed", the model uses the word "Director" (identified as a profession) to correctly label the subsequent proper nouns as a Person.

TagDescription
NNNoun
NNPProper Noun
JJAdjective
VMFinite Verb
Table 1: Subset of the Bengali POS Tagset used for feature extraction.

Conclusion & Future Outlook

While the paper successfully demonstrates the effectiveness of MEMM and linguistic rules, it highlights the ongoing struggle with "Resource Scarcity" in Bengali NLP. This work serves as a vital stepping stone for transitioning from simple rule-based systems to more sophisticated machine learning models.

Future research in this domain is likely to move toward Transformer-based models (like BanglaBERT), but the linguistic insights—specifically the use of suffixes and prefix-based gazetteers—will remain essential "Inductive Biases" for improving accuracy in low-resource settings.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning architectures like Bi-LSTM-CRF or Transformers for Bengali Named Entity Recognition to compare against classical MEMM approaches.
  • What are the most comprehensive publicly available Bengali corpora for NER, and how do they address the scarcity of resources mentioned in the ICCA 2020 paper?
  • Explore how the linguistic feature engineering techniques used for Bengali NER have been adapted for other Indo-Aryan languages like Hindi or Marathi in cross-lingual transfer learning studies.
Contents
Advancing Bengali NER: Harnessing MEMM and Rich Linguistic Heuristics
1. TL;DR
2. Background & Motivation: Why is Bengali NER Hard?
3. Methodology: The Power of MEMM
3.1. 1. The Processing Pipeline
3.2. 2. Feature Engineering
4. Technical Deep Dive: Rules and Logic
5. Experimental Insights
6. Conclusion & Future Outlook