Automated Readability for Filipino Literature: Beyond Traditional Formulas

Developing a machine learning-based grade level classifier for Filipino children’s literature

2019-11-01
Joseph Marvin Imperial, Rachel Edita Roxas, Erica Mae Campos, Jemelee Oandasan, Reyniel Caraballo, Ferry Winsley Sabdani, Ani Rosa Almaroi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a machine learning-based approach to classify the readability levels of Filipino children's literature into four distinct grade clusters. By evaluating Multinomial Naive Bayes, K-Nearest Neighbors (KNN), and Random Forest (RF) across various text features, the authors developed a high-performing soft-voting ensemble model that achieves a validation accuracy of 74.19%.

TL;DR

Researchers have developed a machine learning baseline to solve the long-standing problem of objectively grading Filipino children's books. By shifting from static formulas to an ensemble of KNN and Random Forest models, they achieved 74% accuracy in classifying books for appropriate grade levels, discovering that "stop words" are surprisingly vital markers of text difficulty in the Filipino language.

Background: The "Pagbasa...pag-asa" Challenge

In the Philippines, the mantra "In reading, there is hope" underscores the importance of literacy. However, hope requires the right tools. For decades, educators have lacked a standardized scale to measure the readability of Filipino texts. Traditional English-centric formulas like Fry or SMOG consistently yield invalid results when applied to Filipino due to the language’s unique agglutinative nature and different syntactic rules. This study marks a transition from subjective human labeling to objective, data-driven categorization.

Why Traditional Methods Fail

Traditional readability metrics often rely on simple surface features like sentence length or syllable counts. While these work reasonably well for English, they ignore the nuances of the Filipino language where a single long, inflected word (e.g., nagtutulog-tulugan) carries significant morphological complexity that simple character counts cannot fully capture.

Methodology: The Power of Voting

The researchers utilized a dataset of 258 picture books from Adarna House Inc., a leading publisher. The workflow moved through three critical phases:

  1. Feature Extraction: They tested Count Vectors, TF-IDF, and Trigrams. Interestingly, they chose Trigrams over Bigrams because previous research suggests Trigrams better capture the building blocks of the Filipino language.
  2. Algorithm Selection: They experimented with Multinomial Naive Bayes (MNB), K-Nearest Neighbors (KNN), and Random Forest (RF).
  3. The Ensemble Shift: Recognizing that single models had limitations, they implemented a Soft Voting Mechanism. This averaged the predicted probabilities of KNN and RF, allowing the strengths of both (KNN’s local similarity and RF’s decision-tree robustness) to complement each other.

Overall Methodology Flow Fig 1: The architecture of the proposed machine learning pipeline from data collection to classification.

Key Insights: The "Stop Word" Paradox

In most NLP tasks, "stop words" (common connectors like ang, mga, sa) are treated as noise and removed. However, this study found that in Filipino children's literature, these words are top features for determining readability.

  • Lower Levels: Feature simple, repetitive stop words.
  • Higher Levels: Involve more complex inflections and a higher density of specific functional connectors.

Accuracy Table Table: Comparison of Soft vs. Hard Voting mechanisms showing the superiority of the weighted probability approach.

Experimental Results

The Soft Voting model (KNN + RF with Count Vectors) emerged as the SOTA (State of the Art) for this specific task:

  • Training Accuracy: 82.2%
  • Validation Accuracy: 74.19% (on completely unseen data)

The researchers observed that Random Forest was particularly effective at reducing bias and controlling overfitting by aggregating results from multiple decision trees. This suggests that the hierarchical nature of decision trees matches the hierarchical complexity of language acquisition in children.

Critical Analysis & Conclusion

While this baseline is a major step forward, the study highlights typical "low-resource" challenges. The dataset (258 books) is relatively small, which led the authors to cluster grade levels (e.g., grouping Grades 1-3) to improve generalization.

The Takeaway: This research proves that for Philippine languages, automated readability is not just about word length, but about the distribution of common linguistic markers. Future work utilizing Deep Learning (like Transformers) could potentially remove the need for manual feature engineering, provided a larger corpus is curated. For now, the soft-voting ensemble provides a reliable, objective "yardstick" for Filipino educators and publishers.

Find Similar Papers

Try Our Examples

  • Search for recent studies on readability assessment of Southeast Asian languages that utilize deep learning or BERT-based architectures.
  • Which linguistic studies first identified the correlation between function word frequency and syntactic complexity in Austronesian languages?
  • Explore how the soft-voting ensemble method used in this paper has been adapted for readability classification in other morphologically rich, low-resource languages.
Contents
Automated Readability for Filipino Literature: Beyond Traditional Formulas
1. TL;DR
2. Background: The "Pagbasa...pag-asa" Challenge
3. Why Traditional Methods Fail
4. Methodology: The Power of Voting
5. Key Insights: The "Stop Word" Paradox
6. Experimental Results
7. Critical Analysis & Conclusion