Automated Readability for Filipino Literature: Beyond Traditional Formulas
Developing a machine learning-based grade level classifier for Filipino children’s literature
This paper introduces a machine learning-based approach to classify the readability levels of Filipino children's literature into four distinct grade clusters. By evaluating Multinomial Naive Bayes, K-Nearest Neighbors (KNN), and Random Forest (RF) across various text features, the authors developed a high-performing soft-voting ensemble model that achieves a validation accuracy of 74.19%.
TL;DR
Researchers have developed a machine learning baseline to solve the long-standing problem of objectively grading Filipino children's books. By shifting from static formulas to an ensemble of KNN and Random Forest models, they achieved 74% accuracy in classifying books for appropriate grade levels, discovering that "stop words" are surprisingly vital markers of text difficulty in the Filipino language.
Background: The "Pagbasa...pag-asa" Challenge
In the Philippines, the mantra "In reading, there is hope" underscores the importance of literacy. However, hope requires the right tools. For decades, educators have lacked a standardized scale to measure the readability of Filipino texts. Traditional English-centric formulas like Fry or SMOG consistently yield invalid results when applied to Filipino due to the language’s unique agglutinative nature and different syntactic rules. This study marks a transition from subjective human labeling to objective, data-driven categorization.
Why Traditional Methods Fail
Traditional readability metrics often rely on simple surface features like sentence length or syllable counts. While these work reasonably well for English, they ignore the nuances of the Filipino language where a single long, inflected word (e.g., nagtutulog-tulugan) carries significant morphological complexity that simple character counts cannot fully capture.
Methodology: The Power of Voting
The researchers utilized a dataset of 258 picture books from Adarna House Inc., a leading publisher. The workflow moved through three critical phases:
- Feature Extraction: They tested Count Vectors, TF-IDF, and Trigrams. Interestingly, they chose Trigrams over Bigrams because previous research suggests Trigrams better capture the building blocks of the Filipino language.
- Algorithm Selection: They experimented with Multinomial Naive Bayes (MNB), K-Nearest Neighbors (KNN), and Random Forest (RF).
- The Ensemble Shift: Recognizing that single models had limitations, they implemented a Soft Voting Mechanism. This averaged the predicted probabilities of KNN and RF, allowing the strengths of both (KNN’s local similarity and RF’s decision-tree robustness) to complement each other.
Fig 1: The architecture of the proposed machine learning pipeline from data collection to classification.
Key Insights: The "Stop Word" Paradox
In most NLP tasks, "stop words" (common connectors like ang, mga, sa) are treated as noise and removed. However, this study found that in Filipino children's literature, these words are top features for determining readability.
- Lower Levels: Feature simple, repetitive stop words.
- Higher Levels: Involve more complex inflections and a higher density of specific functional connectors.
Table: Comparison of Soft vs. Hard Voting mechanisms showing the superiority of the weighted probability approach.
Experimental Results
The Soft Voting model (KNN + RF with Count Vectors) emerged as the SOTA (State of the Art) for this specific task:
- Training Accuracy: 82.2%
- Validation Accuracy: 74.19% (on completely unseen data)
The researchers observed that Random Forest was particularly effective at reducing bias and controlling overfitting by aggregating results from multiple decision trees. This suggests that the hierarchical nature of decision trees matches the hierarchical complexity of language acquisition in children.
Critical Analysis & Conclusion
While this baseline is a major step forward, the study highlights typical "low-resource" challenges. The dataset (258 books) is relatively small, which led the authors to cluster grade levels (e.g., grouping Grades 1-3) to improve generalization.
The Takeaway: This research proves that for Philippine languages, automated readability is not just about word length, but about the distribution of common linguistic markers. Future work utilizing Deep Learning (like Transformers) could potentially remove the need for manual feature engineering, provided a larger corpus is curated. For now, the soft-voting ensemble provides a reliable, objective "yardstick" for Filipino educators and publishers.
