Predictive Analytics in Green Accounting: Decoding Environmental Reporting in Greece
Applying Machine Learning Techniques for Environmental Reporting
This paper investigates environmental reporting levels in Greek listed companies using machine learning. It employs an ensemble of regression models (LR, M5, LWR, SMOreg) to predict reporting quality based on financial and visibility variables, achieving a correlation coefficient of 0.7416.
TL;DR
This study bridges the gap between environmental accounting and data science by analyzing the disclosure habits of Greek listed companies. By applying machine learning techniques like RRELIEF for feature selection and Ensemble Regression, the researchers proved that a company's visibility and regulatory alignment (EMAS) are far better predictors of "green transparency" than their actual stock market performance.
Background & Motivation
While environmental reporting is mature in the US and UK, Greece represents a unique case study of a late adopter navigating the transition to International Accounting Standards (IAS). The core challenge is the "Information Position": why do some companies disclose more? Is it because they are profitable, or because the media is watching? Traditional linear models often oversimplify these motivations.
Methodology: Beyond Simple Linear Regression
The authors didn't just look for correlations; they treated reporting as a regression problem.
- Feature Selection (RRELIEF): Before modeling, the researchers used the RRELIEF algorithm to rank variables. Unlike standard filters, RRELIEF estimates the quality of attributes by how well they distinguish between "near" instances in a continuous space.
- The Ensemble Approach: The study compared four distinct learners:
- M5 Model Trees: Bridging decision trees and linear regression.
- Locally Weighted Regression (LWR): An expensive but flexible instance-based learner.
- SMOreg: A Support Vector Machine variant for regression.
- Linear Regression (LR): The baseline.
Architecture & Variable mapping
The model follows the functional form:
Environmental Reporting = f(Information Cost, Proprietary Cost, Media Visibility, Control Variables)

Key Insights from the Results
The experiment highlights a significant disparity in Greek industries. Sector leaders like AGET and S&B achieved near-perfect reporting scores (3.0), while many others remained at zero, indicating a highly fragmented landscape.
Performance Metrics
The Averaging Ensemble proved superior, balancing the biases of individual models:
- Correlation Coefficient: 0.7416 (vs. 0.7258 for standard LR).
- Mean Absolute Error (MAE): 0.4313.

The RRELIEF scores (Table 3 in the paper) provided the most striking "Why":
- Media Visibility (0.126) and EMAS (0.07) were the top influencers.
- Annual Stock Market Return (-0.005) had virtually no impact.
- Insight: Companies do not report because they are doing well financially; they report because they are being watched or are trying to comply with specific European frameworks.
Critical Analysis & Conclusion
This work demonstrates that Ensemble Learning provides a more nuanced lens for social sciences than traditional econometrics.
Limitations: The sample size (44 companies over 2 years) is small by modern ML standards. Furthermore, the "Media Visibility" variable is treated as a binary (web page presence), which lacks the depth of modern sentiment analysis.
Future Outlook: Integrating Natural Language Processing (NLP) to automatically score the text of these financial statements, rather than relying on manual Wiseman indexing, would be the logical next step in automating environmental auditing.
