Proceedings of the 2nd Workshop on Ecology, Environment, and Natural Language Processing
Francesca Grasso, Valerio Basile, Cristina Bosco, Muhammad Okky Ibrohim, Maria Skeppstedt, Manfred Stede (Editors)
- Anthology ID:
- 2026.nlp4ecology-1
- Month:
- May
- Year:
- 2026
- Address:
- Palma de Mallorca, Spain
- Venues:
- NLP4Ecology | WS
- Events:
- Fifteenth Language Resources and Evaluation Conference | Workshop on Ecology, Environment, and Natural Language Processing (2026) | Other Workshops and Events (2026)
- SIG:
- Publisher:
- European Language Resources Association
- URL:
- https://aclanthology.org/2026.nlp4ecology-1/
- DOI:
- 10.63317/2vhvk7rbcds2
- PDF:
- https://aclanthology.org/2026.nlp4ecology-1.pdf
Proceedings of the 2nd Workshop on Ecology, Environment, and Natural Language Processing
Francesca Grasso | Valerio Basile | Cristina Bosco | Muhammad Okky Ibrohim | Maria Skeppstedt | Manfred Stede
Francesca Grasso | Valerio Basile | Cristina Bosco | Muhammad Okky Ibrohim | Maria Skeppstedt | Manfred Stede
Retrieving Climate Change Disinformation by Narrative
Max Upravitelev | Veronika Solopova | Charlott Jakob | Premtim Sahitaj | Sebastian Möller | Vera Schmitt
Max Upravitelev | Veronika Solopova | Charlott Jakob | Premtim Sahitaj | Sebastian Möller | Vera Schmitt
Climate disinformation evolves faster than the fixed taxonomies used to detect it. Thus, we re-frame narrative detection as a retrieval task: given a narrative’s core message as a query, rank texts from a corpus by alignment with that narrative. This formulation requires no predefined label set and can accommodate emerging narratives. We repurpose three climate disinformation datasets (CARDS, Climate Obstruction, climate change subset of PolyNarrative) for retrieval evaluation and propose SpecFi, a framework that generates hypothetical documents to bridge the gap between abstract narrative descriptions and their concrete textual instantiations. SpecFi uses community summaries from graph-based community detection as few-shot examples for generation, achieving a MAP of 0.505 on CARDS without access to narrative labels. We further introduce narrative variance, an embedding-based difficulty metric, and show via partial correlation analysis that standard retrieval degrades on high-variance narratives (BM25 loses 63.4% of MAP), while SpecFi-CS remains robust (32.7% loss). Our analysis also reveals that unsupervised community summaries converge on descriptions close to expert-crafted taxonomies, suggesting that graph-based methods can surface narrative structure from unlabeled text.
Unsupervised GRI-TCFD Alignment with LLM-Assisted Validation for Climate Disclosure and Greenwashing Risk Analysis
Seyed Alireza Mousavian Anaraki | Danilo Croce | Roberta Costa | Luigi Tiburzi | Armando Calabrese | Roberto Basili
Seyed Alireza Mousavian Anaraki | Danilo Croce | Roberta Costa | Luigi Tiburzi | Armando Calabrese | Roberto Basili
Climate-related corporate disclosures play a central role in sustainable finance and regulatory supervision, but remain difficult to analyze due to their length, unstructured format, and strategic language. While existing NLP approaches have been applied to ESG scoring and greenwashing detection, most operate at the document level and lack explicit alignment with formal reporting standards. We propose a scalable paragraph-level framework for aligning sustainability disclosures with the Global Reporting Initiative (GRI) indicators and the Task Force on Climate-related Financial Disclosures (TCFD) pillars. Our approach combines weak supervision, climate-focused GRI-TCFD mapping, embedding-based semantic similarity, and LLM validation for climate detection. In parallel, we introduce a paragraph-level greenwashing proxy based on commitment intensity, claim specificity, and sentiment polarity. This proxy complements regulatory alignment by capturing linguistic signals associated with potentially symbolic climate communication. The resulting augmented dataset is used to fine-tune ClimateBERT models in both single-task and multi-task settings. Experimental results show that weakly supervised dataset augmentation improves robustness and generalization compared to purely manual training, with further gains in the multi-task configuration. By integrating regulatory semantics, domain-adapted language models, and scalable annotation strategies, this study advances standard-aligned climate disclosure analysis and provides tools directly relevant to climate-related financial risk assessment.
Towards Empowering Consumers through Sentence-level Readability Scoring in German ESG Reports
Benjamin Josef Schüßler | Jakob Prange
Benjamin Josef Schüßler | Jakob Prange
With the ever-growing urgency of sustainability in the economy and society, and the massive stream of information that comes with it, consumers need reliable access to that information. To address this need, companies began publishing so called Environmental, Social, and Governance (ESG) reports, both voluntarily and forced by law. To serve the public, these reports must be addressed not only to financial experts but also to non-expert audiences. But are they written clearly enough? In this work, we extend an existing sentence-level dataset of German ESG reports with crowdsourced readability annotations. We find that, in general, native speakers perceive sentences in ESG reports as easy to read, but also that readability is subjective. We apply various readability scoring methods and evaluate them regarding their prediction error and correlation with human rankings. Our analysis shows that, while LLM prompting has potential for distinguishing clear from hard-to-read sentences, a small finetuned transformer predicts human readability with the lowest error. Averaging predictions of multiple models can slightly improve the performance at the cost of slower inference.
Disambiguating Geographic Names in Biodiversity Occurrence Data: A Retrieval-Augmented Generation Approach
Yanni Jose C. Ella | Monica Ashley R. Laviste | John Michael L. Lastimoso | Wilfred John E. Santiañez | Riza Batista-Navarro | Roselyn Santos Gabud
Yanni Jose C. Ella | Monica Ashley R. Laviste | John Michael L. Lastimoso | Wilfred John E. Santiañez | Riza Batista-Navarro | Roselyn Santos Gabud
The availability of georeferenced coordinates is essential for biodiversity research, as it enables species distribution modeling and supports conservation planning. However, datasets often contain ambiguous or inconsistent geographic names that reduce spatial accuracy and underscore the need for methods that resolve geographic name ambiguity. While traditional named entity linking strategies are well established, they remain limited in low-resource domains, e.g., in biodiversity contexts, due to the scarcity of annotated training data and high lexical ambiguity of local geographic names. This study proposes a Retrieval-Augmented Generation (RAG) framework to automatically disambiguate Philippine seaweed-related geographic names in databases and literature. This approach utilizes a custom knowledge base of gazetteers to support large language models (LLMs) in the task of geospatial disambiguation. With a disambiguation accuracy of 87.8% within a 5 km distance error threshold, our evaluation shows that the RAG-enabled pipeline significantly outperforms standard LLM baselines (Accuracy@5km = 0%), demonstrating the need for external knowledge to resolve geospatial ambiguity.
Recent advances in generative AI have enabled the large-scale production of environmental imagery and descriptions, yet questions remain regarding how such content represents emotion, agency, and responsibility. This study examines how human evaluators respond to AI-generated environmental representations, focusing on sentiment, stance, and argumentation as dimensions of qualitative evaluation. Data were collected from 81 multilingual secondary-school EFL learners in Cyprus, who engaged with AI-generated environmental images and accompanying AI-written descriptions through a sequence of structured tasks. Using qualitative discourse analysis informed by sentiment- and stance-oriented frameworks, the study analyses learner-produced texts to identify affective evaluations, moral positioning, and alignment with or challenge to AI-generated discourse. Findings indicate that participants consistently moved beyond surface-level description to articulate emotional engagement, assign responsibility, and critique omissions in AI-generated content, particularly regarding the representation of human-environment relations. The study contributes to research on human-centered AI evaluation by demonstrating the value of sentiment and stance analysis for assessing AI-generated environmental language, and highlights the potential of educational contexts as sites for examining human interpretive responses to automated discourse.
What Stories Do Language Models Tell About Nature? A Multi Layer Evaluation Framework for Ecological Alignment
Jorge Vallego | Eleanor Tiernan | Mah Rukh | Mariana Roccia | Sabina Fiebig Lord
Jorge Vallego | Eleanor Tiernan | Mah Rukh | Mariana Roccia | Sabina Fiebig Lord
Large language models increasingly generate environmental discourse, yet there is no standardised framework for evaluating the ecological narratives they produce. We introduce a structured prompt corpus and a reproducible multi layer evaluation framework grounded in ecolinguistic theory, operationalising five dimensions of ecological alignment: anthropocentrism, agency attribution, erasure of non human impacts, evaluation of growth, and responsibility framing. The framework integrates human judgement, an ecosophy aligned model judge, and automated semantic metrics, and is applied to outputs from ChatGPT, DeepSeek, and Ecophora, our ecosophy guided model. Ecophora achieves the highest alignment across all layers, with near ceiling judge scores of 159/160 and 142/160, together with the strongest automated composite performance. Divergences between automated metrics and holistic judgement indicate that ecological vocabulary alone does not guarantee ecological reasoning. The proposed framework provides a scalable methodology for benchmarking ecological alignment and assessing narrative shifts in language models.
Ecological Discourse Modeling in a Low-Resource Setting: A Longitudinal Vietnamese Climate Corpus with Comparative Topic Modeling
Huyen Phuong Nguyen
Huyen Phuong Nguyen
Climate change discourse has expanded substantially in recent decades, yet computational analyses remain concentrated on high-resource languages. In this paper, we construct a longitudinal Vietnamese climate news corpus and examine thematic structure and temporal evolution in a lower-resource setting. The corpus comprises 10,401 articles published between 2004 and 2026 and is systematically preprocessed using linguistically informed word segmentation. To ensure domestic relevance, we apply transformer-based Named Entity Recognition and construct a geographically grounded subset of 4,501 Vietnam-focused documents. We analyze this dataset using both Latent Dirichlet Allocation and BERTopic. Results reveal stable thematic dimensions alongside longitudinal shifts from event-driven pollution reporting toward governance- and energy-centered narratives. Embedding-based modeling achieves higher semantic coherence while maintaining comparable topic diversity. The main contribution of this work is thus the compilation of a structured Vietnamese climate corpus and a systematic analysis of discourse evolution in an underrepresented language context.
Greench-v1: distilling SLMs on Greenwashing Detection
Federico Raspanti | Alessandro Pietro Bardelli Bardelli | Simona Scala | İrem Demirtaş | Marilena Di Bari | Michele Filannino
Federico Raspanti | Alessandro Pietro Bardelli Bardelli | Simona Scala | İrem Demirtaş | Marilena Di Bari | Michele Filannino
Validating greenwashing claims in environmental, social, and governance (ESG) reports relies heavily on costly and inconsistent manual review. To address this, this paper introduces Greench-v1, a low-latency small language model (based on Qwen3-4B) that screens ESG text at the paragraph level. The model outputs a three-way classification (Greenwashing Alert, No Greenwashing, Not Relevant) paired with a concise, paragraph-grounded rationale to assist human auditors in triage and validation. The system was trained on a custom dataset of roughly 2,000 paragraphs, adapted from the ClimateBERT corpus. This dataset mitigates class imbalance through controlled paraphrasing of rare positive instances and uses GPT-4o to generate evidence-based justifications. Four training regimes were evaluated: (i) Hard distillation: Supervised fine-tuning on teacher-generated outputs. (ii) Soft distillation: Training the student to match the temperature-scaled logits of a domain-specialized Qwen3-14B teacher. (iii) Group Relative Policy Optimization (GRPO): Reward-based updates driven by exact-match alert generation. (iv) Hybrid GRPO: GRPO initialized from the hard-distilled checkpoint. Distillation and efficient policy optimization significantly improved performance over untuned baselines. Soft distillation and GRPO achieved the strongest results, increasing the “Greenwashing Alert” weighted F1-score by 36.7% and 49.0%, respectively, resulting in a deployable tool for screening large volumes of ESG narratives.
Analyzing Environmental Discourse through Construction-Based Pattern Extraction
Elisa Chierchiello | Eliana Di Palma | Ludovica Pannitto | Cristina Bosco
Elisa Chierchiello | Eliana Di Palma | Ludovica Pannitto | Cristina Bosco
Environmental issues are at the centre of a debate currently taking place across all communication channels. This paper provides an analysis of texts in which these issues are discussed, with the novelty of applying a methodology that enables the extraction and comparison of different narratives and points of view. The texts used in this study are the English Living Planet Reports published biennially by the WWF from 2014 to 2024. The methodology is based on the extraction of constructions – patterns collected in the English constructicon CASA – which allow us to identify differences in the presentation of the issues discussed in the analysed texts. Our results show that this methodology can be very helpful in the comparative analysis of texts to reveal different perspectives, for example, to observe diachronic variations.
Mapping the Historical Ecology of the Cyclades: A Diachronic Natural Language Processing Analysis of Travel Narratives (1700–1920)
Aikaterini Christopoulou | Vassilis Detsis | Basilis Gatos
Aikaterini Christopoulou | Vassilis Detsis | Basilis Gatos
Historical texts can be valuable for the study of a place’s ecological history but reading and extracting information from them can be a tedious and time-consuming task. Natural Language Processing can help in order to extract the most important information of the text in a quick, effective and reproducible way. In this study, travel narratives for the Cyclades Islands from 4 different time periods (1700-1920) have been chosen for analysis. The first step, the quantitative part, includes the semi-automatic detection of geographical entities in the texts and their connection to predefined keywords in order to enable temporal and spatial statistical analysis. The output of this procedure is then inserted in a Retrieval-Augmented Generative Synthesis pipeline in which the text segments with the connected place and keyword are processed by a locally orchestrated Large Language Model. The final output is used for the understanding and interpretation of the original text. Even though the study focuses mainly on the coherence and repeatability of the workflow, an effort is made to interpret and connect the results to the past ecological profile of these islands. The dataset/supplementary material is provided via an open access repository.
Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News
Brielen Madureira | Mariana Madruga de Brito | Andreas Niekler
Brielen Madureira | Mariana Madruga de Brito | Andreas Niekler
In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. Still, enough annotated data to train accurate deep learning-based classifiers from scratch is often not available. Topic Models have the advantage of being both unsupervised and interpretable, but are typically used only for exploratory analysis or data characterisation. In this study, we investigate how to employ Topic Models as binary classifiers for refining the retrieval of relevant news about seven types of extreme climate events in the German media. Our method relies on the posterior distributions estimated by Topic Models to select relevant documents, without modifying their training procedure. Using an annotated sample to guide the evaluation, we show that the probabilities assigned to keywords used to query news databases can also be informative for selecting relevant topics and improve sample precision. We compare our results to a fine-tuned text embedding classifier and an open-weight LLM, discussing observed trade-offs, e.g. the LLM’s lowest precision. Moreover, we show that results are hazard-dependent, which speaks against considering climate events as a single category in NLP tasks.
Why Is This Green? LLM-Based Explanations of Implicit Green Practices in Social Media
Anna Glazkova | Olga Zakharova | Daria Lebedeva
Anna Glazkova | Olga Zakharova | Daria Lebedeva
Identifying green practices in social media is not merely a matter of lexical matching. Many green practices are expressed implicitly, rely on shared background knowledge, or are embedded in broader contextual narratives. In this paper, we investigate how large language models (LLMs) explain expert annotations of green waste management practices and how they rationalize classification errors made by a fine-tuned model (mBART) on a Russian social media corpus (GreenRu). We analyze explanations generated by two LLMs (T-lite and GigaChat) in two settings: (1) explaining gold expert-assigned labels and (2) interpreting erroneous model predictions. Our qualitative and micro-quantitative analysis shows that green practices are frequently inferred through contextual reasoning rather than explicit terminology. Error patterns of mBART reveal overgeneralization, associative misinterpretation (e.g., linking food sharing to waste recycling), and detection of practices where none are present. We further compare explanatory strategies of the two LLMs. T-lite tends to rely on lexical cues and surface markers that may create an impression of a practice, while GigaChat more often reconstructs broader contextual interpretations. Expert feedback highlights limitations of formal textual analysis, sensitivity to missing contextual knowledge, and difficulties in aligning model reasoning with expert conceptual boundaries. Our findings suggest that explanation-based analysis is a productive tool for diagnosing classification errors and refining annotation guidelines. More broadly, the study demonstrates that modeling implicit sustainability discourse requires contextual grounding and deeper semantic integration beyond keyword-based approaches.
Introducing a Green Leaderboard for Sustainable Risk Prediction in Streaming NLP Shared Tasks.
Alba María Mármol-Romero | Adrián Moreno Muñoz | Arturo Montejo-Raez
Alba María Mármol-Romero | Adrián Moreno Muñoz | Arturo Montejo-Raez
Current NLP shared-task evaluations predominantly rank systems by predictive performance, overlooking computational efficiency and environmental impact. This limitation is particularly critical in streaming and early risk detection scenarios, where models operate continuously, and resource consumption accumulates over time. We propose a sustainability-aware evaluation framework for streaming NLP tasks by introducing the Green Early Detection Score (GED), which integrates classification performance, detection timeliness, and carbon emissions. We also present an energy-based variant tailored to on-device early risk detection settings where energy consumption per inference is a key constraint. Applying these metrics to three editions (2023-2025) of the MentalRiskES shared task, we construct the first Green Leaderboard for early risk detection. Our results show that sustainability-aware ranking substantially reshapes system positions, highlighting efficient models that remain undervalued under performance-only evaluation.
Not Everything Is Greenwashing: Limitations of Automatic Analysis of Sustainability Reports, and a Proposal
Maria Pilar Uribe Silva | Rik van Noord | Malvina Nissim
Maria Pilar Uribe Silva | Rik van Noord | Malvina Nissim
Sustainability reports (SRs) are essential for holding companies accountable, and they are required by law. They also serve as a key communication tool through which companies shape their image and disclose non-financial information. However, the rapid growth of these reports, their lack of standardisation, and the frequent use of strategically ambiguous language make it difficult for stakeholders to evaluate whether sustainability claims are genuine or deceptive. Previous work has focused on extracting misleading climate-related content and identifying greenwashing. We argue that this is not enough, because deception does not only appear in overtly false or misleading green claims, and it often emerges through a variety of subtle linguistic strategies. We therefore propose the development of a framework based on deception theories to examine how deceptive language operates in SRs, and we outline the challenges that should be seen as an invitation for future research. Keywords: Deception, Deceptive Language, Sustainability Reports, Greenwashing