Feng Chen
Author directoryPapers on this page may belong to the following people: Feng Chen, Feng Chen, Feng Chen, Feng Chen, Feng Chen
2026
Automated Detection and Classification of Delusion-related Content in Naturalistic Audio Diaries Using Multi-Agent Language Models
Feng Chen | Justin Tauscher | Changye Li | Meliha Yetisgen | Alex Cohen | Adam Kuczynski | Angelina Tsai | Benjamin Buck | Dror Ben-Zeev | Trevor Cohen
Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026)
Feng Chen | Justin Tauscher | Changye Li | Meliha Yetisgen | Alex Cohen | Adam Kuczynski | Angelina Tsai | Benjamin Buck | Dror Ben-Zeev | Trevor Cohen
Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026)
Speech monologues recorded in naturalistic settings provide opportunities to characterize mental illness phenomenology and detect symptom exacerbation. Large language models (LLMs) offer new possibilities for automating this process, as they require annotated data primarily for evaluation rather than training. In this paper, we present a novel automated, multi-agent LLM pipeline for the fine-grained, multi-label extraction of language suggestive of delusional beliefs, associated affective responses, and behavioral responses from transcripts of naturalistic audio diaries collected from people with moderate persecutory ideation. Evaluating an ensemble of three foundation models, we demonstrate that detailed diagnostic prompt instructions successfully reduce false positives for delusional theme classification, but also constrain the interpretation of affective or behavioral responses. Furthermore, comparing multi-agent adjudication frameworks reveals a critical divergence from standard NLP benchmarks: complex conversational debate between agents diminishes accuracy on clinically ambiguous text by inducing premature consensus. Instead, majority voting establishes robust performance (Micro F1 of 0.872 and 0.779 for delusion detection and classification respectively). This work provides a validated and scalable pipeline for the automated detection and characterization of content suggesting delusional beliefs in naturalistic speech.
Uncertainty-Aware Test-Time Search for Optimization Problem Solving
Linlin Yu | Xujiang Zhao | Dong Li | Yanchi Liu | Wei Cheng | Zhengzhang Chen | Chen Zhao | Feng Chen | Haifeng Chen
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Linlin Yu | Xujiang Zhao | Dong Li | Yanchi Liu | Wei Cheng | Zhengzhang Chen | Chen Zhao | Feng Chen | Haifeng Chen
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Automatically solving optimization problems from natural language descriptions with both efficiency and reliability is highly desirable but remains challenging. Language model hallucinations and the limited availability of labeled datasets often result in misaligned formulations, code errors, and feasibility failures We propose UMCTS, an Uncertainty-aware Monte Carlo Tree Search framework that combines the language understanding capability of large language models with the reliability of well-established solvers. UMCTS structures the solution process into four stages: global instruction, assumptions, mathematical formulation, and solver code generation. It employs Monte Carlo Tree Search with semantic-equivalence pruning, prior-guided exploration, and solver-based feasibility checks. An LLM judge provides numerical reward signals, qualitative error information, and uncertainty estimates. These signals are backpropagated to guide the search and flag unreliable outputs. Across six public benchmarks, UMCTS achieves state-of-the-art solution accuracy, improves efficiency by reducing token usage.
2025
Let The Jury Decide: Fair Demonstration Selection for In-Context Learning through Incremental Greedy Evaluation
Sadaf Md Halim | Chen Zhao | Xintao Wu | Latifur Khan | Christan Grant | Fariha Ishrat Rahman | Feng Chen
Findings of the Association for Computational Linguistics: ACL 2025
Sadaf Md Halim | Chen Zhao | Xintao Wu | Latifur Khan | Christan Grant | Fariha Ishrat Rahman | Feng Chen
Findings of the Association for Computational Linguistics: ACL 2025
Large Language Models (LLMs) are powerful in-context learners, achieving strong performance with just a few high-quality demonstrations. However, fairness concerns arise in many in-context classification tasks, especially when predictions involve sensitive attributes. To address this, we propose JUDGE—a simple yet effective framework for selecting fair and representative demonstrations that improve group fairness in In-Context Learning. JUDGE constructs the demonstration set iteratively using a greedy approach, guided by a small, carefully selected jury set. Our method remains robust across varying LLM architectures and datasets, ensuring consistent fairness improvements. We evaluate JUDGE on four datasets using four LLMs, comparing it against seven baselines. Results show that JUDGE consistently improves fairness metrics without compromising accuracy.
Large Language Models with Reinforcement Learning from Human Feedback Approach for Enhancing Explainable Sexism Detection
Ali Riahi Samani | Tianhao Wang | Kangshuo Li | Feng Chen
Proceedings of the 31st International Conference on Computational Linguistics
Ali Riahi Samani | Tianhao Wang | Kangshuo Li | Feng Chen
Proceedings of the 31st International Conference on Computational Linguistics
Recent advancements in natural language processing, driven by Large Language Models (LLMs), have significantly improved text comprehension, enabling these models to handle complex tasks with greater efficiency. A key feature of LLMs is their ability to engage in contextual learning, which allows them to understand and apply instructions given in natural language to new scenarios without requiring additional training. This capability is particularly valuable in social media, where LLMs can be crucial in addressing challenges in explainable sexism detection. We hypothesize that by leveraging contextual learning capabilities, LLMs can provide clear, explainable insights into why certain content is flagged as problematic, thus enhancing transparency in the sexism detection process. To this end, we propose a Reinforcement Learning from Human Feedback (RLHF) based fine-tuning framework for sexism detection. We studied two well-known LLMs, Mistral-7B and LLaMA-3-8B, in zero-shot, supervised fine-tuning, and RLHF scenarios to conclude the superior ability of LLMs in sexism detection. The experimental results reported in this work, based on three tasks of Explainable Detection of Online Sexism (EDOS), highlight the importance of RLHF for building explainable systems in online discourse. Furthermore, we found that the LLaMA-3-8B model achieves the best results using the RLHF approach, scoring 0.8681 on Task A (binary sexism detection), 0.6829 on Task B (category classification of sexism), and 0.4722 on Task C (fine-grained sexism vectors) test sets.
2024
Uncertainty Estimation on Sequential Labeling via Uncertainty Transmission
Jianfeng He | Linlin Yu | Shuo Lei | Chang-Tien Lu | Feng Chen
Findings of the Association for Computational Linguistics: NAACL 2024
Jianfeng He | Linlin Yu | Shuo Lei | Chang-Tien Lu | Feng Chen
Findings of the Association for Computational Linguistics: NAACL 2024
Sequential labeling is a task predicting labels for each token in a sequence, such as Named Entity Recognition (NER). NER tasks aim to extract entities and predict their labels given a text, which is important in information extraction. Although previous works have shown great progress in improving NER performance, uncertainty estimation on NER (UE-NER) is still underexplored but essential. This work focuses on UE-NER, which aims to estimate uncertainty scores for the NER predictions. Previous uncertainty estimation models often overlook two unique characteristics of NER: the connection between entities (i.e., one entity embedding is learned based on the other ones) and wrong span cases in the entity extraction subtask. Therefore, we propose a Sequential Labeling Posterior Network (SLPN) to estimate uncertainty scores for the extracted entities, considering uncertainty transmitted from other tokens. Moreover, we have defined an evaluation strategy to address the specificity of wrong-span cases. Our SLPN has achieved significant improvements on three datasets, such as a 5.54-point improvement in AUPR on the MIT-Restaurant dataset. Our code is available at https://github.com/he159ok/UncSeqLabeling_SLPN.
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
Jianfeng He | Runing Yang | Linlin Yu | Changbin Li | Ruoxi Jia | Feng Chen | Ming Jin | Chang-Tien Lu
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Jianfeng He | Runing Yang | Linlin Yu | Changbin Li | Ruoxi Jia | Feng Chen | Ming Jin | Chang-Tien Lu
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Text summarization, a key natural language generation (NLG) task, is vital in various domains. However, the high cost of inaccurate summaries in risk-critical applications, particularly those involving human-in-the-loop decision-making, raises concerns about the reliability of uncertainty estimation on text summarization (UE-TS) evaluation methods. This concern stems from the dependency of uncertainty model metrics on diverse and potentially conflicting NLG metrics. To address this issue, we introduce a comprehensive UE-TS benchmark incorporating 31 NLG metrics across four dimensions. The benchmark evaluates the uncertainty estimation capabilities of two large language models and one pre-trained language model on three datasets, with human-annotation analysis incorporated where applicable. We also assess the performance of 14 common uncertainty estimation methods within this benchmark. Our findings emphasize the importance of considering multiple uncorrelated NLG metrics and diverse uncertainty estimation methods to ensure reliable and efficient evaluation of UE-TS techniques. Our code and data are available: https://github.com/he159ok/Benchmark-of-Uncertainty-Estimation-Methods-in-Text-Summarization.
2021
Boosting Cross-Lingual Transfer via Self-Learning with Uncertainty Estimation
Liyan Xu | Xuchao Zhang | Xujiang Zhao | Haifeng Chen | Feng Chen | Jinho D. Choi
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Liyan Xu | Xuchao Zhang | Xujiang Zhao | Haifeng Chen | Feng Chen | Jinho D. Choi
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Recent multilingual pre-trained language models have achieved remarkable zero-shot performance, where the model is only finetuned on one source language and directly evaluated on target languages. In this work, we propose a self-learning framework that further utilizes unlabeled data of target languages, combined with uncertainty estimation in the process to select high-quality silver labels. Three different uncertainties are adapted and analyzed specifically for the cross lingual transfer: Language Heteroscedastic/Homoscedastic Uncertainty (LEU/LOU), Evidential Uncertainty (EVI). We evaluate our framework with uncertainties on two cross-lingual tasks including Named Entity Recognition (NER) and Natural Language Inference (NLI) covering 40 languages in total, which outperforms the baselines significantly by 10 F1 for NER on average and 2.5 accuracy for NLI.
Search
Fix author
Co-authors
- Linlin Yu 3
- Jianfeng He 2
- Chang-Tien Lu 2
- Chen Zhao 2
- Xujiang Zhao 2
- Dror Ben-Zeev 1
- Benjamin Buck 1
- Haifeng Chen 1
- Haifeng Chen 1
- Zhengzhang Chen 1
- Wei Cheng 1
- Jinho D. Choi 1
- Alex Cohen 1
- Trevor Cohen 1
- Christan Grant 1
- Sadaf Md Halim 1
- Ruoxi Jia 1
- Ming Jin 1
- Latifur Khan 1
- Adam Kuczynski 1
- Shuo Lei 1
- Changbin Li 1
- Changye Li 1
- Dong Li 1
- Kangshuo Li 1
- Yanchi Liu 1
- Fariha Ishrat Rahman 1
- Ali Riahi Samani 1
- Justin Tauscher 1
- Angelina Tsai 1
- Tianhao Wang 1
- Xintao Wu 1
- Liyan Xu 1
- Runing Yang 1
- Meliha Yetisgen-Yildiz 1
- Xuchao Zhang 1