Xuan Luo
Author directoryOther people with similar names: Xuan Luo
Unverified author pages with similar names: Xuan Luo
2026
A Simple and Efficient Learning-Style Prompting for LLM Jailbreaking
Xuan Luo | Yue Wang | Zefeng He | Geng Tu | Jing Li | Ruifeng Xu
Findings of the Association for Computational Linguistics: EACL 2026
Xuan Luo | Yue Wang | Zefeng He | Geng Tu | Jing Li | Ruifeng Xu
Findings of the Association for Computational Linguistics: EACL 2026
This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses.The learning-style queries are constructed by a novel reframing paradigm: HILL (Hiding Intention by Learning from LLMs). The deterministic, model-agnostic reframing framework is composed of 4 conceptual components: 1) key concept, 2) exploratory transformation, 3) detail-oriented inquiry, and optionally 4) hypotheticality.Further, new metrics are introduced to thoroughly evaluate the efficiency and harmfulness of jailbreak methods.Experiments on the AdvBench dataset across a wide range of models demonstrate HILL’s strong generalizability. It achieves top attack success rates on the majority of models and across malicious categories while maintaining high efficiency with concise prompts. On the other hand, results of various defense methods show the robustness of HILL, with most defenses having mediocre effects or even increasing the attack success rates. In addition, the assessment of defenses on the constructed safe prompts reveals inherent limitations of LLMs’ safety mechanisms and flaws in the defense methods. This work exposes significant vulnerabilities of safety measures against learning-style elicitation, highlighting a critical challenge of fulfilling both helpfulness and safety alignments.
AEQ-Bench: Measuring Empathy of Omni-Modal Large Models
Xuan Luo | Lewei Yao | Lanqing Hong | Kai Chen | Dehua Tao | Daxin Tan | Yukun Deng | Ruifeng Xu | Jing Li
Findings of the Association for Computational Linguistics: ACL 2026
Xuan Luo | Lewei Yao | Lanqing Hong | Kai Chen | Dehua Tao | Daxin Tan | Yukun Deng | Ruifeng Xu | Jing Li
Findings of the Association for Computational Linguistics: ACL 2026
While the automatic evaluation of omni-modal large models (OLMs) is essential, assessing empathy remains a significant challenge due to its inherent affectivity. To investigate this challenge, we introduce AEQ-Bench (Audio Empathy Quotient Benchmark), a novel benchmark to systematically assess two core empathetic capabilities of OLMs: (i) generating empathetic responses by comprehending affective cues from multi-modal inputs (audio + text), and (ii) judging the empathy of audio responses without relying on text transcription. Compared to existing benchmarks, AEQ-Bench incorporates two novel settings that vary in context specificity and speech tone. Comprehensive assessment across linguistic and paralinguistic metrics reveals that (1) OLMs trained with audio output capabilities generally outperformed models with text-only outputs, and (2) while OLMs align with human judgments for coarse-grained quality assessment, they remain unreliable for evaluating fine-grained paralinguistic expressiveness.
2025
Large Language Models as Reader for Bias Detection
Xuan Luo | Jing Li | Zhong Wenzhong | Geng Tu | Ruifeng Xu
Findings of the Association for Computational Linguistics: EMNLP 2025
Xuan Luo | Jing Li | Zhong Wenzhong | Geng Tu | Ruifeng Xu
Findings of the Association for Computational Linguistics: EMNLP 2025
Detecting bias in media content is crucial for maintaining information integrity and promoting inclusivity. Traditional methods analyze text from the writer’s perspective, which analyzes textual features directly from the writer’s intent, leaving the reader’s perspective underexplored. This paper investigates whether Large Language Models (LLMs) can be leveraged as readers for bias detection by generating reader-perspective comments. Experiments are conducted on the BASIL (news bias) and BeyondGender (gender bias) datasets with LLMs Gemma-7B, Phi-3-3.8B, Llama3.1-8B, Llama3.1-70B, and GPT4. The results demonstrate the effectiveness of reader-perspective comments for open-source LLMs, achieving performance comparable to GPT4’s. The findings highlight the significance of emotion-related comments, which are generally more beneficial than value-related ones in bias detection. In addition, experiments on Llamas show that comment selection ensures consistent performance regardless of model sizes and comment combinations. This study is particularly beneficial for small-size open-source LLMs.