Haiyan Zhao
Author directoryOther people with similar names: Haiyan Zhao
Unverified author pages with similar names: Haiyan Zhao
2026
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
Haiyan Zhao | Xuansheng Wu | Fan Yang | Bo Shen | Ninghao Liu | Mengnan Du
Findings of the Association for Computational Linguistics: EACL 2026
Haiyan Zhao | Xuansheng Wu | Fan Yang | Bo Shen | Ninghao Liu | Mengnan Du
Findings of the Association for Computational Linguistics: EACL 2026
Linear concept vectors effectively steer LLMs, but existing methods suffer from noisy features in diverse datasets that undermine steering robustness. We propose Sparse Autoencoder-Denoised Concept Vectors (SDCV), which selectively keep the most discriminative SAE latents while reconstructing hidden representations. Our key insight is that concept-relevant signals can be explicitly separated from dataset noise by scaling up activations of top-k latents that best differentiate positive and negative samples. Applied to linear probing and difference-in-mean, SDCV consistently improves steering success rates by 4-16% across six challenging concepts, while maintaining topic relevance.
2025
ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
Shuang Liu | Zelong Li | Ruoyun Ma | Haiyan Zhao | Mengnan Du
Proceedings of the Natural Legal Language Processing Workshop 2025
Shuang Liu | Zelong Li | Ruoyun Ma | Haiyan Zhao | Mengnan Du
Proceedings of the Natural Legal Language Processing Workshop 2025
The potential of large language models (LLMs) in contract legal risk analysis remains underexplored. In response, this paper introduces ContractEval, the first benchmark to thoroughly evaluate whether open-source LLMs could match proprietary LLMs in identifying clause-level legal risks in commercial contracts. Using the Contract Understanding Atticus Dataset (CUAD), we assess 4 proprietary and 15 open-source LLMs. Our results highlight five key findings: (1) Proprietary models outperform open-source models in both correctness and output effectiveness. (2) Larger open-source models generally perform better, though the improvement slows down as models get bigger. (3) Reasoning (“thinking”) mode improves output effectiveness but reduces correctness, likely due to over-complicating simpler tasks. (4) Open-source models generate “no related clause” responses more frequently even when relevant clauses are present. (5) Model quantization speed up inference but at the cost of performance drop, showing the tradeoff between efficiency and accuracy. These findings suggest that while most LLMs perform at a level comparable to junior legal assistants, open-source models require targeted fine-tuning to ensure correctness and effectiveness in high-stakes legal settings. ContractEval offers a solid benchmark to guide future development of legal-domain LLMs.
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
Dong Shu | Haiyan Zhao | Jingyu Hu | Weiru Liu | Ali Payani | Lu Cheng | Mengnan Du
Findings of the Association for Computational Linguistics: EMNLP 2025
Dong Shu | Haiyan Zhao | Jingyu Hu | Weiru Liu | Ali Payani | Lu Cheng | Mengnan Du
Findings of the Association for Computational Linguistics: EMNLP 2025
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information. However, the critical challenge of alignment between visual and textual representations is not fully understood. This survey presents a comprehensive examination of alignment and misalignment in LVLMs through an explainability lens. We first examine the fundamentals of alignment, exploring its representational and behavioral aspects, training methodologies, and theoretical foundations. We then analyze misalignment phenomena across three semantic levels: object, attribute, and relational misalignment. Our investigation reveals that misalignment emerges from challenges at multiple levels: the data level, the model level, and the inference level. We provide a comprehensive review of existing mitigation strategies, categorizing them into parameter-frozen and parameter-tuning approaches. Finally, we outline promising future research directions, emphasizing the need for standardized evaluation protocols and in-depth explainability studies.
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
Dong Shu | Xuansheng Wu | Haiyan Zhao | Daking Rai | Ziyu Yao | Ninghao Liu | Mengnan Du
Findings of the Association for Computational Linguistics: EMNLP 2025
Dong Shu | Xuansheng Wu | Haiyan Zhao | Daking Rai | Ziyu Yao | Ninghao Liu | Mengnan Du
Findings of the Association for Computational Linguistics: EMNLP 2025
Large Language Models (LLMs) have transformed natural language processing, yet their internal mechanisms remain largely opaque. Recently, mechanistic interpretability has attracted significant attention from the research community as a means to understand the inner workings of LLMs. Among various mechanistic interpretability approaches, Sparse Autoencoders (SAEs) have emerged as a promising method due to their ability to disentangle the complex, superimposed features within LLMs into more interpretable components. This paper presents a comprehensive survey of SAEs for interpreting and understanding the internal workings of LLMs. Our major contributions include: (1) exploring the technical framework of SAEs, covering basic architecture, design improvements, and effective training strategies; (2) examining different approaches to explaining SAE features, categorized into input-based and output-based explanation methods; (3) discussing evaluation methods for assessing SAE performance, covering both structural and functional metrics; and (4) investigating real-world applications of SAEs in understanding and manipulating LLM behaviors.
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
Dong Shu | Xuansheng Wu | Haiyan Zhao | Mengnan Du | Ninghao Liu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Dong Shu | Xuansheng Wu | Haiyan Zhao | Mengnan Du | Ninghao Liu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Sparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs). However, conventional approaches to analyzing SAEs typically rely solely on input-side activations, without considering the influence between each latent feature and the model’s output. This work is built on two key hypotheses: (1) activated latents do not contribute equally to the construction of the model’s output, and (2) only latents with high influence are effective for model steering. To validate these hypotheses, we propose Gradient Sparse Autoencoder (GradSAE), a simple yet effective method that identifies the most influential latents by incorporating output-side gradient information.