Haotian Zhang
Papers on this page may belong to the following people: Haotian Zhang, Haotian Zhang, Haotian Zhang
2026
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue
Jonathan Ivey | Shivani Kumar | Jiayu Liu | Hua Shen | Sushrita Rakshit | Rohan Raju | Haotian Zhang | Aparna Ananthasubramaniam | Junghwan Kim | Bowen Yi | Dustin Wright | Abraham Israeli | Anders Giovanni Møller | Lechen Zhang | David Jurgens
Findings of the Association for Computational Linguistics: ACL 2026
Jonathan Ivey | Shivani Kumar | Jiayu Liu | Hua Shen | Sushrita Rakshit | Rohan Raju | Haotian Zhang | Aparna Ananthasubramaniam | Junghwan Kim | Bowen Yi | Dustin Wright | Abraham Israeli | Anders Giovanni Møller | Lechen Zhang | David Jurgens
Findings of the Association for Computational Linguistics: ACL 2026
Building datasets for dialogue tasks is expensive and time-consuming, requiring recruitment, training, and data collection from study participants. In response, much recent work has sought to use large language models (LLMs) to simulate both human-human and human-LLM interactions, as they have been shown to generate convincingly human-like text in many settings. However, how well do LLM-based simulations reflect real human dialogue? In this work, we answer this question by generating a large-scale dataset of 100,000 paired LLM-LLM and human-LLM dialogues from the WildChat dataset and quantifying how well the LLM simulations align with their human counterparts. Overall, we find relatively low alignment between simulations and human interactions, with systematic differences in multiple textual properties, including style and conversational dynamics. Further, we find that models perform similarly in simulating English, Chinese, and Russian dialogues. Our results also suggest that LLMs only simulate a narrow range of the overall distribution of human dialogue, as they perform better on the subset of humans who write similarly to the LLM’s own style.
2025
Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?
Qingyuan Liang | Zhao Zhang | Zeyu Sun | Zheng Lin | Qi Luo | Yueyi Xiao | Yizhou Chen | Yuqun Zhang | Haotian Zhang | Lu Zhang | Bin Chen | Yingfei Xiong
Findings of the Association for Computational Linguistics: ACL 2025
Qingyuan Liang | Zhao Zhang | Zeyu Sun | Zheng Lin | Qi Luo | Yueyi Xiao | Yizhou Chen | Yuqun Zhang | Haotian Zhang | Lu Zhang | Bin Chen | Yingfei Xiong
Findings of the Association for Computational Linguistics: ACL 2025
Grammar serves as a cornerstone in programming languages and software engineering, providing frameworks to define the syntactic space and program structure. Existing research demonstrates the effectiveness of grammar-based code representations in small-scale models, showing their ability to reduce syntax errors and enhance performance. However, as language models scale to the billion level or beyond, syntax-level errors become rare, making it unclear whether grammar information still provides performance benefits. To explore this, we develop a series of billion-scale GrammarCoder models, incorporating grammar rules in the code generation process. Experiments on HumanEval (+) and MBPP (+) demonstrate a notable improvement in code generation accuracy. Further analysis shows that grammar-based representations enhance LLMs’ ability to discern subtle code differences, reducing semantic errors caused by minor variations. These findings suggest that grammar-based code representations remain valuable even in billion-scale models, not only by maintaining syntax correctness but also by improving semantic differentiation.
OASIS: Order-Augmented Strategy for Improved Code Search
Zuchen Gao | Zizheng Zhan | Xianming Li | Erxin Yu | Haotian Zhang | Bin Chen | Yuqun Zhang | Jing Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Zuchen Gao | Zizheng Zhan | Xianming Li | Erxin Yu | Haotian Zhang | Bin Chen | Yuqun Zhang | Jing Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Code embeddings capture the semantic representations of code and are crucial for various code-related large language model (LLM) applications, such as code search. Previous training primarily relies on optimizing the InfoNCE loss by comparing positive natural language (NL)-code pairs with in-batch negatives. However, due to the sparse nature of code contexts, training solely by comparing the major differences between positive and negative pairs may fail to capture deeper semantic nuances. To address this issue, we propose a novel order-augmented strategy for improved code search (OASIS). It leverages order-based similarity labels to train models to capture subtle differences in similarity among negative pairs. Extensive benchmark evaluations demonstrate that our OASIS model significantly outperforms previous state-of-the-art models focusing solely on major positive-negative differences. It underscores the value of exploiting subtle differences among negative pairs with order labels for effective code embedding training.
Towards Generating Controllable and Solvable Geometry Problem by Leveraging Symbolic Deduction Engine
Zhuoxuan Jiang | Tianyang Zhang | Peiyan Peng | Jing Chen | Yinong Xun | Haotian Zhang | Lichi Li | Yong Li | Shaohua Zhang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Zhuoxuan Jiang | Tianyang Zhang | Peiyan Peng | Jing Chen | Yinong Xun | Haotian Zhang | Lichi Li | Yong Li | Shaohua Zhang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Generating high-quality geometry problems is both an important and challenging task in education. Compared to math word problems, geometry problems further emphasize multi-modal formats and the translation between informal and formal languages. In this paper, we introduce a novel task for geometry problem generation and propose a new pipeline method: the Symbolic Deduction Engine-based Geometry Problem Generation framework (SDE-GPG). The framework leverages a symbolic deduction engine and contains four main steps: (1) searching a predefined mapping table from knowledge points to extended definitions, (2) sampling extended definitions and performing symbolic deduction, (3) filtering out unqualified problems, and (4) generating textual problems and diagrams. Specifically, our method supports to avoid inherent biases in translating natural language into formal language by designing the mapping table, and guarantees to control the generated problems in terms of knowledge points and difficulties by an elaborate checking function. With obtained formal problems, they are translated to natural language and the accompanying diagrams are automatically drew by rule-based methods. We conduct experiments using real-world combinations of knowledge points from two public datasets. The results demonstrate that the SDE-GPG can effectively generate readable, solvable and controllable geometry problems.
2020
Recurrent Inference in Text Editing
Ning Shi | Ziheng Zeng | Haotian Zhang | Yichen Gong
Findings of the Association for Computational Linguistics: EMNLP 2020
Ning Shi | Ziheng Zeng | Haotian Zhang | Yichen Gong
Findings of the Association for Computational Linguistics: EMNLP 2020
In neural text editing, prevalent sequence-to-sequence based approaches directly map the unedited text either to the edited text or the editing operations, in which the performance is degraded by the limited source text encoding and long, varying decoding steps. To address this problem, we propose a new inference method, Recurrence, that iteratively performs editing actions, significantly narrowing the problem space. In each iteration, encoding the partially edited text, Recurrence decodes the latent representation, generates an action of short, fixed-length, and applies the action to complete a single edit. For a comprehensive comparison, we introduce three types of text editing tasks: Arithmetic Operators Restoration (AOR), Arithmetic Equation Simplification (AES), Arithmetic Equation Correction (AEC). Extensive experiments on these tasks with varying difficulties demonstrate that Recurrence achieves improvements over conventional inference methods.
2019
Applying BERT to Document Retrieval with Birch
Zeynep Akkalyoncu Yilmaz | Shengjin Wang | Wei Yang | Haotian Zhang | Jimmy Lin
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations
Zeynep Akkalyoncu Yilmaz | Shengjin Wang | Wei Yang | Haotian Zhang | Jimmy Lin
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations
We present Birch, a system that applies BERT to document retrieval via integration with the open-source Anserini information retrieval toolkit to demonstrate end-to-end search over large document collections. Birch implements simple ranking models that achieve state-of-the-art effectiveness on standard TREC newswire and social media test collections. This demonstration focuses on technical challenges in the integration of NLP and IR capabilities, along with the design rationale behind our approach to tightly-coupled integration between Python (to support neural networks) and the Java Virtual Machine (to support document retrieval using the open-source Lucene search library). We demonstrate integration of Birch with an existing search interface as well as interactive notebooks that highlight its capabilities in an easy-to-understand manner.
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval
Zeynep Akkalyoncu Yilmaz | Wei Yang | Haotian Zhang | Jimmy Lin
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
Zeynep Akkalyoncu Yilmaz | Wei Yang | Haotian Zhang | Jimmy Lin
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
This paper applies BERT to ad hoc document retrieval on news articles, which requires addressing two challenges: relevance judgments in existing test collections are typically provided only at the document level, and documents often exceed the length that BERT was designed to handle. Our solution is to aggregate sentence-level evidence to rank documents. Furthermore, we are able to leverage passage-level relevance judgments fortuitously available in other domains to fine-tune BERT models that are able to capture cross-domain notions of relevance, and can be directly used for ranking news articles. Our simple neural ranking models achieve state-of-the-art effectiveness on three standard test collections.
2015
Lexical Comparison Between Wikipedia and Twitter Corpora by Using Word Embeddings
Luchen Tan | Haotian Zhang | Charles Clarke | Mark Smucker
Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)
Luchen Tan | Haotian Zhang | Charles Clarke | Mark Smucker
Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)
Search
Fix author
Co-authors
- Zeynep Akkalyoncu Yilmaz 2
- Bin Chen 2
- Jimmy Lin 2
- Wei Yang 2
- Yuqun Zhang 2
- Aparna Ananthasubramaniam 1
- Jing Chen 1
- Yizhou Chen 1
- Charles L. A. Clarke 1
- Zuchen Gao 1
- Yichen Gong 1
- Abraham Israeli 1
- Jonathan Ivey 1
- Zhuoxuan Jiang 1
- David Jurgens 1
- Junghwan Kim 1
- Shivani Kumar 1
- Jing Li 1
- Lichi Li 1
- Xianming Li 1
- Yong Li 1
- Qingyuan Liang 1
- Zheng Lin 1
- Jiayu Liu 1
- Qi Luo 1
- Anders Giovanni Møller 1
- Peiyan Peng 1
- Rohan Raju 1
- Sushrita Rakshit 1
- Hua Shen 1
- Ning Shi 1
- Mark Smucker 1
- Zeyu Sun 1
- Luchen Tan 1
- Shengjin Wang 1
- Dustin Wright 1
- Yueyi Xiao 1
- Yingfei Xiong 1
- Yinong Xun 1
- Bowen Yi 1
- Erxin Yu 1
- Ziheng Zeng 1
- Zizheng Zhan 1
- Lechen Zhang 1
- Lu Zhang 1
- Shaohua Zhang 1
- Tianyang Zhang 1
- Zhao Zhang 1