Xuanming Zhang
Author directoryOther people with similar names: Xuanming Zhang
Unverified author pages with similar names: Xuanming Zhang
2026
Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data
Xuanming Zhang | Shwan Ashrafi | Aziza Mirsaidova | Amir H. Rezaeian | Miguel Ballesteros | Lydia Chilton | Zhou Yu | Dan Roth
Findings of the Association for Computational Linguistics: ACL 2026
Xuanming Zhang | Shwan Ashrafi | Aziza Mirsaidova | Amir H. Rezaeian | Miguel Ballesteros | Lydia Chilton | Zhou Yu | Dan Roth
Findings of the Association for Computational Linguistics: ACL 2026
We study the reasoning behavior of large language models (LLMs) under limited computation budgets. In such settings, producing useful partial solutions quickly is often more practical than exhaustive reasoning, which incurs high inference costs. Many real-world tasks, such as trip planning, require models to deliver the best possible output within a fixed reasoning budget. We introduce an anytime reasoning framework and the Anytime Index, a metric that quantifies how effectively solution quality improves as reasoning tokens increase. To further enhance efficiency, we propose an inference-time self-improvement method using LLM-synthesized preference data, where models learn from their own reasoning comparisons to produce better intermediate solutions. Experiments on NaturalPlan (Trip), AIME, and GPQA datasets show consistent gains across Grok-3, GPT-oss, GPT-4.1/4o, and LLaMA models, improving both reasoning quality and efficiency under budget constraints.
2025
Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants’ Question-Answering in Asynchronous Learning Environments
Li Siyan | Zhen Xu | Vethavikashini Chithrra Raghuram | Xuanming Zhang | Renzhe Yu | Zhou Yu
Findings of the Association for Computational Linguistics: EMNLP 2025
Li Siyan | Zhen Xu | Vethavikashini Chithrra Raghuram | Xuanming Zhang | Renzhe Yu | Zhou Yu
Findings of the Association for Computational Linguistics: EMNLP 2025
Virtual Teaching Assistants (VTAs) can reduce the workload of teaching teams in Asynchronous Learning Environments (ALEs) where timely, personalized support is often limited. As VTA systems grow more capable, rigorous and pedagogically sound evaluation becomes essential. Existing assessments often rely on surface-level metrics and lack sufficient grounding in educational theory, making it difficult to meaningfully compare the pedagogical effectiveness of VTA systems. To bridge this gap, we propose a pedagogically-oriented evaluation framework that is rooted in learning sciences and tailored to asynchronous forum discussions, a common VTA deployment context in ALE. We construct classifiers using expert annotations of VTA responses on a diverse set of forum posts. We evaluate the effectiveness of our classifiers, identifying approaches that improve accuracy as well as challenges that hinder generalization. Our work establishes a foundation for theory-driven evaluation of VTA systems, paving the way for more pedagogically effective AI in education.