Jiaxin Zhang
Author directoryOther people with similar names: Jiaxin Zhang, Jiaxin Zhang, Jiaxin Zhang
Unverified author pages with similar names: Jiaxin Zhang
2026
From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models
Jiaxin Zhang | Wendi Cui | Zhuohang Li | Lifu Huang | Bradley A. Malin | Caiming Xiong | Chien-Sheng Wu
Findings of the Association for Computational Linguistics: ACL 2026
Jiaxin Zhang | Wendi Cui | Zhuohang Li | Lifu Huang | Bradley A. Malin | Caiming Xiong | Chien-Sheng Wu
Findings of the Association for Computational Linguistics: ACL 2026
While Large Language Models (LLMs) show remarkable capabilities, their unreliability remains a critical barrier to deployment in high-stakes domains. This survey charts a functional evolution in addressing this challenge: the evolution of uncertainty from a passive diagnostic metric to an active control signal guiding real-time model behavior. We demonstrate how uncertainty is leveraged as an active control signal across three frontiers: in advanced reasoning to optimize computation and trigger self-correction; in autonomous agents to govern metacognitive decisions about tool use and information seeking; and in reinforcement learning to mitigate reward hacking and enable self-improvement via intrinsic rewards. By grounding these advancements in emerging theoretical frameworks like Bayesian methods and Conformal Prediction, we provide a unified perspective on this transformative trend. This survey provides a comprehensive overview, critical analysis, and practical design patterns, arguing that mastering the new trend of uncertainty is essential for building the next generation of scalable, reliable, and trustworthy AI.
Don’t Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
Prafulla Kumar Choubey | Kung-Hsiang Huang | Pranav Narayanan Venkit | Jiaxin Zhang | Vaibhav Vats | Yu Li | Xiangyu Peng | Chien-Sheng Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Prafulla Kumar Choubey | Kung-Hsiang Huang | Pranav Narayanan Venkit | Jiaxin Zhang | Vaibhav Vats | Yu Li | Xiangyu Peng | Chien-Sheng Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)
Enterprise deep research often fails to produce decision-ready reports due to uneven information coverage, context explosion, and premature stopping. We propose a scalable Enterprise Deep Research (EDR) architecture to address these failures. Our system (i) decomposes requests into coverage-driven objectives via outline generation with reflection, (ii) localizes context with dependency-guided execution and explicit information sharing, and (iii) enforces evidence-based completion criteria so agents iteratively collect information until sufficiency conditions are met. We evaluate on an internal sales enablement task and the public DeepResearch Bench benchmark, where our proposed system design achieves the strongest overall performance compared with competitive deep-research baselines. The results show that dependency-controlled context and explicit evidence sufficiency criteria reduce premature stopping and improve the consistency and depth of enterprise research outputs.
2025
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation
Yu Wang | Jiaxin Zhang | Xiang Gao | Wendi Cui | Peng Li | Kamalika Das
Findings of the Association for Computational Linguistics: NAACL 2025
Yu Wang | Jiaxin Zhang | Xiang Gao | Wendi Cui | Peng Li | Kamalika Das
Findings of the Association for Computational Linguistics: NAACL 2025
In tasks such as summarization and open-book question answering (QA), Large Language Models (LLMs) frequently experience “contextual hallucination”, where they generate irrelevant or incorrect responses despite having access to accurate information in the input. This issue often stems from the models’ propensity to prioritize self-generated content over input context, leading to a disregard for pertinent details. To address this challenge, we introduce, Guided Attention Map Editing (GAME), an innovative approach that dynamically adjusts attention maps to enhance contextual relevance. During inference, GAME employs a trained classifier to identify attention maps likely to induce hallucinations and implements targeted interventions. These interventions, guided by gradient-informed “edit directions”, strategically redistribute attention weights across various heads to efficiently mitigate hallucination. Extensive evaluations on challenging summarization and open-book QA tasks demonstrate that GAME consistently and significantly reduces hallucinations across diverse open-source models, thereby improving the reliability and applicability of LLMs.
Heuristic-based Search Algorithm in Automatic Instruction-focused Prompt Optimization: A Survey
Wendi Cui | Jiaxin Zhang | Zhuohang Li | Hao Sun | Damien Lopez | Kamalika Das | Bradley A. Malin | Sricharan Kumar
Findings of the Association for Computational Linguistics: ACL 2025
Wendi Cui | Jiaxin Zhang | Zhuohang Li | Hao Sun | Damien Lopez | Kamalika Das | Bradley A. Malin | Sricharan Kumar
Findings of the Association for Computational Linguistics: ACL 2025
Recent advances in Large Language Models(LLMs) have led to remarkable achievements across a variety of Natural Language Processing(NLP) tasks, making prompt engineering increasingly central to guiding model outputs. While manual methods (e.g., “chain-of-thought,” “step-by-step” prompts) can be effective, they typically rely on intuition and do not automatically refine prompts over time. In contrast, automatic prompt optimization employing heuristic-based search algorithms can systematically explore and improve prompts with minimal human oversight. This survey proposes a comprehensive taxonomy of these methods, categorizing them by where optimization occurs, what is optimized, what criteria drive the optimization, which operators generate new prompts, and which iterative search algorithms are applied. We further highlight specialized datasets and tools that support and accelerate automated prompt refinement. We conclude by discussing key open challenges, pointing toward future opportunities for more robust and versatile LLM applications.
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
Kaijie Chen | Zihao Lin | Zhiyang Xu | Ying Shen | Yuguang Yao | Joy Rimchala | Jiaxin Zhang | Lifu Huang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Kaijie Chen | Zihao Lin | Zhiyang Xu | Ying Shen | Yuguang Yao | Joy Rimchala | Jiaxin Zhang | Lifu Huang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating “a bitten apple that has been left in the air for more than a week” necessitates understanding temporal decay and commonsense concepts. While recent T2I models have made impressive progress in producing photorealistic images, their reasoning capability remains underdeveloped and insufficiently evaluated. To bridge this gap, we introduce R2I-Bench, a comprehensive benchmark specifically designed to rigorously assess reasoning-driven T2I generation. R2I-Bench comprises 3068 meticulously curated data instances, spanning 7 core reasoning categories, including commonsense, mathematical, logical, compositional, numerical, causal, and concept mixing. To facilitate fine-grained evaluation, we design R2IScore, a QA-style metric based on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality. Extensive experiments with 16 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems.
Towards Statistical Factuality Guarantee for Large Vision-Language Models
Zhuohang Li | Chao Yan | Nicholas J Jackson | Wendi Cui | Bo Li | Jiaxin Zhang | Bradley A. Malin
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Zhuohang Li | Chao Yan | Nicholas J Jackson | Wendi Cui | Bo Li | Jiaxin Zhang | Bradley A. Malin
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Advancements in Large Vision-Language Models (LVLMs) have demonstrated impressive performance in image-conditioned text generation; however, hallucinated outputs–text that misaligns with the visual input–pose a major barrier to their use in safety-critical applications. We introduce ConfLVLM, a conformal-prediction-based framework that achieves finite-sample distribution-free statistical guarantees to the factuality of LVLM output. Taking each generated detail as a hypothesis, ConfLVLM statistically tests factuality via efficient heuristic uncertainty measures to filter out unreliable claims. We conduct extensive experiments covering three representative application domains: general scene understanding, medical radiology report generation, and document understanding. Remarkably, ConfLVLM reduces the error rate of claims generated by LLaVa-1.5 for scene descriptions from 87.8% to 10.0% by filtering out erroneous claims with a 95.3% true positive rate. Our results further show that ConfLVLM is highly flexible, and can be applied to any black-box LVLMs paired with any uncertainty measure for any image-conditioned free-form text generation task while providing a rigorous guarantee on controlling hallucination risk.
Confidence-Aware Reasoning: Optimizing Self-Guided Thinking Trajectories in Large Reasoning Models
Jiaxin Zhang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Jiaxin Zhang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Chain-of-thought enables large reasoning models (LRMs) to reason through multi-step problems but often leads to unnecessarily long or redundant reasoning traces, a phenomenon known as overthinking. This results in inflated inference costs and potential degradation in answer quality. To address these challenges, we propose Confidence-Aware Reasoning (), an inference-time framework that optimizes reasoning trajectories by selectively pruning low-utility reasoning blocks and halting early when sufficient confidence has been achieved. is theoretically grounded in Bayesian optimal experimental design, treating each reasoning block as a sequential decision whose utility is approximated by its marginal contribution to reducing final answer uncertainty. We introduce a lightweight implementation that leverages token-level confidence to dynamically modulate reasoning depth without additional supervision. Evaluations on multiple benchmarks, including AMC, AIME, GPQA-Diamond, and MATH-500 show that improves answer accuracy by up to +13.3%, while reducing average reasoning length by 40%–50%. Our findings demonstrate that information-theoretic insights can effectively control self-guided reasoning and enable LRMs to “think just enough” at test time.
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
Wendi Cui | Jiaxin Zhang | Zhuohang Li | Hao Sun | Damien Lopez | Kamalika Das | Bradley A. Malin | Sricharan Kumar
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Wendi Cui | Jiaxin Zhang | Zhuohang Li | Hao Sun | Damien Lopez | Kamalika Das | Bradley A. Malin | Sricharan Kumar
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Designing optimal prompts for Large Language Models (LLMs) is a complex and resource-intensive task, often requiring substantial human expertise. Existing approaches typically separate the optimization of prompt instructions and in-context learning examples, leading to incohesive, suboptimal results. To overcome this limitation, we propose a novel Cohesive In-Context Prompt Optimization framework that refines both prompt instructions and examples. In our formulation, coherence refers to the degree to which instructions and examples work synergistically to improve task performance—emerging as a byproduct of performance-driven optimization. However, formulating such an optimization in the discrete and high-dimensional space of natural language poses significant challenges in both convergence and computational efficiency. To address these issues, we introduce SEE, a scalable and efficient prompt optimization framework that adopts metaheuristic optimization principles and strategically balances exploration and exploitation to enhance optimization performance and achieve efficient convergence. SEE features a quad-phased design that alternates between global traversal (exploration) and local optimization (exploitation) and adaptively chooses LLM operators during the optimization process. We have conducted a comprehensive evaluation across 35 benchmark tasks, and SEE significantly outperforms state-of-the-art baseline methods by a large margin, achieving an average performance gain of 13.94 while reducing computational costs by 58.67%.
Search
Fix author
Co-authors
- Wendi Cui 5
- Zhuohang Li 4
- Bradley A. Malin 4
- Kamalika Das 3
- Lifu Huang 2
- Sricharan Kumar 2
- Damien Lopez 2
- Hao Sun 2
- Chien-Sheng Wu 2
- Kaijie Chen 1
- Prafulla Kumar Choubey 1
- Xiang Gao 1
- Kung-Hsiang Huang 1
- Nicholas J Jackson 1
- Bo Li 1
- Peng Li 1
- Yu Li 1
- Zihao Lin 1
- Xiangyu Peng 1
- Joy Rimchala 1
- Ying Shen 1
- Vaibhav Vats 1
- Pranav Narayanan Venkit 1
- Yu Wang 1
- Caiming Xiong 1
- Zhiyang Xu 1
- Chao Yan 1
- Yuguang Yao 1