Jingyu Zhang
Author directoryOther people with similar names: Jingyu Zhang
Unverified author pages with similar names: Jingyu Zhang
2026
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
Hoang Phan | Xianjun Yang | Yuanshun Yao | Jingyu Zhang | Shengjie Bi | Xiaocheng Tang | Madian Khabsa | Lijuan Liu | Deren Lei
Findings of the Association for Computational Linguistics: ACL 2026
Hoang Phan | Xianjun Yang | Yuanshun Yao | Jingyu Zhang | Shengjie Bi | Xiaocheng Tang | Madian Khabsa | Lijuan Liu | Deren Lei
Findings of the Association for Computational Linguistics: ACL 2026
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models. However, the RLVR recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies. We empirically confirm this concern, observing that open-source reasoning models suffer performance degradation on core capabilities such as perception and faithfulness. While imposing regularization terms like KL divergence can help prevent deviation from the base model, these terms are calculated on the current task, thus they do not guarantee broader knowledge. Meanwhile, commonly used experience replay across heterogeneous domains makes it nontrivial to decide how much training focus each objective should receive. To address this, we propose RECAP—a replay strategy with dynamic objective reweighting for general knowledge preservation. Our reweighting mechanism adapts in an online manner using short-horizon signals of convergence and instability, shifting the post-training focus away from saturated objectives and toward underperforming or volatile ones. Our method is end-to-end and readily applicable to existing RLVR pipelines without training additional models or heavy tuning. Extensive experiments on benchmarks based on Qwen2.5-VL-3B and Qwen2.5-VL-7B demonstrate the effectiveness of our method, which not only preserves general capabilities but also improves reasoning by enabling more flexible trade-offs among in-task rewards.
2025
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
Jingyu Zhang | Marc Marone | Tianjian Li | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Jingyu Zhang | Marc Marone | Tianjian Li | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
To trust the fluent generations of large language models (LLMs), humans must be able to verify their correctness against trusted, external sources. Recent efforts, such as providing citations via retrieved documents or post-hoc provenance, enhance verifiability but provide no guarantees on their correctness. To address these limitations, we tackle the verifiability goal with a different philosophy: _trivializing the verification process by developing models that quote verbatim statements from trusted sources in their pre-training data._We propose Quote-Tuning, which demonstrates the feasibility of aligning models to quote. The core of Quote-Tuning is a fast membership inference function that efficiently verifies text against trusted corpora. We leverage this tool to design a reward function to quantify quotes in model responses, and curate datasets for preference learning. Experiments show that Quote-Tuning significantly increases verbatim quotes from high-quality documents by up to 130% relative to base models while maintaining response quality. Quote-Tuning is applicable in different tasks, generalizes to out-of-domain data and diverse model families, and provides additional benefits to truthfulness. Our method not only serves as a hassle-free method to increase quoting but also opens up avenues for improving LLM trustworthiness through better verifiability.
TurkingBench: A Challenge Benchmark for Web Agents
Kevin Xu | Yeganeh Kordi | Tanay Nayak | Adi Asija | Yizhong Wang | Kate Sanders | Adam Byerly | Jingyu Zhang | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Kevin Xu | Yeganeh Kordi | Tanay Nayak | Adi Asija | Yizhong Wang | Kate Sanders | Adam Byerly | Jingyu Zhang | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-tasks within web-based environments.Building on this idea, we present TurkingBench, a benchmark consisting of tasks presented as web pages with textual instructions and multi-modal contexts. Unlike previous approaches that rely on artificially synthesized web pages, our benchmark uses natural HTML pages originally designed for crowdsourcing workers to perform various annotation tasks. Each task’s HTML instructions are instantiated with different values derived from crowdsourcing tasks, creating diverse instances. This benchmark includes 32.2K instances spread across 158 tasks.To support the evaluation of TurkingBench, we have developed a framework that links chatbot responses to actions on web pages (e.g., modifying a text box, selecting a radio button). We assess the performance of cutting-edge private and open-source models, including language-only and vision-language models (such as GPT4 and InternVL), on this benchmark. Our results show that while these models outperform random chance, there is still significant room for improvement. We hope that this benchmark will drive progress in the evaluation and development of web-based agents.
Jailbreak Distillation: Renewable Safety Benchmarking
Jingyu Zhang | Ahmed Elgohary | Xiawei Wang | A S M Iftekhar | Ahmed Magooda | Benjamin Van Durme | Daniel Khashabi | Kyle Jackson
Findings of the Association for Computational Linguistics: EMNLP 2025
Jingyu Zhang | Ahmed Elgohary | Xiawei Wang | A S M Iftekhar | Ahmed Magooda | Benjamin Van Durme | Daniel Khashabi | Kyle Jackson
Findings of the Association for Computational Linguistics: EMNLP 2025
Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a novel benchmark construction framework that “distills” jailbreak attacks into high-quality and easily-updatable safety benchmarks. JBDistill utilizes a small set of development models and existing jailbreak attack algorithms to create a candidate prompt pool, then employs prompt selection algorithms to identify an effective subset of prompts as safety benchmarks. JBDistill addresses challenges in existing safety evaluation: the use of consistent evaluation prompts across models ensures fair comparisons and reproducibility. It requires minimal human effort to rerun the JBDistill pipeline and produce updated benchmarks, alleviating concerns on saturation and contamination. Extensive experiments demonstrate our benchmarks generalize robustly to 13 diverse evaluation models held out from benchmark construction, including proprietary, specialized, and newer-generation LLMs, significantly outperforming existing safety benchmarks in effectiveness while maintaining high separability and diversity. Our framework thus provides an effective, sustainable, and adaptable solution for streamlining safety evaluation.
Core: Robust Factual Precision with Informative Sub-Claim Identification
Zhengping Jiang | Jingyu Zhang | Nathaniel Weir | Seth Ebner | Miriam Wanner | Kate Sanders | Daniel Khashabi | Anqi Liu | Benjamin Van Durme
Findings of the Association for Computational Linguistics: ACL 2025
Zhengping Jiang | Jingyu Zhang | Nathaniel Weir | Seth Ebner | Miriam Wanner | Kate Sanders | Daniel Khashabi | Anqi Liu | Benjamin Van Durme
Findings of the Association for Computational Linguistics: ACL 2025
Hallucinations pose a challenge to the application of large language models (LLMs) thereby motivating the development of metrics to evaluate factual precision. We observe that popular metrics using the Decompose-Then-Verify framework, such as FActScore, can be manipulated by adding obvious or repetitive subclaims to artificially inflate scores. This observation motivates our new customizable plug-and-play subclaim selection component called Core, which filters down individual subclaims according to their uniqueness and informativeness. We show that many popular factual precision metrics augmented by Core are substantially more robust on a wide range of knowledge domains. We release an evaluation framework supporting easy and modular use of Core and various decomposition strategies, which we recommend adoption by the community. We also release an expansion of the FActScore biography dataset to facilitate further studies of decomposition-based factual precision evaluation.
Certified Mitigation of Worst-Case LLM Copyright Infringement
Jingyu Zhang | Jiacan Yu | Marc Marone | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jingyu Zhang | Jiacan Yu | Marc Marone | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
The exposure of large language models (LLMs) to copyrighted material during pre-training raises concerns about unintentional copyright infringement post deployment. This has driven the development of “copyright takedown” methods—post-training approaches aimed at preventing models from generating content substantially similar to copyrighted ones. While current mitigation approaches are somewhat effective for average-case risks, we demonstrate that they overlook worst-case copyright risks exhibited by the existence of long, verbatim quotes from copyrighted sources. We propose BloomScrub, a remarkably simple yet highly effective inference-time approach that provides certified copyright takedown. Our method repeatedly interleaves quote detection with rewriting techniques to transform potentially infringing segments. By leveraging efficient data sketches (Bloom filters), our approach enables scalable copyright screening—even for large-scale real-world corpora. When quotes beyond a length threshold cannot be removed, the system can abstain from responding, offering certified risk reduction. Experimental results show that BloomScrub reduces infringement risk, preserves utility, and accommodates different levels of enforcement stringency with adaptive abstention. Our results suggest that lightweight, inference-time methods can be surprisingly effective for copyright prevention.
RATIONALYST: Pre-training Process-Supervision for Improving Reasoning
Dongwei Jiang | Guoxuan Wang | Yining Lu | Andrew Wang | Jingyu Zhang | Chuyu Liu | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Dongwei Jiang | Guoxuan Wang | Yining Lu | Andrew Wang | Jingyu Zhang | Chuyu Liu | Benjamin Van Durme | Daniel Khashabi
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
The reasoning steps generated by LLMs might be incomplete, as they mimic logical leaps common in everyday communication found in their pre-training data: underlying rationales are frequently left implicit (unstated). To address this challenge, we introduce RATIONALYST, a model for process-supervision of reasoning based on pre-training on a vast collection of rationale annotations extracted from unlabeled data. We extract 79k rationales from web-scale unlabelled dataset (the Pile) and a combination of reasoning datasets with minimal human intervention. This web-scale pre-training for reasoning allows RATIONALYST to consistently generalize across diverse reasoning tasks, including mathematical, commonsense, scientific, and logical reasoning. Fine-tuned from LLaMa-3-8B, RATIONALYST improves the accuracy of reasoning by an average of 3.9% on 7 representative reasoning benchmarks. It also demonstrates superior performance compared to significantly larger verifiers like GPT-4 and similarly sized models fine-tuned on matching training sets.
Search
Fix author
Co-authors
- Daniel Khashabi 6
- Benjamin Van Durme 6
- Marc Marone 2
- Kate Sanders 2
- Adi Asija 1
- Shengjie Bi 1
- Adam Byerly 1
- Seth Ebner 1
- Ahmed Elgohary 1
- A S M Iftekhar 1
- Kyle Jackson 1
- Dongwei Jiang 1
- Zheng Ping Jiang 1
- Madian Khabsa 1
- Yeganeh Kordi 1
- Deren Lei 1
- Tianjian Li 1
- Anqi Liu 1
- Chuyu Liu 1
- Lijuan Liu 1
- Yining Lu 1
- Ahmed Magooda 1
- Tanay Nayak 1
- Hoang Phan 1
- Xiaocheng Tang 1
- Andrew Wang 1
- Guoxuan Wang 1
- Xiawei Wang 1
- Yizhong Wang 1
- Miriam Wanner 1
- Nathaniel Weir 1
- Kevin Xu 1
- Xianjun Yang 1
- Yuanshun Yao 1
- Jiacan Yu 1