Yun Zhu
Author directoryOther people with similar names: Yun Zhu
Unverified author pages with similar names: Yun Zhu
2026
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
Yu Li | Xiaoran Shang | Qizhi Pei | Yun Zhu | Xin Gao | Honglin Lin | Zhanping Zhong | Zhuoshi Pan | Zheng Liu | Xiaoyang Wang | Conghui He | Dahua Lin | Feng Zhao | Lijun Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yu Li | Xiaoran Shang | Qizhi Pei | Yun Zhu | Xin Gao | Honglin Lin | Zhanping Zhong | Zhuoshi Pan | Zheng Liu | Xiaoyang Wang | Conghui He | Dahua Lin | Feng Zhao | Lijun Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifacts, overlooking the systemic connections that underlie their evolution. To disentangle these complex relationships, we introduce the concept of data lineage to the LLM ecosystem and propose an automated multi-agent framework to reconstruct the evolutionary graph of dataset development. Through large-scale lineage analysis, we characterize domain-specific structural patterns, such as vertical refinement in Math-oriented datasets and horizontal aggregation in General-domain corpora. Moreover, we uncover pervasive systemic issues, including structural redundancy induced by implicit dataset intersections and the propagation of benchmark contamination along lineage paths. To demonstrate the practical value of lineage analysis for data construction, we leverage the reconstructed lineage graph to create a lineage-aware diversity-oriented dataset. By anchoring instruction sampling at upstream leaf sources, this approach mitigates downstream homogenization and hidden redundancy, yielding a more diverse post-training corpus. We further highlight lineage-centric analysis as an efficient and robust topological alternative to sample-level dataset comparison for large-scale data ecosystems. By grounding data construction in explicit lineage structures, our work advances post-training data curation toward a more systematic and controllable paradigm.
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
Zheng Liu | Honglin Lin | Xiaoyang Wang | Xin Gao | Yu Li | Mengzhang Cai | Yun Zhu | Zhanping Zhong | Qizhi Pei | Zhuoshi Pan | Xiaoran Shang | Conghui He | Bin Cui | Wentao Zhang | Lijun Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Zheng Liu | Honglin Lin | Xiaoyang Wang | Xin Gao | Yu Li | Mengzhang Cai | Yun Zhu | Zhanping Zhong | Qizhi Pei | Zhuoshi Pan | Xiaoran Shang | Conghui He | Bin Cui | Wentao Zhang | Lijun Wu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chart reasoning is a critical capability for Vision Language Models (VLMs). However, the development of open-source models is severely hindered by the lack of high-quality training data. Existing datasets suffer from a dual challenge: synthetic charts are often simplistic and repetitive, while the associated QA pairs are prone to hallucinations and lack the reasoning depth required for complex tasks. To bridge this gap, we propose ChartVerse, a scalable framework designed to synthesize complex charts and reliable reasoning data from scratch. (1) To address the bottleneck of simple patterns, we first introduce Rollout Posterior Entropy (RPE), a novel metric that quantifies chart complexity. Guided by RPE, we develop complexity-aware chart coder to autonomously synthesize diverse, high-complexity charts via executable programs. (2) To guarantee reasoning rigor, we develop truth-anchored inverse QA synthesis. Diverging from standard generation, we adopt an answer-first paradigm: we extract deterministic answers directly from the source code, generate questions conditional on these anchors, and enforce strict consistency verification. To further elevate difficulty and reasoning depth, we filter samples based on model fail-rate and distill high-quality Chain-of-Thought (CoT) reasoning. We curate ChartVerse-SFT-600K and ChartVerse-RL-40K using Qwen3-VL-30B-A3B-Thinking as the teacher. Experimental results demonstrate that ChartVerse-8B achieves state-of-the-art performance, notably surpassing its teacher and rivaling the stronger Qwen3-32B-Thinking.
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
Haoyu Zheng | Yun Zhu | Yuqian Yuan | Bo Yuan | Wenqiao Zhang | Siliang Tang | Jun Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Haoyu Zheng | Yun Zhu | Yuqian Yuan | Bo Yuan | Wenqiao Zhang | Siliang Tang | Jun Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Strategic planning is critical for multi-step reasoning, yet compact Language Language Models (LLMs) often lack the capacity to formulate global strategies, leading to error propagation in long-horizon tasks. Our analysis reveals that LLMs possess latent reasoning capabilities that can be unlocked when conditioned on explicit plans from a teacher model; however, runtime reliance on external guidance is often impractical due to latency and availability constraints. To bridge this gap, we propose PILOT (Planning via Internalized Latent Optimization Trajectories), a non-invasive framework designed to internalize the strategic oversight of large models into intrinsic Latent Guidance. Instead of altering backbone weights, PILOT employs a lightweight Hyper-Network to synthesize a query-conditioned Latent Guidance. This vector acts as an internal steering mechanism, guiding the model’s representations toward optimal reasoning paths. Extensive experiments on mathematical and coding benchmarks demonstrate that PILOT effectively stabilizes reasoning trajectories, consistently outperforming strong baselines (e.g., +8.9% on MATH500) with negligible inference latency. Our code is available at: https://anonymous.4open.science/r/PILOT-B266
COSMOS: Connectivity-Oriented Submodular Maximization for Optimal Subgraph Retrieval
Boci Peng | Xiao Liu | Boren Hu | Yun Zhu | Xuanbo Fan | Yanwei Yue | Chunyu Yang | Yan Zhang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Boci Peng | Xiao Liu | Boren Hu | Yun Zhu | Xuanbo Fan | Yanwei Yue | Chunyu Yang | Yan Zhang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Retrieving coherent evidence subgraphs is critical for Knowledge Base Question Answering (KBQA). Existing paradigms often treat facts independently, rely on biased heuristics, or employ myopic search, failing to optimize collective subgraph utility. In this paper, we propose COSMOS (Connectivity-Oriented Submodular Maximization for Optimal Subgraph Retrieval), a unified framework that formalizes evidence retrieval as a constrained submodular maximization problem. This formulation mathematically captures the trade-off between information relevance and structural complexity. To tractably solve this combinatorial challenge, COSMOS employs a decompose-and-conquer strategy, which first performs a seed-guided greedy expansion to maximize local semantic utility, followed by a topology-aware component aggregation to bridge disjoint evidence clusters via Maximum Spanning Tree aggregation. Guided by theoretical bounds, we introduce Structure-Aware Contrastive Tuning to align semantic space with KG topology. Experimental results on WebQSP, CWQ, and M3GQA benchmarks demonstrate that COSMOS achieves state-of-the-art performance.
2025
Meta-Reflection: A Feedback-Free Reflection Learning Framework
Yaoke Wang | Yun Zhu | Xintong Bao | Wenqiao Zhang | Suyang Dai | Kehan Chen | Wenqiang Li | Gang Huang | Siliang Tang | Yueting Zhuang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yaoke Wang | Yun Zhu | Xintong Bao | Wenqiao Zhang | Suyang Dai | Kehan Chen | Wenqiang Li | Gang Huang | Siliang Tang | Yueting Zhuang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Despite the remarkable capabilities of large language models (LLMs) in natural language understanding and reasoning, they often display undesirable behaviors, such as generating hallucinations and unfaithful reasoning. A prevalent strategy to mitigate these issues is the use of reflection, which refines responses through an iterative process. However, while promising, reflection heavily relies on high-quality external feedback and requires iterative multi-agent inference processes, thus hindering its practical application. In this paper, we propose Meta-Reflection, a novel feedback-free reflection mechanism that necessitates only a single inference pass without external feedback. Motivated by the human ability to remember and retrieve reflections from past experiences when encountering similar problems, Meta-Reflection integrates reflective insights into a codebook, allowing the historical insights to be stored, retrieved, and used to guide LLMs in problem-solving. To thoroughly investigate and evaluate the practicality of Meta-Reflection in real-world scenarios, we introduce an industrial e-commerce benchmark named E-commerce Customer Intent Detection. Extensive experiments conducted on both public datasets and the ECID benchmark highlight the effectiveness and efficiency of our proposed approach. Project is available at https://github.com/DCDmllm/Meta-Reflection
M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering Benchmark
Boci Peng | Yongchao Liu | Xiaohe Bo | Jiaxin Guo | Yun Zhu | Xuanbo Fan | Chuntao Hong | Yan Zhang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Boci Peng | Yongchao Liu | Xiaohe Bo | Jiaxin Guo | Yun Zhu | Xuanbo Fan | Chuntao Hong | Yan Zhang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Recently, GraphRAG systems have achieved remarkable progress in enhancing the performance and reliability of large language models (LLMs). However, most previous benchmarks are template-based and primarily focus on few-entity queries, which are monotypic and simplistic, failing to offer comprehensive and robust assessments. Besides, the lack of ground-truth reasoning paths also hinders the assessments of different components in GraphRAG systems. To address these limitations, we propose M³GQA, a complex, diverse, and high-quality GraphRAG benchmark focusing on multi-entity queries, with six distinct settings for comprehensive evaluation. In order to construct diverse data with semantically correct ground-truth reasoning paths, we introduce a novel reasoning-driven four-step data construction method, including tree sampling, reasoning path backtracking, query creation, and multi-stage refinement and filtering. Extensive experiments demonstrate that M³GQA effectively reflects the capabilities of GraphRAG methods, offering valuable insights into the model performance and reliability. By pushing the boundaries of current methods, M³GQA establishes a comprehensive, robust, and reliable benchmark for advancing GraphRAG research.
2024
Bridging Local Details and Global Context in Text-Attributed Graphs
Yaoke Wang | Yun Zhu | Wenqiao Zhang | Yueting Zhuang | Yunfei Li | Siliang Tang
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Yaoke Wang | Yun Zhu | Wenqiao Zhang | Yueting Zhuang | Yunfei Li | Siliang Tang
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Representation learning on text-attributed graphs (TAGs) is vital for real-world applications, as they combine semantic textual and contextual structural information. Research in this field generally consist of two main perspectives: local-level encoding and global-level aggregating, respectively refer to textual node information unification (e.g., using Language Models) and structure-augmented modeling (e.g., using Graph Neural Networks). Most existing works focus on combining different information levels but overlook the interconnections, i.e., the contextual textual information among nodes, which provides semantic insights to bridge local and global levels. In this paper, we propose GraphBridge, a multi-granularity integration framework that bridges local and global perspectives by leveraging contextual textual information, enhancing fine-grained understanding of TAGs. Besides, to tackle scalability and efficiency challenges, we introduce a graph-aware token reduction module. Extensive experiments across various models and datasets show that our method achieves state-of-the-art performance, while our graph-aware token reduction module significantly enhances efficiency and solves scalability issues. Codes are available at https://github.com/wykk00/GraphBridge.
Search
Fix author
Co-authors
- Siliang Tang 3
- Wenqiao Zhang 3
- Xuanbo Fan 2
- Xin Gao 2
- Conghui He 2
- Yu Li 2
- Honglin Lin 2
- Zheng Liu 2
- Zhuoshi Pan 2
- Qizhi Pei 2
- Boci Peng 2
- Xiaoran Shang 2
- Xiaoyang Wang 2
- Yaoke Wang 2
- Lijun Wu 2
- Yan Zhang 2
- Zhanping Zhong 2
- Yueting Zhuang 2
- Xintong Bao 1
- Xiaohe Bo 1
- Mengzhang Cai 1
- Kehan Chen 1
- Bin Cui 1
- Suyang Dai 1
- Jiaxin Guo 1
- Chuntao Hong 1
- Boren Hu 1
- Gang Huang 1
- Wenqiang Li 1
- Yunfei Li 1
- Dahua Lin 1
- Xiao Liu 1
- Yongchao Liu 1
- Jun Xiao 1
- Chunyu Yang 1
- Bo Yuan 1
- Yuqian Yuan 1
- Yanwei Yue 1
- Wentao Zhang 1
- Feng Zhao 1
- Haoyu Zheng 1