Yu Sun
Author directoryOther people with similar names: Yu Sun, Yu Sun, Yu Sun
Unverified author pages with similar names: Yu Sun
2026
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
Yishu Lei | Shuwei He | Hu Jing | Dan Zhang | Xianlong Luo | Danxiang Zhu | Shikun Feng | Rui Liu | Jingzhou HE | Yu Sun | Hua Wu | Haifeng Wang
Findings of the Association for Computational Linguistics: ACL 2026
Yishu Lei | Shuwei He | Hu Jing | Dan Zhang | Xianlong Luo | Danxiang Zhu | Shikun Feng | Rui Liu | Jingzhou HE | Yu Sun | Hua Wu | Haifeng Wang
Findings of the Association for Computational Linguistics: ACL 2026
Extending the input modality of Large Language Models (LLMs) to the audio domain is essential for achieving comprehensive multimodal perception. However, it is well-known that acoustic information is intrinsically heterogeneous, entangling attributes such as speech, music, and environmental context. Existing research is limited to a dense, parameter-shared adapter to model these diverse patterns, which induces gradient conflict during optimization, as parameter updates required for distinct attributes contradict each other. To address this limitation, we introduce the MoE-Adapter, a sparse Mixture-of-Experts (MoE) architecture designed to decouple acoustic information. Specifically, it employs a dynamic gating mechanism that routes audio tokens to specialized experts capturing complementary feature subspaces while retaining shared experts for global context, thereby mitigating gradient conflicts and enabling fine-grained feature learning. Comprehensive experiments show that the MoE-Adapter achieves superior performance on both audio semantic and paralinguistic tasks, consistently outperforming dense linear baselines with comparable computational costs. To facilitate future research, our code are publicly available at https://github.com/Alittleegg/Eureka-Audio.
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
Yao Chen | Yilong Chen | Yinqi Yang | Junyuan Shang | Zhenyu Zhang | Zefeng Zhang | Shuaiyi Nie | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang | Tingwen Liu
Findings of the Association for Computational Linguistics: ACL 2026
Yao Chen | Yilong Chen | Yinqi Yang | Junyuan Shang | Zhenyu Zhang | Zefeng Zhang | Shuaiyi Nie | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang | Tingwen Liu
Findings of the Association for Computational Linguistics: ACL 2026
Existing approaches to increasing the effective depth of Transformers predominantly rely on parameter reuse, extending computation through recursive execution.Under this paradigm, the network structure remains static along the training timeline, and additional computational depth is uniformly assigned to entire blocks at the parameter level.This rigidity across training time and parameter space leads to substantial computational redundancy during training.In contrast, we argue that depth allocation during training should not be a static preset, but rather a progressively growing structural process. Our systematic analysis reveals a deep-to-shallow maturation trajectory across layers, where high-entropy attention heads play a crucial role in semantic integration. Motivated by this observation, we introduce the Sparse Growing Transformer (SGT).SGT is a training-time sparse depth allocation framework that progressively extends recurrence from deeper to shallower layers via targeted attention looping on informative heads. This mechanism induces structural sparsity by selectively increasing depth only for a small subset of parameters as training evolves.Extensive experiments across multiple parameter scales demonstrate that SGT consistently outperforms training-time static block-level looping baselines under comparable settings, while reducing the additional training FLOPs overhead from approximately 16–20% to only 1–3% relative to a standard Transformer backbone.
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
Hu Jing | Danxiang Zhu | Xianlong Luo | Dan Zhang | Shuwei He | Yishu Lei | Shikun Feng | Hai-Tao Zheng | Jingzhou HE | Yu Sun | Hua Wu | Haifeng Wang
Findings of the Association for Computational Linguistics: ACL 2026
Hu Jing | Danxiang Zhu | Xianlong Luo | Dan Zhang | Shuwei He | Yishu Lei | Shikun Feng | Hai-Tao Zheng | Jingzhou HE | Yu Sun | Hua Wu | Haifeng Wang
Findings of the Association for Computational Linguistics: ACL 2026
Large Audio Language Models (LALMs) have garnered significant research interest. Despite being built upon text-based large language models (LLMs), LALMs frequently exhibit a degradation in knowledge and reasoning capabilities. We hypothesize that this limitation stems from the failure of current training paradigms to effectively bridge the acoustic-semantic gap within the feature representation space. To address this challenge, we propose CORD, a unified alignment framework that performs online cross-modal self-distillation. Specifically, it aligns audio-conditioned reasoning with its text-conditioned counterpart within a unified model. Leveraging the text modality as an internal teacher, CORD performs multi-granularity alignment throughout the audio rollout process. At the token level, it employs on-policy reverse KL divergence with importance-aware weighting to prioritize early and semantically critical tokens. At the sequence level, CORD introduces a judge-based global reward to optimize complete reasoning trajectories via Group Relative Policy Optimization (GRPO). Empirical results across multiple benchmarks demonstrate that CORD consistently enhances audio-conditioned reasoning and substantially bridges the audio–text performance gap with only 80k synthetic training samples, validating the efficacy and data efficiency of our on-policy, multi-level cross-modal alignment approach.
AttnPO: Attention-Guided Process Supervision for Efficient Reasoning
Shuaiyi Nie | Dingsiyu | Wenyuan Zhang | Linhao Yu | Tianmeng Yang | Yao Chen | Weichong Yin | Yu Sun | Hua Wu | Tingwen Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Shuaiyi Nie | Dingsiyu | Wenyuan Zhang | Linhao Yu | Tianmeng Yang | Yao Chen | Weichong Yin | Yu Sun | Hua Wu | Tingwen Liu
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large reasoning models trained with reinforcement learning and verifiable rewards (RLVR) achieve strong performance on complex reasoning tasks, yet often overthink, generating redundant reasoning without performance gains. Existing trajectory-level length penalties often fail to effectively shorten reasoning length and degrade accuracy, as they uniformly treat all reasoning steps and lack fine-grained signals to distinguish redundancy from necessity. Meanwhile, process-supervised methods are typically resource-intensive and suffer from inaccurate credit assignment. To address these issues, we propose ATTNPO, a low-overhead process-supervised RL framework that leverages the model’s intrinsic attention signals for step-level credit assignment. We first identify a set of special attention heads that naturally focus on essential steps while suppressing redundant ones. By leveraging the attention scores of these heads, We then employ two sub-strategies to mitigate overthinking by discouraging redundant steps while preserving accuracy by reducing penalties on essential steps. Experimental results show that ATTNPO substantially reduces reasoning length while significantly improving performance across 9 benchmarks.
Uncertainty-Aware Routing for Principled Alignment with MoE Dynamics
Yilong Chen | Junyuan Shang | Yuchen Feng | Zhenyu Zhang | Naibin Gu | Ziqi Wang | Tingwen Liu | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yilong Chen | Junyuan Shang | Yuchen Feng | Zhenyu Zhang | Naibin Gu | Ziqi Wang | Tingwen Liu | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Mixture-of-Experts (MoE) is a cornerstone for scaling LLMs, yet its training dynamics remain poorly understood, often leading to sub-optimal specialization. Moving beyond static routing, we present a systematic study of the MoE lifecycle using Helmholtz Free Energyand Router Entropy. We identify a universal Three-Stage Phase Transition—Exploration, Symmetry Breaking, and Stabilization—marked by an Energy Climb and Plateau. This reflects Frustrated Exploration, caused by structural interference between specialization drives and uniformity constraints. To address this, we propose Uncertainty-Aware Routing (UAR), which aligns routing with the model’s epistemic state via: (1) Evidence-Triggered Expansion, increasing active experts for high-energy tokens, and (2) Epistemic Masking, applying load-balancing only in high-uncertainty regimes to shield mature experts. Experiments confirm UAR reduces perplexity and improves expert distinctiveness, offering a principled path toward thermodynamically aligned computation.
2025
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
Tingfeng Hui | Zhenyu Zhang | Shuohuan Wang | Yu Sun | Hua Wu | Sen Su
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tingfeng Hui | Zhenyu Zhang | Shuohuan Wang | Yu Sun | Hua Wu | Sen Su
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Mixture-of-Experts (MoE) shines brightly in large language models (LLMs) and demonstrates outstanding performance in plentiful natural language processing tasks. However, existing methods transforming LLMs from dense to MoE face significant data requirements and typically rely on large-scale post-training. In this paper, we propose Upcycling Instruction Tuning (UpIT), a data-efficient approach for tuning a dense pre-trained model into a MoE instruction model. Specifically, we first point out that intermediate checkpoints during instruction tuning of the dense model are naturally suitable for specialized experts, and then propose an expert expansion stage to flexibly achieve models with flexible numbers of experts, where genetic algorithm and parameter merging are introduced to ensure sufficient diversity of new extended experts. To ensure that each specialized expert in the MoE model works as expected, we select a small amount of seed data that each expert excels to pre-optimize the router. Extensive experiments with various data scales and upcycling settings demonstrate the outstanding performance and data efficiency of UpIT, as well as stable improvement in expert or data scaling. Further analysis reveals the importance of ensuring expert diversity in upcycling.
HFT: Half Fine-Tuning for Large Language Models
Tingfeng Hui | Zhenyu Zhang | Shuohuan Wang | Weiran Xu | Yu Sun | Hua Wu
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tingfeng Hui | Zhenyu Zhang | Shuohuan Wang | Weiran Xu | Yu Sun | Hua Wu
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large language models (LLMs) with one or more fine-tuning phases have become necessary to unlock various capabilities, enabling LLMs to follow natural language instructions and align with human preferences. However, it carries the risk of catastrophic forgetting during sequential training, the parametric knowledge or the ability learned in previous stages may be overwhelmed by incoming training data. This paper finds that LLMs can restore some original knowledge by regularly resetting partial parameters. Inspired by this, we introduce Half Fine-Tuning (HFT) for LLMs, as a substitute for full fine-tuning (FFT), to mitigate the forgetting issues, where half of the parameters are selected to learn new tasks. In contrast, the other half are frozen to retain previous knowledge. We provide a feasibility analysis from the optimization perspective and interpret the parameter selection operation as a regularization term. HFT could be seamlessly integrated into existing fine-tuning frameworks without changing the model architecture. Extensive experiments and analysis on supervised fine-tuning, direct preference optimization, and continual learning consistently demonstrate the effectiveness, robustness, and efficiency of HFT. Compared with FFT, HFT not only significantly alleviates the forgetting problem, but also achieves the best performance in a series of downstream benchmarks, with an approximately 30% reduction in training time.
BeamLoRA: Beam-Constraint Low-Rank Adaptation
Naibin Gu | Zhenyu Zhang | Xiyu Liu | Peng Fu | Zheng Lin | Shuohuan Wang | Yu Sun | Hua Wu | Weiping Wang | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Naibin Gu | Zhenyu Zhang | Xiyu Liu | Peng Fu | Zheng Lin | Shuohuan Wang | Yu Sun | Hua Wu | Weiping Wang | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Due to the demand for efficient fine-tuning of large language models, Low-Rank Adaptation (LoRA) has been widely adopted as one of the most effective parameter-efficient fine-tuning methods. Nevertheless, while LoRA improves efficiency, there remains room for improvement in accuracy. Herein, we adopt a novel perspective to assess the characteristics of LoRA ranks. The results reveal that different ranks within the LoRA modules not only exhibit varying levels of importance but also evolve dynamically throughout the fine-tuning process, which may limit the performance of LoRA. Based on these findings, we propose BeamLoRA, which conceptualizes each LoRA module as a beam where each rank naturally corresponds to a potential sub-solution, and the fine-tuning process becomes a search for the optimal sub-solution combination. BeamLoRA dynamically eliminates underperforming sub-solutions while expanding the parameter space for promising ones, enhancing performance with a fixed rank. Extensive experiments across three base models and 12 datasets spanning math reasoning, code generation, and commonsense reasoning demonstrate that BeamLoRA consistently enhances the performance of LoRA, surpassing the other baseline methods.
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
Yilong Chen | Junyuan Shang | Zhenyu Zhang | Yanxi Xie | Jiawei Sheng | Tingwen Liu | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yilong Chen | Junyuan Shang | Zhenyu Zhang | Yanxi Xie | Jiawei Sheng | Tingwen Liu | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large language models (LLMs) face inherent performance bottlenecks under parameter constraints, particularly in processing critical tokens that demand complex reasoning. Empirical analysis reveals challenging tokens induce abrupt gradient spikes across layers, exposing architectural stress points in standard Transformers. Building on this insight, we propose Inner Thinking Transformer (ITT), which reimagines layer computations as implicit thinking steps. ITT dynamically allocates computation through Adaptive Token Routing, iteratively refines representations via Residual Thinking Connections, and distinguishes reasoning phases using Thinking Step Encoding. ITT enables deeper processing of critical tokens without parameter expansion. Evaluations across 162M-466M parameter models show ITT achieves 96.5% performance of a 466M Transformer using only 162M parameters, reduces training data by 43.2%, and outperforms Transformer/Loop variants in 11 benchmarks. By enabling elastic computation allocation during inference, ITT balances performance and efficiency through architecture-aware optimization of implicit thinking pathways.
Curiosity-Driven Reinforcement Learning from Human Feedback
Haoran Sun | Yekun Chai | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Haoran Sun | Yekun Chai | Shuohuan Wang | Yu Sun | Hua Wu | Haifeng Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Reinforcement learning from human feedback (RLHF) has proven effective in aligning large language models (LLMs) with human preferences, but often at the cost of reduced output diversity. This trade-off between diversity and alignment quality remains a significant challenge. Drawing inspiration from curiosity-driven exploration in reinforcement learning, we introduce curiosity-driven RLHF (CD-RLHF), a framework that incorporates intrinsic rewards for novel states, alongside traditional sparse extrinsic rewards, to optimize both output diversity and alignment quality. We demonstrate the effectiveness of CD-RLHF through extensive experiments on a range of tasks, including text summarization and instruction following. Our approach achieves significant gains in diversity on multiple diversity-oriented metrics while maintaining alignment with human preferences comparable to standard RLHF. We will make our code publicly available.
Search
Fix author
Co-authors
- Hua Wu (吴华) 10
- Haifeng Wang 7
- Shuohuan Wang 7
- Zhenyu Zhang 6
- Tingwen Liu 4
- Yilong Chen 3
- Junyuan Shang 3
- Yao Chen 2
- Shikun Feng 2
- Naibin Gu 2
- Jingzhou HE 2
- Shuwei He 2
- Tingfeng Hui 2
- Hu Jing 2
- Yishu Lei 2
- Xianlong Luo 2
- Shuaiyi Nie 2
- Dan Zhang 2
- Danxiang Zhu 2
- Yekun Chai 1
- Dingsiyu 1
- Yuchen Feng 1
- Peng Fu 1
- Zheng Lin 1
- Rui Liu 1
- Xiyu Liu 1
- Jiawei Sheng 1
- Sen Su 1
- Haoran Sun 1
- Weiping Wang 1
- Ziqi Wang 1
- Yanxi Xie 1
- Weiran Xu 1
- Tianmeng Yang 1
- Yinqi Yang 1
- Weichong Yin 1
- Linhao Yu 1
- Wenyuan Zhang 1
- Zefeng Zhang 1
- Hai-Tao Zheng 1