Tiancheng Hu
Other people with similar names: Tiancheng Hu
Unverified author pages with similar names: Tiancheng Hu
2026
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
Tiancheng Hu | Benjamin Minixhofer | Nigel Collier
Findings of the Association for Computational Linguistics: ACL 2026
Tiancheng Hu | Benjamin Minixhofer | Nigel Collier
Findings of the Association for Computational Linguistics: ACL 2026
The “alignment tax” of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We demonstrate that this trade-off can be navigated effectively via a simple post-hoc intervention: interpolating between a model’s weights before and after alignment. Crucially, this is not a strict trade-off. We find that the process consistently reveals Pareto-optimal interpolations—models that improve accuracy beyond both parents while substantially recovering the calibration lost during alignment. Our work demonstrates that simple model merging provides a computationally efficient method for mitigating the full scope of the alignment tax, yielding models that are more capable and more reliable.
Confidence Estimation for LLMs in Multi-turn Interactions
Caiqi Zhang | Ruihan Yang | Xiaochen Zhu | Chengzu Li | Tiancheng Hu | Yijiang River Dong | Deqing Yang | Nigel Collier
Findings of the Association for Computational Linguistics: ACL 2026
Caiqi Zhang | Ruihan Yang | Xiaochen Zhu | Chengzu Li | Tiancheng Hu | Yijiang River Dong | Deqing Yang | Nigel Collier
Findings of the Association for Computational Linguistics: ACL 2026
While confidence estimation is a promising direction for mitigating hallucinations in Large Language Models (LLMs), current research overwhelmingly focuses on single-turn settings. The dynamics of model confidence in multi-turn conversations, where context accumulates and ambiguity is progressively resolved, remain largely unexplored. This work presents the first systematic study of confidence estimation in multi-turn interactions, establishing a formal evaluation framework grounded in two key desiderata: per-turn calibration and monotonicity of confidence as more information becomes available. To facilitate this, we introduce novel metrics, including a length-normalized Expected Calibration Error (InfoECE), and a new “Hinter-Guesser” paradigm for generating controlled evaluation datasets. Our experiments reveal that widely-used confidence techniques struggle with calibration and monotonicity in multi-turn dialogues. In contrast, a novel logit-based probe we introduce, P(Sufficient), proves comparatively more effective, robustly tracking evidence accumulation and distinguishing it from conversational filler. Our work provides a foundational methodology for developing more reliable and trustworthy conversational agents.
Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective
Beiduo Chen | Tiancheng Hu | Caiqi Zhang | Robert Litschko | Anna Korhonen | Barbara Plank
Findings of the Association for Computational Linguistics: ACL 2026
Beiduo Chen | Tiancheng Hu | Caiqi Zhang | Robert Litschko | Anna Korhonen | Barbara Plank
Findings of the Association for Computational Linguistics: ACL 2026
Reasoning-tuned LLMs utilizing long Chain-of-Thought (CoT) excel at single-answer tasks, yet their ability to model Human Label Variation—which requires capturing probabilistic ambiguity rather than resolving it—remains underexplored. We investigate this through systematic disentanglement experiments on distribution-based tasks, employing Cross-CoT experiments to isolate the effect of reasoning text from intrinsic model priors. We observe a distinct “decoupled mechanism”: while CoT improves distributional alignment, final accuracy is dictated by CoT content (99% variance contribution), whereas distributional ranking is governed by model priors (over 80%). Step-wise analysis further shows that while CoT’s influence on accuracy grows monotonically during the reasoning process, distributional structure is largely determined by LLM’s intrinsic priors. These findings suggest that long CoT serves as a decisive LLM decision-maker for the top option but fails to function as a granular distribution calibrator for ambiguous tasks.
TRACE: A Corpus of Team Creative Discussions
Yixuan Jiang | Tiancheng Hu | Jose Hernandez-Orallo | David Stillwell | Luning Sun
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yixuan Jiang | Tiancheng Hu | Jose Hernandez-Orallo | David Stillwell | Luning Sun
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Understanding how discussion dynamics shape team creativity has been limited by the difficulty of measuring process at scale. We introduce Trace, a corpus of 309 group discussions from 103 teams (460 participants) across six creative problem-solving tasks. The dataset follows an input-process-output framework, integrating team composition (demographics, personalities), full discussion transcripts, and creativity outcomes. Using sentence embeddings and factor analysis, we identify four interpretable discussion dimensions: Coherence, Exploration, Convergence, and Participation. Analysis reveals a depth-breadth trade-off: coherent idea development inversely relates to semantic exploration. Larger teams explore more broadly but converge less effectively while team diversity shapes participation patterns more than discussion content. Novelty and usefulness in the creativity outcomes follow distinct pathways: Exploration and Convergence predict novelty, whereas Coherence predicts usefulness. These findings ground our understanding of how teams talk their way to creative solutions and provide guidance for designing multiagent systems.
Value of Information: A Framework for Human–Agent Communication
Yijiang River Dong | Tiancheng Hu | Zheng Hui | Caiqi Zhang | Ivan Vulić | Andreea Bobu | Nigel Collier
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yijiang River Dong | Tiancheng Hu | Zheng Hui | Caiqi Zhang | Ivan Vulić | Andreea Bobu | Nigel Collier
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large Language Model (LLM) agents deployed for real-world tasks face a fundamental dilemma: user requests are underspecified, yet agents must decide whether to act on incomplete information or interrupt users for clarification. Existing approaches either rely on brittle confidence thresholds that require task-specific tuning, or fail to account for the varying stakes of different decisions. We introduce a decision-theoretic framework that resolves this trade-off through the Value of Information (VoI), enabling agents to dynamically weigh the expected utility gain from asking questions against the cognitive cost imposed on users. Our inference-time method requires no hyperparameter tuning and adapts seamlessly across contexts—from casual games to medical diagnosis. Experiments across four diverse domains (20 Questions, medical diagnosis, flight booking, and e-commerce) show that VoI consistently matches or exceeds the best manually-tuned baselines, achieving up to 1.36 utility points higher in high-cost settings. This work provides a parameter-free framework for adaptive agent communication that explicitly balances task risk, query ambiguity, and user effort.
2025
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
Yijiang River Dong | Tiancheng Hu | Yinhong Liu | Ahmet Üstün | Nigel Collier
Findings of the Association for Computational Linguistics: EMNLP 2025
Yijiang River Dong | Tiancheng Hu | Yinhong Liu | Ahmet Üstün | Nigel Collier
Findings of the Association for Computational Linguistics: EMNLP 2025
While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences across users, overlooking diverse human values and minority viewpoints.Although personalized preference learning addresses this by tailoring separate preferences for individual users, the field lacks standardized methods to assess its effectiveness. We present a multi-faceted evaluation framework that measures not only performance but also fairness, unintended effects, and adaptability across varying levels of preference divergence. Through extensive experiments comparing eight personalization methods across three preference datasets, we demonstrate that performance differences between methods could reach 36% when users strongly disagree, and personalization can introduce up to 20% safety misalignment. These findings highlight the critical need for holistic evaluation approaches to advance the development of more effective and inclusive preference learning systems.
Personalization up to a Point: Why Personalized Content Moderation Needs Boundaries, and How We Can Enforce Them
Emanuele Moscato | Tiancheng Hu | Matthias Orlikowski | Paul Röttger | Debora Nozza
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Emanuele Moscato | Tiancheng Hu | Matthias Orlikowski | Paul Röttger | Debora Nozza
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Personalized content moderation can protect users from harm while facilitating free expression by tailoring moderation decisions to individual preferences rather than enforcing universal rules. However, content moderation that is fully personalized to individual preferences, no matter what these preferences are, may lead to even the most hazardous types of content being propagated on social media. In this paper, we explore this risk using hate speech as a case study. Certain types of hate speech are illegal in many countries. We show that, while fully personalized hate speech detection models increase overall user welfare (as measured by user-level classification performance), they also make predictions that violate such legal hate speech boundaries, especially when tailored to users who tolerate highly hateful content. To address this problem, we enforce legal boundaries in personalized hate speech detection by overriding predictions from personalized models with those from a boundary classifier. This approach significantly reduces legal violations while minimally affecting overall user welfare. Our findings highlight both the promise and the risks of personalized moderation, and offer a practical solution to balance user preferences with legal and ethical obligations.
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
Ona de Gibert | Joseph Attieh | Teemu Vahtola | Mikko Aulamo | Zihao Li | Raúl Vázquez | Tiancheng Hu | Jörg Tiedemann
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Ona de Gibert | Joseph Attieh | Teemu Vahtola | Mikko Aulamo | Zihao Li | Raúl Vázquez | Tiancheng Hu | Jörg Tiedemann
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
We investigate the potential of LLM-generated synthetic data for improving low-resource Machine Translation (MT). Focusing on seven diverse target languages, we construct a document-level synthetic corpus from English Europarl, and extend it via pivoting to 147 additional language pairs. Automatic and human evaluation confirm its overall high quality. We study its practical application by (i) identifying effective training regimes, (ii) comparing our data with the HPLT dataset, (iii) studying the effect of varying training data size, and (iiii) testing its utility beyond English-centric MT. Finally, we introduce SynOPUS, a public repository for synthetic parallel datasets. Our findings show that LLM-generated synthetic data, even when noisy, can substantially improve MT performance for low-resource languages.
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
Tiancheng Hu | Nigel Collier
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tiancheng Hu | Nigel Collier
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Understanding how individuals perceive and react to information is fundamental for advancing social and behavioral sciences and developing human-centered AI systems. Current approaches often lack the granular data needed to model these personalized responses, relying instead on aggregated labels that obscure the rich variability driven by individual differences. We introduce iNews, a novel large-scale dataset specifically designed to facilitate the modeling of personalized affective responses to news content. Our dataset comprises annotations from 291 demographically diverse UK participants across 2,899 multimodal Facebook news posts from major UK outlets, with an average of 5.18 annotators per sample. For each post, annotators provide multifaceted labels including valence, arousal, dominance, discrete emotions, content relevance judgments, sharing likelihood, and modality importance ratings. Crucially, we collect comprehensive annotator persona information covering demographics, personality, media trust, and consumption patterns, which explain 15.2% of annotation variance - substantially higher than existing NLP datasets. Incorporating this information yields a 7% accuracy gain in zero-shot prediction and remains beneficial even with 32-shot in-context learning.
Search
Fix author
Co-authors
- Nigel Collier 5
- Yijiang River Dong 3
- Caiqi Zhang 3
- Joseph Attieh 1
- Mikko Aulamo 1
- Andreea Bobu 1
- Beiduo Chen 1
- Jose Hernandez-Orallo 1
- Zheng Hui 1
- Yixuan Jiang 1
- Anna Korhonen 1
- Chengzu Li 1
- Zihao Li 1
- Robert Litschko 1
- Yinhong Liu 1
- Benjamin Minixhofer 1
- Emanuele Moscato 1
- Debora Nozza 1
- Matthias Orlikowski 1
- Barbara Plank 1
- Paul Röttger 1
- David Stillwell 1
- Luning Sun 1
- Jörg Tiedemann 1
- Teemu Vahtola 1
- Ivan Vulić 1
- Raúl Vázquez 1
- Deqing Yang 1
- Ruihan Yang 1
- Xiaochen Zhu 1
- Ona de Gibert 1
- Ahmet Üstün 1