Jin Huang
Author directoryOther people with similar names: Jin Huang
Unverified author pages with similar names: Jin Huang
2026
Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks
Xinhe Wang | Jin Huang | Xingjian Zhang | Tianhao Wang | Jiaqi W. Ma
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xinhe Wang | Jin Huang | Xingjian Zhang | Tianhao Wang | Jiaqi W. Ma
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Reasoning benchmarks such as the Abstraction and Reasoning Corpus (ARC) and ARC-AGI are widely used to assess progress in artificial intelligence and are often interpreted as probes of core, so-called “fluid” reasoning abilities. Despite their apparent simplicity for humans, these tasks remain challenging for frontier vision-language models (VLMs), a gap commonly attributed to deficiencies in machine reasoning. We challenge this interpretation and hypothesize that the gap arises primarily from limitations in visual perception rather than from shortcomings in inductive reasoning.To verify this hypothesis, we introduce a two-stage experimental pipeline that explicitly separates perception and reasoning. In the perception stage, each image is independently converted into a natural-language description, while in the reasoning stage a model induces and applies rules using these descriptions. This design prevents leakage of cross-image inductive signals and isolates reasoning from perception bottlenecks. Across three ARC-style datasets, Mini-ARC, ACRE, and Bongard-LOGO, we show that the perception capability is the dominant factor underlying the observed performance gap by comparing the two-stage pipeline with against standard end-to-end one-stage evaluation. Manual inspection of reasoning traces in the VLM outputs further reveals that approximately 80 percent of model failures stem from perception errors. Together, these results demonstrate that ARC-style benchmarks conflate perceptual and reasoning challenges and that observed performance gaps may overstate deficiencies in machine reasoning. Our findings underscore the need for evaluation protocols that disentangle perception from reasoning when assessing progress in machine intelligence.
2025
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
Xingjian Zhang | Yutong Xie | Jin Huang | Jinge Ma | Zhaoying Pan | Qijia Liu | Ziyang Xiong | Tolga Ergen | Dongsub Shim | Honglak Lee | Qiaozhu Mei
Findings of the Association for Computational Linguistics: NAACL 2025
Xingjian Zhang | Yutong Xie | Jin Huang | Jinge Ma | Zhaoying Pan | Qijia Liu | Ziyang Xiong | Tolga Ergen | Dongsub Shim | Honglak Lee | Qiaozhu Mei
Findings of the Association for Computational Linguistics: NAACL 2025
Scientific innovation relies on detailed workflows, which include critical steps such as contextualizing literature, generating ideas, validating ideas, interpreting results, and planning new research. Scientific publications that document these workflows are extensive and unstructured, making it difficult to effectively navigate and explore the space of scientific innovation. To meet this challenge, we introduce MASSW, a comprehensive dataset of Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 152,000 peer-reviewed publications from 17 leading computer science conferences spanning the past 50 years. Using Large Language Models (LLMs), we automatically extract five core aspects from these publications – context, key idea, method, outcome, and projected impact – which correspond to five key steps in a research workflow. We show that these LLM-extract summaries have a comparable quality to human annotations, and they facilitate a variety of downstream tasks, corresponding to different types of predictions and recommendations along the scientific workflow. Overall, MASSW demonstrates decent utility as a pre-computed and trustful resource for the AI4Science community to create and benchmark a wide-range of new AI methods for optimizing scientific workflows and fostering scientific innovation. Our code and datasets are made available anonymously: link.
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
Hengxing Cai | Xiaochen Cai | Junhan Chang | Sihang Li | Lin Yao | Wang Changxin | Zhifeng Gao | Hongshuai Wang | Li Yongge | Mujie Lin | Shuwen Yang | Jiankun Wang | Mingjun Xu | Jin Huang | Xi Fang | Jiaxi Zhuang | Yuqi Yin | Yaqi Li | Changhong Chen | Zheng Cheng | Zifeng Zhao | Linfeng Zhang | Guolin Ke
Findings of the Association for Computational Linguistics: NAACL 2025
Hengxing Cai | Xiaochen Cai | Junhan Chang | Sihang Li | Lin Yao | Wang Changxin | Zhifeng Gao | Hongshuai Wang | Li Yongge | Mujie Lin | Shuwen Yang | Jiankun Wang | Mingjun Xu | Jin Huang | Xi Fang | Jiaxi Zhuang | Yuqi Yin | Yaqi Li | Changhong Chen | Zheng Cheng | Zifeng Zhao | Linfeng Zhang | Guolin Ke
Findings of the Association for Computational Linguistics: NAACL 2025
Recent breakthroughs in Large Language Models (LLMs) have revolutionized scientific literature analysis. However, existing benchmarks fail to adequately evaluate the proficiency of LLMs in this domain, particularly in scenarios requiring higher-level abilities beyond mere memorization and the handling of multimodal data.In response to this gap, we introduce SciAssess, a benchmark specifically designed for the comprehensive evaluation of LLMs in scientific literature analysis. It aims to thoroughly assess the efficacy of LLMs by evaluating their capabilities in Memorization (L1), Comprehension (L2), and Analysis & Reasoning (L3). It encompasses a variety of tasks drawn from diverse scientific fields, including biology, chemistry, material, and medicine.To ensure the reliability of SciAssess, rigorous quality control measures have been implemented, ensuring accuracy, anonymization, and compliance with copyright standards. SciAssess evaluates 11 LLMs, highlighting their strengths and areas for improvement. We hope this evaluation supports the ongoing development of LLM applications in scientific literature analysis.SciAssess and its resources are available at https://github.com/sci-assess/SciAssess.
Search
Fix author
Co-authors
- Xingjian Zhang 2
- Hengxing Cai 1
- Xiaochen Cai 1
- Junhan Chang 1
- Wang Changxin 1
- Changhong Chen 1
- Zheng Cheng 1
- Tolga Ergen 1
- Xi Fang 1
- Zhifeng Gao 1
- Guolin Ke 1
- Honglak Lee 1
- Sihang Li 1
- Yaqi Li 1
- Mujie Lin 1
- Qijia Liu 1
- Jiaqi W. Ma 1
- Jinge Ma 1
- Qiaozhu Mei 1
- Zhaoying Pan 1
- Dongsub Shim 1
- Hongshuai Wang 1
- Jiankun Wang 1
- Tianhao Wang 1
- Xinhe Wang 1
- Yutong Xie 1
- Ziyang Xiong 1
- Mingjun Xu 1
- Shuwen Yang 1
- Lin Yao 1
- Yuqi Yin 1
- Li Yongge 1
- Linfeng Zhang 1
- Zifeng Zhao 1
- Jiaxi Zhuang 1