Wei Wang
Other people with similar names: Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang, Wei Wang
Unverified author pages with similar names: Wei Wang
2026
DataSciBench: An LLM Agent Benchmark for Data Science
Dan Zhang | Sining Zhoubian | Min Cai | Fengzu Li | Lekang Yang | Wei Wang | Tianjiao Dong | Ziniu Hu | Jie Tang | Yisong Yue
Findings of the Association for Computational Linguistics: ACL 2026
Dan Zhang | Sining Zhoubian | Min Cai | Fengzu Li | Lekang Yang | Wei Wang | Tianjiao Dong | Ziniu Hu | Jie Tang | Yisong Yue
Findings of the Association for Computational Linguistics: ACL 2026
This paper presents DataSciBench, a comprehensive benchmark for evaluating Large Language Models (LLMs) in data science. Unlike existing benchmarks limited to single task, simple evaluation metrics, and readily available ground truth (GT), DataSciBench is built on curated, natural, and challenging prompts with complex evaluation criteria and uncertain GT. To bridge the gap, we develop a semi-automated GT generation pipeline, integrating LLM-based self-consistency and human verification to ensure accuracy, predefined task types, and aggregate functions (metrics). Furthermore, we introduce an innovative Intention-Function-Code (IFC) framework, assessing code execution outcomes through metrics and programmatic rules. Evaluating 26 models (8 API-based, 8 open-source general, 9 code generation, and 1 agentic models), our approach offers rigorous insights into LLM strengths and weaknesses. Experimental results show API-based models outperform open-source counterparts across all metrics, with DeepAnalyze-8B leading among open-sourced models. We release all code and data at https://github.com/THUDM/DataSciBench.
2025
IntelliCockpitBench: A Comprehensive Benchmark to Evaluate VLMs for Intelligent Cockpit
Liang Lin | Siyuan Chai | Jiahao Wu | Hongbing Hu | Xiaotao Gu | Hao Hu | Fan Zhang | Wei Wang | Dan Zhang
Findings of the Association for Computational Linguistics: ACL 2025
Liang Lin | Siyuan Chai | Jiahao Wu | Hongbing Hu | Xiaotao Gu | Hao Hu | Fan Zhang | Wei Wang | Dan Zhang
Findings of the Association for Computational Linguistics: ACL 2025
The integration of sophisticated Vision-Language Models (VLMs) in vehicular systems is revolutionizing vehicle interaction and safety, performing tasks such as Visual Question Answering (VQA). However, a critical gap persists due to the lack of a comprehensive benchmark for multimodal VQA models in vehicular scenarios. To address this, we propose IntelliCockpitBench, a benchmark that encompasses diverse automotive scenarios. It includes images from front, side, and rear cameras, various road types, weather conditions, and interior views, integrating data from both moving and stationary states. Notably, all images and queries in the benchmark are verified for high levels of authenticity, ensuring the data accurately reflects real-world conditions. A sophisticated scoring methodology combining human and model-generated assessments enhances reliability and consistency. Our contributions include a diverse and authentic dataset for automotive VQA and a robust evaluation metric aligning human and machine assessments. All code and data can be found at https://github.com/Lane315/IntelliCockpitBench.