Fernando De la Torre
Author directory2026
LLA MADRS: Evaluating Open-Source LLMs on Real Clinical Interviews—To Reason or Not to Reason?
Gaoussou Youssouf Kebe | Jeffrey M. Girard | Einat Liebenthal | Justin Baker | Fernando De la Torre | Louis-Philippe Morency
Transactions of the Association for Computational Linguistics, Volume 14
Gaoussou Youssouf Kebe | Jeffrey M. Girard | Einat Liebenthal | Justin Baker | Fernando De la Torre | Louis-Philippe Morency
Transactions of the Association for Computational Linguistics, Volume 14
Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored. We present LLAMADRS, a benchmark for structured clinical assessment from dialogue built on the CAMI corpus of psychiatric interviews, comprising 5,804 expert annotations across 541 sessions. We evaluate 25 open-source models (standard and reasoning-augmented; 0.6B–400B parameters) and generate over 400,000 predictions. Our results demonstrate that strong open-source LLMs achieve item-level accuracy with residual error below clinically substantial thresholds. Additionally, an Item-then-Sum (ITS) strategy, assessing symptoms individually through discrete LLM calls before synthesizing final scores, significantly reduces error relative to Direct Total Score (DTS) prediction across most model architectures and scales, despite reasoning models attempting similar decomposition in the reasoning traces of their DTS predictions. In fact, we find that performance gains attributed to “reasoning” depend fundamentally on prompt design: standard models equipped with structured task definitions and examples match reasoning-augmented counterparts. Among the latter, longer reasoning traces correlate with reduced error; while higher model scale does across both architectures. Our results clarify when and why reasoning helps and offer actionable guidance for deploying LLMs in semi-structured clinical assessment.
2025
On the Fine-Grained Planning Abilities of VLM Web Agents
Surgan Jandial | Yinong Oliver Wang | Andrea Bajcsy | Fernando De la Torre
Findings of the Association for Computational Linguistics: EMNLP 2025
Surgan Jandial | Yinong Oliver Wang | Andrea Bajcsy | Fernando De la Torre
Findings of the Association for Computational Linguistics: EMNLP 2025
Vision-Language Models (VLMs) have shown promise as web agents, yet their planning—the ability to devise strategies or action sequences to complete tasks—remains understudied. While prior works focus on VLM’s perception and overall success rates (i.e., goal completion), fine-grained investigation of their planning has been overlooked. To address this gap, we examine VLMs’ capability to (1) understand temporal relationships within web contexts, and (2) assess plans of actions across diverse scenarios. We design four simple yet effective tests to delve into these nuanced aspects around planning. Our results across nineteen VLMs reveal that these models exhibit limited performance in the aforementioned skills and are not reliable to function as web agents. To facilitate future work, we release our planning evaluations and data, providing a foundation for advancing the future research in this area.