David W. Eyre
Also published as: David W Eyre
2026
BDI at MEDIQA-EVAL 2026: A ReAct-Style Multimodal Agent for Fine-Grained Medical Response Assessment
Justin Xu | Zizheng Zhang | Augustine Luk | Benjamin Khong | Haochen Cui | Samuel Hwang | Alyssa Pradhan | Kevin Yuan | David W. Eyre
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Justin Xu | Zizheng Zhang | Augustine Luk | Benjamin Khong | Haochen Cui | Samuel Hwang | Alyssa Pradhan | Kevin Yuan | David W. Eyre
Proceedings of the 8th Workshop on Clinical Natural Language Processing (Clinical NLP) @ LREC 2026
Free-text evaluation of multimodal clinical question answering (QA) systems remains a central challenge in medical NLP due to the complexity of medical knowledge, the necessity of integrating visual and textual information, and the limitations of existing automatic evaluation metrics for open-ended outputs. In this work, we present a training-free, agentic evaluation framework that formulates response scoring as evidence-guided orchestration of components rather than a task requiring conventional end-to-end fine-tuning of underlying LLMs/VLMs. Our ReAct-style evaluator combines (i) structured reasoning, (ii) multimodal retrieval of similar encounters, (iii) auxiliary explainable feature-based regression models that provide numeric priors and human-interpretable signals, (iv) VLM-generated visual QA references for comparison, and (v) optional image augmentation tools. Unlike standard LLM-as-a-judge approaches that rely on direct generative scoring, our agent decomposes evaluation into modular stages of evidence acquisition, structured feature modeling, and integrative reasoning. We apply this architecture to the MEDIQA-EVAL shared task - a multimodal, multilingual clinical evaluation challenge that assesses system-generated answers for patient queries paired with images along multiple clinical quality dimensions. We report results across both English and Chinese tracks, comparing against baseline prompting methods, and discuss the feasibility and limitations of lightweight agentic systems for clinical QA evaluation.
2025
Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning
Justin Xu | Yiming Li | Zizheng Zhang | Augustine Yui Hei Luk | Mayank Jobanputra | Samarth Oza | Ashley Murray | Meghana Reddy Kasula | Andrew Parker | David W Eyre
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Justin Xu | Yiming Li | Zizheng Zhang | Augustine Yui Hei Luk | Mayank Jobanputra | Samarth Oza | Ashley Murray | Meghana Reddy Kasula | Andrew Parker | David W Eyre
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Large language models (LLMs) can produce fluent but factually incorrect outputs and often have limited ability to attribute their claims to source material. This undermines their reliability, particularly in multi-hop and high-stakes domains such as medicine. We propose Tree-of-Quote (ToQ) prompting, a structured framework that decomposes complex questions into subquestions, generates quotes to support each step without retrieval, and selectively advances reasoning based on quote quality. We also introduce FQ-Score, a unified metric that captures answer correctness, attribution fidelity, and reasoning quality. Experiments on StrategyQA, 2WikiMultiHopQA, MuSiQue, MoreHopQA, and MedQA demonstrate that ToQ improves factuality and attribution over standard prompting baselines. To validate FQ-Score as a proxy for human judgment, we conduct two reader studies with clinicians on medical questions, and observe strong correlations. Both clinician scores and FQ-Scores also indicate a preference for ToQ over baselines due to a combination of greater correctness, completeness, and logical flow. Our results suggest ToQ is a promising approach for building more trustworthy and auditable LLM systems.
RadEval: A framework for radiology text evaluation
Justin Xu | Xi Zhang | Javid Abderezaei | Julie Bauml | Roger Boodoo | Fatemeh Haghighi | Ali Ganjizadeh | Eric Brattain | Dave Van Veen | Zaiqiao Meng | David W Eyre | Jean-Benoit Delbrouck
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
Justin Xu | Xi Zhang | Javid Abderezaei | Julie Bauml | Roger Boodoo | Fatemeh Haghighi | Ali Ganjizadeh | Eric Brattain | Dave Van Veen | Zaiqiao Meng | David W Eyre | Jean-Benoit Delbrouck
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
We introduce RadEval, a unified, open-source framework for evaluating radiology texts. RadEval consolidates a diverse range of metrics - from classic n‐gram overlap (BLEU, ROUGE) and contextual measures (BERTScore) to clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1) and advanced LLM‐based evaluators (GREEN). We refine and standardize implementations, extend GREEN to support multiple imaging modalities with a more lightweight model, and pretrain a domain-specific radiology encoder - demonstrating strong zero-shot retrieval performance. We also release a richly annotated expert dataset with over 450 clinically significant error labels and show how different metrics correlate with radiologist judgment. Finally, RadEval provides statistical testing tools and baseline model evaluations across multiple publicly available datasets, facilitating reproducibility and robust benchmarking in radiology report generation.
Search
Fix author
Co-authors
- Justin Xu 3
- Zizheng Zhang 2
- Javid Abderezaei 1
- Julie Bauml 1
- Roger Boodoo 1
- Eric Brattain 1
- Haochen Cui 1
- Jean-Benoit Delbrouck 1
- Ali Ganjizadeh 1
- Fatemeh Haghighi 1
- Samuel Hwang 1
- Mayank Jobanputra 1
- Meghana Reddy Kasula 1
- Benjamin Khong 1
- Yiming Li 1
- Augustine Luk 1
- Augustine Yui Hei Luk 1
- Zaiqiao Meng 1
- Ashley Murray 1
- Samarth Oza 1
- Andrew Parker 1
- Alyssa Pradhan 1
- Dave Van Veen 1
- Kevin Yuan 1
- Xi Zhang 1