Alexander M. Fichtl
2026
MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
Pia Chouayfati | Alexander M. Fichtl | Miriam Anschütz | George Doumat | Georg Groh
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Pia Chouayfati | Alexander M. Fichtl | Miriam Anschütz | George Doumat | Georg Groh
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical PatientVector schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the PatientVector schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
Challenging Quadratic Attention - A Holistic View On the Rise of Alternative Language Model Architectures
Alexander M. Fichtl | Jeremias Bohn | Josefin Kelber | Edoardo Mosca | Georg Groh
Proceedings of The Big Picture v2: Crafting a Research Narrative
Alexander M. Fichtl | Jeremias Bohn | Josefin Kelber | Edoardo Mosca | Georg Groh
Proceedings of The Big Picture v2: Crafting a Research Narrative
Transformers have dominated sequence processing tasks for the past seven years—most notably language modeling. However, the inherent quadratic complexity of their attention mechanism remains a significant bottleneck as context length increases. We review and distill the recent efforts to overcome this bottleneck, including advances in (sub-quadratic) attention variants, recurrent neural networks, state space models, and hybrid architectures. We critically analyze approaches regarding compute and memory complexity, benchmark results, and fundamental limitations to assess whether the dominance of pure-attention transformers may soon be challenged, which we consider possible, particularly in domain-specific and edge-device applications.