Annette Hautli-Janisz
Papers on this page may belong to the following people: Annette Hautli, Annette Hautli-Janisz
2026
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Lucas Anastasiou | Katarina Boland | Anna De Liddo | Neele Falk | Annette Hautli-Janisz | Gabriella Lapesa | Julia Romberg
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Lucas Anastasiou | Katarina Boland | Anna De Liddo | Neele Falk | Annette Hautli-Janisz | Gabriella Lapesa | Julia Romberg
Proceedings of The 2nd Workshop on Language-driven Deliberation Technology
Proceedings of the 13th Workshop on Argument Mining and Reasoning
Mohamed Elaraby | Annette Hautli-Janisz | Julia Romberg | Elena Musi | Federico Ruggeri | John Lawrence
Proceedings of the 13th Workshop on Argument Mining and Reasoning
Mohamed Elaraby | Annette Hautli-Janisz | Julia Romberg | Elena Musi | Federico Ruggeri | John Lawrence
Proceedings of the 13th Workshop on Argument Mining and Reasoning
Probing Bias Formation in Medical LLMs through Activation Steering
Bayram Ayadi | Annette Hautli-Janisz
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Bayram Ayadi | Annette Hautli-Janisz
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Large Language Models specialized for the medical domain achieve high performance on static benchmarks, but remain vulnerable to sycophantic confabulation, where the models generate medically spurious rationales to justify incorrect user hints. This robustness gap poses severe risks in clinical environments, as models may prioritize contextual faithfulness to a biased prompt over their internal parametric medical knowledge. This study introduces a mechanistic approach to identify and mitigate these failures in MedGemma-27B, isolating hint integration circuits using Sparse Autoencoders and geometric manifold analysis. Our findings reveal that sycophantic bias is a highly distributed and polymorphic concept, with biased reasoning routed through shifting dimensions across transformer layers. We identify the optimal layer for intervention and demonstrate that cluster-conditioned dynamic steering tailored to the geometric subspace of the prompt outperforms static global interventions, though it reveals a fundamental tension between bias resilience and the retention of internal parametric knowledge. This work proposes a principled framework toward clinical AI systems that are more robust and aligned with expert medical logic, demonstrating the potential of cluster-conditioned geometric interventions while characterizing the inherent trade-offs in clinical knowledge retention.