Markus Löffler
2026
Developing the German Medical Text Corpus (GeMTeX): Legal Compliance and Semantic Enrichment
Justin Hofenbitzer | Christina Lohr | Andrea Riedel | Rebekka Kiser | Aliaksandra Shutsko | Abanoub Abdelmalak | Peter Klügl | Jutta Romberg | Sarah Riepenhausen | Miriam Schechner | Jakob Faller | Frank Meineke | Luise Modersohn | Markus Löffler | Juliane Fluck | Udo Hahn | Stefan Schulz | Martin Boeker
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Justin Hofenbitzer | Christina Lohr | Andrea Riedel | Rebekka Kiser | Aliaksandra Shutsko | Abanoub Abdelmalak | Peter Klügl | Jutta Romberg | Sarah Riepenhausen | Miriam Schechner | Jakob Faller | Frank Meineke | Luise Modersohn | Markus Löffler | Juliane Fluck | Udo Hahn | Stefan Schulz | Martin Boeker
Proceedings of the Fifteenth Language Resources and Evaluation Conference
GeMTeX is a large-scale German Medical Text Corpus project with the goal to publish a clinical national reference corpus. The resource is currently under construction and comprises, as of February 2026, more than 15k clinical documents (20M tokens) from six German university hospitals. When building GeMTeX, attention was paid to comply with European regulatory requirements. In phase I, patients were asked to allow reuse of their clinical documents based on the legal foundation of an “informed consent”. In phase II, consented documents from six major clinical sites in Germany underwent a thorough de-identification process. In phase III, we currently enrich this unlocked dataset with semantic information from the clinical domain. This annotation process is guided by Snomed CT, which supports to directly ground expressions within clinical documents in a worldwide shared medical documentation and ontology standard. The resource is currently under active development and is accessible upon request under controlled access conditions. We refer interested researchers to visit https://kiinformatik.mri.tum.de/en/gemtex or reach out via gemtex.mi@mh.tum.de.
The German Medical Text Corpus: Early 2026 Update
Justin Hofenbitzer | Christina Lohr | Frank Meineke | Markus Löffler | Martin Boeker
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Justin Hofenbitzer | Christina Lohr | Frank Meineke | Markus Löffler | Martin Boeker
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Clinical text resources are a central component for the study of medical language, as well as the training and evaluation of large language models, chatbots, and artificial intelligence systems supporting clinical routines. With the German Medical Text Corpus (GeMTeX), we are currently working on the largest shareable clinical document dataset in German. The multi-centric project ensures diversity across different university hospitals, clinical domains, and text sorts. After a thorough de-identification process, the clinical texts are semantically annotated using Snomed CT, a language-independent, standardized medical ontology. While the corpus is still under active development, it is accessible upon request under controlled access conditions. As of February 2026, GeMTeX comprises more than 15k documents and 20M tokens. We refer researchers interested in the resource to visit https://kiinformatik.mri.tum.de/en/gemtex or reach out to us via gemtex.mi@mh.tum.de.