Matthijs S. Berends

Author directory

2026

Clinical language models are typically pretrained with self-supervised objectives whose geometry reflects linguistic co-occurrence rather than clinical knowledge structure. For downstream tasks that operate directly on the representation space, without task-specific fine-tuning, this gap limits what the model can do. We introduce DOKTERBERT (Dutch Ontology-grounded Knowledge-injected Text Encoder for Representations using BERT), a Dutch clinical language model pre-trained with a structure-aware contrastive objective that aligns contextual span representations to SNOMED concept anchors, with negative pressure weighted by graph distance in the SNOMED hierarchy. We evaluate DOKTERBERT against three Dutch baselines (RobBERT, MedRoBERTa.nl, and MedRoBERTa.nl-SapBERT) through supervised named entity recognition on MultiClinNER-nl and an unsupervised representation analysis spanning retrieval, clustering, entity linking, and concept-level separation. On supervised NER, all four models perform comparably; on the representational evaluations, DOKTERBERT separates from every baseline. Standard fine-tuning evaluation obscures pre-training-level differences in representation quality that representation analysis exposes, and these differences matter for clinical applications that depend on embedding geometry.