CoreELM: An Open-Source Framework for Aligning Large Language Models to Embedding Spaces

Brian David Ondov; Chia-Hsuan Chang; Yujia Zhou; Mauro Giuffrè; Hua Xu

CoreELM: An Open-Source Framework for Aligning Large Language Models to Embedding Spaces

Brian Ondov, Chia-Hsuan Chang, Yujia Zhou, Mauro Giuffrè, Hua Xu

Abstract

Text embeddings have become an essential part of a variety of language applications. However, methods for interpreting, exploring and reversing embedding spaces are limited, reducing transparency and precluding potentially valuable generative use cases. In this work, we develop an open-source, domain-agnostic framework for aligning Large Language Models to embedding spaces using the recently reported Embedding Language Model (ELM) method. We demonstrate our framework by training models to recover, summarize, and compare clinical trial abstracts from embeddings alone. In addition to inverting embeddings back to text more reliably than existing methods, our models can decode novel, interpolated embeddings into new clinical trial abstracts that human experts cannot distinguish from real ones. We further show that these generated abstracts are responsive to moving embeddings along concept vectors for age and sex of study subjects. Our public ELM implementation and experimental results will aid the alignment of Large Language Models to embedding spaces in the biomedical domain and beyond.

Anthology ID:: 2026.bionlp-1.15
Volume:: BioNLP 2026
Month:: July
Year:: 2026
Address:: San Diego, California
Editors:: Dina Demner-Fushman, Sophia Ananiadou, Kirk Roberts, Junichi Tsujii
Venues:: BioNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 156–180
Language:
URL:: https://aclanthology.org/2026.bionlp-1.15/
DOI:
Bibkey:
Cite (ACL):: Brian Ondov, Chia-Hsuan Chang, Yujia Zhou, Mauro Giuffrè, and Hua Xu. 2026. CoreELM: An Open-Source Framework for Aligning Large Language Models to Embedding Spaces. In BioNLP 2026, pages 156–180, San Diego, California. Association for Computational Linguistics.
Cite (Informal):: CoreELM: An Open-Source Framework for Aligning Large Language Models to Embedding Spaces (Ondov et al., BioNLP 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.bionlp-1.15.pdf

PDF Cite Search Fix data