Can’t Hide Behind the API: Stealing Black-Box Commercial Embedding Models

Manveer Singh Tamber; Jasper Xian; Jimmy Lin

doi:10.18653/v1/2025.findings-naacl.104

Can’t Hide Behind the API: Stealing Black-Box Commercial Embedding Models

Manveer Singh Tamber, Jasper Xian, Jimmy Lin

Abstract

Embedding models that generate dense vector representations of text are widely used and hold significant commercial value. Companies such as OpenAI and Cohere offer proprietary embedding models via paid APIs, but despite being “hidden” behind APIs, these models are not protected from theft. We present, to our knowledge, the first effort to “steal” these models for retrieval by training thief models on text–embedding pairs obtained from the APIs. Our experiments demonstrate that it is possible to replicate the retrieval effectiveness of commercial embedding models with a cost of under $300. Notably, our methods allow for distilling from multiple teachers into a single robust student model, and for distilling into presumably smaller models with fewer dimension vectors, yet competitive retrieval effectiveness. Our findings raise important considerations for deploying commercial embedding models and suggest measures to mitigate the risk of model theft.

Anthology ID:: 2025.findings-naacl.104
Volume:: Findings of the Association for Computational Linguistics: NAACL 2025
Month:: April
Year:: 2025
Address:: Albuquerque, New Mexico
Editors:: Luis Chiruzzo, Alan Ritter, Lu Wang
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1958–1969
Language:
URL:: https://aclanthology.org/2025.findings-naacl.104/
DOI:: 10.18653/v1/2025.findings-naacl.104
Bibkey:
Cite (ACL):: Manveer Singh Tamber, Jasper Xian, and Jimmy Lin. 2025. Can’t Hide Behind the API: Stealing Black-Box Commercial Embedding Models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 1958–1969, Albuquerque, New Mexico. Association for Computational Linguistics.
Cite (Informal):: Can’t Hide Behind the API: Stealing Black-Box Commercial Embedding Models (Tamber et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-naacl.104.pdf

PDF Cite Search Fix data