Orlando Amaral Cejas


2026

Short Message Service is a fundamental communication channel in modern telecom networks, yet its ubiquity continues to be exploited for large-scale attacks. Recent advances in multilingual embeddings and large language models have shown strong performance on general text classification tasks. However, their effectiveness and efficiency for multilingual SMS fraud detection under constraints of real telecom environments, remain underexplored. In particular, existing studies largely focus on monolingual datasets, cloud-based inference, or overlook the constraints imposed by real telecom environments. In this work, we investigate large-scale multilingual SMS classification for HAM, SPAM, and SMISHING detection using embedding models. We construct and curate a proprietary multilingual SMS dataset and conduct a systematic evaluation of four different multilingual embedding models. Using the obtained dataset, we fine-tune the models, demonstrating that domain-adapted embeddings significantly improve SMS classification across several languages. Overall, this study addresses the gap between embedding-centric NLP research and real-world telecom requirements, providing empirical and practical insights for effective and efficient deployment of multilingual SMS fraud detection systems. We open source all non-proprietary material at: https://figshare.com/s/1df3ca8d08a4eb4d2712.