Renato O. Miyaji
Author directory2026
Tail Smoothing and Cross-Lingual Volatility: Evaluating Estimative Uncertainty in Large Language Models for Brazilian Portuguese
Renato O. Miyaji | Pedro L. P. Corrêa
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
Renato O. Miyaji | Pedro L. P. Corrêa
Proceedings of the 17th Brazilian Symposium in Information and Human Language Technology
This study evaluates how LLMs interpret Words of Estimative Probability (WEPs) in Brazilian Portuguese compared to English. We translated an English benchmark and compared multilingual models (GPT-5.1, Gemini 3 Flash) against a region-specific model (Sabiá 4). Our findings reveal a “tail smoothing” phenomenon, where models systematically compress extreme probabilities. Notably, while Gemini 3 Flash demonstrated remarkable cross-lingual stability, GPT-5.1 exhibited significant calibration degradation. Counterintuitively, when measured against the English human baseline, Sabiá 4 displayed the most aggressive distribution compression. These results suggest a complex dynamic: either linguistic fine-tuning does not ensure alignment, or it actively captures a culturally specific pragmatic ambiguity in Brazilian Portuguese. This exposes the vulnerabilities and nuances of deploying LLMs in nuanced semantic tasks in Brazilian Portuguese.