Unipa-GPT: A Framework to Assess Open-source Alternatives to Chat-GPT for Italian Chat-bots

Irene Siragusa, Roberto Pirrone


Abstract
This paper illustrates the implementation of Open Unipa-GPT, an open source version of the Unipa-GPT chatbot that leverages on open-source Large Language Models for embeddings and text generation. The system relies on a Retrieval Augmented Generation approach, thus mitigating hallucination errors in the generation phase. A detailed comparison between different models is reported to illustrate their performance as regards embedding generation, retrieval, and text generation. In the last case, models were tested in simple inference setup after a fine-tuning procedure. Experiments demonstrate that an open-source LLMs can be efficiently used for embedding generation, but noon of the models does reach the performances obtained by closed models, such as gpt-3.5-turbo in generating answers.
Anthology ID:
2024.clicit-1.100
Volume:
Proceedings of the 10th Italian Conference on Computational Linguistics (CLiC-it 2024)
Month:
December
Year:
2024
Address:
Pisa, Italy
Editors:
Felice Dell'Orletta, Alessandro Lenci, Simonetta Montemagni, Rachele Sprugnoli
Venue:
CLiC-it
SIG:
Publisher:
CEUR Workshop Proceedings
Note:
Pages:
929–939
Language:
URL:
https://aclanthology.org/2024.clicit-1.100/
DOI:
Bibkey:
Cite (ACL):
Irene Siragusa and Roberto Pirrone. 2024. Unipa-GPT: A Framework to Assess Open-source Alternatives to Chat-GPT for Italian Chat-bots. In Proceedings of the 10th Italian Conference on Computational Linguistics (CLiC-it 2024), pages 929–939, Pisa, Italy. CEUR Workshop Proceedings.
Cite (Informal):
Unipa-GPT: A Framework to Assess Open-source Alternatives to Chat-GPT for Italian Chat-bots (Siragusa & Pirrone, CLiC-it 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.clicit-1.100.pdf