Fatou Sow


2026

The lack of sufficiently large French resources linking discourse connectives to the relations they express in context hinders the training of models that map connectives to their discourse relations. We address this gap by introducing two complementary datasets: a large-scale, semi-automatically annotated corpus and a manually validated dataset. Both resources annotate connectives and their discourse relations according to the French lexicon LEXCONN. Relying on the large-scale corpus, we train Relex, a CamemBERT-based model fine-tuned to predict, among 19 relation types, the relation expressed by a connective. Despite being trained on a fixed inventory of connectives, Relex extends to previously unseen connectives and achieves an F1 score of 0.59.

2025

Les marqueurs discursifs sont des éléments linguistiques qui peuvent être employés pour construire la cohérence d’un discours car ils expriment les relations entre les unités discursives. Ils constituent ainsi des indices utiles pour la résolution de problèmes de traitement de langue en rapport avec la sémantique du texte, le discours ou la compréhension de systèmes. Dans cet article, nous présentons un état de l’art des marqueurs discursifs en traitement automatique des langues (TAL). Nous introduisons les représentations textuelles des marqueurs discursifs puis nous nous intéressons à la détection des marqueurs et l’utilisation de leurs sens pour améliorer ou évaluer des tâches de TAL.

2024

One of the biggest hurdles for the effective analysis of data collected on social platforms is the need for deeper insights on the content and meaning of this data. Emotion annotation can bring new perspectives on this issue and can enable the identification of content–specific features. This study aims at investigating the ways in which variation in online content can be explored through emotion annotation and corpus-based analysis. The paper describes the emotion annotation of three data sets in French composed of extremist, sexist and hateful messages respectively. To this end, first a fine-grained, corpus annotation scheme was used to annotate the data sets and then several empirical studies were carried out to characterize the content in the light of emotional categories. Results suggest that emotion annotations can provide new insights for online content analysis and stronger empirical background for automatic content detection.