Elena Morandini


2026

This paper proposes a context-first NLP detection pipeline for automated anti-language identification in RICO wiretap transcripts. Current threat-detection classifiers fail on organized crime discourse because criminal intent is encoded through implicature and relexicalization rather than explicit lexical markers. The pipeline addresses this architectural mismatch by formalizing van Dijk’s (2011) socio-cognitive CDA framework as a sequential seven-step decision tree mapped to concrete NLP subtasks: from speaker-role classification and genre detection to deontic feature extraction and ensemble scoring. A six-feature micro-level vector (F1–F6), validated against a 14,072-word corpus of authenticated Mafia communications, operationalizes the ideological square as a two-axis feature space that measures discursive distance between the ingroup and the outgroup. Preliminary evaluation confirms statistically significant patterns (χ² = 90.82, p < 0.001 for pragmatic divergence; 4:1 deontic saturation ratio) consistent with anti-language characteristics. The pipeline enables three LEA applications: automated flagging, context-sensitive decoding, and communication network analysis. Ethical considerations regarding false positives, privacy, and evidentiary standards are discussed.
Search
Co-authors
    Venues
    Fix author