Priscilla Adenuga


2026

This paper presents a manually annotated morpho-syntactic corpus of Ògè, an under-resourced indigenous language spoken in Nigeria. The corpus consists of ten folk narratives (approximately 4,667 tokens) collected for the investigation of nominal structure. Annotation is expert-driven and includes token-level part-of-speech tagging together with a structured Determiner Phrase (DP) classification framework designed to capture language-specific nominal configurations. The scheme distinguishes between bare nouns and modified noun phrases, reflecting a central structural property of Ògè: noun forms remain morphologically stable across contexts, while modifiers exhibit formal and positional variation contributing to reference, specificity, and discourse prominence. The DP classification layer encodes both simple and complex nominal constructions, enabling systematic analysis of internal phrase structure. Designed as a reusable digital resource, the corpus supports morphosyntactic tagging, noun phrase boundary detection, and modeling of nominal structure in low-resource NLP settings. The annotated dataset will be made publicly available through the SADiLaR repository. This work demonstrates how descriptive linguistic analysis can inform annotation design and provides a replicable framework for developing structured resources for under-resourced African languages. Keywords: Ògè, low-resource NLP, annotated corpus, nominal structure, African languages
Large Language Models (LLMs) are increasingly deployed across a wide range of applications, from conversational assistants to decision support systems. However, these systems remain vulnerable to prompt injection attacks, in which carefully crafted inputs manipulate model behavior and circumvent intended safeguards. While existing research has largely approached prompt injection as a technical or security problem, the linguistic mechanisms through which such attacks operate remain insufficiently understood. In this paper, we argue that prompt injection attacks are fundamentally linguistic in nature, exploiting discourse structure, pragmatic framing, and instruction hierarchies encoded in natural language prompts. Drawing on concepts from speech act theory, discourse analysis, and pragmatics, we propose a typology of four linguistic strategies used to manipulate Large Language Models: instruction override, role framing, hypothetical framing, and procedural prompting. Through detailed linguistic analysis of representative examples, we demonstrate how each strategy exploits identifiable properties of natural language interaction to reshape model interpretation and influence output generation. For each strategy, we also discuss implications for detection and mitigation, arguing that effective safeguards must attend to discourse-level patterns in prompts rather than relying solely on surface-level keyword filtering. Our findings contribute to emerging research at the intersection of computational linguistics and AI security and highlight the importance of integrating linguistic expertise into the design of more robust and reliable language-based AI systems.