Antony McCabe


2026

Abstract Automating export control screening with large language models (LLMs) introduces risks of contextual sensitivity to non-technical signals, including partner country identity. This paper presents a reproducible counterfactual framework for diagnosing how partner country information influences LLM compliance decisions across five instruction-tuned models and 249 jurisdictions. Across four controlled experiments, partner identity systematically alters export control classifications even when project content is held constant or removed entirely, with BRICS and GCC partners attracting consistently higher control rates than EU and Five Eyes collaborations. Under causal partner controls, classification deltas become small and statistically non-significant, suggesting that observed disparities reflect contextual cue sensitivity rather than fixed geopolitical bias, though residual associative priors cannot be fully excluded. Model-level analysis reveals qualitatively distinct behavioural patterns, including directionally opposite responses to identical inputs across architectures, indicating that model selection carries consequences beyond accuracy in compliance-sensitive deployments. A structured two-stage prompting protocol reduces decision volatility by 60–80% while preserving interpretability. The proposed framework provides an auditable diagnostic method for evaluating contextual sensitivity in regulated AI systems and offers practical guidance for accountable compliance automation.