Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges

Daniel Schlör, Marius Bohn, Maximilian Wolf, Kevin Bergner, Christian Goldschmied, Andreas Hotho


Abstract
Evaluating Large Language Model (LLM) agents on complex multi-step cybersecurity tasks requires structured, reproducible evaluation rubrics. We present BraceGreen, a framework that formalizes Capture-the-Flag (CTF) attack paths as knowledge graphs and uses them as gold-standard rubrics for agentic LLM evaluation. Each node in our knowledge graphs represents an attack step annotated with MITRE ATT&CK tactics, goals, commands, expected outputs, and semantic outcomes, while edges encode prerequisites, dependencies, and alternative paths. Our LangGraph-based evaluation workflow employs LLM-as-judge with chain-of-thought reasoning to semantically compare agent predictions against knowledge graph-encoded alternatives. We contribute a benchmark of 7 CTF machines with knowledge graph annotations, three evaluation modes (command prediction, goal inference, anticipated result), and integration with live machine infrastructure via virtual machines and a MCP server. Our approach bridges the gap between unstructured CTF writeups and graph-structured evaluation rubrics.
Anthology ID:
2026.kallm-1.15
Volume:
Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Gilles Sérasset, Katerina Gkirtzou, Michael Cochez, Jan-Christoph Kalo
Venues:
KaLLM | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
144–154
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-kgllm-15
DOI:
10.63317/4ne74wm45iti
Bibkey:
Cite (ACL):
Daniel Schlör, Marius Bohn, Maximilian Wolf, Kevin Bergner, Christian Goldschmied, and Andreas Hotho. 2026. Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges. In Proceedings of the Knowledge Graphs and Large Language Models Workshop (KG-LLM) @ LREC26, pages 144–154, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Towards Knowledge Graph-Grounded Evaluation of Agentic LLMs on Cybersecurity Capture-the-Flag Challenges (Schlör et al., KaLLM 2026)
Copy Citation: