Barbora Štěpánková
Author directoryOther people with similar names: Barbora Štěpánková
Unverified author pages with similar names: Barbora Štěpánková
2026
Command-Line Obfuscation Detection in Real-World Telemetry under Extreme Class Imbalance
Vojtěch Outrata | Barbora Štěpánková | Michael Adam Polák | Martin Kopp
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Vojtěch Outrata | Barbora Štěpánková | Michael Adam Polák | Martin Kopp
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
To avoid detection by endpoint security tools, adversaries employ command-line obfuscation to alter syntax while preserving functionality. This paper proposes a scalable detection method specifically for command-line data, centered on a custom-trained, small transformer-based model optimized for low-latency inference across massive data streams. We demonstrate the method’s efficacy through a two-phase evaluation: first, by benchmarking the model on a controlled dataset simulating realistic command-line telemetry with extreme class imbalance, where it outperforms previous approaches. Second, we evaluate the model against multiple days of high-volume telemetry from diverse real-world environments. Our results show that this approach provides the high precision and computational efficiency required to handle large-scale command-line logs while effectively reducing analyst workload.
Semantic Clustering of Obfuscated Command-Line Detections for Alert Reduction
Barbora Štěpánková | Vojtěch Outrata | Martin Kopp
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Barbora Štěpánková | Vojtěch Outrata | Martin Kopp
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
We propose a post-processing method for grouping large volumes of command-line detections into semantically coherent cluster-level alerts. The approach combines embedding-based clustering with LLM-based cluster-level filtering: command-lines are first encoded using Sentence-BERT embeddings and grouped via hierarchical agglomerative clustering, after which a large language model evaluates cluster representatives to reduce the number of false positive alerts. We evaluate the method on real-world telemetry from a commercial endpoint protection system, applying it to detections produced by an obfuscation detection model. On one week of data, the pipeline reduces alert volume by approximately 98% while maintaining high cluster purity and semantic coherence, demonstrating its effectiveness as a scalable post-processing step in high-volume detection settings.
2025
Song Lyrics Adaptations: Computational Interpretation of the Pentathlon Principle
Barbora Štěpánková | Rudolf Rosa
Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities
Barbora Štěpánková | Rudolf Rosa
Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities
Songs are an integral part of human culture, and they often resonate the most when we can sing them in our native language. However, translating song lyrics presents a unique challenge: maintaining singability, naturalness, and semantic fidelity. In this work, we computationally interpret Low’s Pentathlon Principle of singable translations to be able to properly measure the quality of adapted lyrics, breaking it down into five measurable metrics that reflect the key aspects of singable translations. Building on this foundation, we introduce a text-to-text song lyrics translation system based on generative large language models, designed to meet the Pentathlon Principle’s criteria, without relying on melodies or bilingual training data.We experiment on the English-Czech language pair: we collect a dataset of English-to-Czech bilingual song lyrics and identify the desirable values of the five Pentathlon Principle metrics based on the values achieved by human translators. Through detailed human assessment of automatically generated lyric translations, we confirm the appropriateness of the proposed metrics as well as the general validity of the Pentathlon Principle, with some insights into the variation in people’s individual preferences. All code and data are available at https://github.com/stepankovab/Computational-Interpretation-of-the-Pentathlon-Principle.