Jan Štěpánek
2026
Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies
Marie Mikulová | Barbora Štěpánková | Daniel Zeman | Jan Štěpánek | Milan Straka | Jan Hajič
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Marie Mikulová | Barbora Štěpánková | Daniel Zeman | Jan Štěpánek | Milan Straka | Jan Hajič
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the “Prague Dependency Treebank-Consolidated” (PDT-C). In comparison to the original PDT, PDT-C is more than twice as large, but it is also much more diverse in terms of genres and domains. In this paper, we describe the conversion of the new resource to Universal Dependencies. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion. We argue that while PDT is less “universal” and more tightly bound to one language, its multi-layer annotation is rich and provides all information needed for basic UD trees, and much more.
Prague Dependency Treebank - Consolidated 2.0: Enriching a Complex Annotation Scheme
Marie Mikulová | Jiří Mírovský | Milan Straka | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Marie Mikulová | Jiří Mírovský | Milan Straka | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
Proceedings of the Fifteenth Language Resources and Evaluation Conference
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relation. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
Semantic-pragmatic Annotations in the Prague Dependency Treebank
Marie Mikulová | Eva Hajicova | Jiří Mírovský | Anna Nedoluzhko | Michal Novák | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
Findings of the Association for Computational Linguistics: ACL 2026
Marie Mikulová | Eva Hajicova | Jiří Mírovský | Anna Nedoluzhko | Michal Novák | Pavlína Synková | Jan Štěpánek | Barbora Štěpánková | Jan Hajič
Findings of the Association for Computational Linguistics: ACL 2026
We present semantic-pragmatic specification and annotation (ellipsis, coreference, bridging and discourse relations, information structure, scope of negation) in the multi-layer, genre-diversified, 3+ million-token Prague Dependency Treebank – Consolidated 2. 0. While morphology and syntax work almost exclusively on sentence level, the semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context. In the contribution, we describe these phenomena from both the linguistic perspective (form of expression, relation to syntax and morphology) and the cognitive perspective (relation to context, real world knowledge, as well as to the related processes such as thinking or reasoning) – classifying the possible relations between the semantic-pragmatic units into cognitively plausible, distinguishable, and human-understandable categories. We have applied our results to the corpus, by annotating it in its entirety. The resulting dataset is publicly and freely available, to serve for verification and further investigation of (not only) these phenomena.
Towards Consistent UMR Annotation of Deverbal Nouns: Evidence from Czech and Latin
Hana Hledíková | Federica Gamba | Marketa Lopatkova | Jan Štěpánek
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Hana Hledíková | Federica Gamba | Marketa Lopatkova | Jan Štěpánek
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Deverbal nouns pose challenges for semantic annotation frameworks that aim to represent event structures consistently across lexical categories. This paper examines problematic phenomena in the annotation of deverbal nouns in Czech and Latin within the Universal Meaning Representation (UMR) framework, addressing both manual graph construction and rule-based automatic conversion from existing resources. Current UMR guidelines lack operational criteria for deciding when a noun should be treated as an eventive concept, particularly in the absence of a PropBank-like lexicon with sufficient nominal coverage. We therefore propose practical annotation principles: deverbal nouns denoting events (such as učení ‘teaching’), results of events (řešení ‘solution’), or event participants (učitel ‘teacher’) should be related to underlying event concepts (represented as verbs in their particular senses, i.e., učit-001 ‘to teach’, vyřešit-001 ‘to solve’, and učit-001 ‘to teach’, respectively), while other deverbal nouns should remain unrelated to respective events (such as učebna ‘teaching room’). To reduce inter-annotator variation, we further suggest systematic strategies for selecting verbal labels, including the use of light-verb constructions, synonymous verbs, and a preference for imperfective verbs in Czech aspectual pairs. For automatic conversion, we outline a rule-based approach that combines multiple lexical resources and frequency-based heuristics to identify corresponding verb senses. Our findings provide guidelines for more consistent UMR annotation across languages.
Meaning Annotation Experience. A Tribute to Petr Sgall
Marie Mikulová | Jan Štěpánek | Barbora Štěpánková | Jarmila Panevova | Eva Hajicova
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Marie Mikulová | Jan Štěpánek | Barbora Štěpánková | Jarmila Panevova | Eva Hajicova
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
We present ongoing work on annotating fine-grained semantic distinctions for circumstantial meanings, focusing on spatial expressions. We describe our theoretical background, and annotation process, as well as how we evaluate the results obtained. Using multiple independent annotations across the 3-million-token, genre-diverse Prague Dependency Treebank – Consolidated corpus of Czech data, we analyse inter-annotator agreement, recurrent disagreement patterns, and the limits of semantic categorization. Our results highlight the inherent vagueness of linguistic meaning. We also propose strategies for handling disagreement, such as weighted annotations, intermediate labels, and fuzzy labels that preserve annotation nuance. This work builds on the legacy of Petr Sgall and the Functional Generative Description theory that underpins the multi-layer form–meaning framework.
First Shared Task on UMR Parsing
Jan Štěpánek | Daniel Zeman | Marketa Lopatkova | Federica Gamba | Hana Hledíková | Nianwen Xue
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
Jan Štěpánek | Daniel Zeman | Marketa Lopatkova | Federica Gamba | Hana Hledíková | Nianwen Xue
Proceedings of The Seventh International Workshop on Designing Meaning Representations (DMR 2026) @ LREC 2026
The paper presents the first shared task on parsing Uniform Meaning Representation (UMR), a graph-based framework for cross-linguistic semantic annotation of typologically diverse languages. The task requires systems to enrich plain text with sentence-level structure, node–token alignment, and document-level relations. It involves processing data for seven languages from four language families (Indo-European, Sino-Tibetan, Na-Dene, and Algic). Six languages have at least some training data; for one language, data is not available, leading to a zero-shot scenario. The training dataset as well as the gold-standard test set for all seven languages is released and made available for follow-up research. We present the task setup and evaluation methodology, using two graph matching approaches – a traditional, and an alignment-sensitive one, tailored specifically for UMR. Two participating systems are compared, each representing different modeling approaches. Results highlight the challenges of UMR parsing, particularly for alignment prediction and document-level semantics, and reveal substantial variation across languages and annotation conditions.
2025
Label Bias in Symbolic Representation of Meaning
Marie Mikulová | Jan Štěpánek | Jan Hajič
Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025)
Marie Mikulová | Jan Štěpánek | Jan Hajič
Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025)
This paper contributes to the trend of building semantic representations and exploring the relations between a language and the world it represents. We analyse alternative approaches to semantic representation, focusing on methodology of determining meaning categories, their arrangement and granularity, and annotation consistency and reliability. Using the task of semantic classification of circumstantial meanings within the Prague Dependency Treebank framework, we present our principles for analyzing meaning categories. Compared with the discussed projects, the unique aspect of our approach is its focus on how a language, in its structure, reflects reality. We employ a two-level classification: a higher, coarse-grained set of general semantic concepts (defined by questions: where, how, why, etc.) and a fine-grained set of circumstantial meanings based on data-driven analysis, reflecting meanings fixed in the language. We highlight that the inherent vagueness of linguistic meaning is crucial for capturing the limitless variety of the world but it can lead to label biases in datasets. Therefore, besides semantically clear categories, we also use fuzzy meaning categories.
Comparing Manual and Automatic UMRs for Czech and Latin
Jan Štěpánek | Daniel Zeman | Markéta Lopatková | Federica Gamba | Hana Hledíková
Proceedings of the Sixth International Workshop on Designing Meaning Representations
Jan Štěpánek | Daniel Zeman | Markéta Lopatková | Federica Gamba | Hana Hledíková
Proceedings of the Sixth International Workshop on Designing Meaning Representations
Uniform Meaning Representation (UMR) is a semantic framework designed to represent the meaning of texts in a structured and interpretable manner. In this paper, we evaluate the results of the automatic conversion of existing resources to UMR, focusing on Czech (PDT-C treebank) and Latin (LDT treebank). We present both quantitative and qualitative evaluations based on a comparison between manually and automatically generated UMR structures for a sample of Czech and Latin sentences. The findings indicate comparable results of the automatic conversion for both languages. The key challenges prove to be the higher level of semantic abstraction required by UMR and the fact that UMR allows for capturing semantic structure in multiple ways, potentially with varying levels of granularity.
From Form to Meaning: The Case of Particles within the Prague Dependency Treebank Annotation Scheme
Marie Mikulova | Barbora Štěpánková | Jan Štěpánek
Proceedings of the 31st International Conference on Computational Linguistics
Marie Mikulova | Barbora Štěpánková | Jan Štěpánek
Proceedings of the 31st International Conference on Computational Linguistics
In the last decades, computational linguistics has become increasingly interested in annotation schemes that aim at an adequate description of the meaning of the sentences and texts. Discussions are ongoing on an appropriate annotation scheme for a large and complex amount of diverse information. In this contribution devoted to description of polyfunctional uninflected words (namely particles), i.e. words which, although having only one paradigmatic form, can have several different syntactic functions and even express relatively different semantic distinctions, we argue that it is the multi-layer system (linked from meaning to text) that allows a comprehensive description of the relations between morphological properties, syntactic function and expressed meaning, and thus contributes to greater accuracy in the description of the phenomena concerned and to the overall consistency of the annotated data. These aspects are demonstrated within the Prague Dependency Treebank annotation scheme, whose pioneering proposal can be found in the first COLING proceedings from 1965 (Sgall 1965), and to this day, the concept has proved to be sound and serves very well for complex annotation.
2022
Quality and Efficiency of Manual Annotation: Pre-annotation Bias
Marie Mikulová | Milan Straka | Jan Štěpánek | Barbora Štěpánková | Jan Hajic
Proceedings of the Thirteenth Language Resources and Evaluation Conference
Marie Mikulová | Milan Straka | Jan Štěpánek | Barbora Štěpánková | Jan Hajic
Proceedings of the Thirteenth Language Resources and Evaluation Conference
This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task - dependency syntax annotation. It compares the annotation efforts made by annotators using a pre-annotated version (with a high-accuracy parser) and those made by fully manual annotation. The aim of the experiment is to judge the final annotation quality when pre-annotation is used. In addition, it evaluates the effect of automatic linguistically-based (rule-formulated) checks and another annotation on the same data available to the annotators, and their influence on annotation quality and efficiency. The experiment confirmed that the pre-annotation is an efficient tool for faster manual syntactic annotation which increases the consistency of the resulting annotation without reducing its quality.
2020
Prague Dependency Treebank - Consolidated 1.0
Jan Hajič | Eduard Bejček | Jaroslava Hlavacova | Marie Mikulová | Milan Straka | Jan Štěpánek | Barbora Štěpánková
Proceedings of the Twelfth Language Resources and Evaluation Conference
Jan Hajič | Eduard Bejček | Jaroslava Hlavacova | Marie Mikulová | Milan Straka | Jan Štěpánek | Barbora Štěpánková
Proceedings of the Twelfth Language Resources and Evaluation Conference
We present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0 (PDT-C 1.0), the purpose of which is - as it always been the case for the family of the Prague Dependency Treebanks - to serve both as a training data for various types of NLP tasks as well as for linguistically-oriented research. PDT-C 1.0 contains four different datasets of Czech, uniformly annotated using the standard PDT scheme (albeit not everything is annotated manually, as we describe in detail here). The texts come from different sources: daily newspaper articles, Czech translation of the Wall Street Journal, transcribed dialogs and a small amount of user-generated, short, often non-standard language segments typed into a web translator. Altogether, the treebank contains around 180,000 sentences with their morphological, surface and deep syntactic annotation. The diversity of the texts and annotations should serve well the NLP applications as well as it is an invaluable resource for linguistic research, including comparative studies regarding texts of different genres. The corpus is publicly and freely available.
2016
Searching in the Penn Discourse Treebank Using the PML-Tree Query
Jiří Mírovský | Lucie Poláková | Jan Štěpánek
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Jiří Mírovský | Lucie Poláková | Jan Štěpánek
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
The PML-Tree Query is a general, powerful and user-friendly system for querying richly linguistically annotated treebanks. The paper shows how the PML-Tree Query can be used for searching for discourse relations in the Penn Discourse Treebank 2.0 mapped onto the syntactic annotation of the Penn Treebank.
2013
Coordination Structures in Dependency Treebanks
Martin Popel | David Mareček | Jan Štěpánek | Daniel Zeman | Zdeněk Žabokrtský
Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Martin Popel | David Mareček | Jan Štěpánek | Daniel Zeman | Zdeněk Žabokrtský
Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2012
Prague Markup Language Framework
Jirka Hana | Jan Štěpánek
Proceedings of the Sixth Linguistic Annotation Workshop
Jirka Hana | Jan Štěpánek
Proceedings of the Sixth Linguistic Annotation Workshop
Announcing Prague Czech-English Dependency Treebank 2.0
Jan Hajič | Eva Hajičová | Jarmila Panevová | Petr Sgall | Ondřej Bojar | Silvie Cinková | Eva Fučíková | Marie Mikulová | Petr Pajas | Jan Popelka | Jiří Semecký | Jana Šindlerová | Jan Štěpánek | Josef Toman | Zdeňka Urešová | Zdeněk Žabokrtský
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Jan Hajič | Eva Hajičová | Jarmila Panevová | Petr Sgall | Ondřej Bojar | Silvie Cinková | Eva Fučíková | Marie Mikulová | Petr Pajas | Jan Popelka | Jiří Semecký | Jana Šindlerová | Jan Štěpánek | Josef Toman | Zdeňka Urešová | Zdeněk Žabokrtský
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
We introduce a substantial update of the Prague Czech-English Dependency Treebank, a parallel corpus manually annotated at the deep syntactic layer of linguistic representation. The English part consists of the Wall Street Journal (WSJ) section of the Penn Treebank. The Czech part was translated from the English source sentence by sentence. This paper gives a high level overview of the underlying linguistic theory (the so-called tectogrammatical annotation) with some details of the most important features like valency annotation, ellipsis reconstruction or coreference.
HamleDT: To Parse or Not to Parse?
Daniel Zeman | David Mareček | Martin Popel | Loganathan Ramasamy | Jan Štěpánek | Zdeněk Žabokrtský | Jan Hajič
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Daniel Zeman | David Mareček | Martin Popel | Loganathan Ramasamy | Jan Štěpánek | Zdeněk Žabokrtský | Jan Hajič
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
We propose HamleDT ― HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. While the license terms prevent us from directly redistributing the corpora, most of them are easily acquirable for research purposes. What we provide instead is the software that normalizes tree structures in the data obtained by the user from their original providers.
Prague Dependency Treebank 2.5 – a Revisited Version of PDT 2.0
Eduard Bejček | Jarmila Panevová | Jan Popelka | Pavel Straňák | Magda Ševčíková | Jan Štěpánek | Zdeněk Žabokrtský
Proceedings of COLING 2012
Eduard Bejček | Jarmila Panevová | Jan Popelka | Pavel Straňák | Magda Ševčíková | Jan Štěpánek | Zdeněk Žabokrtský
Proceedings of COLING 2012
2010
Ways of Evaluation of the Annotators in Building the Prague Czech-English Dependency Treebank
Marie Mikulová | Jan Štěpánek
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Marie Mikulová | Jan Štěpánek
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
In this paper, we present several ways to measure and evaluate the annotation and annotators, proposed and used during the building of the Czech part of the Prague Czech-English Dependency Treebank. At first, the basic principles of the treebank annotation project are introduced (division to three layers: morphological, analytical and tectogrammatical). The main part of the paper describes in detail one of the important phases of the annotation process: three ways of evaluation of the annotators - inter-annotator agreement, error rate and performance. The measuring of the inter-annotator agreement is complicated by the fact that the data contain added and deleted nodes, making the alignment between annotations non-trivial. The error rate is measured by a set of automatic checking procedures that guard the validity of some invariants in the data. The performance of the annotators is measured by a booking web application. All three measures are later compared and related to each other.
Querying Diverse Treebanks in a Uniform Way
Jan Štěpánek | Petr Pajas
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Jan Štěpánek | Petr Pajas
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
This paper presents a system for querying treebanks in a uniform way. The system is able to work with both dependency and constituency based treebanks in any language. We demonstrate its abilities on 11 different treebanks. The query language used by the system provides many features not available in other existing systems while still keeping the performance efficient. The paper also describes the conversion of ten treebanks into a common XML-based format used by the system, touching the question of standards and formats. The paper then shows several examples of linguistically interesting questions that the system is able to answer, for example browsing verbal clauses without subjects or extraposed relative clauses, generating the underlying grammar in a constituency treebank, searching for non-projective edges in a dependency treebank, or word-order typology of a language based on the treebank. The performance of several implementations of the system is also discussed by measuring the time requirements of some of the queries.
2009
The CoNLL-2009 Shared Task: Syntactic and Semantic Dependencies in Multiple Languages
Jan Hajič | Massimiliano Ciaramita | Richard Johansson | Daisuke Kawahara | Maria Antònia Martí | Lluís Màrquez | Adam Meyers | Joakim Nivre | Sebastian Padó | Jan Štěpánek | Pavel Straňák | Mihai Surdeanu | Nianwen Xue | Yi Zhang
Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL 2009): Shared Task
Jan Hajič | Massimiliano Ciaramita | Richard Johansson | Daisuke Kawahara | Maria Antònia Martí | Lluís Màrquez | Adam Meyers | Joakim Nivre | Sebastian Padó | Jan Štěpánek | Pavel Straňák | Mihai Surdeanu | Nianwen Xue | Yi Zhang
Proceedings of the Thirteenth Conference on Computational Natural Language Learning (CoNLL 2009): Shared Task
System for Querying Syntactically Annotated Corpora
Petr Pajas | Jan Štěpánek
Proceedings of the ACL-IJCNLP 2009 Software Demonstrations
Petr Pajas | Jan Štěpánek
Proceedings of the ACL-IJCNLP 2009 Software Demonstrations
2008
Search
Fix author
Co-authors
- Marie Mikulová 10
- Jan Hajic 9
- Daniel Zeman 5
- Petr Pajas 4
- Milan Straka 4
- Barbora Štěpánková 4
- Zdeněk Žabokrtský 4
- Federica Gamba 3
- Eva Hajicova 3
- Hana Hledíková 3
- Marketa Lopatkova 3
- Jiří Mírovský 3
- Jarmila Panevová 3
- Barbora Štěpánková 3
- Eduard Bejček 2
- David Mareček 2
- Martin Popel 2
- Jan Popelka 2
- Pavel Straňák 2
- Pavlína Synková 2
- Nianwen Xue 2
- Ondřej Bojar 1
- Massimiliano Ciaramita 1
- Silvie Cinkova 1
- Eva Fucikova 1
- Jirka Hana 1
- Jaroslava Hlaváčová 1
- Richard Johansson 1
- Daisuke Kawahara 1
- M. Antònia Martí 1
- Adam Meyers 1
- Lluís Màrquez 1
- Anna Nedoluzhko 1
- Joakim Nivre 1
- Michal Novák 1
- Sebastian Padó 1
- Lucie Polakova 1
- Loganathan Ramasamy 1
- Jiří Semecký 1
- Petr Sgall 1
- Mihai Surdeanu 1
- Josef Toman 1
- Zdenka Uresova 1
- Yi Zhang 1
- Magda Ševčíková 1
- Jana Šindlerová 1