Ana García-Serrano
Also published as: Ana M. García-Serrano, Ana Garcia-Serrano
2026
Data Matters: Looking for High-Quality Corpora to Build Robust and Reliable Models for Humanists
Jaione Macicior-Mitxelena | Ana García-Serrano
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Jaione Macicior-Mitxelena | Ana García-Serrano
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
The digitization of Spanish historical newspapers poses significant challenges due to low scan quality, typographical diversity, complex layouts and linguistic variation from contemporary Spanish. While advances in Optical Character Recognition (OCR) and layout-aware models offer promising results, their effectiveness strongly depends on the quality and consistency of the underlying training corpora. This work focuses on corpus construction and evaluation for historical document processing. Two experiments were conducted. In the first corpus los101 was used, a manually curated and structurally annotated subcorpus derived from historical Spanish newspapers, designed to ensure coherent ground truth under heterogeneous real-world conditions. This corpus enables systematic experimentation across OCR and document layout analysis tasks. In a second experimental phase, we apply an additional layout-focused corpus characterized by structural regularity and consistent page organization, allowing us to isolate the impact of layout homogeneity on segmentation performance. State-of-the-art OCR models and a layout detection model are evaluated as validation instruments to assess corpus adequacy rather than as primary contributions. Quantitative and qualitative analyses based on (1) relationship between annotation quality, (2) structural variability, and (3) model behavior, show that heterogeneous corpora challenge both transcription and segmentation stability, while layout-consistent data significantly improves structural detection reliability.
The 7th Financial Narrative Processing Workshop
Mo El-Haj | Antonio Moreno Sandoval | Ana Garcia-Serrano | Chung-Chi Chen | Paul Rayson | Yanco Amor Torterolo Orta | Paloma Martinez | Jordi Porta
The 7th Financial Narrative Processing Workshop
Mo El-Haj | Antonio Moreno Sandoval | Ana Garcia-Serrano | Chung-Chi Chen | Paul Rayson | Yanco Amor Torterolo Orta | Paloma Martinez | Jordi Porta
The 7th Financial Narrative Processing Workshop
2010
Q-WordNet: Extracting Polarity from WordNet Senses
Rodrigo Agerri | Ana García-Serrano
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Rodrigo Agerri | Ana García-Serrano
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
This paper presents Q-WordNet, a lexical resource consisting of WordNet senses automatically annotated by positive and negative polarity. Polarity classification amounts to decide whether a text (sense, sentence, etc.) may be associated to positive or negative connotations. Polarity classification is becoming important within the fields of Opinion Mining and Sentiment Analysis for determining opinions about commercial products, on companies reputation management, brand monitoring, or to track attitudes by mining online forums, blogs, etc. Inspired by work on classification of word senses by polarity (e.g., SentiWordNet), and taking WordNet as a starting point, we build Q-WordNet. Instead of applying external tools such as supervised classifiers to annotated WordNet synsets by polarity, we try to effectively maximize the linguistic information contained in WordNet, thereby taking advantage of the human effort put by lexicographers and annotators. The resulting resource is a subset of WordNet senses classified as positive or negative. In this approach, neutral polarity is seen as the absence of positive or negative polarity. The evaluation of Q-WordNet shows an improvement with respect to previous approaches. We believe that Q-WordNet can be used as a starting point for data-driven approaches in sentiment analysis.
2002
Natural Language Dialogue in a Virtual Assistant Interface
Ana M. García-Serrano | Luis Rodrigo-Aguado | Javier Calle
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Ana M. García-Serrano | Luis Rodrigo-Aguado | Javier Calle
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Integrating Spanish Linguistic Resources in a Web Site Assistant
Paloma Martínez | Ana García-Serrano | Alberto Ruiz-Cristina
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)
Paloma Martínez | Ana García-Serrano | Alberto Ruiz-Cristina
Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02)