Other Workshops and Events (2012)
Volumes
- Proceedings of the EACL 2012 Joint Workshop of LINGVIS & UNCLH 17 papers
- Proceedings of the Second Workshop on Computational Linguistics and Writing (CL&W 2012): Linguistic and Cognitive Aspects of Document Creation and Document Engineering 7 papers
- Proceedings of the Workshop on Computational Approaches to Deception Detection 16 papers
- Proceedings of the Workshop on Innovative Hybrid Approaches to the Processing of Textual Data 16 papers
- Proceedings of the Workshop on Semantic Analysis in Social Media 9 papers
- Proceedings of the Joint Workshop on Unsupervised and Semi-Supervised Learning in NLP 8 papers
- Proceedings of the Workshop on Computational Models of Language Acquisition and Loss 13 papers
- JEP-TALN-RECITAL 2012, Workshop DEFT 2012: DÉfi Fouille de Textes (DEFT 2012 Workshop: Text Mining Challenge) 10 papers
- JEP-TALN-RECITAL 2012, Workshop DEGELS 2012: Défi GEste Langue des Signes (DEGELS 2012: Gestures and Sign Language Challenge) 9 papers
- JEP-TALN-RECITAL 2012, Workshop TALAf 2012: Traitement Automatique des Langues Africaines (TALAf 2012: African Language Processing) 11 papers
- JEP-TALN-RECITAL 2012, Workshop ILADI 2012: Interactions Langagières pour personnes Agées Dans les habitats Intelligents (ILADI 2012: Language Interaction for Elderly in Smart Homes) 7 papers
- NAACL-HLT Workshop on Future directions and needs in the Spoken Dialog Community: Tools and Data (SDCTD 2012) 20 papers
- Proceedings of the NAACL-HLT Workshop on the Induction of Linguistic Structure 16 papers
- Proceedings of the Second Workshop on Language in Social Media 10 papers
- Proceedings of Workshop on Evaluation Metrics and System Comparison for Automatic Summarization 7 papers
- Proceedings of the NAACL-HLT 2012 Workshop: Will We Ever Really Replace the N-gram Model? On the Future of Language Modeling for HLT 8 papers
- Proceedings of the Second Workshop on Semantic Interpretation in an Actionable Context 3 papers
- Proceedings of the ACL-2012 Special Workshop on Rediscovering 50 Years of Discoveries 14 papers
- Proceedings of the 1st Workshop on Speech and Multimodal Interaction in Assistive Environments 8 papers
- Proceedings of the First Workshop on Multilingual Modeling 5 papers
- Proceedings of the 3rd Workshop on the People’s Web Meets NLP: Collaboratively Constructed Semantic Resources and their Applications to NLP 7 papers
- Proceedings of the Workshop on Detecting Structure in Scholarly Discourse 7 papers
- Proceedings of the Workshop on Advances in Discourse Analysis and its Computational Aspects 6 papers
- Proceedings of the Second Workshop on Advances in Text Input Methods 11 papers
- Proceedings of the First Workshop on Eye-tracking and Natural Language Processing 7 papers
- Proceedings of the 2nd Workshop on Sentiment Analysis where AI meets Psychology 13 papers
- Proceedings of the Workshop on Information Extraction and Entity Analytics on Social Media Data 5 papers
- Proceedings of the Workshop on Machine Translation and Parsing in Indian Languages 21 papers
- Proceedings of the Second Workshop on Applying Machine Learning Techniques to Optimise the Division of Labour in Hybrid MT 10 papers
- Proceedings of the Workshop on Speech and Language Processing Tools in Education 13 papers
- Proceedings of the Workshop on Reordering for Statistical Machine Translation 7 papers
- Proceedings of the Workshop on Question Answering for Complex Domains 7 papers
- Proceedings of the First International Workshop on Optimization Techniques for Human Language Technology 9 papers
- Workshop on Post-Editing Technology and Practice 11 papers
- Fourth Workshop on Computational Approaches to Arabic-Script-based Languages 11 papers
- Workshop on Monolingual Machine Translation 8 papers
- Proceedings of the Joint Workshop on Exploiting Synergies between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra) HyTra 18 papers
- Proceedings of the Workshop on Applications of Tree Automata Techniques in Natural Language Processing ATANLP 5 papers
- Proceedings of the 6th Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities LaTeCH 17 papers
- Proceedings of the 3rd Workshop on Cognitive Modeling and Computational Linguistics (CMCL 2012) CMCL 9 papers
- Proceedings of the Seventh Workshop on Building Educational Applications Using NLP BEA 40 papers
- Proceedings of the First Workshop on Predicting and Improving Text Readability for target reader populations PITR 9 papers
- Proceedings of the Twelfth Meeting of the Special Interest Group on Computational Morphology and Phonology SIGMORPHON 10 papers
- BioNLP: Proceedings of the 2012 Workshop on Biomedical Natural Language Processing BioNLP 31 papers
- Proceedings of the NAACL-HLT 2012 Workshop on Computational Linguistics for Literature CLFL 15 papers
- Proceedings of the Third Workshop on Speech and Language Processing for Assistive Technologies SLPAT 11 papers
- Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction (AKBC-WEKEX) AKBC 24 papers
- Proceedings of the Seventh Workshop on Statistical Machine Translation WMT 61 papers
- Proceedings of the ACL 2012 Joint Workshop on Statistical Parsing and Semantic Processing of Morphologically Rich Languages SPMRL 13 papers
- Proceedings of the Sixth Linguistic Annotation Workshop LAW 27 papers
- Proceedings of the 3rd Workshop in Computational Approaches to Subjectivity and Sentiment Analysis WASSA 18 papers
- Proceedings of the Workshop on Extra-Propositional Aspects of Meaning in Computational Linguistics EXprom 11 papers
- Workshop Proceedings of TextGraphs-7: Graph-based Methods for Natural Language Processing TextGraphs 9 papers
- Proceedings of the Sixth Workshop on Syntax, Semantics and Structure in Statistical Translation SSST 14 papers
- Proceedings of the 4th Named Entity Workshop (NEWS) 2012 NEWS 13 papers
- Proceedings of the 11th International Workshop on Tree Adjoining Grammars and Related Formalisms (TAG+11) TAG+ 28 papers
- Proceedings of the 3rd Workshop on South and Southeast Asian Natural Language Processing WSSANLP 23 papers
- Proceedings of the 3rd Workshop on Cognitive Aspects of the Lexicon CogALex 18 papers
- Proceedings of the 10th Workshop on Asian Language Resources ALR 15 papers
- Proceedings of the 10th International Workshop on Finite State Methods and Natural Language Processing FSMNLP 20 papers
- Proceedings of the Second CIPS-SIGHAN Joint Conference on Chinese Language Processing SIGHAN 42 papers
- Proceedings of the Third International Workshop on Free/Open-Source Rule-Based Machine Translation FreeOpMT 8 papers
- Proceedings of the 9th International Workshop on Spoken Language Translation: Keynotes IWSLT 3 papers
- Proceedings of the 9th International Workshop on Spoken Language Translation: Evaluation Campaign IWSLT 20 papers
- Proceedings of the 9th International Workshop on Spoken Language Translation: Papers IWSLT 20 papers
- Proceedings of the Australasian Language Technology Association Workshop 2012 ALTA 21 papers
- Proceedings of ACL 2012 Student Research Workshop ACL 13 papers
- Joint Conference on EMNLP and CoNLL - Shared Task EMNLP CoNLL 17 papers
- INLG 2012 Proceedings of the Seventh International Natural Language Generation Conference INLG 29 papers
- Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue SIGDIAL 44 papers
up
Proceedings of the EACL 2012 Joint Workshop of LINGVIS & UNCLH
Proceedings of the EACL 2012 Joint Workshop of LINGVIS & UNCLH
Miriam Butt | Sheelagh Carpendale | Gerald Penn | Jelena Prokić | Michael Cysouw
Miriam Butt | Sheelagh Carpendale | Gerald Penn | Jelena Prokić | Michael Cysouw
Lexical Semantics and Distribution of Suffixes - A Visual Analysis
Christian Rohrdantz | Andreas Niekler | Annette Hautli | Miriam Butt | Daniel A. Keim
Christian Rohrdantz | Andreas Niekler | Annette Hautli | Miriam Butt | Daniel A. Keim
Looking at word meaning. An interactive visualization of Semantic Vector Spaces for Dutch synsets
Kris Heylen | Dirk Speelman | Dirk Geeraerts
Kris Heylen | Dirk Speelman | Dirk Geeraerts
First steps in checking and comparing Princeton WordNet and Estonian Wordnet
Ahti Lohk | Kadri Vare | Leo Võhandu
Ahti Lohk | Kadri Vare | Leo Võhandu
Visualising Typological Relationships: Plotting WALS with Heat Maps
Richard Littauer | Rory Turnbull | Alexis Palmer
Richard Littauer | Rory Turnbull | Alexis Palmer
Automating Second Language Acquisition Research: Integrating Information Visualisation and Machine Learning
Helen Yannakoudakis | Ted Briscoe | Theodora Alexopoulou
Helen Yannakoudakis | Ted Briscoe | Theodora Alexopoulou
Visualising Linguistic Evolution in Academic Discourse
Verena Lyding | Ekaterina Lapshinova-Koltunski | Stefania Degaetano-Ortlieb | Henrik Dittmann | Chris Culy
Verena Lyding | Ekaterina Lapshinova-Koltunski | Stefania Degaetano-Ortlieb | Henrik Dittmann | Chris Culy
Estimating and visualizing language similarities using weighted alignment and force-directed graph layout
Gerhard Jäger
Gerhard Jäger
Tracking the dynamics of kinship and social category terms with AustKin II
Patrick McConvell | Laurent Dousset
Patrick McConvell | Laurent Dousset
up
Proceedings of the Second Workshop on Computational Linguistics and Writing (CL&W 2012): Linguistic and Cognitive Aspects of Document Creation and Document Engineering
Proceedings of the Second Workshop on Computational Linguistics and Writing (CL&W 2012): Linguistic and Cognitive Aspects of Document Creation and Document Engineering
Michael Piotrowski | Cerstin Mahlow | Robert Dale
Michael Piotrowski | Cerstin Mahlow | Robert Dale
From Character to Word Level: Enabling the Linguistic Analyses of Inputlog Process Data
Mariëlle Leijten | Lieve Macken | Veronique Hoste | Eric Van Horenbeeck | Luuk Van Waes
Mariëlle Leijten | Lieve Macken | Veronique Hoste | Eric Van Horenbeeck | Luuk Van Waes
From Drafting Guideline to Error Detection: Automating Style Checking for Legislative Texts
Stefan Höfler | Kyoko Sugisaki
Stefan Höfler | Kyoko Sugisaki
up
Proceedings of the Workshop on Computational Approaches to Deception Detection
Proceedings of the Workshop on Computational Approaches to Deception Detection
Eileen Fitzpatrick | Joan Bachenko | Tommaso Fornaciari
Eileen Fitzpatrick | Joan Bachenko | Tommaso Fornaciari
Linguistic Cues to Deception Assessed by Computer Programs: A Meta-Analysis
Valerie Hauch | Iris Blandón-Gitlin | Jaume Masip | Siegfried Ludwig Sporer
Valerie Hauch | Iris Blandón-Gitlin | Jaume Masip | Siegfried Ludwig Sporer
“I Don’t Know Where He is Not”: Does Deception Research yet Offer a Basis for Deception Detectives?
Anna Vartapetiance | Lee Gillam
Anna Vartapetiance | Lee Gillam
Seeing through Deception: A Computational Approach to Deceit Detection in Written Communication
Ángela Almela | Rafael Valencia-García | Pascual Cantos
Ángela Almela | Rafael Valencia-García | Pascual Cantos
In Search of a Gold Standard in Studies of Deception
Stephanie Gokhman | Jeff Hancock | Poornima Prabhu | Myle Ott | Claire Cardie
Stephanie Gokhman | Jeff Hancock | Poornima Prabhu | Myle Ott | Claire Cardie
On the Use of Homogenous Sets of Subjects in Deceptive Language Analysis
Tommaso Fornaciari | Massimo Poesio
Tommaso Fornaciari | Massimo Poesio
Invited Talk: Current and Future Needs for Deception Detection in a Government Screening Environment
Daniel Baxter
Daniel Baxter
The Voice and Eye Gaze Behavior of an Imposter: Automated Interviewing and Detection for Rapid Screening at the Border
Aaron Elkins | Douglas Derrick | Monica Gariup
Aaron Elkins | Douglas Derrick | Monica Gariup
Argument Formation in the Reasoning Process: Toward a Generic Model of Deception Detection
Deqing Li | Eugene Santos
Deqing Li | Eugene Santos
Pastiche Detection Based on Stopword Rankings. Exposing Impersonators of a Romanian Writer
Liviu P. Dinu | Vlad Niculae | Maria-Octavia Sulea
Liviu P. Dinu | Vlad Niculae | Maria-Octavia Sulea
Making the Subjective Objective? Computer-Assisted Quantification of Qualitative Content Cues to Deception
Siegfried Ludwig Sporer
Siegfried Ludwig Sporer
up
Proceedings of the Workshop on Innovative Hybrid Approaches to the Processing of Textual Data
Proceedings of the Workshop on Innovative Hybrid Approaches to the Processing of Textual Data
Natalia Grabar | Marie Dupuch | Amandine Périnet | Thierry Hamon
Natalia Grabar | Marie Dupuch | Amandine Périnet | Thierry Hamon
Experiments on Hybrid Corpus-Based Sentiment Lexicon Acquisition
Goran Glavaš | Jan Šnajder | Bojana Dalbelo Bašić
Goran Glavaš | Jan Šnajder | Bojana Dalbelo Bašić
A Study of Hybrid Similarity Measures for Semantic Relation Extraction
Alexander Panchenko | Olga Morozova
Alexander Panchenko | Olga Morozova
Hybrid Combination of Constituency and Dependency Trees into an Ensemble Dependency Parser
Nathan Green | Zdeněk Žabokrtský
Nathan Green | Zdeněk Žabokrtský
An Unsupervised and Data-Driven Approach for Spell Checking in Vietnamese OCR-scanned Texts
Cong Duy Vu Hoang | Ai Ti Aw
Cong Duy Vu Hoang | Ai Ti Aw
Contrasting Objective and Subjective Portuguese Texts from Heterogeneous Sources
Michel Généreux | William Martinez
Michel Généreux | William Martinez
A Joint Named Entity Recognition and Entity Linking System
Rosa Stern | Benoît Sagot | Frédéric Béchet
Rosa Stern | Benoît Sagot | Frédéric Béchet
Collaborative Annotation of Dialogue Acts: Application of a New ISO Standard to the Switchboard Corpus
Alex C. Fang | Harry Bunt | Jing Cao | Xiaoyue Liu
Alex C. Fang | Harry Bunt | Jing Cao | Xiaoyue Liu
Coupling Knowledge-Based and Data-Driven Systems for Named Entity Recognition
Damien Nouvel | Jean-Yves Antoine | Nathalie Friburger | Arnaud Soulet
Damien Nouvel | Jean-Yves Antoine | Nathalie Friburger | Arnaud Soulet
A Random Forest System Combination Approach for Error Detection in Digital Dictionaries
Michael Bloodgood | Peng Ye | Paul Rodrigues | David Zajic | David Doermann
Michael Bloodgood | Peng Ye | Paul Rodrigues | David Zajic | David Doermann
Methods Combination and ML-based Re-ranking of Multiple Hypothesis for Question-Answering Systems
Arnaud Grappy | Brigitte Grau | Sophie Rosset
Arnaud Grappy | Brigitte Grau | Sophie Rosset
up
Proceedings of the Workshop on Semantic Analysis in Social Media
Unsupervised Part-of-Speech Tagging in Noisy and Esoteric Domains With a Syntactic-Semantic Bayesian HMM
William M. Darling | Michael J. Paul | Fei Song
William M. Darling | Michael J. Paul | Fei Song
Towards Scalable Speech Act Recognition in Twitter: Tackling Insufficient Training Data
Renxian Zhang | Dehong Gao | Wenjie Li
Renxian Zhang | Dehong Gao | Wenjie Li
A User and NLP-Assisted Strategic Workflow for a Social Semantic OWL 2-Based Knowledge Platform
Jinan El-Hachem | Volker Haarslev
Jinan El-Hachem | Volker Haarslev
up
Proceedings of the Joint Workshop on Unsupervised and Semi-Supervised Learning in NLP
Proceedings of the Joint Workshop on Unsupervised and Semi-Supervised Learning in NLP
Omri Abend | Chris Biemann | Anna Korhonen | Ari Rappoport | Roi Reichart | Anders Søgaard
Omri Abend | Chris Biemann | Anna Korhonen | Ari Rappoport | Roi Reichart | Anders Søgaard
Fast Unsupervised Dependency Parsing with Arc-Standard Transitions
Mohammad Sadegh Rasooli | Heshaam Faili
Mohammad Sadegh Rasooli | Heshaam Faili
Dependency-Based Open Information Extraction
Pablo Gamallo | Marcos Garcia | Santiago Fernández-Lanza
Pablo Gamallo | Marcos Garcia | Santiago Fernández-Lanza
Improving Distantly Supervised Extraction of Drug-Drug and Protein-Protein Interactions
Tamara Bobić | Roman Klinger | Philippe Thomas | Martin Hofmann-Apitius
Tamara Bobić | Roman Klinger | Philippe Thomas | Martin Hofmann-Apitius
up
Proceedings of the Workshop on Computational Models of Language Acquisition and Loss
Proceedings of the Workshop on Computational Models of Language Acquisition and Loss
Robert Berwick | Anna Korhonen | Thierry Poibeau | Aline Villavicencio
Robert Berwick | Anna Korhonen | Thierry Poibeau | Aline Villavicencio
Distinguishing Contact-Induced Change from Language Drift in Genetically Related Languages
T. Mark Ellison | Luisa Miceli
T. Mark Ellison | Luisa Miceli
A Morphologically Annotated Hebrew CHILDES Corpus
Aviad Albert | Brian MacWhinney | Bracha Nir | Shuly Wintner
Aviad Albert | Brian MacWhinney | Bracha Nir | Shuly Wintner
An annotated English child language database
Aline Villavicencio | Beracah Yankama | Rodrigo Wilkens | Marco Idiart | Robert Berwick
Aline Villavicencio | Beracah Yankama | Rodrigo Wilkens | Marco Idiart | Robert Berwick
Unseen features. Collecting semantic data from congenital blind subjects
Alessandro Lenci | Marco Baroni | Giovanna Marotta
Alessandro Lenci | Marco Baroni | Giovanna Marotta
I say have you say tem: profiling verbs in children data in English and Portuguese
Rodrigo Wilkens | Aline Villavicencio
Rodrigo Wilkens | Aline Villavicencio
up
JEP-TALN-RECITAL 2012, Workshop DEFT 2012: DÉfi Fouille de Textes (DEFT 2012 Workshop: Text Mining Challenge)
JEP-TALN-RECITAL 2012, Workshop DEFT 2012: DÉfi Fouille de Textes (DEFT 2012 Workshop: Text Mining Challenge)
Cyril Grouin | Dominic Forest | Gilles Sérasset
Cyril Grouin | Dominic Forest | Gilles Sérasset
Indexation libre et contrôlée d’articles scientifiques. Présentation et résultats du défi fouille de textes DEFT2012 (Controlled and free indexing of scientific papers. Presentation and results of the DEFT2012 text-mining challenge) [in French]
Patrick Paroubek | Pierre Zweigenbaum | Dominic Forest | Cyril Grouin
Patrick Paroubek | Pierre Zweigenbaum | Dominic Forest | Cyril Grouin
Acquisition terminologique pour identifier les mots-clés d’articles scientifiques (Terminological acquisition for identifying keywords of scientific articles) [in French]
Thierry Hamon
Thierry Hamon
Indexation à base des syntagmes nominaux (Nominal-chunk based indexing) [in French]
Amine Amri | Maroua Mbarek | Chedi Bechikh | Chiraz Latiri | Hatem Haddad
Amine Amri | Maroua Mbarek | Chedi Bechikh | Chiraz Latiri | Hatem Haddad
Détection de mots-clés par approches au grain caractère et au grain mot (Keywords extraction by repeated string analysis) [in French]
Gaëlle Doualan | Mathieu Boucher | Romain Brixtel | Gaël Lejeune | Gaël Dias
Gaëlle Doualan | Mathieu Boucher | Romain Brixtel | Gaël Lejeune | Gaël Dias
Participation de l’IRISA à DeFT2012 : recherche d’information et apprentissage pour la génération de mots-clés (IRISA participation to DeFT2012: information retrieval and machine-learning for keyword generation) [in French]
Vincent Claveau | Christian Raymond
Vincent Claveau | Christian Raymond
Participation du LINA à DEFT2012 (LINA at DEFT2012) [in French]
Florian Boudin | Amir Hazem | Nicolas Hernandez | Prajol Shrestha
Florian Boudin | Amir Hazem | Nicolas Hernandez | Prajol Shrestha
up
JEP-TALN-RECITAL 2012, Workshop DEGELS 2012: Défi GEste Langue des Signes (DEGELS 2012: Gestures and Sign Language Challenge)
JEP-TALN-RECITAL 2012, Workshop DEGELS 2012: Défi GEste Langue des Signes (DEGELS 2012: Gestures and Sign Language Challenge)
Annelies Braffort | Leïla Boutora | Gilles Sérasset
Annelies Braffort | Leïla Boutora | Gilles Sérasset
Défi d’annotation DEGELS2012 : la segmentation (DEGELS2012 annotation challenge: Segmentation) [in French]
Annelies Braffort | Leïla Boutora
Annelies Braffort | Leïla Boutora
Critères de segmentation de la gestualité co-verbale (Segmentation criteria for the annotation of co-speech gestures) [in French]
Gaëlle Ferré
Gaëlle Ferré
Par où couper pour aller à la plage ? (Where do you switch to get the band ?) [in French]
Dominique Boutet | Karine Martel | Marion Blondel
Dominique Boutet | Karine Martel | Marion Blondel
Segmentation et annotation du geste : Méthodologie pour travailler en équipe (Gesture segmentation and coding: a methodology for team working) [in French]
Marion Tellier | Brahim Azaoui | Jorane Saubesty
Marion Tellier | Brahim Azaoui | Jorane Saubesty
Segmenter et annoter le discours d’un locuteur de LSF : permanence formelle et variablité fonctionnelle des unités (Segment and annotate a discourse of a French Sign Language speaker: formal continuity and fonctional variation of units) [in French]
Agnès Millet | Isabelle Estève
Agnès Millet | Isabelle Estève
Influence de la segmentation temporelle sur la caractérisation de signes (Influence of the temporal segmentation on the sign characterization) [in French]
François Lefebvre-Albaret | Jérémie Segouat
François Lefebvre-Albaret | Jérémie Segouat
up
JEP-TALN-RECITAL 2012, Workshop TALAf 2012: Traitement Automatique des Langues Africaines (TALAf 2012: African Language Processing)
JEP-TALN-RECITAL 2012, Workshop TALAf 2012: Traitement Automatique des Langues Africaines (TALAf 2012: African Language Processing)
Chantal Enguehard | Mathieu Mangeot | Gilles Sérasset
Chantal Enguehard | Mathieu Mangeot | Gilles Sérasset
Mbochi : corpus oral, traitement automatique et exploration phonologique (Mboshi: oral corpus, automatic processing & phonological mining) [in French]
Annie Rialland | Martial Embanga Aborobongui | Martine Adda-Decker | Lori Lamel
Annie Rialland | Martial Embanga Aborobongui | Martine Adda-Decker | Lori Lamel
Élaboration d’un dictionnaire bilingue kanouri-français (Construction of the Kanuri-French bilingual dictionary) [in French]
Chérif Ari Abdoulkarim | Arimi Boukar | Kevin Anthony Jarrett | Maï Moussa Maï | Manoua Djibir | Taweye Aïchéta Chégou Koré
Chérif Ari Abdoulkarim | Arimi Boukar | Kevin Anthony Jarrett | Maï Moussa Maï | Manoua Djibir | Taweye Aïchéta Chégou Koré
Vers l’informatisation de quelques langues d’Afrique de l’Ouest (Towards the computerization of some west-african languages) [in French]
Chantal Enguehard | Soumana Kané | Mathieu Mangeot | Issouf Modi | Mamadou Lamine Sanogo
Chantal Enguehard | Soumana Kané | Mathieu Mangeot | Issouf Modi | Mamadou Lamine Sanogo
La transcription phonétique au bout des doigts, claviers et polices ergonomiques pour la transcription en API (Phonetic transcription at fingertips, ergonomics keyboards and fonts) [in French]
Bernard Gautheron | Antonia Simon-Colazo
Bernard Gautheron | Antonia Simon-Colazo
Analyse des performances de modèles de langage sub-lexicale pour des langues peu-dotées à morphologie riche (Performance analysis of sub-word language modeling for under-resourced languages with rich morphology: case study on Swahili and Amharic) [in French]
Hadrien Gelas | Solomon Teferra Abate | Laurent Besacier | François Pellegrino
Hadrien Gelas | Solomon Teferra Abate | Laurent Besacier | François Pellegrino
Règles de formation des noms en hausa (Formation rules of names in Hausa) [in French]
Abdou Mijinguini | Harouna Naroua
Abdou Mijinguini | Harouna Naroua
Vers un analyseur syntaxique du wolof (Towards a syntactic analyzer of Wolof) [in French]
Mar Ndiaye | Cherif Mbodj
Mar Ndiaye | Cherif Mbodj
Formalisation de l’amazighe standard avec NooJ (Formalization of the standard Amazigh with NooJ) [in French]
Fatima Zahra Nejme | Siham Boulaknadel
Fatima Zahra Nejme | Siham Boulaknadel
up
JEP-TALN-RECITAL 2012, Workshop ILADI 2012: Interactions Langagières pour personnes Agées Dans les habitats Intelligents (ILADI 2012: Language Interaction for Elderly in Smart Homes)
JEP-TALN-RECITAL 2012, Workshop ILADI 2012: Interactions Langagières pour personnes Agées Dans les habitats Intelligents (ILADI 2012: Language Interaction for Elderly in Smart Homes)
François Portet | Michel Vacher | Gilles Sérasset
François Portet | Michel Vacher | Gilles Sérasset
Conférence invitée: Nouveaux paradigmes et technologies pour la santé et l’autonomie (Invited Conference: New Paradigms and Technologies for Health and Autonomy) [in French]
Alain Franco
Alain Franco
Les technologies de la parole et du TALN pour l’assistance à domicile des personnes âgées : un rapide tour d’horizon (Quick tour of NLP and speech technologies for ambient assisted living) [in French]
François Portet | Michel Vacher | Solange Rossato
François Portet | Michel Vacher | Solange Rossato
Interactions sonores et vocales dans l’habitat (Acoustic Interaction At Home) [in French]
Pierrick Milhorat | Dan Istrate | Jérôme Boudy | Gérard Chollet
Pierrick Milhorat | Dan Istrate | Jérôme Boudy | Gérard Chollet
Reconnaissance d’ordres domotiques en conditions bruitées pour l’assistance à domicile (Recognition of Voice Commands by Multisource ASR and Noise Cancellation in a Smart Home Environment) [in French]
Benjamin Lecouteux | Michel Vacher | François Portet
Benjamin Lecouteux | Michel Vacher | François Portet
up
NAACL-HLT Workshop on Future directions and needs in the Spoken Dialog Community: Tools and Data (SDCTD 2012)
NAACL-HLT Workshop on Future directions and needs in the Spoken Dialog Community: Tools and Data (SDCTD 2012)
Maxine Eskenazi | Alan Black | David Traum
Maxine Eskenazi | Alan Black | David Traum
Up from Limited Dialog Systems!
Giuseppe Riccardi | Philipp Cimiano | Alexandros Potamianos | Christina Unger
Giuseppe Riccardi | Philipp Cimiano | Alexandros Potamianos | Christina Unger
Position Paper: Towards Standardized Metrics and Tools for Spoken and Multimodal Dialog System Evaluation
Sebastian Möller | Klaus-Peter Engelbrecht | Florian Kretzschmar | Stefan Schmidt | Benjamin Weiss
Sebastian Möller | Klaus-Peter Engelbrecht | Florian Kretzschmar | Stefan Schmidt | Benjamin Weiss
Dialogue Systems Using Online Learning: Beyond Empirical Methods
Heriberto Cuayáhuitl | Nina Dethlefs
Heriberto Cuayáhuitl | Nina Dethlefs
Statistical User Simulation for Spoken Dialogue Systems: What for, Which Data, Which Future?
Olivier Pietquin
Olivier Pietquin
The Future of Spoken Dialogue Systems is in their Past: Long-Term Adaptive, Conversational Assistants
David Schlangen
David Schlangen
Future Directions in Spoken Dialog Systems: A Community of Possibilities
Alan W. Black | Maxine Eskenazi
Alan W. Black | Maxine Eskenazi
Framework for the Development of Spoken Dialogue System based on Collaboratively Constructed Semantic Resources
Masahiro Araki | Daisuke Takegoshi
Masahiro Araki | Daisuke Takegoshi
A Simulation-based Framework for Spoken Language Understanding and Action Selection in Situated Interaction
David Cohen | Ian Lane
David Cohen | Ian Lane
Mining Search Query Logs for Spoken Language Understanding
Dilek Hakkani-Tür | Gokhan Tür | Asli Celikyilmaz
Dilek Hakkani-Tür | Gokhan Tür | Asli Celikyilmaz
HRItk: The Human-Robot Interaction ToolKit Rapid Development of Speech-Centric Interactive Systems in ROS
Ian Lane | Vinay Prasad | Gaurav Sinha | Arlette Umuhoza | Shangyu Luo | Akshay Chandrashekaran | Antoine Raux
Ian Lane | Vinay Prasad | Gaurav Sinha | Arlette Umuhoza | Shangyu Luo | Akshay Chandrashekaran | Antoine Raux
up
Proceedings of the NAACL-HLT Workshop on the Induction of Linguistic Structure
Proceedings of the NAACL-HLT Workshop on the Induction of Linguistic Structure
Trevor Cohn | Phil Blunsom | Joao Graca
Trevor Cohn | Phil Blunsom | Joao Graca
Unsupervised Induction of Frame-Semantic Representations
Ashutosh Modi | Ivan Titov | Alexandre Klementiev
Ashutosh Modi | Ivan Titov | Alexandre Klementiev
Transferring Frames: Utilization of Linked Lexical Resources
Lars Borin | Markus Forsberg | Richard Johansson | Kristiina Muhonen | Tanja Purtonen | Kaarlo Voionmaa
Lars Borin | Markus Forsberg | Richard Johansson | Kristiina Muhonen | Tanja Purtonen | Kaarlo Voionmaa
Capitalization Cues Improve Dependency Grammar Induction
Valentin I. Spitkovsky | Hiyan Alshawi | Daniel Jurafsky
Valentin I. Spitkovsky | Hiyan Alshawi | Daniel Jurafsky
Toward Tree Substitution Grammars with Latent Annotations
Francis Ferraro | Benjamin Van Durme | Matt Post
Francis Ferraro | Benjamin Van Durme | Matt Post
Nudging the Envelope of Direct Transfer Methods for Multilingual Named Entity Recognition
Oscar Täckström
Oscar Täckström
Unsupervised Dependency Parsing using Reducibility and Fertility features
David Mareček | Zdeněk Žabokrtský
David Mareček | Zdeněk Žabokrtský
Induction of Linguistic Structure with Combinatory Categorial Grammars
Yonatan Bisk | Julia Hockenmaier
Yonatan Bisk | Julia Hockenmaier
up
Proceedings of the Second Workshop on Language in Social Media
Proceedings of the Second Workshop on Language in Social Media
Sara Owsley Sood | Meenakshi Nagarajan | Michael Gamon
Sara Owsley Sood | Meenakshi Nagarajan | Michael Gamon
Analyzing Urdu Social Media for Sentiments using Transfer Learning with Controlled Translations
Smruthi Mukund | Rohini Srihari
Smruthi Mukund | Rohini Srihari
Detecting Distressed and Non-distressed Affect States in Short Forum Texts
Michael Thaul Lehrman | Cecilia Ovesdotter Alm | Rubén A. Proaño
Michael Thaul Lehrman | Cecilia Ovesdotter Alm | Rubén A. Proaño
A Demographic Analysis of Online Sentiment during Hurricane Irene
Benjamin Mandel | Aron Culotta | John Boulahanis | Danielle Stark | Bonnie Lewis | Jeremy Rodrigue
Benjamin Mandel | Aron Culotta | John Boulahanis | Danielle Stark | Bonnie Lewis | Jeremy Rodrigue
Detecting Influencers in Written Online Conversations
Or Biran | Sara Rosenthal | Jacob Andreas | Kathleen McKeown | Owen Rambow
Or Biran | Sara Rosenthal | Jacob Andreas | Kathleen McKeown | Owen Rambow
up
Proceedings of Workshop on Evaluation Metrics and System Comparison for Automatic Summarization
Proceedings of Workshop on Evaluation Metrics and System Comparison for Automatic Summarization
John M. Conroy | Hoa Trang Dang | Ani Nenkova | Karolina Owczarzak
John M. Conroy | Hoa Trang Dang | Ani Nenkova | Karolina Owczarzak
An Assessment of the Accuracy of Automatic Evaluation in Summarization
Karolina Owczarzak | John M. Conroy | Hoa Trang Dang | Ani Nenkova
Karolina Owczarzak | John M. Conroy | Hoa Trang Dang | Ani Nenkova
Using the Omega Index for Evaluating Abstractive Community Detection
Gabriel Murray | Giuseppe Carenini | Raymond Ng
Gabriel Murray | Giuseppe Carenini | Raymond Ng
Ecological Validity and the Evaluation of Speech Summarization Quality
Anthony McCallum | Cosmin Munteanu | Gerald Penn | Xiaodan Zhu
Anthony McCallum | Cosmin Munteanu | Gerald Penn | Xiaodan Zhu
up
Proceedings of the NAACL-HLT 2012 Workshop: Will We Ever Really Replace the N-gram Model? On the Future of Language Modeling for HLT
Proceedings of the NAACL-HLT 2012 Workshop: Will We Ever Really Replace the N-gram Model? On the Future of Language Modeling for HLT
Bhuvana Ramabhadran | Sanjeev Khudanpur | Ebru Arisoy
Bhuvana Ramabhadran | Sanjeev Khudanpur | Ebru Arisoy
Measuring the Influence of Long Range Dependencies with Neural Network Language Models
Hai Son Le | Alexandre Allauzen | François Yvon
Hai Son Le | Alexandre Allauzen | François Yvon
Large, Pruned or Continuous Space Language Models on a GPU for Statistical Machine Translation
Holger Schwenk | Anthony Rousseau | Mohammed Attik
Holger Schwenk | Anthony Rousseau | Mohammed Attik
Deep Neural Network Language Models
Ebru Arisoy | Tara N. Sainath | Brian Kingsbury | Bhuvana Ramabhadran
Ebru Arisoy | Tara N. Sainath | Brian Kingsbury | Bhuvana Ramabhadran
Unsupervised Vocabulary Adaptation for Morph-based Language Models
André Mansikkaniemi | Mikko Kurimo
André Mansikkaniemi | Mikko Kurimo
up
Proceedings of the Second Workshop on Semantic Interpretation in an Actionable Context
Proceedings of the Second Workshop on Semantic Interpretation in an Actionable Context
Dan Goldwasser | Regina Barzilay | Dan Roth
Dan Goldwasser | Regina Barzilay | Dan Roth
up
Proceedings of the ACL-2012 Special Workshop on Rediscovering 50 Years of Discoveries
Proceedings of the ACL-2012 Special Workshop on Rediscovering 50 Years of Discoveries
Rafael E. Banchs
Rafael E. Banchs
Rediscovering ACL Discoveries Through the Lens of ACL Anthology Network Citing Sentences
Dragomir Radev | Amjad Abu-Jbara
Dragomir Radev | Amjad Abu-Jbara
Towards a Computational History of the ACL: 1980-2008
Ashton Anderson | Dan Jurafsky | Daniel A. McFarland
Ashton Anderson | Dan Jurafsky | Daniel A. McFarland
Discovering Factions in the Computational Linguistics Community
Yanchuan Sim | Noah A. Smith | David A. Smith
Yanchuan Sim | Noah A. Smith | David A. Smith
Extracting glossary sentences from scholarly articles: A comparative evaluation of pattern bootstrapping and deep analysis
Melanie Reiplinger | Ulrich Schäfer | Magdalena Wolska
Melanie Reiplinger | Ulrich Schäfer | Magdalena Wolska
Towards an ACL Anthology Corpus with Logical Document Structure. An Overview of the ACL 2012 Contributed Task
Ulrich Schäfer | Jonathon Read | Stephan Oepen
Ulrich Schäfer | Jonathon Read | Stephan Oepen
Towards High-Quality Text Stream Extraction from PDF. Technical Background to the ACL 2012 Contributed Task
Øyvind Raddum Berg | Stephan Oepen | Jonathon Read
Øyvind Raddum Berg | Stephan Oepen | Jonathon Read
up
Proceedings of the 1st Workshop on Speech and Multimodal Interaction in Assistive Environments
Proceedings of the 1st Workshop on Speech and Multimodal Interaction in Assistive Environments
Dimitra Anastasiou | Desislava Zhekova | Cui Jian | Robert Ross
Dimitra Anastasiou | Desislava Zhekova | Cui Jian | Robert Ross
Multimodal Human-Machine Interaction for Service Robots in Home-Care Environments
Stefan Goetze | Sven Fischer | Niko Moritz | Jens-E. Appell | Frank Wallhoff
Stefan Goetze | Sven Fischer | Niko Moritz | Jens-E. Appell | Frank Wallhoff
Integration of Multimodal Interaction as Assistance in Virtual Environments
Kiran Pala | Ram Naresh | Sachin Joshi | Suryakanth V Ganagshetty
Kiran Pala | Ram Naresh | Sachin Joshi | Suryakanth V Ganagshetty
Toward a Virtual Assistant for Vulnerable Users: Designing Careful Interaction
Ramin Yaghoubzadeh | Stefan Kopp
Ramin Yaghoubzadeh | Stefan Kopp
Speech and Gesture Interaction in an Ambient Assisted Living Lab
Dimitra Anastasiou | Cui Jian | Desislava Zhekova
Dimitra Anastasiou | Cui Jian | Desislava Zhekova
Reduction of Non-stationary Noise for a Robotic Living Assistant using Sparse Non-negative Matrix Factorization
Benjamin Cauchi | Stefan Goetze | Simon Doclo
Benjamin Cauchi | Stefan Goetze | Simon Doclo
up
Proceedings of the First Workshop on Multilingual Modeling
Proceedings of the First Workshop on Multilingual Modeling
Jagadeesh Jagarlamudi | Sujith Ravi | Xiaojun Wan | Hal Daume III
Jagadeesh Jagarlamudi | Sujith Ravi | Xiaojun Wan | Hal Daume III
Implementing a Language-Independent MT Methodology
Sokratis Sofianopoulos | Marina Vassiliou | George Tambouratzis
Sokratis Sofianopoulos | Marina Vassiliou | George Tambouratzis
Language Independent Named Entity Identification using Wikipedia
Mahathi Bhagavatula | Santosh GSK | Vasudeva Varma
Mahathi Bhagavatula | Santosh GSK | Vasudeva Varma
up
Proceedings of the 3rd Workshop on the People’s Web Meets NLP: Collaboratively Constructed Semantic Resources and their Applications to NLP
Proceedings of the 3rd Workshop on the People’s Web Meets NLP: Collaboratively Constructed Semantic Resources and their Applications to NLP
Iryna Gurevych | Nicoletta Calzolari Zamorani | Jungi Kim
Iryna Gurevych | Nicoletta Calzolari Zamorani | Jungi Kim
Sentiment Analysis Using a Novel Human Computation Game
Claudiu-Cristian Musat | Alireza Ghasemi | Boi Faltings
Claudiu-Cristian Musat | Alireza Ghasemi | Boi Faltings
Collaboratively Building Language Resources while Localising the Web
Asanka Wasala | Reinhard Schäler | Ruvan Weerasinghe | Chris Exton
Asanka Wasala | Reinhard Schäler | Ruvan Weerasinghe | Chris Exton
up
Proceedings of the Workshop on Detecting Structure in Scholarly Discourse
Proceedings of the Workshop on Detecting Structure in Scholarly Discourse
Antal Van Den Bosch | Hagit Shatkay
Antal Van Den Bosch | Hagit Shatkay
Identifying Comparative Claim Sentences in Full-Text Scientific Articles
Dae Hoon Park | Catherine Blake
Dae Hoon Park | Catherine Blake
Open-domain Anatomical Entity Mention Detection
Tomoko Ohta | Sampo Pyysalo | Jun’ichi Tsujii | Sophia Ananiadou
Tomoko Ohta | Sampo Pyysalo | Jun’ichi Tsujii | Sophia Ananiadou
up
Proceedings of the Workshop on Advances in Discourse Analysis and its Computational Aspects
Proceedings of the Workshop on Advances in Discourse Analysis and its Computational Aspects
Eva Hajičová | Lucie Poláková | Jiří Mírovský
Eva Hajičová | Lucie Poláková | Jiří Mírovský
Exploiting Discourse Relations between Sentences for Text Clustering
Nik Adilah Hanin Binti Zahri | Fumiyo Fukumoto | Suguru Matsuyoshi
Nik Adilah Hanin Binti Zahri | Fumiyo Fukumoto | Suguru Matsuyoshi
up
Proceedings of the Second Workshop on Advances in Text Input Methods
Proceedings of the Second Workshop on Advances in Text Input Methods
Kalika Bali | Monojit Choudhury | Yoh Okuno
Kalika Bali | Monojit Choudhury | Yoh Okuno
An Ensemble Model of Word-based and Character-based Models for Japanese and Chinese Input Method
Yoh Okuno | Shinsuke Mori
Yoh Okuno | Shinsuke Mori
Multi-objective Optimization for Efficient Brahmic Keyboards
Albert Brouillette | Devraj Sarmah | Jugal Kalita
Albert Brouillette | Devraj Sarmah | Jugal Kalita
Using Collocations and K-means Clustering to Improve the N-pos Model for Japanese IME
Long Chen | Xianchao Wu | Jingzhou He
Long Chen | Xianchao Wu | Jingzhou He
phloat : Integrated Writing Environment for ESL learners
Yuta Hayashibe | Masato Hagiwara | Satoshi Sekine
Yuta Hayashibe | Masato Hagiwara | Satoshi Sekine
Bangla Phonetic Input Method with Foreign Words Handling
Khan Md. Anwarus Salam | Setsuo Yamada | Tetsuro Nishino
Khan Md. Anwarus Salam | Setsuo Yamada | Tetsuro Nishino
up
Proceedings of the First Workshop on Eye-tracking and Natural Language Processing
Proceedings of the First Workshop on Eye-tracking and Natural Language Processing
Michael Carl | Pushpak Bhattacharyya | Kamal Kumar Choudhary
Michael Carl | Pushpak Bhattacharyya | Kamal Kumar Choudhary
Identifying instances of processing effort in translation through heat maps: an eye-tracking study using multiple input sources
Fabio Alves | José Luiz Gonçalves | Karina Szpak
Fabio Alves | José Luiz Gonçalves | Karina Szpak
Robustness and processing difficulty models. A pilot study for eye-tracking data on the French Treebank
Stéphane Rauzy | Philippe Blache
Stéphane Rauzy | Philippe Blache
Scanpaths in reading are informative about sentence processing
Titus von der Malsburg | Shravan Vasishth | Reinhold Kliegl
Titus von der Malsburg | Shravan Vasishth | Reinhold Kliegl
up
Proceedings of the 2nd Workshop on Sentiment Analysis where AI meets Psychology
Proceedings of the 2nd Workshop on Sentiment Analysis where AI meets Psychology
Sivaji Bandyopadhyay | Manabu Okumura
Sivaji Bandyopadhyay | Manabu Okumura
Categorical Probability Proportion Difference (CPPD): A Feature Selection Method for Sentiment Classification
Basant Agarwal | Namita Mittal
Basant Agarwal | Namita Mittal
Classification of Interviews - A Case Study on Cancer Patients
Braja Gopal Patra | Amitava Kundu | Dipankar Das | Sivaji Bandyopadhyay
Braja Gopal Patra | Amitava Kundu | Dipankar Das | Sivaji Bandyopadhyay
Analyzing Sentiment Word Relations with Affect, Judgment, and Appreciation
Alena Neviarouskaya | Masaki Aono
Alena Neviarouskaya | Masaki Aono
Emotiphons: Emotion Markers in Conversational Speech - Comparison across Indian Languages
Nandini Bondale | Thippur Sreenivas
Nandini Bondale | Thippur Sreenivas
up
Proceedings of the Workshop on Information Extraction and Entity Analytics on Social Media Data
Proceedings of the Workshop on Information Extraction and Entity Analytics on Social Media Data
Sriram Raghavan | Ganesh Ramakrishnan
Sriram Raghavan | Ganesh Ramakrishnan
up
Proceedings of the Workshop on Machine Translation and Parsing in Indian Languages
Proceedings of the Workshop on Machine Translation and Parsing in Indian Languages
Dipti Misra Sharma | Prashanth Mannem | Joseph vanGenabith | Sobha Lalitha Devi | Radhika Mamidi | Ranjani Parthasarathi
Dipti Misra Sharma | Prashanth Mannem | Joseph vanGenabith | Sobha Lalitha Devi | Radhika Mamidi | Ranjani Parthasarathi
Sublexical Translations for Low-Resource Language
Khan Md. Anwarus Salam | Setsuo Yamada | Tetsuro Nishino
Khan Md. Anwarus Salam | Setsuo Yamada | Tetsuro Nishino
A Diagnostic Evaluation Approach Targeting MT Systems for Indian Languages
Renu Balyan | Sudip Kumar Naskar | Antonio Toral | Niladri Chatterjee
Renu Balyan | Sudip Kumar Naskar | Antonio Toral | Niladri Chatterjee
An Approach to Discourse Parsing using Sangati and Rhetorical Structure Theory
Subalalitha C.N. | Ranjani Parthasarathi
Subalalitha C.N. | Ranjani Parthasarathi
Clause Boundary Identification for Malayalam Using CRF
Lakshmi S. | Vijay Sundar Ram R | Sobha Lalitha Devi
Lakshmi S. | Vijay Sundar Ram R | Sobha Lalitha Devi
Disambiguation of pre/post positions in English - Malayalam Text Translation
Jayan V | Sunil R | Bhadran V K
Jayan V | Sunil R | Bhadran V K
Morphological Processing for English-Tamil Statistical Machine Translation
Loganathan Ramasamy | Ondřej Bojar | Zdeněk Žabokrtský
Loganathan Ramasamy | Ondřej Bojar | Zdeněk Žabokrtský
Dative Case in Telugu: A Parsing Perspective
Umamaheshwar Rao Garapati | Rajyarama Koppaka | Srinivas Addanki
Umamaheshwar Rao Garapati | Rajyarama Koppaka | Srinivas Addanki
Two-stage Approach for Hindi Dependency Parsing Using MaltParser
Naman Jain | Karan Singla | Aniruddha Tammewar | Sambhav Jain
Naman Jain | Karan Singla | Aniruddha Tammewar | Sambhav Jain
Hindi Dependency Parsing using a combined model of Malt and MST
B. Venkata Seshu Kumari | Rajeswara Rao Ramisetty
B. Venkata Seshu Kumari | Rajeswara Rao Ramisetty
up
Proceedings of the Second Workshop on Applying Machine Learning Techniques to Optimise the Division of Labour in Hybrid MT
Proceedings of the Second Workshop on Applying Machine Learning Techniques to Optimise the Division of Labour in Hybrid MT
Josef van Genabith | Toni Badia | Christian Federmann | Maite Melero | Marta R. Costa-jussà | Tsuyoshi Okita
Josef van Genabith | Toni Badia | Christian Federmann | Maite Melero | Marta R. Costa-jussà | Tsuyoshi Okita
Hybrid Adaptation of Named Entity Recognition for Statistical Machine Translation
Vassilina Nikoulina | Agnes Sandor | Marc Dymetman
Vassilina Nikoulina | Agnes Sandor | Marc Dymetman
Confusion Network Based System Combination for Chinese Translation Output: Word-Level or Character-Level?
Maoxi Li | MingWen Wang
Maoxi Li | MingWen Wang
Using Cross-Lingual Explicit Semantic Analysis for Improving Ontology Translation
Kartik Asooja | Jorge Gracia | Nitish Aggarwal | Asunción Gómez Pérez
Kartik Asooja | Jorge Gracia | Nitish Aggarwal | Asunción Gómez Pérez
System Combination with Extra Alignment Information
Xiaofeng Wu | Tsuyoshi Okita | Josef van Genabith | Qun Liu
Xiaofeng Wu | Tsuyoshi Okita | Josef van Genabith | Qun Liu
Topic Modeling-based Domain Adaptation for System Combination
Tsuyoshi Okita | Antonio Toral | Josef van Genabith
Tsuyoshi Okita | Antonio Toral | Josef van Genabith
up
Proceedings of the Workshop on Speech and Language Processing Tools in Education
Proceedings of the Workshop on Speech and Language Processing Tools in Education
Radhika Mamidi | Kishore Prahallad
Radhika Mamidi | Kishore Prahallad
Effective Mentor Suggestion System for Online Collaboration Platform
Advait Raut | Upasana Gaikwad | Ramakrishna Bairi | Ganesh Ramakrishnan
Advait Raut | Upasana Gaikwad | Ramakrishna Bairi | Ganesh Ramakrishnan
Automatic pronunciation assessment for language learners with acoustic-phonetic features
Vaishali Patil | Preeti Rao
Vaishali Patil | Preeti Rao
An Issue-oriented Syllabus Retrieval System using Terminology-based Syllabus Structuring and Visualization
Hideki Mima
Hideki Mima
Real-Time Tone Recognition in A Computer-Assisted Language Learning System for German Learners of Mandarin
Hussein Hussein | Hansjörg Mixdorff | Rüdiger Hoffmann
Hussein Hussein | Hansjörg Mixdorff | Rüdiger Hoffmann
Textbook Construction from Lecture Transcripts
Aliabbas Petiwala | Kannan Moudgalya | Pushpak Bhattacharyya
Aliabbas Petiwala | Kannan Moudgalya | Pushpak Bhattacharyya
Enriching An Academic knowledge base using Linked Open Data
Chetana Gavankar | Ashish Kulkarni | Yuan Fang Li | Ganesh Ramakrishnan
Chetana Gavankar | Ashish Kulkarni | Yuan Fang Li | Ganesh Ramakrishnan
Automatic Pronunciation Scoring And Mispronunciation Detection Using CMUSphinx
Ronanki Srikanth | Bo Li | James Salsman
Ronanki Srikanth | Bo Li | James Salsman
A template matching approach for detecting pronunciation mismatch
Lavanya Prahallad | Radhika Mamidi | Kishore Prahallad
Lavanya Prahallad | Radhika Mamidi | Kishore Prahallad
up
Proceedings of the Workshop on Reordering for Statistical Machine Translation
Proceedings of the Workshop on Reordering for Statistical Machine Translation
Karthik Visweswariah | Ananthakrishnan Ramanathan | Mitesh M. Khapra
Karthik Visweswariah | Ananthakrishnan Ramanathan | Mitesh M. Khapra
Whitepaper for Shared Task on Learning Reordering from Word Alignments at RSMT 2012
Mitesh M. Khapra | Ananthakrishnan Ramanathan | Karthik Visweswariah
Mitesh M. Khapra | Ananthakrishnan Ramanathan | Karthik Visweswariah
Report of the Shared Task on Learning Reordering from Word Alignments at RSMT 2012
Mitesh M. Khapra | Ananthakrishnan Ramanathan | Karthik Visweswariah
Mitesh M. Khapra | Ananthakrishnan Ramanathan | Karthik Visweswariah
Building a reordering system using tree-to-string hierarchical model
Jacob Dlougach | Irina Galinskaya
Jacob Dlougach | Irina Galinskaya
up
Proceedings of the Workshop on Question Answering for Complex Domains
Proceedings of the Workshop on Question Answering for Complex Domains
Nanda Kambhatla | Sachindra Joshi | Ganesh Ramakrishnan | Kiran Kate | Priyanka Agrawal
Nanda Kambhatla | Sachindra Joshi | Ganesh Ramakrishnan | Kiran Kate | Priyanka Agrawal
Question Classification and Answering from Procedural Text in English
Somnath Banerjee | Sivaji Bandyopadhyay
Somnath Banerjee | Sivaji Bandyopadhyay
Structured and Logical Representations of Assamese Text for Question-Answering System
Shikhar Kr. Sarma | Rita Chakraborty
Shikhar Kr. Sarma | Rita Chakraborty
Towards a thematic role based target identification model for question answering
Rivindu Perera | Udayangi Perera
Rivindu Perera | Udayangi Perera
up
Proceedings of the First International Workshop on Optimization Techniques for Human Language Technology
Proceedings of the First International Workshop on Optimization Techniques for Human Language Technology
Pushpak Bhattacharyya | Asif Ekbal | Sriparna Saha | Mark Johnson | Diego Molla-Aliod | Mark Dras
Pushpak Bhattacharyya | Asif Ekbal | Sriparna Saha | Mark Johnson | Diego Molla-Aliod | Mark Dras
BioPOS: Biologically Inspired Algorithms for POS Tagging
Ana Paula Silva | Arlindo Silva | Irene Rodrigues
Ana Paula Silva | Arlindo Silva | Irene Rodrigues
Optimization for Efficient Determination of Chunk in Automatic Evaluation for Machine Translation
Hiroshi Echizen’ya | Kenji Araki | Eduard Hovy
Hiroshi Echizen’ya | Kenji Araki | Eduard Hovy
Optimizing Transliteration for Hindi/Marathi to English Using only Two Weights
Manikrao Dhore | Shantanu Dixit | Ruchi Dhore
Manikrao Dhore | Shantanu Dixit | Ruchi Dhore
Selection of Discriminative Features for Translation Texts
Kuo-Ming Tang | Chien-Kang Huang | Chia-Ming Lee
Kuo-Ming Tang | Chien-Kang Huang | Chia-Ming Lee
Semi-supervised Learning of Naive Bayes Classifier with feature constraints
Nagesh Bhattu Sristy | D.V.L.N Somayajulu
Nagesh Bhattu Sristy | D.V.L.N Somayajulu
Optimization and Sampling for NLP from a Unified Viewpoint
Marc Dymetman | Guillaume Bouchard | Simon Carter
Marc Dymetman | Guillaume Bouchard | Simon Carter
up
Workshop on Post-Editing Technology and Practice
This paper introduces a publicly available database of recorded translation sessions for Translation Process Research (TPR). User activity data (UAD) of translators behavior was collected over the past 5 years in several translation studies with Translog 1 , a data acquisition software which logs keystrokes and gaze data during text reception and production. The database compiles this data into a consistent format which can be processed by various visualization and analysis tools.
Post-editing time as a measure of cognitive effort
Maarit Koponen | Wilker Aziz | Luciana Ramos | Lucia Specia
Maarit Koponen | Wilker Aziz | Luciana Ramos | Lucia Specia
Post-editing machine translations has been attracting increasing attention both as a common practice within the translation industry and as a way to evaluate Machine Translation (MT) quality via edit distance metrics between the MT and its post-edited version. Commonly used metrics such as HTER are limited in that they cannot fully capture the effort required for post-editing. Particularly, the cognitive effort required may vary for different types of errors and may also depend on the context. We suggest post-editing time as a way to assess some of the cognitive effort involved in post-editing. This paper presents two experiments investigating the connection between post-editing time and cognitive effort. First, we examine whether sentences with long and short post-editing times involve edits of different levels of difficulty. Second, we study the variability in post-editing time and other statistics among editors.
Average Pause Ratio as an Indicator of Cognitive Effort in Post-Editing: A Case Study
Isabel Lacruz | Gregory M. Shreve | Erik Angelone
Isabel Lacruz | Gregory M. Shreve | Erik Angelone
Pauses are known to be good indicators of cognitive demand in monolingual language production and in translation. However, a previous effort by O’Brien (2006) to establish an analogous relationship in post-editing did not produce the expected result. In this case study, we introduce a metric for pause activity, the average pause ratio, which is sensitive to both the number and duration of pauses. We measured cognitive effort in a segment by counting the number of complete editing events. We found that the average pause ratio was higher for less cognitively demanding segments than for more cognitively demanding segments. Moreover, this effect became more pronounced as the minimum threshold for pause length was shortened.
Reliably Assessing the Quality of Post-edited Translation Based on Formalized Structured Translation Specifications
Alan K. Melby | Jason Housley | Paul J. Fields | Emily Tuioti
Alan K. Melby | Jason Housley | Paul J. Fields | Emily Tuioti
Post-editing of machine translation has become more common in recent years. This has created the need for a formal method of assessing the performance of post-editors in terms of whether they are able to produce post-edited target texts that follow project specifications. This paper proposes the use of formalized structured translation specifications (FSTS) as a basis for post-editor assessment. To determine if potential evaluators are able to reliably assess the quality of post-edited translations, an experiment used texts representing the work of five fictional post-editors. Two software applications were developed to facilitate the assessment: the Ruqual Specifications Writer, which aids in establishing post-editing project specifications; and Ruqual Rubric Viewer, which provides a graphical user interface for constructing a rubric in a machine-readable format. Seventeen non-experts rated the translation quality of each simulated post-edited text. Intraclass correlation analysis showed evidence that the evaluators were highly reliable in evaluating the performance of the post-editors. Thus, we assert that using FSTS specifications applied through the Ruqual software tools provides a useful basis for evaluating the quality of post-edited texts.
Learning to Automatically Post-Edit Dropped Words in MT
Jacob Mundt | Kristen Parton | Kathleen McKeown
Jacob Mundt | Kristen Parton | Kathleen McKeown
Automatic post-editors (APEs) can improve adequacy of MT output by detecting and reinserting dropped content words, but the location where these words are inserted is critical. In this paper, we describe a probabilistic approach for learning reinsertion rules for specific languages and MT systems, as well as a method for synthesizing training data from reference translations. We test the insertion logic on MT systems for Chinese to English and Arabic to English. Our adaptive APE is able to insert within 3 words of the best location 73% of the time (32% in the exact location) in Arabic-English MT output, and 67% of the time in Chinese-English output (30% in the exact location), and delivers improved performance on automated adequacy metrics over a previous rule-based approach to insertion. We consider how particular aspects of the insertion problem make it particularly amenable to machine learning solutions.
It is a well-known fact that the amount of content which is available to be translated and localized far outnumbers the current amount of translation resources. Automation in general and Machine Translation (MT) in particular are one of the key technologies which can help improve this situation. However, a tool that integrates all of the components needed for the localization process is still missing, and MT is still out of reach for most localisation professionals. In this paper we present an online translation environment which empowers users with MT by enabling engines to be created from their data, without a need for technical knowledge or special hardware requirements and at low cost. Documents in a variety of formats can then be post-edited after being processed with their Translation Memories, MT engines and glossaries. We give an overview of the tool and present a case study of a project for a large games company, showing the applicability of our tool.
To post-edit or not to post-edit? Estimating the benefits of MT post-editing for a European organization
Alexandros Poulis | David Kolovratnik
Alexandros Poulis | David Kolovratnik
In the last few years the European Parliament has witnessed a significant increase in translation demand. Although Translation Memory (TM) tools, terminology databases and bilingual concordancers have provided significant leverage in terms of quality and productivity the European Parliament is in need for advanced language technology to keep facing successfully the challenge of multilingualism. This paper describes an ongoing large-scale machine translation post-editing evaluation campaign the purpose of which is to estimate the business benefits from the use of machine translation for the European Parliament. This paper focuses mainly on the design, the methodology and the tools used by the evaluators but it also presents some preliminary results for the following language pairs: Polish-English, Danish-English, Lithuanian-English, English-German and English-French.
How Good Is Crowd Post-Editing? Its Potential and Limitations
Midori Tatsumi | Takako Aikawa | Kentaro Yamamoto | Hitoshi Isahara
Midori Tatsumi | Takako Aikawa | Kentaro Yamamoto | Hitoshi Isahara
This paper is a partial report of a research effort on evaluating the effect of crowd-sourced post-editing. We first discuss the emerging trend of crowd-sourced post-editing of machine translation output, along with its benefits and drawbacks. Second, we describe the pilot study we have conducted on a platform that facilitates crowd-sourced post-editing. Finally, we provide our plans for further studies to have more insight on how effective crowd-sourced post-editing is.
Error Detection for Post-editing Rule-based Machine Translation
Justina Valotkaite | Munshi Asadullah
Justina Valotkaite | Munshi Asadullah
The increasing role of post-editing as a way of improving machine translation output and a faster alternative to translating from scratch has lately attracted researchers’ attention and various attempts have been proposed to facilitate the task. We experiment with a method to provide support for the post-editing task through error detection. A deep linguistic error analysis was done of a sample of English sentences translated from Portuguese by two Rule-based Machine Translation systems. We designed a set of rules to deal with various systematic translation errors and implemented a subset of these rules covering the errors of tense and number. The evaluation of these rules showed a satisfactory performance. In addition, we performed an experiment with human translators which confirmed that highlighting translation errors during the post-editing can help the translators perform the post-editing task up to 12 seconds per error faster and improve their efficiency by minimizing the number of missed errors.
In this paper, we present the Moses-based infrastructure we developed and use as a productivity tool for the localisation of software documentation and user interface (UI) strings at Autodesk into twelve languages. We describe the adjustments we have made to the machine translation (MT) training workflow to suit our needs and environment, our server environment and the MT Info Service that handles all translation requests and allows the integration of MT in our various localisation systems. We also present the results of our latest post-editing productivity test, where we measured the productivity gain for translators post-editing MT output versus translating from scratch. Our analysis of the data indicates the presence of a strong correlation between the amount of editing applied to the raw MT output by the translators and their productivity gain. In addition, within the last calendar year our system has processed over thirteen million tokens of documentation content of which we have a record of the performed post-editing. This has allowed us to evaluate the performance of our MT engines for the different languages across our product portfolio, as well as spotlight potential issues with MT in the localisation process.
up
Fourth Workshop on Computational Approaches to Arabic-Script-based Languages
Fourth Workshop on Computational Approaches to Arabic-Script-based Languages
Ali Farghaly | Farhad Oroumchian
Ali Farghaly | Farhad Oroumchian
Translating English Discourse Connectives into Arabic: a Corpus-based Analysis and an Evaluation Metric
Najeh Hajlaoui | Andrei Popescu-Belis
Najeh Hajlaoui | Andrei Popescu-Belis
Discourse connectives can often signal multiple discourse relations, depending on their context. The automatic identification of the Arabic translations of seven English discourse connectives shows how these connectives are differently translated depending on their actual senses. Automatic labelling of English source connectives can help a machine translation system to translate them more correctly. The corpus-based analysis of Arabic translations also enables the definition of a connective-specific evaluation metric for machine translation, which is here validated by human judges on sample English/Arabic translation data.
Idiomatic MWEs and Machine Translation A Retrieval and Representation Model: the AraMWE Project
Giuliano Lancioni | Marco Boella
Giuliano Lancioni | Marco Boella
A preliminary implementation of AraMWE, a hybrid project that includes a statistical component and a CCG symbolic component to extract and treat MWEs and idioms in Arabic and Eng- lish parallel texts is presented, together with a general sketch of the system, a thorough description of the statistical component and a proof of concept of the CCG component.
Developing an Open-domain English-Farsi Translation System Using AFEC: Amirkabir Bilingual Farsi-English Corpus
Fattaneh Jabbari | Somayeh Bakshaei | Seyyed Mohammad Mohammadzadeh Ziabary | Shahram Khadivi
Fattaneh Jabbari | Somayeh Bakshaei | Seyyed Mohammad Mohammadzadeh Ziabary | Shahram Khadivi
The translation quality of Statistical Machine Translation (SMT) depends on the amount of input data especially for morphologically rich languages. Farsi (Persian) language is such a language which has few NLP resources. It also suffers from the non-standard written characters which causes a large variety in the written form of each character. Moreover, the structural difference between Farsi and English results in long range reorderings which cannot be modeled by common SMT reordering models. Here, we try to improve the existing English-Farsi SMT system focusing on these challenges first by expanding our bilingual limited-domain corpus to an open-domain one. Then, to alleviate the character variations, a new text normalization algorithm is offered. Finally, some hand-crafted rules are applied to reduce the structural differences. Using the new corpus, the experimental results showed 8.82% BLEU improvement by applying new normalization method and 9.1% BLEU when rules are used.
In this paper, we study the problem of finding named entities in the Arabic text. For this task we present the development of our pipeline software for Arabic named entity recognition (ARNE), which includes tokenization, morphological analysis, Buckwalter transliteration, part of speech tagging and named entity recognition of person, location and organisation named entities. In our first attempt to recognize named entites, we have used a simple, fast and language independent gazetteer lookup approach. In our second attempt, we have used the morphological analysis provided by our pipeline to remove affixes and observed hence an improvement in our performance. The pipeline presented in this paper, can be used in future as a basis for a named entity recognition system that recognized named entites not only using gazetteers, but also making use of morphological information and part of speech tagging.
Approaches to Arabic Name Transliteration and Matching in the DataFlux Quality Knowledge Base
Brant N. Kay | Brian C. Rineer
Brant N. Kay | Brian C. Rineer
This paper discusses a hybrid approach to transliterating and matching Arabic names, as implemented in the DataFlux Quality Knowledge Base (QKB), a knowledge base used by data management software systems from SAS Institute, Inc. The approach to transliteration relies on a lexicon of names with their corresponding transliterations as its primary method, and falls back on PERL regular expression rules to transliterate any names that do not exist in the lexicon. Transliteration in the QKB is bi-directional; the technology transliterates Arabic names written in the Arabic script to the Latin script, and transliterates Arabic names written in the Latin script to Arabic. Arabic name matching takes a similar approach and relies on a lexicon of Arabic names and their corresponding transliterations, falling back on phonetic transliteration rules to transliterate names into the Latin script. All names are ultimately rendered in the Latin script before matching takes place. Thus, the technology is capable of matching names across the Arabic and Latin scripts, as well as within the Arabic script or within the Latin script. The goal of the authors of this paper was to build a software system capable of transliterating and matching Arabic names across scripts with an accuracy deemed to be acceptable according to internal software quality standards.
Using Arabic Transliteration to Improve Word Alignment from French- Arabic Parallel Corpora
Houda Saadane | Ouafa Benterki | Nasredine Semmar | Christian Fluhr
Houda Saadane | Ouafa Benterki | Nasredine Semmar | Christian Fluhr
In this paper, we focus on the use of Arabic transliteration to improve the results of a linguistics-based word alignment approach from parallel text corpora. This approach uses, on the one hand, a bilingual lexicon, named entities, cognates and grammatical tags to align single words, and on the other hand, syntactic dependency relations to align compound words. We have evaluated the word aligner integrating Arabic transliteration using two methods: A manual evaluation of the alignment quality and an evaluation of the impact of this alignment on the translation quality by using the Moses statistical machine translation system. The obtained results show that Arabic transliteration improves the quality of both alignment and translation.
Research done on Arabic sentiment analysis is considered very limited almost in its early steps compared to other languages like English whether at document-level or sentence-level. In this paper, we test the effect of preprocessing (normalization, stemming, and stop words removal) on the performance of an Arabic sentiment analysis system using Arabic tweets from twitter. The sentiment (positive or negative) of the crawled tweets is analyzed to interpret the attitude of the public with regards to topic of interest. Using Twitter as the main source of data reflects the importance of the system for the Middle East region, which mostly speaks Arabic.
Rescoring N-Best Hypotheses for Arabic Speech Recognition: A Syntax- Mining Approach
Dia AbuZeina | Moustafa Elshafei | Husni Al-Muhtaseb | Wasfi Al-Khatib
Dia AbuZeina | Moustafa Elshafei | Husni Al-Muhtaseb | Wasfi Al-Khatib
Improving speech recognition accuracy through linguistic knowledge is a major research area in automatic speech recognition systems. In this paper, we present a syntax-mining approach to rescore N-Best hypotheses for Arabic speech recognition systems. The method depends on a machine learning tool (WEKA-3-6-5) to extract the N-Best syntactic rules of the Baseline tagged transcription corpus which was tagged using Stanford Arabic tagger. The proposed method was tested using the Baseline system that contains a pronunciation dictionary of 17,236 vocabularies (28,682 words and variants) from 7.57 hours pronunciation corpus of modern standard Arabic (MSA) broadcast news. Using Carnegie Mellon University (CMU) PocketSphinx speech recognition engine, the Baseline system achieved a Word Error Rate (WER) of 16.04 % on a test set of 400 utterances ( about 0.57 hours) containing 3585 diacritized words. Even though there were enhancements in some tested files, we found that this method does not lead to significant enhancement (for Arabic). Based on this research work, we conclude this paper by introducing a new design for language models to account for longer-distance constrains, instead of a few proceeding words.
We annotate a small corpus of religious Arabic with morphological segmentation boundaries and fine-grained segment-based part of speech tags. Experiments on both segmentation and POS tagging show that the religious corpus-trained segmenter and POS tagger outperform the Arabic Treebak-trained ones although the latter is 21 times as big, which shows the need for building religious Arabic linguistic resources. The small corpus we annotate improves segmentation accuracy by 5% absolute (from 90.84% to 95.70%), and POS tagging by 9% absolute (from 82.22% to 91.26) when using gold standard segmentation, and by 9.6% absolute (from 78.62% to 88.22) when using automatic segmentation.
Exploiting Wikipedia as a Knowledge Base for the Extraction of Linguistic Resources: Application on Arabic-French Comparable Corpora and Bilingual Lexicons
Rahma Sellami | Fatiha Sadat | Lamia Hadrich Belguith
Rahma Sellami | Fatiha Sadat | Lamia Hadrich Belguith
We present simple and effective methods for extracting comparable corpora and bilingual lexicons from Wikipedia. We shall exploit the large scale and the structure of Wikipedia articles to extract two resources that will be very useful for natural language applications. We build a comparable corpus from Wikipedia using categories as topic restrictions and we extract bilingual lexicons from inter-language links aligned with statistical method or a combined statistical and linguistic method.
up
Workshop on Monolingual Machine Translation
Improving English to Spanish Out-of-Domain Translations by Morphology Generalization and Generation
Lluís Formiga | Adolfo Hernández | José B. Mariño | Enric Monte
Lluís Formiga | Adolfo Hernández | José B. Mariño | Enric Monte
This paper presents a detailed study of a method for morphology generalization and generation to address out-of-domain translations in English-to-Spanish phrase-based MT. The paper studies whether the morphological richness of the target language causes poor quality translation when translating out-of-domain. In detail, this approach first translates into Spanish simplified forms and then predicts the final inflected forms through a morphology generation step based on shallow and deep-projected linguistic information available from both the source and target-language sentences. Obtained results highlight the importance of generalization, and therefore generation, for dealing with out-of-domain data.
Monolingual Data Optimisation for Bootstrapping SMT Engines
Jie Jiang | Andy Way | Nelson Ng | Rejwanul Haque | Mike Dillinger | Jun Lu
Jie Jiang | Andy Way | Nelson Ng | Rejwanul Haque | Mike Dillinger | Jun Lu
Content localisation via machine translation (MT) is a sine qua non, especially for international online business. While most applications utilise rule-based solutions due to the lack of suitable in-domain parallel corpora for statistical MT (SMT) training, in this paper we investigate the possibility of applying SMT where huge amounts of monolingual content only are available. We describe a case study where an analysis of a very large amount of monolingual online trading data from eBay is conducted by ALS with a view to reducing this corpus to the most representative sample in order to ensure the widest possible coverage of the total data set. Furthermore, minimal yet optimal sets of sentences/words/terms are selected for generation of initial translation units for future SMT system-building.
Shallow and Deep Paraphrasing for Improved Machine Translation Parameter Optimization
Dennis N. Mehay | Michael White
Dennis N. Mehay | Michael White
String comparison methods such as BLEU (Papineni et al., 2002) are the de facto standard in MT evaluation (MTE) and in MT system parameter tuning (Och, 2003). It is difficult for these metrics to recognize legitimate lexical and grammatical paraphrases, which is important for MT system tuning (Madnani, 2010). We present two methods to address this: a shallow lexical substitution technique and a grammar-driven paraphrasing technique. Grammatically precise paraphrasing is novel in the context of MTE, and demonstrating its usefulness is a key contribution of this paper. We use these techniques to paraphrase a single reference, which, when used for parameter tuning, leads to superior translation performance over baselines that use only human-authored references.
Two stage Machine Translation System using Pattern-based MT and Phrase-based SMT
Jin’ichi Murakami | Takuya Nishimura | Masoto Tokuhisa
Jin’ichi Murakami | Takuya Nishimura | Masoto Tokuhisa
We have developed a two-stage machine translation (MT) system. The first stage consists of an automatically created pattern-based machine translation system (PBMT), and the second stage consists of a standard phrase-based statistical machine translation (SMT) system. We studied for the Japanese-English simple sentence task. First, we obtained English sentences from Japanese sentences using an automatically created Japanese-English pattern-based machine translation. We call the English sentences obtained in this way as “English”. Second, we applied a standard SMT (Moses) to the results. This means that we translated the “English” sentences into English by SMT. We also conducted ABX tests (Clark, 1982) to compare the outputs by the standard SMT (Moses) with those by the proposed system for 100 sentences. The experimental results indicated that 30 sentences output by the proposed system were evaluated as being better than those outputs by the standard SMT system, whereas 9 sentences output by the standard SMT system were thought to be better than those outputs by the proposed system. This means that our proposed system functioned effectively in the Japanese-English simple sentence task.
This paper presents a method to improve a word alignment model in a phrase-based Statistical Machine Translation system for a low-resourced language using a string similarity approach. Our method captures similar words that can be seen as semi-monolingual across languages, such as numbers, named entities, and adapted/loan words. We use several string similarity metrics to measure the monolinguality of the words, such as Longest Common Subsequence Ratio (LCSR), Minimum Edit Distance Ratio (MEDR), and we also use a modified BLEU Score (modBLEU). Our approach is to add intersecting alignment points for word pairs that are orthographically similar, before applying a word alignment heuristic, to generate a better word alignment. We demonstrate this approach on Indonesian-to-English translation task, where the languages share many similar words that are poorly aligned given a limited training data. This approach gives a statistically significant improvement by up to 0.66 in terms of BLEU score.
Addressing some Issues of Data Sparsity towards Improving English- Manipuri SMT using Morphological Information
Thoudam Doren Singh
Thoudam Doren Singh
The performance of an SMT system heavily depends on the availability of large parallel corpora. Unavailability of these resources in the required amount for many language pair is a challenging issue. The required size of the resource involving morphologically rich and highly agglutinative language is essentially much more for the SMT systems. This paper investigates on some of the issues on enriching the resource for this kind of languages. Handling of inflectional and derivational morphemes of the morphologically rich target language plays important role in the enrichment process. Mapping from the source to the target side is carried out for the English-Manipuri SMT task using factored model. The SMT system developed shows improvement in the performance both in terms of the automatic scoring and subjective evaluation over the baseline system.
We aim to use statistical machine translation technology to correct grammar errors and style issues in monolingual text. Here, as a feasibility test, we focus on depassivization in German and we abstract from surface forms to parts of speech. Our results are not yet satisfactory but yield useful insights into directions for improvement.
up
Proceedings of the Joint Workshop on Exploiting Synergies between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra)
Proceedings of the Joint Workshop on Exploiting Synergies between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra)
Marta R. Costa-jussà | Patrik Lambert | Rafael E. Banchs | Reinhard Rapp | Bogdan Babych
Marta R. Costa-jussà | Patrik Lambert | Rafael E. Banchs | Reinhard Rapp | Bogdan Babych
Measuring Comparability of Documents in Non-Parallel Corpora for Efficient Extraction of (Semi-)Parallel Translation Equivalents
Fangzhong Su | Bogdan Babych
Fangzhong Su | Bogdan Babych
An Empirical Evaluation of Stop Word Removal in Statistical Machine Translation
Tze Yuang Chong | Rafael Banchs | Eng Siong Chng
Tze Yuang Chong | Rafael Banchs | Eng Siong Chng
Natural Language Descriptions of Visual Scenes Corpus Generation and Analysis
Muhammad Usman Ghani Khan | Rao Muhammad Adeel Nawab | Yoshihiko Gotoh
Muhammad Usman Ghani Khan | Rao Muhammad Adeel Nawab | Yoshihiko Gotoh
Combining EBMT, SMT, TM and IR Technologies for Quality and Scale
Sandipan Dandapat | Sara Morrissey | Andy Way | Josef van Genabith
Sandipan Dandapat | Sara Morrissey | Andy Way | Josef van Genabith
PRESEMT: Pattern Recognition-based Statistically Enhanced MT
George Tambouratzis | Marina Vassiliou | Sokratis Sofianopoulos
George Tambouratzis | Marina Vassiliou | Sokratis Sofianopoulos
ATLAS - Human Language Technologies integrated within a Multilingual Web Content Management System
Svetla Koeva
Svetla Koeva
Were the clocks striking or surprising? Using WSD to improve MT performance
Špela Vintar | Darja Fišer | Aljoša Vrščaj
Špela Vintar | Darja Fišer | Aljoša Vrščaj
Design of a hybrid high quality machine translation system
Bogdan Babych | Kurt Eberle | Johanna Geiß | Mireia Ginestí-Rosell | Anthony Hartley | Reinhard Rapp | Serge Sharoff | Martin Thomas
Bogdan Babych | Kurt Eberle | Johanna Geiß | Mireia Ginestí-Rosell | Anthony Hartley | Reinhard Rapp | Serge Sharoff | Martin Thomas
Can Machine Learning Algorithms Improve Phrase Selection in Hybrid Machine Translation?
Christian Federmann
Christian Federmann
up
Proceedings of the Workshop on Applications of Tree Automata Techniques in Natural Language Processing
Proceedings of the Workshop on Applications of Tree Automata Techniques in Natural Language Processing
Frank Drewes | Marco Kuhlmann
Frank Drewes | Marco Kuhlmann
Preservation of Recognizability for Weighted Linear Extended Top-Down Tree Transducers
Nina Seemann | Daniel Quernheim | Fabienne Braune | Andreas Maletti
Nina Seemann | Daniel Quernheim | Fabienne Braune | Andreas Maletti
Deciding the Twins Property for Weighted Tree Automata over Extremal Semifields
Matthias Büchse | Anja Fischer
Matthias Büchse | Anja Fischer
up
Proceedings of the 6th Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities
Proceedings of the 6th Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities
Kalliopi Zervanou | Antal van den Bosch
Kalliopi Zervanou | Antal van den Bosch
Lexicon Construction and Corpus Annotation of Historical Language with the CoBaLT Editor
Tom Kenter | Tomaž Erjavec | Maja Žorga Dulmin | Darja Fišer
Tom Kenter | Tomaž Erjavec | Maja Žorga Dulmin | Darja Fišer
A high speed transcription interface for annotating primary linguistic data
Mark Dingemanse | Jeremy Hammond | Herman Stehouwer | Aarthy Somasundaram | Sebastian Drude
Mark Dingemanse | Jeremy Hammond | Herman Stehouwer | Aarthy Somasundaram | Sebastian Drude
BAD: An Assistant tool for making verses in Basque
Manex Agirrezabal | Iñaki Alegria | Bertol Arrieta | Mans Hulden
Manex Agirrezabal | Iñaki Alegria | Bertol Arrieta | Mans Hulden
Toward Language Independent Methodology for Generating Artwork Descriptions – Exploring FrameNet Information
Dana Dannélls | Lars Borin
Dana Dannélls | Lars Borin
Harvesting Indices to Grow a Controlled Vocabulary: Towards Improved Access to Historical Legal Texts
Michael Piotrowski | Cathrin Senn
Michael Piotrowski | Cathrin Senn
Ontology-Based Incremental Annotation of Characters in Folktales
Thierry Declerck | Nikolina Koleva | Hans-Ulrich Krieger
Thierry Declerck | Nikolina Koleva | Hans-Ulrich Krieger
Advanced Visual Analytics Methods for Literature Analysis
Daniela Oelke | Dimitrios Kokkinakis | Mats Malm
Daniela Oelke | Dimitrios Kokkinakis | Mats Malm
Distributional techniques for philosophical enquiry
Aurélie Herbelot | Eva von Redecker | Johanna Müller
Aurélie Herbelot | Eva von Redecker | Johanna Müller
Linguistically-Adapted Structural Query Annotation for Digital Libraries in the Social Sciences
Caroline Brun | Vassilina Nikoulina | Nikolaos Lagos
Caroline Brun | Vassilina Nikoulina | Nikolaos Lagos
Parsing the Past - Identification of Verb Constructions in Historical Text
Eva Pettersson | Beáta Megyesi | Joakim Nivre
Eva Pettersson | Beáta Megyesi | Joakim Nivre
Computing Similarity between Cultural Heritage Items using Multimodal Features
Nikolaos Aletras | Mark Stevenson
Nikolaos Aletras | Mark Stevenson
Enabling the Discovery of Digital Cultural Heritage Objects through Wikipedia
Mark Michael Hall | Oier Lopez de Lacalle | Aitor Soroa Etxabe | Paul Clough | Eneko Agirre
Mark Michael Hall | Oier Lopez de Lacalle | Aitor Soroa Etxabe | Paul Clough | Eneko Agirre
up
Proceedings of the 3rd Workshop on Cognitive Modeling and Computational Linguistics (CMCL 2012)
Proceedings of the 3rd Workshop on Cognitive Modeling and Computational Linguistics (CMCL 2012)
David Reitter | Roger Levy
David Reitter | Roger Levy
Semi-supervised learning for automatic conceptual property extraction
Colin Kelly | Barry Devereux | Anna Korhonen
Colin Kelly | Barry Devereux | Anna Korhonen
Why long words take longer to read: the role of uncertainty about word length
Klinton Bicknell | Roger Levy
Klinton Bicknell | Roger Levy
Fractal Unfolding: A Metamorphic Approach to Learning to Parse Recursive Structure
Whitney Tabor | Pyeong Whan Cho | Emily Szkudlarek
Whitney Tabor | Pyeong Whan Cho | Emily Szkudlarek
Sequential vs. Hierarchical Syntactic Models of Human Incremental Sentence Processing
Victoria Fossum | Roger Levy
Victoria Fossum | Roger Levy
up
Proceedings of the Seventh Workshop on Building Educational Applications Using NLP
Proceedings of the Seventh Workshop on Building Educational Applications Using NLP
Joel Tetreault | Jill Burstein | Claudia Leacock
Joel Tetreault | Jill Burstein | Claudia Leacock
Question Ranking and Selection in Tutorial Dialogues
Lee Becker | Martha Palmer | Sarel van Vuuren | Wayne Ward
Lee Becker | Martha Palmer | Sarel van Vuuren | Wayne Ward
Identifying science concepts and student misconceptions in an interactive essay writing tutor
Steven Bethard | Ifeyinwa Okoye | Md. Arafat Sultan | Haojie Hang | James H. Martin | Tamara Sumner
Steven Bethard | Ifeyinwa Okoye | Md. Arafat Sultan | Haojie Hang | James H. Martin | Tamara Sumner
Automatic Grading of Scientific Inquiry
Avirup Sil | Angela Shelton | Diane Jass Ketelhut | Alexander Yates
Avirup Sil | Angela Shelton | Diane Jass Ketelhut | Alexander Yates
Exploring Grammatical Error Correction with Not-So-Crummy Machine Translation
Nitin Madnani | Joel Tetreault | Martin Chodorow
Nitin Madnani | Joel Tetreault | Martin Chodorow
HOO 2012: A Report on the Preposition and Determiner Error Correction Shared Task
Robert Dale | Ilya Anisimoff | George Narroway
Robert Dale | Ilya Anisimoff | George Narroway
Measuring the Use of Factual Information in Test-Taker Essays
Beata Beigman Klebanov | Derrick Higgins
Beata Beigman Klebanov | Derrick Higgins
PREFER: Using a Graph-Based Approach to Generate Paraphrases for Language Learning
Mei-Hua Chen | Shi-Ting Huang | Chung-Chi Huang | Hsien-Chin Liou | Jason S. Chang
Mei-Hua Chen | Shi-Ting Huang | Chung-Chi Huang | Hsien-Chin Liou | Jason S. Chang
Using an Ontology for Improved Automated Content Scoring of Spontaneous Non-Native Speech
Miao Chen | Klaus Zechner
Miao Chen | Klaus Zechner
Predicting Learner Levels for Online Exercises of Hebrew
Markus Dickinson | Sandra Kübler | Anthony Meyer
Markus Dickinson | Sandra Kübler | Anthony Meyer
On using context for automatic correction of non-word misspellings in student essays
Michael Flor | Yoko Futagi
Michael Flor | Yoko Futagi
Judging Grammaticality with Count-Induced Tree Substitution Grammars
Francis Ferraro | Matt Post | Benjamin Van Durme
Francis Ferraro | Matt Post | Benjamin Van Durme
Developing ARET: An NLP-based Educational Tool Set for Arabic Reading Enhancement
Mohammed Maamouri | Wajdi Zaghouani | Violetta Cavalli-Sforza | Dave Graff | Mike Ciul
Mohammed Maamouri | Wajdi Zaghouani | Violetta Cavalli-Sforza | Dave Graff | Mike Ciul
A Comparison of Greedy and Optimal Assessment of Natural Language Student Input Using Word-to-Word Similarity Metrics
Vasile Rus | Mihai Lintean
Vasile Rus | Mihai Lintean
On Improving the Accuracy of Readability Classification using Insights from Second Language Acquisition
Sowmya Vajjala | Detmar Meurers
Sowmya Vajjala | Detmar Meurers
An Interactive Analytic Tool for Peer-Review Exploration
Wenting Xiong | Diane Litman | Jingtao Wang | Christian Schunn
Wenting Xiong | Diane Litman | Jingtao Wang | Christian Schunn
Vocabulary Profile as a Measure of Vocabulary Sophistication
Su-Youn Yoon | Suma Bhat | Klaus Zechner
Su-Youn Yoon | Suma Bhat | Klaus Zechner
Short Answer Assessment: Establishing Links Between Research Strands
Ramon Ziai | Niels Ott | Detmar Meurers
Ramon Ziai | Niels Ott | Detmar Meurers
Detection and Correction of Preposition and Determiner Errors in English: HOO 2012
Pinaki Bhaskar | Aniruddha Ghosh | Santanu Pal | Sivaji Bandyopadhyay
Pinaki Bhaskar | Aniruddha Ghosh | Santanu Pal | Sivaji Bandyopadhyay
Informing Determiner and Preposition Error Correction with Hierarchical Word Clustering
Adriane Boyd | Marion Zepf | Detmar Meurers
Adriane Boyd | Marion Zepf | Detmar Meurers
Precision Isn’t Everything: A Hybrid Approach to Grammatical Error Detection
Michael Heilman | Aoife Cahill | Joel Tetreault
Michael Heilman | Aoife Cahill | Joel Tetreault
HOO 2012 Error Recognition and Correction Shared Task: Cambridge University Submission Report
Ekaterina Kochmar | Øistein Andersen | Ted Briscoe
Ekaterina Kochmar | Øistein Andersen | Ted Briscoe
A Naive Bayes classifier for automatic correction of preposition and determiner errors in ESL text
Gerard Lynch | Erwan Moreau | Carl Vogel
Gerard Lynch | Erwan Moreau | Carl Vogel
KU Leuven at HOO-2012: A Hybrid Approach to Detection and Correction of Determiner and Preposition Errors in Non-native English Text
Li Quan | Oleksandr Kolomiyets | Marie-Francine Moens
Li Quan | Oleksandr Kolomiyets | Marie-Francine Moens
The UI System in the HOO 2012 Shared Task on Error Correction
Alla Rozovskaya | Mark Sammons | Dan Roth
Alla Rozovskaya | Mark Sammons | Dan Roth
NAIST at the HOO 2012 Shared Task
Keisuke Sakaguchi | Yuta Hayashibe | Shuhei Kondo | Lis Kanashiro | Tomoya Mizumoto | Mamoru Komachi | Yuji Matsumoto
Keisuke Sakaguchi | Yuta Hayashibe | Shuhei Kondo | Lis Kanashiro | Tomoya Mizumoto | Mamoru Komachi | Yuji Matsumoto
Helping Our Own: NTHU NLPLAB System Description
Jian-Cheng Wu | Joseph Chang | Yi-Chun Chen | Shih-Ting Huang | Mei-Hua Chen | Jason S. Chang
Jian-Cheng Wu | Joseph Chang | Yi-Chun Chen | Shih-Ting Huang | Mei-Hua Chen | Jason S. Chang
Crowdsourced Comprehension: Predicting Prerequisite Structure in Wikipedia
Partha Talukdar | William Cohen
Partha Talukdar | William Cohen
up
Proceedings of the First Workshop on Predicting and Improving Text Readability for target reader populations
Proceedings of the First Workshop on Predicting and Improving Text Readability for target reader populations
Sandra Williams | Advaith Siddharthan | Ani Nenkova
Sandra Williams | Advaith Siddharthan | Ani Nenkova
Toward Determining the Comprehensibility of Machine Translations
Tucker Maney | Linda Sibert | Dennis Perzanowski | Kalyan Gupta | Astrid Schmidt-Nielsen
Tucker Maney | Linda Sibert | Dennis Perzanowski | Kalyan Gupta | Astrid Schmidt-Nielsen
Towards Automatic Lexical Simplification in Spanish: An Empirical Study
Biljana Drndarević | Horacio Saggion
Biljana Drndarević | Horacio Saggion
Offline Sentence Processing Measures for testing Readability with Users
Advaith Siddharthan | Napoleon Katsos
Advaith Siddharthan | Napoleon Katsos
Graphical Schemes May Improve Readability but Not Understandability for People with Dyslexia
Luz Rello | Horacio Saggion | Ricardo Baeza-Yates | Eduardo Graells
Luz Rello | Horacio Saggion | Ricardo Baeza-Yates | Eduardo Graells
Building Readability Lexicons with Unannotated Corpora
Julian Brooke | Vivian Tsang | David Jacob | Fraser Shein | Graeme Hirst
Julian Brooke | Vivian Tsang | David Jacob | Fraser Shein | Graeme Hirst
up
Proceedings of the Twelfth Meeting of the Special Interest Group on Computational Morphology and Phonology
Proceedings of the Twelfth Meeting of the Special Interest Group on Computational Morphology and Phonology
Lynne Cahill | Adam Albright
Lynne Cahill | Adam Albright
A Regularized Compression Method to Unsupervised Word Segmentation
Ruey-Cheng Chen | Chiung-Min Tsai | Jieh Hsiang
Ruey-Cheng Chen | Chiung-Min Tsai | Jieh Hsiang
Bounded copying is subsequential: Implications for metathesis and reduplication
Jane Chandlee | Jeffrey Heinz
Jane Chandlee | Jeffrey Heinz
An approximation approach to the problem of the acquisition of phonotactics in Optimality Theory
Giorgio Magri
Giorgio Magri
up
BioNLP: Proceedings of the 2012 Workshop on Biomedical Natural Language Processing
BioNLP: Proceedings of the 2012 Workshop on Biomedical Natural Language Processing
Kevin B. Cohen | Dina Demner-Fushman | Sophia Ananiadou | Bonnie Webber | Jun’ichi Tsujii | John Pestian
Kevin B. Cohen | Dina Demner-Fushman | Sophia Ananiadou | Bonnie Webber | Jun’ichi Tsujii | John Pestian
Graph-based alignment of narratives for automated neurological assessment
Emily Prud’hommeaux | Brian Roark
Emily Prud’hommeaux | Brian Roark
Bootstrapping Biomedical Ontologies for Scientific Text using NELL
Dana Movshovitz-Attias | William W. Cohen
Dana Movshovitz-Attias | William W. Cohen
Semantic distance and terminology structuring methods for the detection of semantically close terms
Marie Dupuch | Laëtitia Dupuch | Thierry Hamon | Natalia Grabar
Marie Dupuch | Laëtitia Dupuch | Thierry Hamon | Natalia Grabar
Analyzing Patient Records to Establish If and When a Patient Suffered from a Medical Condition
James Cogley | Nicola Stokes | Joe Carthy | John Dunnion
James Cogley | Nicola Stokes | Joe Carthy | John Dunnion
Alignment-HMM-based Extraction of Abbreviations from Biomedical Text
Dana Movshovitz-Attias | William W. Cohen
Dana Movshovitz-Attias | William W. Cohen
Medical diagnosis lost in translation – Analysis of uncertainty and negation expressions in English and Swedish clinical texts
Danielle L Mowery | Sumithra Velupillai | Wendy W Chapman
Danielle L Mowery | Sumithra Velupillai | Wendy W Chapman
A Hybrid Stepwise Approach for De-identifying Person Names in Clinical Documents
Oscar Ferrández | Brett South | Shuying Shen | Stéphane Meystre
Oscar Ferrández | Brett South | Shuying Shen | Stéphane Meystre
PubMed-Scale Event Extraction for Post-Translational Modifications, Epigenetics and Protein Structural Relations
Jari Björne | Sofie Van Landeghem | Sampo Pyysalo | Tomoko Ohta | Filip Ginter | Yves Van de Peer | Sophia Ananiadou | Tapio Salakoski
Jari Björne | Sofie Van Landeghem | Sampo Pyysalo | Tomoko Ohta | Filip Ginter | Yves Van de Peer | Sophia Ananiadou | Tapio Salakoski
New Resources and Perspectives for Biomedical Event Extraction
Sampo Pyysalo | Pontus Stenetorp | Tomoko Ohta | Jin-Dong Kim | Sophia Ananiadou
Sampo Pyysalo | Pontus Stenetorp | Tomoko Ohta | Jin-Dong Kim | Sophia Ananiadou
Combining Compositionality and Pagerank for the Identification of Semantic Relations between Biomedical Words
Thierry Hamon | Christopher Engström | Mounira Manser | Zina Badji | Natalia Grabar | Sergei Silvestrov
Thierry Hamon | Christopher Engström | Mounira Manser | Zina Badji | Natalia Grabar | Sergei Silvestrov
Domain Adaptation of Coreference Resolution for Radiology Reports
Emilia Apostolova | Noriko Tomuro | Pattanasak Mongkolwat | Dina Demner-Fushman
Emilia Apostolova | Noriko Tomuro | Pattanasak Mongkolwat | Dina Demner-Fushman
What can NLP tell us about BioNLP?
Attapol Thamrongrattanarit | Michael Shafir | Michael Crivaro | Bensiin Borukhov | Marie Meteer
Attapol Thamrongrattanarit | Michael Shafir | Michael Crivaro | Bensiin Borukhov | Marie Meteer
A Prototype Tool Set to Support Machine-Assisted Annotation
Brett South | Shuying Shen | Jianwei Leng | Tyler Forbush | Scott DuVall | Wendy Chapman
Brett South | Shuying Shen | Jianwei Leng | Tyler Forbush | Scott DuVall | Wendy Chapman
MedLingMap: A growing resource mapping the Bio-Medical NLP field
Marie Meteer | Bensiin Borukhov | Mike Crivaro | Michael Shafir | Attapol Thamrongrattanarit
Marie Meteer | Bensiin Borukhov | Mike Crivaro | Michael Shafir | Attapol Thamrongrattanarit
Exploring Label Dependency in Active Learning for Phenotype Mapping
Shefali Sharma | Leslie Lange | Jose Luis Ambite | Yigal Arens | Chun-Nan Hsu
Shefali Sharma | Leslie Lange | Jose Luis Ambite | Yigal Arens | Chun-Nan Hsu
Evaluating Joint Modeling of Yeast Biology Literature and Protein-Protein Interaction Networks
Ramnath Balasubramanyan | Kathryn Rivard | William W. Cohen | Jelena Jakovljevic | John L. Woolford
Ramnath Balasubramanyan | Kathryn Rivard | William W. Cohen | Jelena Jakovljevic | John L. Woolford
RankPref: Ranking Sentences Describing Relations between Biomedical Entities with an Application
Catalina Oana Tudor | K Vijay-Shanker
Catalina Oana Tudor | K Vijay-Shanker
Finding small molecule and protein pairs in scientific literature using a bootstrapping method
Ying Yan | Jee-Hyub Kim | Samuel Croset | Dietrich Rebholz-Schuhmann
Ying Yan | Jee-Hyub Kim | Samuel Croset | Dietrich Rebholz-Schuhmann
Classifying Gene Sentences in Biomedical Literature by Combining High-Precision Gene Identifiers
Sun Kim | Won Kim | Don Comeau | W. John Wilbur
Sun Kim | Won Kim | Don Comeau | W. John Wilbur
Effect of small sample size on text categorization with support vector machines
Pawel Matykiewicz | John Pestian
Pawel Matykiewicz | John Pestian
Using Natural Language Processing to Extract Drug-Drug Interaction Information from Package Inserts
Richard Boyce | Gregory Gardner | Henk Harkema
Richard Boyce | Gregory Gardner | Henk Harkema
Automatic Approaches for Gene-Drug Interaction Extraction from Biomedical Text: Corpus and Comparative Evaluation
Nate Sutton | Laura Wojtulewicz | Neel Mehta | Graciela Gonzalez
Nate Sutton | Laura Wojtulewicz | Neel Mehta | Graciela Gonzalez
up
Proceedings of the NAACL-HLT 2012 Workshop on Computational Linguistics for Literature
Proceedings of the NAACL-HLT 2012 Workshop on Computational Linguistics for Literature
David Elson | Anna Kazantseva | Rada Mihalcea | Stan Szpakowicz
David Elson | Anna Kazantseva | Rada Mihalcea | Stan Szpakowicz
Computational Analysis of Referring Expressions in Narratives of Picture Books
Choonkyu Lee | Smaranda Muresan | Karin Stromswold
Choonkyu Lee | Smaranda Muresan | Karin Stromswold
A Computational Analysis of Style, Affect, and Imagery in Contemporary Poetry
Justine Kao | Dan Jurafsky
Justine Kao | Dan Jurafsky
Unsupervised Stylistic Segmentation of Poetry with Change Curves and Extrinsic Features
Julian Brooke | Adam Hammond | Graeme Hirst
Julian Brooke | Adam Hammond | Graeme Hirst
Digitizing 18th-Century French Literature: Comparing transcription methods for a critical edition text
Ann Irvine | Laure Marcellesi | Afra Zomorodian
Ann Irvine | Laure Marcellesi | Afra Zomorodian
up
Proceedings of the Third Workshop on Speech and Language Processing for Assistive Technologies
Proceedings of the Third Workshop on Speech and Language Processing for Assistive Technologies
Jan Alexandersson | Peter Ljunglöf | Kathleen F. McCoy | Brian Roark | Annalu Waller
Jan Alexandersson | Peter Ljunglöf | Kathleen F. McCoy | Brian Roark | Annalu Waller
A free and open-source tool that reads movie subtitles aloud
Peter Ljunglöf | Sandra Derbring | Maria Olsson
Peter Ljunglöf | Sandra Derbring | Maria Olsson
WinkTalk: a demonstration of a multimodal speech synthesis platform linking facial expressions to expressive synthetic voices
Éva Székely | Zeeshan Ahmed | João P. Cabral | Julie Carson-Berndsen
Éva Székely | Zeeshan Ahmed | João P. Cabral | Julie Carson-Berndsen
Applying Prediction Techniques to Phoneme-based AAC Systems
Ha Trinh | Annalu Waller | Keith Vertanen | Per Ola Kristensson | Vicki L. Hanson
Ha Trinh | Annalu Waller | Keith Vertanen | Per Ola Kristensson | Vicki L. Hanson
Assisting Social Conversation between Persons with Alzheimer’s Disease and their Conversational Partners
Nancy L. Green | Curry Guinn | Ronnie Smith
Nancy L. Green | Curry Guinn | Ronnie Smith
Communication strategies for a computerized caregiver for individuals with Alzheimer’s disease
Frank Rudzicz | Rozanne Wilson | Alex Mihailidis | Elizabeth Rochon | Carol Leonard
Frank Rudzicz | Rozanne Wilson | Alex Mihailidis | Elizabeth Rochon | Carol Leonard
Generating Situated Assisting Utterances to Facilitate Tactile-Map Understanding: A Prototype System
Kris Lohmann | Ole Eichhorn | Timo Baumann
Kris Lohmann | Ole Eichhorn | Timo Baumann
up
Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction (AKBC-WEKEX)
Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction (AKBC-WEKEX)
James Fan | Raphael Hoffman | Aditya Kalyanpur | Sebastian Riedel | Fabian Suchanek | Partha Pratim Talukdar
James Fan | Raphael Hoffman | Aditya Kalyanpur | Sebastian Riedel | Fabian Suchanek | Partha Pratim Talukdar
Towards Distributed MCMC Inference in Probabilistic Knowledge Bases
Mathias Niepert | Christian Meilicke | Heiner Stuckenschmidt
Mathias Niepert | Christian Meilicke | Heiner Stuckenschmidt
Collectively Representing Semi-Structured Data from the Web
Bhavana Dalvi | William Cohen | Jamie Callan
Bhavana Dalvi | William Cohen | Jamie Callan
Automatic Evaluation of Relation Extraction Systems on Large-scale
Mirko Bronzi | Zhaochen Guo | Filipe Mesquita | Denilson Barbosa | Paolo Merialdo
Mirko Bronzi | Zhaochen Guo | Filipe Mesquita | Denilson Barbosa | Paolo Merialdo
Relabeling Distantly Supervised Training Data for Temporal Knowledge Base Population
Suzanne Tamang | Heng Ji
Suzanne Tamang | Heng Ji
Web Based Collection and Comparison of Cognitive Properties in English and Chinese
Bin Li | Jiajun Chen | Yingjie Zhang
Bin Li | Jiajun Chen | Yingjie Zhang
Population of a Knowledge Base for News Metadata from Unstructured Text and Web Data
Rosa Stern | Benoît Sagot
Rosa Stern | Benoît Sagot
Real-time Population of Knowledge Bases: Opportunities and Challenges
Ndapandula Nakashole | Gerhard Weikum
Ndapandula Nakashole | Gerhard Weikum
Adding Distributional Semantics to Knowledge Base Entities through Web-scale Entity Linking
Matt Gardner
Matt Gardner
A Context-Aware Approach to Entity Linking
Veselin Stoyanov | James Mayfield | Tan Xu | Douglas Oard | Dawn Lawrie | Tim Oates | Tim Finin
Veselin Stoyanov | James Mayfield | Tan Xu | Douglas Oard | Dawn Lawrie | Tim Oates | Tim Finin
Constructing a Textual KB from a Biology TextBook
Peter Clark | Phil Harrison | Niranjan Balasubramanian | Oren Etzioni
Peter Clark | Phil Harrison | Niranjan Balasubramanian | Oren Etzioni
Human-Machine Cooperation: Supporting User Corrections to Automatically Constructed KBs
Michael Wick | Karl Schultz | Andrew McCallum
Michael Wick | Karl Schultz | Andrew McCallum
Rel-grams: A Probabilistic Model of Relations in Text
Niranjan Balasubramanian | Stephen Soderland | Mausam | Oren Etzioni
Niranjan Balasubramanian | Stephen Soderland | Mausam | Oren Etzioni
Automatic Knowledge Base Construction using Probabilistic Extraction, Deductive Reasoning, and Human Feedback
Daisy Zhe Wang | Yang Chen | Sean Goldberg | Christan Grant | Kun Li
Daisy Zhe Wang | Yang Chen | Sean Goldberg | Christan Grant | Kun Li
up
Proceedings of the Seventh Workshop on Statistical Machine Translation
Proceedings of the Seventh Workshop on Statistical Machine Translation
Chris Callison-Burch | Philipp Koehn | Christof Monz | Matt Post | Radu Soricut | Lucia Specia
Chris Callison-Burch | Philipp Koehn | Christof Monz | Matt Post | Radu Soricut | Lucia Specia
Findings of the 2012 Workshop on Statistical Machine Translation
Chris Callison-Burch | Philipp Koehn | Christof Monz | Matt Post | Radu Soricut | Lucia Specia
Chris Callison-Burch | Philipp Koehn | Christof Monz | Matt Post | Radu Soricut | Lucia Specia
TerrorCat: a Translation Error Categorization-based MT Quality Metric
Mark Fishel | Rico Sennrich | Maja Popović | Ondřej Bojar
Mark Fishel | Rico Sennrich | Maja Popović | Ondřej Bojar
Quality estimation for Machine Translation output using linguistic analysis and decoding features
Eleftherios Avramidis
Eleftherios Avramidis
PRHLT Submission to the WMT12 Quality Estimation Task
Jesús González Rubio | Alberto Sanchis | Francisco Casacuberta
Jesús González Rubio | Alberto Sanchis | Francisco Casacuberta
Tree Kernels for Machine Translation Quality Estimation
Christian Hardmeier | Joakim Nivre | Jörg Tiedemann
Christian Hardmeier | Joakim Nivre | Jörg Tiedemann
LORIA System for the WMT12 Quality Estimation Shared Task
David Langlois | Sylvain Raybaud | Kamel Smaïli
David Langlois | Sylvain Raybaud | Kamel Smaïli
Quality Estimation: an experimental study using unsupervised similarity measures
Erwan Moreau | Carl Vogel
Erwan Moreau | Carl Vogel
The UPC Submission to the WMT 2012 Shared Task on Quality Estimation
Daniele Pighin | Meritxell González | Lluís Màrquez
Daniele Pighin | Meritxell González | Lluís Màrquez
Morpheme- and POS-based IBM1 and language model scores for translation quality estimation
Maja Popović
Maja Popović
DCU-Symantec Submission for the WMT 2012 Quality Estimation Task
Raphael Rubino | Jennifer Foster | Joachim Wagner | Johann Roturier | Rasul Samad Zadeh Kaljahi | Fred Hollowood
Raphael Rubino | Jennifer Foster | Joachim Wagner | Johann Roturier | Rasul Samad Zadeh Kaljahi | Fred Hollowood
The SDL Language Weaver Systems in the WMT12 Quality Estimation Shared Task
Radu Soricut | Nguyen Bach | Ziyuan Wang
Radu Soricut | Nguyen Bach | Ziyuan Wang
Combining Quality Prediction and System Selection for Improved Automatic Translation Output
Radu Soricut | Sushant Narsale
Radu Soricut | Sushant Narsale
Match without a Referee: Evaluating MT Adequacy without Reference Translations
Yashar Mehdad | Matteo Negri | Marcello Federico
Yashar Mehdad | Matteo Negri | Marcello Federico
Review of Hypothesis Alignment Algorithms for MT System Combination via Confusion Network Decoding
Antti-Veikko Rosti | Xiaodong He | Damianos Karakos | Gregor Leusch | Yuan Cao | Markus Freitag | Spyros Matsoukas | Hermann Ney | Jason Smith | Bing Zhang
Antti-Veikko Rosti | Xiaodong He | Damianos Karakos | Gregor Leusch | Yuan Cao | Markus Freitag | Spyros Matsoukas | Hermann Ney | Jason Smith | Bing Zhang
On Hierarchical Re-ordering and Permutation Parsing for Phrase-based Decoding
Colin Cherry | Robert C. Moore | Chris Quirk
Colin Cherry | Robert C. Moore | Chris Quirk
CCG Syntactic Reordering Models for Phrase-based Machine Translation
Dennis Nolan Mehay | Christopher Hardie Brew
Dennis Nolan Mehay | Christopher Hardie Brew
Using Categorial Grammar to Label Translation Rules
Jonathan Weese | Chris Callison-Burch | Adam Lopez
Jonathan Weese | Chris Callison-Burch | Adam Lopez
Using Syntactic Head Information in Hierarchical Phrase-Based Translation
Junhui Li | Zhaopeng Tu | Guodong Zhou | Josef van Genabith
Junhui Li | Zhaopeng Tu | Guodong Zhou | Josef van Genabith
Formemes in English-Czech Deep Syntactic MT
Ondřej Dušek | Zdeněk Žabokrtský | Martin Popel | Martin Majliš | Michal Novák | David Mareček
Ondřej Dušek | Zdeněk Žabokrtský | Martin Popel | Martin Majliš | Michal Novák | David Mareček
The TALP-UPC phrase-based translation systems for WMT12: Morphology simplification and domain adaptation
Lluís Formiga | Carlos A. Henríquez Q. | Adolfo Hernández | José B. Mariño | Enric Monte | José A. R. Fonollosa
Lluís Formiga | Carlos A. Henríquez Q. | Adolfo Hernández | José B. Mariño | Enric Monte | José A. R. Fonollosa
Joshua 4.0: Packing, PRO, and Paraphrases
Juri Ganitkevitch | Yuan Cao | Jonathan Weese | Matt Post | Chris Callison-Burch
Juri Ganitkevitch | Yuan Cao | Jonathan Weese | Matt Post | Chris Callison-Burch
QCRI at WMT12: Experiments in Spanish-English and German-English Machine Translation of News Text
Francisco Guzmán | Preslav Nakov | Ahmed Thabet | Stephan Vogel
Francisco Guzmán | Preslav Nakov | Ahmed Thabet | Stephan Vogel
The RWTH Aachen Machine Translation System for WMT 2012
Matthias Huck | Stephan Peitz | Markus Freitag | Malte Nuhn | Hermann Ney
Matthias Huck | Stephan Peitz | Markus Freitag | Malte Nuhn | Hermann Ney
Towards Effective Use of Training Data in Statistical Machine Translation
Philipp Koehn | Barry Haddow
Philipp Koehn | Barry Haddow
Joint WMT 2012 Submission of the QUAERO Project
Markus Freitag | Stephan Peitz | Matthias Huck | Hermann Ney | Jan Niehues | Teresa Herrmann | Alex Waibel | Hai-son Le | Thomas Lavergne | Alexandre Allauzen | Bianka Buschbeck | Josep Maria Crego | Jean Senellart
Markus Freitag | Stephan Peitz | Matthias Huck | Hermann Ney | Jan Niehues | Teresa Herrmann | Alex Waibel | Hai-son Le | Thomas Lavergne | Alexandre Allauzen | Bianka Buschbeck | Josep Maria Crego | Jean Senellart
LIMSI @ WMT12
Hai-Son Le | Thomas Lavergne | Alexandre Allauzen | Marianna Apidianaki | Li Gong | Aurélien Max | Artem Sokolov | Guillaume Wisniewski | François Yvon
Hai-Son Le | Thomas Lavergne | Alexandre Allauzen | Marianna Apidianaki | Li Gong | Aurélien Max | Artem Sokolov | Guillaume Wisniewski | François Yvon
The Karlsruhe Institute of Technology Translation Systems for the WMT 2012
Jan Niehues | Yuqi Zhang | Mohammed Mediani | Teresa Herrmann | Eunah Cho | Alex Waibel
Jan Niehues | Yuqi Zhang | Mohammed Mediani | Teresa Herrmann | Eunah Cho | Alex Waibel
Kriya - The SFU System for Translation Task at WMT-12
Majid Razmara | Baskaran Sankaran | Ann Clifton | Anoop Sarkar
Majid Razmara | Baskaran Sankaran | Ann Clifton | Anoop Sarkar
DEPFIX: A System for Automatic Correction of Czech MT Outputs
Rudolf Rosa | David Mareček | Ondřej Dušek
Rudolf Rosa | David Mareček | Ondřej Dušek
LIUM’s SMT Machine Translation Systems for WMT 2012
Christophe Servan | Patrik Lambert | Anthony Rousseau | Holger Schwenk | Loïc Barrault
Christophe Servan | Patrik Lambert | Anthony Rousseau | Holger Schwenk | Loïc Barrault
Selecting Data for English-to-Czech Machine Translation
Aleš Tamchyna | Petra Galuščáková | Amir Kamran | Miloš Stanojević | Ondřej Bojar
Aleš Tamchyna | Petra Galuščáková | Amir Kamran | Miloš Stanojević | Ondřej Bojar
Constructing Parallel Corpora for Six Indian Languages via Crowdsourcing
Matt Post | Chris Callison-Burch | Miles Osborne
Matt Post | Chris Callison-Burch | Miles Osborne
Twitter Translation using Translation-Based Cross-Lingual Retrieval
Laura Jehl | Felix Hieber | Stefan Riezler
Laura Jehl | Felix Hieber | Stefan Riezler
Evaluating the Learning Curve of Domain Adaptive Statistical Machine Translation Systems
Nicola Bertoldi | Mauro Cettolo | Marcello Federico | Christian Buck
Nicola Bertoldi | Mauro Cettolo | Marcello Federico | Christian Buck
Phrase Model Training for Statistical Machine Translation with Word Lattices of Preprocessing Alternatives
Joern Wuebker | Hermann Ney
Joern Wuebker | Hermann Ney
up
Proceedings of the ACL 2012 Joint Workshop on Statistical Parsing and Semantic Processing of Morphologically Rich Languages
Proceedings of the ACL 2012 Joint Workshop on Statistical Parsing and Semantic Processing of Morphologically Rich Languages
Marianna Apidianaki | Ido Dagan | Jennifer Foster | Yuval Marton | Djamé Seddah | Reut Tsarfaty
Marianna Apidianaki | Ido Dagan | Jennifer Foster | Yuval Marton | Djamé Seddah | Reut Tsarfaty
Probabilistic Lexical Generalization for French Dependency Parsing
Enrique Henestroza Anguiano | Marie Candito
Enrique Henestroza Anguiano | Marie Candito
Unsupervised frame based Semantic Role Induction: application to French and English
Alejandra Lorenzo | Christophe Cerisara
Alejandra Lorenzo | Christophe Cerisara
Machine Learning of Syntactic Attachment from Morphosyntactic and Semantic Co-occurrence Statistics
Szymon Acedański | Adam Slaski | Adam Przepiórkowski
Szymon Acedański | Adam Slaski | Adam Przepiórkowski
Combining Rule-Based and Statistical Syntactic Analyzers
Iakes Goenaga | Koldobika Gojenola | María Jesús Aranzabe | Arantza Díaz de Ilarraza | Kepa Bengoetxea
Iakes Goenaga | Koldobika Gojenola | María Jesús Aranzabe | Arantza Díaz de Ilarraza | Kepa Bengoetxea
Statistical Parsing of Spanish and Data Driven Lemmatization
Joseph Le Roux | Benoît Sagot | Djamé Seddah
Joseph Le Roux | Benoît Sagot | Djamé Seddah
Assigning Deep Lexical Types Using Structured Classifier Features for Grammatical Dependencies
João Silva | António Branco
João Silva | António Branco
up
Proceedings of the Sixth Linguistic Annotation Workshop
The Role of Linguistic Models and Language Annotation in Feature Selection for Machine Learning
James Pustejovsky
James Pustejovsky
Who Did What to Whom? A Contrastive Study of Syntacto-Semantic Dependencies
Angelina Ivanova | Stephan Oepen | Lilja Øvrelid | Dan Flickinger
Angelina Ivanova | Stephan Oepen | Lilja Øvrelid | Dan Flickinger
Exploiting naive vs expert discourse annotations: an experiment using lexical cohesion to predict Elaboration / Entity-Elaboration confusions
Clémentine Adam | Marianne Vergez-Couret
Clémentine Adam | Marianne Vergez-Couret
Pair Annotation: Adaption of Pair Programming to Corpus Annotation
Isin Demirşahin | İhsan Yalcinkaya | Deniz Zeyrek
Isin Demirşahin | İhsan Yalcinkaya | Deniz Zeyrek
Structured Named Entities in two distinct press corpora: Contemporary Broadcast News and Old Newspapers
Sophie Rosset | Cyril Grouin | Karën Fort | Olivier Galibert | Juliette Kahn | Pierre Zweigenbaum
Sophie Rosset | Cyril Grouin | Karën Fort | Olivier Galibert | Juliette Kahn | Pierre Zweigenbaum
Intra-Chunk Dependency Annotation : Expanding Hindi Inter-Chunk Annotated Treebank
Prudhvi Kosaraju | Bharat Ram Ambati | Samar Husain | Dipti Misra Sharma | Rajeev Sangal
Prudhvi Kosaraju | Bharat Ram Ambati | Samar Husain | Dipti Misra Sharma | Rajeev Sangal
A GrAF-compliant Indonesian Speech Recognition Web Service on the Language Grid for Transcription Crowdsourcing
Bayu Distiawan | Ruli Manurung
Bayu Distiawan | Ruli Manurung
Towards Adaptation of Linguistic Annotations to Scholarly Annotation Formalisms on the Semantic Web
Karin Verspoor | Kevin Livingston
Karin Verspoor | Kevin Livingston
Intonosyntactic Data Structures: The Rhapsodie Treebank of Spoken French
Kim Gerdes | Sylvain Kahane | Anne Lacheret | Paola Pietandrea | Arthur Truong
Kim Gerdes | Sylvain Kahane | Anne Lacheret | Paola Pietandrea | Arthur Truong
Annotation Schemes to Encode Domain Knowledge in Medical Narratives
Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Rui Li | Jeff B. Pelz | Pengcheng Shi | Anne Haake
Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Rui Li | Jeff B. Pelz | Pengcheng Shi | Anne Haake
Search Result Diversification Methods to Assist Lexicographers
Lars Borin | Markus Forsberg | Karin Friberg Heppin | Richard Johansson | Annika Kjellandsson
Lars Borin | Markus Forsberg | Karin Friberg Heppin | Richard Johansson | Annika Kjellandsson
Simultaneous error detection at two levels of syntactic annotation
Adam Przepiórkowski | Michał Lenart
Adam Przepiórkowski | Michał Lenart
Developing Learner Corpus Annotation for Korean Particle Errors
Sun-Hee Lee | Markus Dickinson | Ross Israel
Sun-Hee Lee | Markus Dickinson | Ross Israel
Annotating Archaeological Texts: An Example of Domain-Specific Annotation in the Humanities
Francesca Bonin | Fabio Cavulli | Aronne Noriller | Massimo Poesio | Egon W. Stemle
Francesca Bonin | Fabio Cavulli | Aronne Noriller | Massimo Poesio | Egon W. Stemle
AlvisAE: a collaborative Web text annotation editor for knowledge acquisition
Frédéric Papazian | Robert Bossy | Claire Nédellec
Frédéric Papazian | Robert Bossy | Claire Nédellec
up
Proceedings of the 3rd Workshop in Computational Approaches to Subjectivity and Sentiment Analysis
Proceedings of the 3rd Workshop in Computational Approaches to Subjectivity and Sentiment Analysis
Alexandra Balahur | Andres Montoyo | Patricio Martinez Barco | Ester Boldrini
Alexandra Balahur | Andres Montoyo | Patricio Martinez Barco | Ester Boldrini
Random Walk Weighting over SentiWordNet for Sentiment Polarity Detection on Twitter
Arturo Montejo-Ráez | Eugenio Martínez-Cámara | M. Teresa Martín-Valdivia | L. Alfonso Ureña-López
Arturo Montejo-Ráez | Eugenio Martínez-Cámara | M. Teresa Martín-Valdivia | L. Alfonso Ureña-López
Mining Sentiments from Tweets
Akshat Bakliwal | Piyush Arora | Senthil Madhappan | Nikhil Kapre | Mukesh Singh | Vasudeva Varma
Akshat Bakliwal | Piyush Arora | Senthil Madhappan | Nikhil Kapre | Mukesh Singh | Vasudeva Varma
SAMAR: A System for Subjectivity and Sentiment Analysis of Arabic Social Media
Muhammad Abdul-Mageed | Sandra Kuebler | Mona Diab
Muhammad Abdul-Mageed | Sandra Kuebler | Mona Diab
Opinum: statistical sentiment analysis for opinion classification
Boyan Bonev | Gema Ramírez-Sánchez | Sergio Ortiz Rojas
Boyan Bonev | Gema Ramírez-Sánchez | Sergio Ortiz Rojas
Sentimantics: Conceptual Spaces for Lexical Sentiment Polarity Representation with Contextuality
Amitava Das | Björn Gambäck
Amitava Das | Björn Gambäck
Unifying Local and Global Agreement and Disagreement Classification in Online Debates
Jie Yin | Nalin Narang | Paul Thomas | Cecile Paris
Jie Yin | Nalin Narang | Paul Thomas | Cecile Paris
Cross-discourse Development of Supervised Sentiment Analysis in the Clinical Domain
Phillip Smith | Mark Lee
Phillip Smith | Mark Lee
POLITICAL-ADS: An annotated corpus for modeling event-level evaluativity
Kevin Reschke | Pranav Anand
Kevin Reschke | Pranav Anand
up
Proceedings of the Workshop on Extra-Propositional Aspects of Meaning in Computational Linguistics
Proceedings of the Workshop on Extra-Propositional Aspects of Meaning in Computational Linguistics
Roser Morante | Caroline Sporleder
Roser Morante | Caroline Sporleder
Disfluencies as Extra-Propositional Indicators of Cognitive Processing
Kathryn Womack | Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Jeff B. Pelz | Pengcheng Shi | Anne Haake
Kathryn Womack | Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Jeff B. Pelz | Pengcheng Shi | Anne Haake
How do Negation and Modality Impact on Opinions?
Farah Benamara | Baptiste Chardon | Yannick Mathieu | Vladimir Popescu | Nicholas Asher
Farah Benamara | Baptiste Chardon | Yannick Mathieu | Vladimir Popescu | Nicholas Asher
Linking Uncertainty in Physicians’ Narratives to Diagnostic Correctness
Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Jeff B. Pelz | Pengcheng Shi | Anne Haake
Wilson McCoy | Cecilia Ovesdotter Alm | Cara Calvelli | Jeff B. Pelz | Pengcheng Shi | Anne Haake
Factuality Detection on the Cheap: Inferring Factuality for Increased Precision in Detecting Negated Events
Erik Velldal | Jonathon Read
Erik Velldal | Jonathon Read
Improving Speculative Language Detection using Linguistic Knowledge
Guillermo Moncecchi | Jean-Luc Minel | Dina Wonsever
Guillermo Moncecchi | Jean-Luc Minel | Dina Wonsever
Bridging the Gap Between Scope-based and Event-based Negation/Speculation Annotations: A Bridge Not Too Far
Pontus Stenetorp | Sampo Pyysalo | Tomoko Ohta | Sophia Ananiadou | Jun’ichi Tsujii
Pontus Stenetorp | Sampo Pyysalo | Tomoko Ohta | Sophia Ananiadou | Jun’ichi Tsujii
Statistical Modality Tagging from Rule-based Annotations and Crowdsourcing
Vinodkumar Prabhakaran | Michael Bloodgood | Mona Diab | Bonnie Dorr | Lori Levin | Christine D. Piatko | Owen Rambow | Benjamin Van Durme
Vinodkumar Prabhakaran | Michael Bloodgood | Mona Diab | Bonnie Dorr | Lori Levin | Christine D. Piatko | Owen Rambow | Benjamin Van Durme
up
Workshop Proceedings of TextGraphs-7: Graph-based Methods for Natural Language Processing
Workshop Proceedings of TextGraphs-7: Graph-based Methods for Natural Language Processing
Irina Matveeva | Ahmed Hassan | Gael Dias
Irina Matveeva | Ahmed Hassan | Gael Dias
A New Parametric Estimation Method for Graph-based Clustering
Javid Ebrahimi | Mohammad Saniee Abadeh
Javid Ebrahimi | Mohammad Saniee Abadeh
up
Proceedings of the Sixth Workshop on Syntax, Semantics and Structure in Statistical Translation
Proceedings of the Sixth Workshop on Syntax, Semantics and Structure in Statistical Translation
Marine Carpuat | Lucia Specia | Dekai Wu
Marine Carpuat | Lucia Specia | Dekai Wu
WSD for n-best reranking and local language modeling in SMT
Marianna Apidianaki | Guillaume Wisniewski | Artem Sokolov | Aurélien Max | François Yvon
Marianna Apidianaki | Guillaume Wisniewski | Artem Sokolov | Aurélien Max | François Yvon
Linguistically-Enriched Models for Bulgarian-to-English Machine Translation
Rui Wang | Petya Osenova | Kiril Simov
Rui Wang | Petya Osenova | Kiril Simov
Enriching Parallel Corpora for Statistical Machine Translation with Semantic Negation Rephrasing
Dominikus Wetzel | Francis Bond
Dominikus Wetzel | Francis Bond
Using Parallel Features in Parsing of Machine-Translated Sentences for Correction of Grammatical Errors
Rudolf Rosa | Ondřej Dušek | David Mareček | Martin Popel
Rudolf Rosa | Ondřej Dušek | David Mareček | Martin Popel
Unsupervised vs. supervised weight estimation for semantic MT evaluation metrics
Chi-kiu Lo | Dekai Wu
Chi-kiu Lo | Dekai Wu
Head Finalization Reordering for Chinese-to-Japanese Machine Translation
Dan Han | Katsuhito Sudoh | Xianchao Wu | Kevin Duh | Hajime Tsukada | Masaaki Nagata
Dan Han | Katsuhito Sudoh | Xianchao Wu | Kevin Duh | Hajime Tsukada | Masaaki Nagata
Extracting Semantic Transfer Rules from Parallel Corpora with SMT Phrase Aligners
Petter Haugereid | Francis Bond
Petter Haugereid | Francis Bond
Towards Probabilistic Acceptors and Transducers for Feature Structures
Daniel Quernheim | Kevin Knight
Daniel Quernheim | Kevin Knight
Using Domain-specific and Collaborative Resources for Term Translation
Mihael Arcan | Christian Federmann | Paul Buitelaar
Mihael Arcan | Christian Federmann | Paul Buitelaar
Improving Statistical Machine Translation through co-joining parts of verbal constructs in English-Hindi translation
Karunesh Kumar Arora | R Mahesh K Sinha
Karunesh Kumar Arora | R Mahesh K Sinha
up
Proceedings of the 4th Named Entity Workshop (NEWS) 2012
Whitepaper of NEWS 2012 Shared Task on Machine Transliteration
Min Zhang | Haizhou Li | A Kumaran | Ming Liu
Min Zhang | Haizhou Li | A Kumaran | Ming Liu
Report of NEWS 2012 Machine Transliteration Shared Task
Min Zhang | Haizhou Li | A Kumaran | Ming Liu
Min Zhang | Haizhou Li | A Kumaran | Ming Liu
Accurate Unsupervised Joint Named-Entity Extraction from Unaligned Parallel Text
Robert Munro | Christopher D. Manning
Robert Munro | Christopher D. Manning
Automatically generated NE tagged corpora for English and Hungarian
Eszter Simon | Dávid Márk Nemeskey
Eszter Simon | Dávid Márk Nemeskey
Rescoring a Phrase-based Machine Transliteration System with Recurrent Neural Network Language Models
Andrew Finch | Paul Dixon | Eiichiro Sumita
Andrew Finch | Paul Dixon | Eiichiro Sumita
Syllable-based Machine Transliteration with Extra Phrase Features
Chunyue Zhang | Tingting Li | Tiejun Zhao
Chunyue Zhang | Tingting Li | Tiejun Zhao
English-Korean Named Entity Transliteration Using Substring Alignment and Re-ranking Methods
Chun-Kai Wu | Yu-Chun Wang | Richard Tzong-Han Tsai
Chun-Kai Wu | Yu-Chun Wang | Richard Tzong-Han Tsai
up
Proceedings of the 11th International Workshop on Tree Adjoining Grammars and Related Formalisms (TAG+11)
Proceedings of the 11th International Workshop on Tree Adjoining Grammars and Related Formalisms (TAG+11)
Giorgio Satta | Chung-Hye Han
Giorgio Satta | Chung-Hye Han
Delayed Tree-Locality, Set-locality, and Clitic Climbing
Joan Chen-Main | Tonia Bleam | Aravind Joshi
Joan Chen-Main | Tonia Bleam | Aravind Joshi
Deriving syntax-semantics mappings: node linking, type shifting and scope ambiguity
Dennis Ryan Storoshenko | Robert Frank
Dennis Ryan Storoshenko | Robert Frank
Describing São Tomense Using a Tree-Adjoining Meta-Grammar
Emmanuel Schang | Denys Duchier | Brunelle Magnana Ekoukou | Yannick Parmentier | Simon Petitjean
Emmanuel Schang | Denys Duchier | Brunelle Magnana Ekoukou | Yannick Parmentier | Simon Petitjean
An Attempt Towards Learning Semantics: Distributional Learning of IO Context-Free Tree Grammars
Ryo Yoshinaka
Ryo Yoshinaka
A Formal Model for Plausible Dependencies in Lexicalized Tree Adjoining Grammar
Laura Kallmeyer | Marco Kuhlmann
Laura Kallmeyer | Marco Kuhlmann
Using FB-LTAG Derivation Trees to Generate Transformation-Based Grammar Exercises
Claire Gardent | Laura Perez-Beltrachini
Claire Gardent | Laura Perez-Beltrachini
PLCFRS Parsing Revisited: Restricting the Fan-Out to Two
Wolfgang Maier | Miriam Kaeshammer | Laura Kallmeyer
Wolfgang Maier | Miriam Kaeshammer | Laura Kallmeyer
State-Split for Hypergraphs with an Application to Tree Adjoining Grammars
Johannes Osterholzer | Torsten Stüber
Johannes Osterholzer | Torsten Stüber
up
Proceedings of the 3rd Workshop on South and Southeast Asian Natural Language Processing
Proceedings of the 3rd Workshop on South and Southeast Asian Natural Language Processing
Virach Sornlertlamvanich | Abbas Malik
Virach Sornlertlamvanich | Abbas Malik
Computational evidence that Hindi and Urdu share a grammar but not the lexicon
K.V.S Prasad | Shafqat Mumtaz Virk
K.V.S Prasad | Shafqat Mumtaz Virk
Semantic Relation Extraction from a Cultural Database
Canasai Kruengkrai | Virach Sornlertlamvanich | Watchira Buranasing | Thatsanee Charoenporn
Canasai Kruengkrai | Virach Sornlertlamvanich | Watchira Buranasing | Thatsanee Charoenporn
Bengali Question Classification: Towards Developing QA System
Somnath Banerjee | Sivaji Bandyopadhyay
Somnath Banerjee | Sivaji Bandyopadhyay
Morphological Analyzer for Kokborok
Khumbar Debbarma | Braja Gopal Patra | Dipankar Das | Sivaji Bandyopadhyay
Khumbar Debbarma | Braja Gopal Patra | Dipankar Das | Sivaji Bandyopadhyay
Comparing Different Criteria for Vietnamese Word Segmentation
Quy T. Nguyen | Ngan L.T. Nguyen | Yusuke Miyao
Quy T. Nguyen | Ngan L.T. Nguyen | Yusuke Miyao
A Light Weight Stemmer for Urdu Language: A Scarce Resourced Language
Sajjad Ahmad Khan | Waqas Anwar | Usama Ijaz Bajwa | Xuan Wang
Sajjad Ahmad Khan | Waqas Anwar | Usama Ijaz Bajwa | Xuan Wang
Domain Based Classification of Punjabi Text Documents using Ontology and Hybrid Based Approach
Nidhi Krail | Vishal Gupta
Nidhi Krail | Vishal Gupta
Using English Acoustic Models for Hindi Automatic Speech Recognition
Anik Dey | Ying Li | Pascale Fung
Anik Dey | Ying Li | Pascale Fung
BIS Annotation Standards With Reference to Konkani Language
Madhavi Sardesai | Jyoti Pawar | Shantaram Walawalikar | Edna Vaz
Madhavi Sardesai | Jyoti Pawar | Shantaram Walawalikar | Edna Vaz
Automatic Extraction of Compound Verbs from Bangla Corpora
Sibanshu Mukhopadhayay | Tirthankar Dasgupta | Manjira Sinha | Anupam Basu
Sibanshu Mukhopadhayay | Tirthankar Dasgupta | Manjira Sinha | Anupam Basu
Bidirectional Bengali Script and Meetei Mayek Transliteration of Web Based Manipuri News Corpus
Thoudam Doren Singh
Thoudam Doren Singh
Rule-based Machine Translation between Indonesian and Malaysian
Raymond Hendy Susanto | Septina Dian Larasati | Francis M. Tyers
Raymond Hendy Susanto | Septina Dian Larasati | Francis M. Tyers
Building Multilingual Search Index using open source framework
Arjun Atreya | Swapnil Chaudhari | Pushpak Bhattacharyya | Ganesh Ramakrishnan
Arjun Atreya | Swapnil Chaudhari | Pushpak Bhattacharyya | Ganesh Ramakrishnan
Error tracking in search engine development
Swapnil Chaudhari | Arjun Atreya V | Pushpak Bhattacharyya | Ganesh Ramakrishnan
Swapnil Chaudhari | Arjun Atreya V | Pushpak Bhattacharyya | Ganesh Ramakrishnan
up
Proceedings of the 3rd Workshop on Cognitive Aspects of the Lexicon
On discriminating fMRI representations of abstract WordNet taxonomic categories
Andrew Anderson | Tao Yuan | Brian Murphy | Massimo Poesio
Andrew Anderson | Tao Yuan | Brian Murphy | Massimo Poesio
Automatic index creation to support navigation in lexical graphs encoding part_of relations
Michael Zock | Debela Tesfaye
Michael Zock | Debela Tesfaye
Modeling Word Meaning: Distributional Semantics and the Corpus Quality-Quantity Trade-Off
Seshadri Sridharan | Brian Murphy
Seshadri Sridharan | Brian Murphy
Verb interpretation for basic action types: annotation, ontology induction and creation of prototypical scenes
Francesca Frontini | Irene De Felice | Fahad Khan | Irene Russo | Monica Monachini | Gloria Gagliardi | Alessandro Panunzi
Francesca Frontini | Irene De Felice | Fahad Khan | Irene Russo | Monica Monachini | Gloria Gagliardi | Alessandro Panunzi
Automatic Construction of a MultiWord Expressions Bilingual Lexicon: A Statistical Machine Translation Evaluation Perspective
Dhouha Bouamor | Nasredine Semmar | Pierre Zweigenbaum
Dhouha Bouamor | Nasredine Semmar | Pierre Zweigenbaum
Hand-Crafting a Lexical Network With a Knowledge-Based Graph Editor
Nabil Gader | Veronika Lux-Pogodalla | Alain Polguère
Nabil Gader | Veronika Lux-Pogodalla | Alain Polguère
A Procedural DTD Project for Dictionary Entry Parsing Described with Parameterized Grammars
Neculai Curteanu | Mihai Alex Moruz
Neculai Curteanu | Mihai Alex Moruz
Automatic Generation of the Universal Word Explanation from UNL Ontology
Khan Md. Anwarus Salam | Hiroshi Uchida | Tetsuro Nishino
Khan Md. Anwarus Salam | Hiroshi Uchida | Tetsuro Nishino
Building Multilingual Lexical Resources using Wordnets: Structure, Design and Implementation
Shikhar Kr. Sarma | Dibyajyoti Sarmah | Biswajit Brahma | Himadri Bharali | Mayashree Mahanta | Utpal Saikia
Shikhar Kr. Sarma | Dibyajyoti Sarmah | Biswajit Brahma | Himadri Bharali | Mayashree Mahanta | Utpal Saikia
A New Semantic Lexicon and Similarity Measure in Bangla
Manjira Sinha | Abhik Jana | Tirthankar Dasgupta | Anupam Basu
Manjira Sinha | Abhik Jana | Tirthankar Dasgupta | Anupam Basu
Where’s the meeting that was cancelled? existential implications of transitive verbs
Patricia Amaral | Valeria de Paiva | Cleo Condoravdi | Annie Zaenen
Patricia Amaral | Valeria de Paiva | Cleo Condoravdi | Annie Zaenen
up
Proceedings of the 10th Workshop on Asian Language Resources
Proceedings of the 10th Workshop on Asian Language Resources
Ruvan Weerasinghe | Sarmad Hussain | Virach Sornlertlamvanich | Rachel Edita O. Roxas
Ruvan Weerasinghe | Sarmad Hussain | Virach Sornlertlamvanich | Rachel Edita O. Roxas
Building Large Scale Text Corpus for Tibetan Natural Language Processing by Extracting Text from Web Pages
Huidan Liu | Minghua Nuo | Jian Wu | Yeping He
Huidan Liu | Minghua Nuo | Jian Wu | Yeping He
A Structured Approach for Building Assamese Corpus: Insights, Applications and Challenges
Shikhar Kr. Sarma | Himadri Bharali | Ambeswar Gogoi | Ratul Deka | Anup Kr. Barman
Shikhar Kr. Sarma | Himadri Bharali | Ambeswar Gogoi | Ratul Deka | Anup Kr. Barman
Corpus Building of Literary Lesser Rich Language-Bodo: Insights and Challenges
Biswajit Brahma | Anup Kr. Barman | Shikhar Kr. Sarma | Bhatima Boro
Biswajit Brahma | Anup Kr. Barman | Shikhar Kr. Sarma | Bhatima Boro
Repairing Bengali Verb Chunks for Improved Bengali to Hindi Machine Translation
Sanjay Chatterji | Nabanita Datta | Arnab Dhar | Biswanath Barik | Sudeshna Sarkar | Anupam Basu
Sanjay Chatterji | Nabanita Datta | Arnab Dhar | Biswanath Barik | Sudeshna Sarkar | Anupam Basu
Constrained Hidden Markov Model for Bilingual Keyword Pairs Alignment
Denny Cahyadi | Fabien Cromieres | Sadao Kurohashi
Denny Cahyadi | Fabien Cromieres | Sadao Kurohashi
N-gram and Gazetteer List Based Named Entity Recognition for Urdu: A Scarce Resourced Language
Faryal Jahangir | Waqas Anwar | Usama Ijaz Bajwa | Xuan Wang
Faryal Jahangir | Waqas Anwar | Usama Ijaz Bajwa | Xuan Wang
up
Proceedings of the 10th International Workshop on Finite State Methods and Natural Language Processing
Proceedings of the 10th International Workshop on Finite State Methods and Natural Language Processing
Iñaki Alegria | Mans Hulden
Iñaki Alegria | Mans Hulden
Effect of Language and Error Models on Efficiency of Finite-State Spell-Checking and Correction
Tommi A Pirinen | Sam Hardwick
Tommi A Pirinen | Sam Hardwick
Integrating Aspectually Relevant Properties of Verbs into a Morphological Analyzer for English
Katina Bontcheva
Katina Bontcheva
Finite-State Technology in a Verse-Making Tool
Manex Agirrezabal | Iñaki Alegria | Bertol Arrieta | Mans Hulden
Manex Agirrezabal | Iñaki Alegria | Bertol Arrieta | Mans Hulden
WFST-Based Grapheme-to-Phoneme Conversion: Open Source tools for Alignment, Model-Building and Decoding
Josef R. Novak | Nobuaki Minematsu | Keikichi Hirose
Josef R. Novak | Nobuaki Minematsu | Keikichi Hirose
Implementation of Replace Rules Using Preference Operator
Senka Drobac | Miikka Silfverberg | Anssi Yli-Jyrä
Senka Drobac | Miikka Silfverberg | Anssi Yli-Jyrä
First Approaches on Spanish Medical Record Classification Using Diagnostic Term to Class Transduction
A. Casillas | A. Díaz de Ilarraza | K. Gojenola | M. Oronoz | Alicia Pérez
A. Casillas | A. Díaz de Ilarraza | K. Gojenola | M. Oronoz | Alicia Pérez
Developing an Open-Source FST Grammar for Verb Chain Transfer in a Spanish-Basque MT System
Aingeru Mayor | Mans Hulden | Gorka Labaka
Aingeru Mayor | Mans Hulden | Gorka Labaka
Conversion of Procedural Morphologies to Finite-State Morphologies: A Case Study of Arabic
Mans Hulden | Younes Samih
Mans Hulden | Younes Samih
A Methodology for Obtaining Concept Graphs from Word Graphs
Marcos Calvo | Jon Ander Gómez | Lluís-F. Hurtado | Emilio Sanchis
Marcos Calvo | Jon Ander Gómez | Lluís-F. Hurtado | Emilio Sanchis
Finite-State Acoustic and Translation Model Composition in Statistical Speech Translation: Empirical Assessment
Alicia Pérez | M. Inés Torres | Francisco Casacuberta
Alicia Pérez | M. Inés Torres | Francisco Casacuberta
up
Proceedings of the Second CIPS-SIGHAN Joint Conference on Chinese Language Processing
A Language Modeling Approach to Identifying Code-Switched Sentences and Words
Liang-Chih Yu | Wei-Cheng He | Wei-Nan Chien
Liang-Chih Yu | Wei-Cheng He | Wei-Nan Chien
The CIPS-SIGHAN CLP 2012 ChineseWord Segmentation onMicroBlog Corpora Bakeoff
Huiming Duan | Zhifang Sui | Ye Tian | Wenjie Li
Huiming Duan | Zhifang Sui | Ye Tian | Wenjie Li
Word Segmentation on Chinese Mirco-Blog Data with a Linear-Time Incremental Model
Kaixu Zhang | Maosong Sun | Changle Zhou
Kaixu Zhang | Maosong Sun | Changle Zhou
Soochow University Word Segmenter for SIGHAN 2012 Bakeoff
Yan Fang | Zhongqing Wang | Shoushan Li | Zhongguo Li | Richen Xu | Leixin Cai
Yan Fang | Zhongqing Wang | Shoushan Li | Zhongguo Li | Richen Xu | Leixin Cai
CRFs-Based Chinese Word Segmentation for Micro-Blog with Small-Scale Data
Longyue Wang | Derek F. Wong | Lidia S. Chao | Junwen Xing
Longyue Wang | Derek F. Wong | Lidia S. Chao | Junwen Xing
A Cascaded Approach for CIPS-SIGHAN Micro-Blog Word Segmentation Bakeoff 2012
Bei Shi | Xianpei Han | Le Sun
Bei Shi | Xianpei Han | Le Sun
Adapting Conventional Chinese Word Segmenter for Segmenting Micro-blog Text: Combining Rule-based and Statistic-based Approaches
Ning Xi | Bin Li | Guangchao Tang | Shujian Huang | Yinggong Zhao | Hao Zhou | Xinyu Dai | Jiajun Chen
Ning Xi | Bin Li | Guangchao Tang | Shujian Huang | Yinggong Zhao | Hao Zhou | Xinyu Dai | Jiajun Chen
Rules-based Chinese Word Segmentation on MicroBlog for CIPS-SIGHAN on CLP2012
Jing Zhang | Degen Huang | Xia Han | Wei Wang
Jing Zhang | Degen Huang | Xia Han | Wei Wang
Micro blogs Oriented Word Segmentation System
Yijia Liu | Meishan Zhang | Wanxiang Che | Ting Liu | Yihe Deng
Yijia Liu | Meishan Zhang | Wanxiang Che | Ting Liu | Yihe Deng
A Comparison of Chinese Word Segmentation on News and Microblog Corpora with a Lexicon Based Method
Yuxiang Jia | Hongying Zan | Ming Fan | Zhimin Wang
Yuxiang Jia | Hongying Zan | Ming Fan | Zhimin Wang
A MMSM-based Hybrid Method for Chinese MicroBlog Word Segmentation
Xiao Sun | Chengcheng Li | Chenyi Tang | Jiaqi Ye
Xiao Sun | Chengcheng Li | Chenyi Tang | Jiaqi Ye
The Task 2 of CIPS-SIGHAN 2012 Named Entity Recognition and Disambiguation in Chinese Bakeoff
Zhengyan He | Houfeng Wang | Sujian Li
Zhengyan He | Houfeng Wang | Sujian Li
SIR-NERD: A Chinese Named Entity Recognition and Disambiguation System using a Two-Stage Method
Zehuan Peng | Le Sun | Xianpei Han
Zehuan Peng | Le Sun | Xianpei Han
A Template Based Hybrid Model for Chinese Personal Name Disambiguation
Hao Zong | Derek F. Wong | Lidia S. Chao
Hao Zong | Derek F. Wong | Lidia S. Chao
Attribute based Chinese Named Entity Recognition and Disambiguation
Wei Han | Guang Liu | Yuzhao Mao | Zhenni Huang
Wei Han | Guang Liu | Yuzhao Mao | Zhenni Huang
Chinese Name Disambiguation Based on Adaptive Clustering with the Attribute Features
Wei Tian | Xiao Pan | Zhengtao Yu | Yantuan Xian | Xiuzhen Yang
Wei Tian | Xiao Pan | Zhengtao Yu | Yantuan Xian | Xiuzhen Yang
Explore Chinese Encyclopedic Knowledge to Disambiguate Person Names
Jie Liu | Ruifeng Xu | Qin Lu | Jian Xu
Jie Liu | Ruifeng Xu | Qin Lu | Jian Xu
A Joint Chinese Named Entity Recognition and Disambiguation System
Longyue Wang | Shuo Li | Derek F. Wong | Lidia S. Chao
Longyue Wang | Shuo Li | Derek F. Wong | Lidia S. Chao
Chinese Personal Name Disambiguation Based on Vector Space Model
Qing-hu Fan | Hong-ying Zan | Yu-mei Chai | Yu-xiang Jia | Gui-ling Niu
Qing-hu Fan | Hong-ying Zan | Yu-mei Chai | Yu-xiang Jia | Gui-ling Niu
Multiple TreeBanks Integration for Chinese Phrase Structure Grammar Parsing Using Bagging
Meishan Zhang | Wanxiang Che | Ting Liu
Meishan Zhang | Wanxiang Che | Ting Liu
A Simplified Chinese Parser with Factored Model
Qiuping Huang | Liangye He | Derek F. Wong | Lidia S. Chao
Qiuping Huang | Liangye He | Derek F. Wong | Lidia S. Chao
Traditional Chinese Parsing Evaluation at SIGHAN Bake-offs 2012
Yuen-Hsien Tseng | Lung-Hao Lee | Liang-Chih Yu
Yuen-Hsien Tseng | Lung-Hao Lee | Liang-Chih Yu
Improving PCFG Chinese Parsing with Context-Dependent Probability Re-estimation
Yu-Ming Hsieh | Ming-Hong Bai | Jason S. Chang | Keh-Jiann Chen
Yu-Ming Hsieh | Ming-Hong Bai | Jason S. Chang | Keh-Jiann Chen
Sentence Parsing with Double Sequential Labeling in Traditional Chinese Parsing Task
Shih-Hung Wu | Hsien-You Hsieh | Liang-Pu Chen
Shih-Hung Wu | Hsien-You Hsieh | Liang-Pu Chen
up
Proceedings of the Third International Workshop on Free/Open-Source Rule-Based Machine Translation
Proceedings of the Third International Workshop on Free/Open-Source Rule-Based Machine Translation
Cristina España-Bonet | Aarne Ranta
Cristina España-Bonet | Aarne Ranta
The GF Eclipse Plugin provides an integrated development environment (IDE) for developing grammars in the Grammatical Framework (GF). Built on top of the Eclipse Platform, it aids grammar writing by providing instant syntax checking, semantic warnings and cross-reference resolution. Inline documentation and a library browser facilitate the use of existing resource libraries, and compilation and testing of grammars is greatly improved through single-click launch configurations and an in-built test case manager for running treebank regression tests. This IDE promotes grammar-based systems by making the tasks of writing grammars and using resource libraries more efficient, and provides powerful tools to reduce the barrier to entry to GF and encourage new users of the framework.
Choosing the correct paradigm for unknown words in rule-based machine translation systems
V. M. Sánchez-Cartagena | M. Esplà-Gomis | F. Sánchez-Martínez | J. A. Pérez-Ortiz
V. M. Sánchez-Cartagena | M. Esplà-Gomis | F. Sánchez-Martínez | J. A. Pérez-Ortiz
Previous work on an interactive system aimed at helping non-expert users to enlarge the monolingual dictionaries of rule-based machine translation (MT) systems worked by discarding those inflection paradigms that cannot generate a set of inflected word forms validated by the user. This method, however, cannot deal with the common case where a set of different paradigms generate exactly the same set of inflected word forms, although with different inflection information attached. In this paper, we propose the use of an n-gram-based model of lexical categories and inflection information to select a single paradigm in cases where more than one paradigm generates the same set of word forms. Results obtained with a Spanish monolingual dictionary show that the correct paradigm is chosen for around 75% of the unknown words, thus making the resulting system (available under an open-source license) of valuable help to enlarge the monolingual dictionaries used in MT involving non-expert users without technical linguistic knowledge.
An open-source toolkit for integrating shallow-transfer rules into phrase-based statistical machine translation
V. M. Sánchez-Cartagena | F. Sánchez-Martínez | J. A. Pérez-Ortiz
V. M. Sánchez-Cartagena | F. Sánchez-Martínez | J. A. Pérez-Ortiz
In this paper, we present an open-source toolkit to enrich a phrase-based statistical machine translation system (Moses) with phrase pairs generated from the linguistic resources of a shallow-transfer rule-based machine translation system (Apertium). A system built with this toolkit was not outperformed by any other participant in the shared translation task of the Sixth Workshop on Statistical Machine Translation (WMT 11) for the Spanish–English language pair.
A rule-based machine translation system from Serbo-Croatian to Macedonian
Hrvoje Peradin | Francis Tyers
Hrvoje Peradin | Francis Tyers
This paper describes the development of a one-way machine translation system from SerboCroatian to Macedonian on the Apertium platform. Details of resources and development methods are given, as well as an evaluation, and general directives for future work.
up
Proceedings of the 9th International Workshop on Spoken Language Translation: Evaluation Campaign
Overview of the IWSLT 2012 evaluation campaign
M. Federico | M. Cettolo | L. Bentivogli | M. Paul | S. Stüker
M. Federico | M. Cettolo | L. Bentivogli | M. Paul | S. Stüker
We report on the ninth evaluation campaign organized by the IWSLT workshop. This year, the evaluation offered multiple tracks on lecture translation based on the TED corpus, and one track on dialog translation from Chinese to English based on the Olympic trilingual corpus. In particular, the TED tracks included a speech transcription track in English, a speech translation track from English to French, and text translation tracks from English to French and from Arabic to English. In addition to the official tracks, ten unofficial MT tracks were offered that required translating TED talks into English from either Chinese, Dutch, German, Polish, Portuguese (Brazilian), Romanian, Russian, Slovak, Slovene, or Turkish. 16 teams participated in the evaluation and submitted a total of 48 primary runs. All runs were evaluated with objective metrics, while runs of the official translation tracks were also ranked by crowd-sourced judges. In particular, subjective ranking for the TED task was performed on a progress test which permitted direct comparison of the results from this year against the best results from the 2011 round of the evaluation campaign.
The NICT ASR system for IWSLT2012
Hitoshi Yamamoto | Youzheng Wu | Chien-Lin Huang | Xugang Lu | Paul R. Dixon | Shigeki Matsuda | Chiori Hori | Hideki Kashioka
Hitoshi Yamamoto | Youzheng Wu | Chien-Lin Huang | Xugang Lu | Paul R. Dixon | Shigeki Matsuda | Chiori Hori | Hideki Kashioka
This paper describes our automatic speech recognition (ASR) system for the IWSLT 2012 evaluation campaign. The target data of the campaign is selected from the TED talks, a collection of public speeches on a variety of topics spoken in English. Our ASR system is based on weighted finite-state transducers and exploits an combination of acoustic models for spontaneous speech, language models based on n-gram and factored recurrent neural network trained with effectively selected corpora, and unsupervised topic adaptation framework utilizing ASR results. Accordingly, the system achieved 10.6% and 12.0% word error rate for the tst2011 and tst2012 evaluation set, respectively.
The KIT translation systems for IWSLT 2012
Mohammed Mediani | Yuqi Zhang | Thanh-Le Ha | Jan Niehues | Eunah Cho | Teresa Herrmann | Rainer Kärgel | Alexander Waibel
Mohammed Mediani | Yuqi Zhang | Thanh-Le Ha | Jan Niehues | Eunah Cho | Teresa Herrmann | Rainer Kärgel | Alexander Waibel
In this paper, we present the KIT systems participating in the English-French TED Translation tasks in the framework of the IWSLT 2012 machine translation evaluation. We also present several additional experiments on the English-German, English-Chinese and English-Arabic translation pairs. Our system is a phrase-based statistical machine translation system, extended with many additional models which were proven to enhance the translation quality. For instance, it uses the part-of-speech (POS)-based reordering, translation and language model adaptation, bilingual language model, word-cluster language model, discriminative word lexica (DWL), and continuous space language model. In addition to this, the system incorporates special steps in the preprocessing and in the post-processing step. In the preprocessing the noisy corpora are filtered by removing the noisy sentence pairs, whereas in the postprocessing the agreement between a noun and its surrounding words in the French translation is corrected based on POS tags with morphological information. Our system deals with speech transcription input by removing case information and punctuation except periods from the text translation model.
The UEDIN systems for the IWSLT 2012 evaluation
Eva Hasler | Peter Bell | Arnab Ghoshal | Barry Haddow | Philipp Koehn | Fergus McInnes | Steve Renals | Pawel Swietojanski
Eva Hasler | Peter Bell | Arnab Ghoshal | Barry Haddow | Philipp Koehn | Fergus McInnes | Steve Renals | Pawel Swietojanski
This paper describes the University of Edinburgh (UEDIN) systems for the IWSLT 2012 Evaluation. We participated in the ASR (English), MT (English-French, German-English) and SLT (English-French) tracks.
The NAIST machine translation system for IWSLT2012
Graham Neubig | Kevin Duh | Masaya Ogushi | Takamoto Kano | Tetsuo Kiso | Sakriani Sakti | Tomoki Toda | Satoshi Nakamura
Graham Neubig | Kevin Duh | Masaya Ogushi | Takamoto Kano | Tetsuo Kiso | Sakriani Sakti | Tomoki Toda | Satoshi Nakamura
This paper describes the NAIST statistical machine translation system for the IWSLT2012 Evaluation Campaign. We participated in all TED Talk tasks, for a total of 11 language-pairs. For all tasks, we use the Moses phrase-based decoder and its experiment management system as a common base for building translation systems. The focus of our work is on performing a comprehensive comparison of a multitude of existing techniques for the TED task, exploring issues such as out-of-domain data filtering, minimum Bayes risk decoding, MERT vs. PRO tuning, word alignment combination, and morphology.
FBK’s machine translation systems for IWSLT 2012’s TED lectures
N. Ruiz | A. Bisazza | R. Cattoni | M. Federico
N. Ruiz | A. Bisazza | R. Cattoni | M. Federico
This paper reports on FBK’s Machine Translation (MT) submissions at the IWSLT 2012 Evaluation on the TED talk translation tasks. We participated in the English-French and the Arabic-, Dutch-, German-, and Turkish-English translation tasks. Several improvements are reported over our last year baselines. In addition to using fill-up combinations of phrase-tables for domain adaptation, we explore the use of corpora filtering based on cross-entropy to produce concise and accurate translation and language models. We describe challenges encountered in under-resourced languages (Turkish) and language-specific preprocessing needs.
The RWTH Aachen speech recognition and machine translation system for IWSLT 2012
Stephan Peitz | Saab Mansour | Markus Freitag | Minwei Feng | Matthias Huck | Joern Wuebker | Malte Nuhn | Markus Nußbaum-Thom | Hermann Ney
Stephan Peitz | Saab Mansour | Markus Freitag | Minwei Feng | Matthias Huck | Joern Wuebker | Malte Nuhn | Markus Nußbaum-Thom | Hermann Ney
In this paper, the automatic speech recognition (ASR) and statistical machine translation (SMT) systems of RWTH Aachen University developed for the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT) 2012 are presented. We participated in the ASR (English), MT (English-French, Arabic-English, Chinese-English, German-English) and SLT (English-French) tracks. For the MT track both hierarchical and phrase-based SMT decoders are applied. A number of different techniques are evaluated in the MT and SLT tracks, including domain adaptation via data selection, translation model interpolation, phrase training for hierarchical and phrase-based systems, additional reordering model, word class language model, various Arabic and Chinese segmentation methods, postprocessing of speech recognition output with an SMT system, and system combination. By application of these methods we can show considerable improvements over the respective baseline systems.
The HIT-LTRC machine translation system for IWSLT 2012
Xiaoning Zhu | Yiming Cui | Conghui Zhu | Tiejun Zhao | Hailong Cao
Xiaoning Zhu | Yiming Cui | Conghui Zhu | Tiejun Zhao | Hailong Cao
In this paper, we describe HIT-LTRC’s participation in the IWSLT 2012 evaluation campaign. In this year, we took part in the Olympics Task which required the participants to translate Chinese to English with limited data. Our system is based on Moses[1], which is an open source machine translation system. We mainly used the phrase-based models to carry out our experiments, and factored-based models were also performed in comparison. All the involved tools are freely available. In the evaluation campaign, we focus on data selection, phrase extraction method comparison and phrase table combination.
This paper reports on the participation of FBK at the IWSLT2012 evaluation campaign on automatic speech recognition: namely in the English ASR track. Both primary and contrastive submissions have been sent for evaluation. The ASR system features acoustic models trained on a portion of the TED talk recordings that was automatically selected according to the fidelity of the provided transcriptions. Three decoding steps are performed interleaved by acoustic feature normalization and acoustic model adaptation. A final rescoring step, based on the usage of an interpolated language model, is applied to word graphs generated in the third decoding step. For the primary submission, language models entering the interpolation are trained on both out-of-domain and in-domain text data, instead the contrastive submission uses both ”general purpose” and auxiliary language models trained only on out-of-domain text data. Despite this fact, similar performance are obtained with the two submissions.
The 2012 KIT and KIT-NAIST English ASR systems for the IWSLT evaluation
Christian Saam | Christian Mohr | Kevin Kilgour | Michael Heck | Matthias Sperber | Keigo Kubo | Sebatian Stüker | Sakriani Sakri | Graham Neubig | Tomoki Toda | Satoshi Nakamura | Alex Waibel
Christian Saam | Christian Mohr | Kevin Kilgour | Michael Heck | Matthias Sperber | Keigo Kubo | Sebatian Stüker | Sakriani Sakri | Graham Neubig | Tomoki Toda | Satoshi Nakamura | Alex Waibel
This paper describes our English Speech-to-Text (STT) systems for the 2012 IWSLT TED ASR track evaluation. The systems consist of 10 subsystems that are combinations of different front-ends, e.g. MVDR based and MFCC based ones, and two different phone sets. The outputs of the subsystems are combined via confusion network combination. Decoding is done in two stages, where the systems of the second stage are adapted in an unsupervised manner on the combination of the first stage outputs using VTLN, MLLR, and cM-LLR.
The KIT-NAIST (contrastive) English ASR system for IWSLT 2012
Michael Heck | Keigo Kubo | Matthias Sperber | Sakriani Sakti | Sebastian Stüker | Christian Saam | Kevin Kilgour | Christian Mohr | Graham Neubig | Tomoki Toda | Satoshi Nakamura | Alex Waibel
Michael Heck | Keigo Kubo | Matthias Sperber | Sakriani Sakti | Sebastian Stüker | Christian Saam | Kevin Kilgour | Christian Mohr | Graham Neubig | Tomoki Toda | Satoshi Nakamura | Alex Waibel
This paper describes the KIT-NAIST (Contrastive) English speech recognition system for the IWSLT 2012 Evaluation Campaign. In particular, we participated in the ASR track of the IWSLT TED task. The system was developed by Karlsruhe Institute of Technology (KIT) and Nara Institute of Science and Technology (NAIST) teams in collaboration within the interACT project. We employ single system decoding with fully continuous and semi-continuous models, as well as a three-stage, multipass system combination framework built with the Janus Recognition Toolkit. On the IWSLT 2010 test set our single system introduced in this work achieves a WER of 17.6%, and our final combination achieves a WER of 14.4%.
EBMT system of Kyoto University in OLYMPICS task at IWSLT 2012
Chenhui Chu | Toshiaki Nakazawa | Sadao Kurohashi
Chenhui Chu | Toshiaki Nakazawa | Sadao Kurohashi
This paper describes the EBMT system of Kyoto University that participated in the OLYMPICS task at IWSLT 2012. When translating very different language pairs such as Chinese-English, it is very important to handle sentences in tree structures to overcome the difference. Many recent studies incorporate tree structures in some parts of translation process, but not all the way from model training (alignment) to decoding. Our system is a fully tree-based translation system where we use the Bayesian phrase alignment model on dependency trees and example-based translation. To improve the translation quality, we conduct some special processing for the IWSLT 2012 OLYMPICS task, including sub-sentence splitting, non-parallel sentence filtering, adoption of an optimized Chinese segmenter and rule-based decoding constraints.
The LIG English to French machine translation system for IWSLT 2012
Laurent Besacier | Benjamin Lecouteux | Marwen Azouzi | Ngoc Quang Luong
Laurent Besacier | Benjamin Lecouteux | Marwen Azouzi | Ngoc Quang Luong
This paper presents the LIG participation to the E-F MT task of IWSLT 2012. The primary system proposed made a large improvement (more than 3 point of BLEU on tst2010 set) compared to our last year participation. Part of this improvment was due to the use of an extraction from the Gigaword corpus. We also propose a preliminary adaptation of the driven decoding concept for machine translation. This method allows an efficient combination of machine translation systems, by rescoring the log-linear model at the N-best list level according to auxiliary systems: the basis technique is essentially guiding the search using one or previous system outputs. The results show that the approach allows a significant improvement in BLEU score using Google translate to guide our own SMT system. We also try to use a confidence measure as an additional log-linear feature but we could not get any improvment with this technique.
The MIT-LL/AFRL IWSLT 2012 MT system
Jennifer Drexler | Wade Shen | Tim Anderson | Raymond Slyh | Brian Ore | Eric Hansen | Terry Gleason
Jennifer Drexler | Wade Shen | Tim Anderson | Raymond Slyh | Brian Ore | Eric Hansen | Terry Gleason
This paper describes the MIT-LL/AFRL statistical MT system and the improvements that were developed during the IWSLT 2012 evaluation campaign. As part of these efforts, we experimented with a number of extensions to the standard phrase-based model that improve performance on the Arabic to English and English to French TED-talk translation task. We also applied our existing ASR system to the TED-talk lecture ASR task, and combined our ASR and MT systems for the TED-talk SLT task. We discuss the architecture of the MIT-LL/AFRL MT system, improvements over our 2011 system, and experiments we ran during the IWSLT-2012 evaluation. Specifically, we focus on 1) cross-domain translation using MAP adaptation, 2) cross-entropy filtering of MT training data, and 3) improved Arabic morphology for MT preprocessing.
Minimum Bayes-risk decoding extended with similar examples: NAIST-NCT at IWSLT 2012
Hiroaki Shimizu | Masao Utiyama | Eiichiro Sumita | Satoshi Nakamura
Hiroaki Shimizu | Masao Utiyama | Eiichiro Sumita | Satoshi Nakamura
This paper describes our methods used in the NAIST-NICT submission to the International Workshop on Spoken Language Translation (IWSLT) 2012 evaluation campaign. In particular, we propose two extensions to minimum bayes-risk decoding which reduces a expected loss.
This paper presents efforts in preparation of the Polish-to-English SMT system for the TED lectures domain that is to be evaluated during the IWSLT 2012 Conference. Our attempts cover systems which use stems and morphological information on Polish words (using two different tools) and stems and POS.
Forest-to-string translation using binarized dependency forest for IWSLT 2012 OLYMPICS task
Hwidong Na | Jong-Hyeok Lee
Hwidong Na | Jong-Hyeok Lee
We participated in the OLYMPICS task in IWSLT 2012 and submitted two formal runs using a forest-to-string translation system. Our primary run achieved better translation quality than our contrastive run, but worse than a phrase-based and a hierarchical system using Moses.
Romanian to English automatic MT experiments at IWSLT12 – system description paper
Ştefan Daniel Dumitrescu | Radu Ion | Dan Ştefănescu | Tiberiu Boroş | Dan Tufiş
Ştefan Daniel Dumitrescu | Radu Ion | Dan Ştefănescu | Tiberiu Boroş | Dan Tufiş
The paper presents the system developed by RACAI for the ISWLT 2012 competition, TED task, MT track, Romanian to English translation. We describe the starting baseline phrase-based SMT system, the experiments conducted to adapt the language and translation models and our post-translation cascading system designed to improve the translation without external resources. We further present our attempts at creating a better controlled decoder than the open-source Moses system offers.
The TÜBİTAK statistical machine translation system for IWSLT 2012
Coşkun Mermer | Hamza Kaya | İlknur Durgar El-Kahlout | Mehmet Uğur Doğan
Coşkun Mermer | Hamza Kaya | İlknur Durgar El-Kahlout | Mehmet Uğur Doğan
WedescribetheTU ̈B ̇ITAKsubmissiontotheIWSLT2012 Evaluation Campaign. Our system development focused on utilizing Bayesian alignment methods such as variational Bayes and Gibbs sampling in addition to the standard GIZA++ alignments. The submitted tracks are the Arabic-English and Turkish-English TED Talks translation tasks.
up
Proceedings of the 9th International Workshop on Spoken Language Translation: Papers
Active error detection and resolution for speech-to-speech translation
Rohit Prasad | Rohit Kumar | Sankaranarayanan Ananthakrishnan | Wei Chen | Sanjika Hewavitharana | Matthew Roy | Frederick Choi | Aaron Challenner | Enoch Kan | Arvid Neelakantan | Prem Natarajan
Rohit Prasad | Rohit Kumar | Sankaranarayanan Ananthakrishnan | Wei Chen | Sanjika Hewavitharana | Matthew Roy | Frederick Choi | Aaron Challenner | Enoch Kan | Arvid Neelakantan | Prem Natarajan
We describe a novel two-way speech-to-speech (S2S) translation system that actively detects a wide variety of common error types and resolves them through user-friendly dialog with the user(s). We present algorithms for detecting out-of-vocabulary (OOV) named entities and terms, sense ambiguities, homophones, idioms, ill-formed input, etc. and discuss novel, interactive strategies for recovering from such errors. We also describe our approach for prioritizing different error types and an extensible architecture for implementing these decisions. We demonstrate the efficacy of our system by presenting analysis on live interactions in the English-to-Iraqi Arabic direction that are designed to invoke different error types for spoken language translation. Our analysis shows that the system can successfully resolve 47% of the errors, resulting in a dramatic improvement in the transfer of problematic concepts.
A method for translation of paralinguistic information
Takatomo Kano | Sakriani Sakti | Shinnosuke Takamichi | Graham Neubig | Tomoki Toda | Satoshi Nakamura
Takatomo Kano | Sakriani Sakti | Shinnosuke Takamichi | Graham Neubig | Tomoki Toda | Satoshi Nakamura
This paper is concerned with speech-to-speech translation that is sensitive to paralinguistic information. From the many different possible paralinguistic features to handle, in this paper we chose duration and power as a first step, proposing a method that can translate these features from input speech to the output speech in continuous space. This is done in a simple and language-independent fashion by training a regression model that maps source language duration and power information into the target language. We evaluate the proposed method on a digit translation task and show that paralinguistic information in input speech appears in output speech, and that this information can be used by target language speakers to detect emphasis.
We present a novel approach for continuous space language models in statistical machine translation by using Restricted Boltzmann Machines (RBMs). The probability of an n-gram is calculated by the free energy of the RBM instead of a feedforward neural net. Therefore, the calculation is much faster and can be integrated into the translation process instead of using the language model only in a re-ranking step. Furthermore, it is straightforward to introduce additional word factors into the language model. We observed a faster convergence in training if we include automatically generated word classes as an additional word factor. We evaluated the RBM-based language model on the German to English and English to French translation task of TED lectures. Instead of replacing the conventional n-gram-based language model, we trained the RBM-based language model on the more important but smaller in-domain data and combined them in a log-linear way. With this approach we could show improvements of about half a BLEU point on the translation task.
This paper describes a method for selecting text data from a corpus with the aim of training auxiliary Language Models (LMs) for an Automatic Speech Recognition (ASR) system. A novel similarity score function is proposed, which allows to score each document belonging to the corpus in order to select those with the highest scores for training auxiliary LMs which are linearly interpolated with the baseline one. The similarity score function makes use of ”similarity models” built from the automatic transcriptions furnished by earlier stages of the ASR system, while the documents selected for training auxiliary LMs are drawn from the same set of data used to train the baseline LM used in the ASR system. In this way, the resulting interpolated LMs are ”focused” towards the output of the recognizer itself. The approach allows to improve word error rate, measured on a task of spontaneous speech, of about 3% relative. It is important to note that a similar improvement has been obtained using an ”in-domain” set of texts data not contained in the sources used to train the baseline LM. In addition, we compared the proposed similarity score function with two other ones based on perplexity (PP) and on TFxIDF (Term Frequency x Inverse Document Frequency) vector space model. The proposed approach provides about the same performance as that based on TFxIDF model but requires both lower computation and occupation memory.
We present a Monte Carlo model to simulate human judgments in machine translation evaluation campaigns, such as WMT or IWSLT. We use the model to compare different ranking methods and to give guidance on the number of judgments that need to be collected to obtain sufficiently significant distinctions between systems.
Semi-supervised transliteration mining from parallel and comparable corpora
Walid Aransa | Holger Schwenk | Loic Barrault
Walid Aransa | Holger Schwenk | Loic Barrault
Transliteration is the process of writing a word (mainly proper noun) from one language in the alphabet of another language. This process requires mapping the pronunciation of the word from the source language to the closest possible pronunciation in the target language. In this paper we introduce a new semi-supervised transliteration mining method for parallel and comparable corpora. The method is mainly based on a new suggested Three Levels of Similarity (TLS) scores to extract the transliteration pairs. The first level calculates the similarity of of all vowel letters and consonants letters. The second level calculates the similarity of long vowels and vowel letters at beginning and end position of the words and consonants letters. The third level calculates the similarity consonants letters only. We applied our method on Arabic-English parallel and comparable corpora. We evaluated the extracted transliteration pairs using a statistical based transliteration system. This system is built using letters instead or words as tokens. The transliteration system achieves an accuracy of 0.50 and a mean F-score 0.8958 when trained on transliteration pairs extracted from a parallel corpus. The accuracy is 0.30 and the mean F-score 0.84 when we used instead a comparable corpus to automatically extract the transliteration pairs. This shows that the proposed semi-supervised transliteration mining algorithm is effective and can be applied to other language pairs. We also evaluated two segmentation techniques and reported the impact on the transliteration performance.
A simple and effective weighted phrase extraction for machine translation adaptation
Saab Mansour | Hermann Ney
Saab Mansour | Hermann Ney
The task of domain-adaptation attempts to exploit data mainly drawn from one domain (e.g. news) to maximize the performance on the test domain (e.g. weblogs). In previous work, weighting the training instances was used for filtering dissimilar data. We extend this by incorporating the weights directly into the standard phrase training procedure of statistical machine translation (SMT). This allows the SMT system to make the decision whether to use a phrase translation pair or not, a more methodological way than discarding phrase pairs completely when using filtering. Furthermore, we suggest a combined filtering and weighting procedure to achieve better results while reducing the phrase table size. The proposed methods are evaluated in the context of Arabicto-English translation on various conditions, where significant improvements are reported when using the suggested weighted phrase training. The weighting method also improves over filtering, and the combined filtering and weighting is better than a standalone filtering method. Finally, we experiment with mixture modeling, where additional improvements are reported when using weighted phrase extraction over a variety of baselines.
Applications of data selection via cross-entropy difference for real-world statistical machine translation
Amittai Axelrod | QingJun Li | William D. Lewis
Amittai Axelrod | QingJun Li | William D. Lewis
We broaden the application of data selection methods for domain adaptation to a larger number of languages, data, and decoders than shown in previous work, and explore comparable applications for both monolingual and bilingual cross-entropy difference methods. We compare domain adapted systems against very large general-purpose systems for the same languages, and do so without a bias to a particular direction. We present results against real-world generalpurpose systems tuned on domain-specific data, which are substantially harder to beat than standard research baseline systems. We show better performance for nearly all domain adapted systems, despite the fact that the domainadapted systems are trained on a fraction of the content of their general domain counterparts. The high performance of these methods suggest applicability to a wide variety of contexts, particularly in scenarios where only small supplies of unambiguously domain-specific data are available, yet it is believed that additional similar data is included in larger heterogenous-content general-domain corpora.
Although statistical machine translation (SMT) has made great progress since it came into being, the translation of numerical and time expressions is still far from satisfactory. Generally speaking, numbers are likely to be out-of-vocabulary (OOV) words due to their non-exhaustive characteristics even when the size of training data is very large, so it is difficult to obtain accurate translation results for the infinite set of numbers only depending on traditional statistical methods. We propose a language-independent framework to recognize and translate numbers more precisely by using a rule-based method. Through designing operators, we succeed to make rules educible and totally separate from codes, thus, we can extend rules to various language-pairs without re-coding, which contributes a lot to the efficient development of an SMT system with good portability. We classify numbers and time expressions into seven types, which are Arabic number, cardinal numbers, ordinal numbers, date, time of day, day of week and figures. A greedy algorithm is developed to deal with rule conflicts. Experiments have shown that our approach can significantly improve the translation performance.
Evaluation of interactive user corrections for lecture transcription
Heinrich Kolkhorst | Kevin Kilgour | Sebastian Stüker | Alex Waibel
Heinrich Kolkhorst | Kevin Kilgour | Sebastian Stüker | Alex Waibel
In this work, we present and evaluate the usage of an interactive web interface for browsing and correcting lecture transcripts. An experiment performed with potential users without transcription experience provides us with a set of example corrections. On German lecture data, user corrections greatly improve the comprehensibility of the transcripts, yet only reduce the WER to 22%. The precision of user edits is relatively low at 77% and errors in inflection, case and compounds were rarely corrected. Nevertheless, characteristic lecture data errors, such as highly specific terms, were typically corrected, providing valuable additional information.
Factored recurrent neural network language model in TED lecture transcription
Youzheng Wu | Hitoshi Yamamoto | Xugang Lu | Shigeki Matsuda | Chiori Hori | Hideki Kashioka
Youzheng Wu | Hitoshi Yamamoto | Xugang Lu | Shigeki Matsuda | Chiori Hori | Hideki Kashioka
In this study, we extend recurrent neural network-based language models (RNNLMs) by explicitly integrating morphological and syntactic factors (or features). Our proposed RNNLM is called a factored RNNLM that is expected to enhance RNNLMs. A number of experiments are carried out on top of state-of-the-art LVCSR system that show the factored RNNLM improves the performance measured by perplexity and word error rate. In the IWSLT TED test data sets, absolute word error rate reductions over RNNLM and n-gram LM are 0.4∼0.8 points.
Incremental adaptation using translation information and post-editing analysis
Frédéric Blain | Holger Schwenk | Jean Senellart
Frédéric Blain | Holger Schwenk | Jean Senellart
It is well known that statistical machine translation systems perform best when they are adapted to the task. In this paper we propose new methods to quickly perform incremental adaptation without the need to obtain word-by-word alignments from GIZA or similar tools. The main idea is to use an automatic translation as pivot to infer alignments between the source sentence and the reference translation, or user correction. We compared our approach to the standard method to perform incremental re-training. We achieve similar results in the BLEU score using less computational resources. Fast retraining is particularly interesting when we want to almost instantly integrate user feed-back, for instance in a post-editing context or machine translation assisted CAT tool. We also explore several methods to combine the translation models.
In this paper, we study the incorporation of statistical machine translation models to automatic speech recognition models in the framework of computer-assisted translation. The system is given a source language text to be translated and it shows the source text to the human translator to translate it orally. The system captures the user speech which is the dictation of the target language sentence. Then, the human translator uses an interactive-predictive process to correct the system generated errors. We show the efficiency of this method by higher human productivity gain compared to the baseline systems: pure ASR system and integrated ASR and MT systems.
MDI adaptation for the lazy: avoiding normalization in LM adaptation for lecture translation
Nick Ruiz | Marcello Federico
Nick Ruiz | Marcello Federico
This paper provides a fast alternative to Minimum Discrimination Information-based language model adaptation for statistical machine translation. We provide an alternative to computing a normalization term that requires computing full model probabilities (including back-off probabilities) for all n-grams. Rather than re-estimating an entire language model, our Lazy MDI approach leverages a smoothed unigram ratio between an adaptation text and the background language model to scale only the n-gram probabilities corresponding to translation options gathered by the SMT decoder. The effects of the unigram ratio are scaled by adding an additional feature weight to the log-linear discriminative model. We present results on the IWSLT 2012 TED talk translation task and show that Lazy MDI provides comparable language model adaptation performance to classic MDI.
Segmentation and punctuation prediction in speech language translation using a monolingual translation system
Eunah Cho | Jan Niehues | Alex Waibel
Eunah Cho | Jan Niehues | Alex Waibel
In spoken language translation (SLT), finding proper segmentation and reconstructing punctuation marks are not only significant but also challenging tasks. In this paper we present our recent work on speech translation quality analysis for German-English by improving sentence segmentation and punctuation. From oracle experiments, we show an upper bound of translation quality if we had human-generated segmentation and punctuation on the output stream of speech recognition systems. In our oracle experiments we gain 1.78 BLEU points of improvements on the lecture test set. We build a monolingual translation system from German to German implementing segmentation and punctuation prediction as a machine translation task. Using the monolingual translation system we get an improvement of 1.53 BLEU points on the lecture test set, which is a comparable performance against the upper bound drawn by the oracle experiments.
Sequence labeling-based reordering model for phrase-based SMT
Minwei Feng | Jan-Thorsten Peter | Hermann Ney
Minwei Feng | Jan-Thorsten Peter | Hermann Ney
For current statistical machine translation system, reordering is still a major problem for language pairs like Chinese-English, where the source and target language have significant word order differences. In this paper, we propose a novel reordering model based on sequence labeling techniques. Our model converts the reordering problem into a sequence labeling problem, i.e. a tagging task. For the given source sentence, we assign each source token a label which contains the reordering information for that token. We also design an unaligned word tag so that the unaligned word phenomenon is automatically implanted in the proposed model. Our reordering model is conditioned on the whole source sentence. Hence it is able to catch the long dependency in the source sentence. Although the learning on large scale task requests notably amounts of computational resources, the decoder makes use of the tagging information as soft constraints. Therefore, the training procedure of our model is computationally expensive for large task while in the test phase (during translation) our model is very efficient. We carried out experiments on five Chinese-English NIST tasks trained with BOLT data. Results show that our model improves the baseline system by 1.32 BLEU 1.53 TER on average.
We present a new approach to domain adaptation for SMT that enriches standard phrase-based models with lexicalised word and phrase pair features to help the model select appropriate translations for the target domain (TED talks). In addition, we show how source-side sentence-level topics can be incorporated to make the features differentiate between more fine-grained topics within the target domain (topic adaptation). We compare tuning our sparse features on a development set versus on the entire in-domain corpus and introduce a new method of porting them to larger mixed-domain models. Experimental results show that our features improve performance over a MIRA baseline and that in some cases we can get additional improvements with topic features. We evaluate our methods on two language pairs, English-French and German-English, showing promising results.
Spoken language translation using automatically transcribed text in training
Stephan Peitz | Simon Wiesler | Markus Nußbaum-Thom | Hermann Ney
Stephan Peitz | Simon Wiesler | Markus Nußbaum-Thom | Hermann Ney
In spoken language translation a machine translation system takes speech as input and translates it into another language. A standard machine translation system is trained on written language data and expects written language as input. In this paper we propose an approach to close the gap between the output of automatic speech recognition and the input of machine translation by training the translation system on automatically transcribed speech. In our experiments we show improvements of up to 0.9 BLEU points on the IWSLT 2012 English-to-French speech translation task.
Towards a better understanding of statistical post-editing
Marion Potet | Laurent Besacier | Hervé Blanchon | Marwen Azouzi
Marion Potet | Laurent Besacier | Hervé Blanchon | Marwen Azouzi
We describe several experiments to better understand the usefulness of statistical post-edition (SPE) to improve phrase-based statistical MT (PBMT) systems raw outputs. Whatever the size of the training corpus, we show that SPE systems trained on general domain data offers no breakthrough to our baseline general domain PBMT system. However, using manually post-edited system outputs to train the SPE led to a slight improvement in the translations quality compared with the use of professional reference translations. We also show that SPE is far more effective for domain adaptation, mainly because it recovers a lot of specific terms unknown to our general PBMT system. Finally, we compare two domain adaptation techniques, post-editing a general domain PBMT system vs building a new domain-adapted PBMT system with two different techniques, and show that the latter outperforms the first one. Yet, when the PBMT is a “black box”, SPE trained with post-edited system outputs remains an interesting option for domain adaptation.
Adaptation for Machine Translation has been studied in a variety of ways, using an ideal scenario where the training data can be split into ”out-of-domain” and ”in-domain” corpora, on which the adaptation is based. In this paper, we consider a more realistic setting which does not assume the availability of any kind of ”in-domain” data, hence the name ”any-text translation”. In this context, we present a new approach to contextually adapt a translation model onthe-fly, and present several experimental results where this approach outperforms conventionaly trained baselines. We also present a document-level contrastive evaluation whose results can be easily interpreted, even by non-specialists.
up
Proceedings of the Australasian Language Technology Association Workshop 2012
Proceedings of the Australasian Language Technology Association Workshop 2012
Paul Cook | Scott Nowson
Paul Cook | Scott Nowson
Using a large annotated historical corpus to study word-specific effects in sound change
Jennifer Hay
Jennifer Hay
Diverse Words, Shared Meanings: Statistical Machine Translation for Paraphrase, Grounding, and Intent
Chris Brockett
Chris Brockett
A Citation Centric Annotation Scheme for Scientific Articles
Angrosh M.A. | Stephen Cranefield | Nigel Stanger
Angrosh M.A. | Stephen Cranefield | Nigel Stanger
Semantic Judgement of Medical Concepts: Combining Syntagmatic and Paradigmatic Information with the Tensor Encoding Model
Michael Symonds | Guido Zuccon | Bevan Koopman | Peter Bruza | Anthony Nguyen
Michael Symonds | Guido Zuccon | Bevan Koopman | Peter Bruza | Anthony Nguyen
Active Learning and the Irish Treebank
Teresa Lynn | Jennifer Foster | Mark Dras | Elaine Uí Dhonnchadha
Teresa Lynn | Jennifer Foster | Mark Dras | Elaine Uí Dhonnchadha
Experimental Evaluation of a Lexicon- and Corpus-based Ensemble for Multi-way Sentiment Analysis
Minh Duc Cao | Ingrid Zukerman
Minh Duc Cao | Ingrid Zukerman
Segmentation and Translation of Japanese Multi-word Loanwords
James Breen | Timothy Baldwin | Francis Bond
James Breen | Timothy Baldwin | Francis Bond
Measurement of Progress in Machine Translation
Yvette Graham | Timothy Baldwin | Aaron Harwood | Alistair Moffat | Justin Zobel
Yvette Graham | Timothy Baldwin | Aaron Harwood | Alistair Moffat | Justin Zobel
Towards Two-step Multi-document Summarisation for Evidence Based Medicine: A Quantitative Analysis
Abeed Sarker | Diego Mollá-Aliod | Cécile Paris
Abeed Sarker | Diego Mollá-Aliod | Cécile Paris
In Your Eyes: Identifying Clichés in Song Lyrics
Alex G. Smith | Christopher X. S. Zee | Alexandra L. Uitdenbogerd
Alex G. Smith | Christopher X. S. Zee | Alexandra L. Uitdenbogerd
Free-text input vs menu selection: exploring the difference with a tutorial dialogue system.
Jenny Mcdonald | Alistair Knott | Richard Zeng
Jenny Mcdonald | Alistair Knott | Richard Zeng
Classification of Study Region in Environmental Science Abstracts
Jared Willett | Timothy Baldwin | David Martinez | Angus Webb
Jared Willett | Timothy Baldwin | David Martinez | Angus Webb
up
Proceedings of ACL 2012 Student Research Workshop
Proceedings of ACL 2012 Student Research Workshop
Jackie C. K. Cheung | Jun Hatori | Carlos Henriquez | Ann Irvine
Jackie C. K. Cheung | Jun Hatori | Carlos Henriquez | Ann Irvine
A Broad Evaluation of Techniques for Automatic Acquisition of Multiword Expressions
Carlos Ramisch | Vitor De Araujo | Aline Villavicencio
Carlos Ramisch | Vitor De Araujo | Aline Villavicencio
Domain Adaptation of a Dependency Parser with a Class-Class Selectional Preference Model
Raphael Cohen | Yoav Goldberg | Michael Elhadad
Raphael Cohen | Yoav Goldberg | Michael Elhadad
up
Joint Conference on EMNLP and CoNLL - Shared Task
Joint Conference on EMNLP and CoNLL - Shared Task
Sameer Pradhan | Alessandro Moschitti | Nianwen Xue
Sameer Pradhan | Alessandro Moschitti | Nianwen Xue
CoNLL-2012 Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes
Sameer Pradhan | Alessandro Moschitti | Nianwen Xue | Olga Uryupina | Yuchen Zhang
Sameer Pradhan | Alessandro Moschitti | Nianwen Xue | Olga Uryupina | Yuchen Zhang
Latent Structure Perceptron with Feature Induction for Unrestricted Coreference Resolution
Eraldo Fernandes | Cícero dos Santos | Ruy Milidiú
Eraldo Fernandes | Cícero dos Santos | Ruy Milidiú
Data-driven Multilingual Coreference Resolution using Resolver Stacking
Anders Björkelund | Richárd Farkas
Anders Björkelund | Richárd Farkas
Combining the Best of Two Worlds: A Hybrid Approach to Multilingual Coreference Resolution
Chen Chen | Vincent Ng
Chen Chen | Vincent Ng
Using Syntactic Dependencies to Solve Coreferences
Marcus Stamborg | Dennis Medved | Peter Exner | Pierre Nugues
Marcus Stamborg | Dennis Medved | Peter Exner | Pierre Nugues
A Mixed Deterministic Model for Coreference Resolution
Bo Yuan | Qingcai Chen | Yang Xiang | Xiaolong Wang | Liping Ge | Zengjian Liu | Meng Liao | Xianbo Si
Bo Yuan | Qingcai Chen | Yang Xiang | Xiaolong Wang | Liping Ge | Zengjian Liu | Meng Liao | Xianbo Si
Simple Maximum Entropy Models for Multilingual Coreference Resolution
Xinxin Li | Xuan Wang | Xingwei Liao
Xinxin Li | Xuan Wang | Xingwei Liao
UBIU for Multilingual Coreference Resolution in OntoNotes
Desislava Zhekova | Sandra Kübler | Joshua Bonner | Marwa Ragheb | Yu-Yin Hsu
Desislava Zhekova | Sandra Kübler | Joshua Bonner | Marwa Ragheb | Yu-Yin Hsu
A Multigraph Model for Coreference Resolution
Sebastian Martschat | Jie Cai | Samuel Broscheit | Éva Mújdricza-Maydt | Michael Strube
Sebastian Martschat | Jie Cai | Samuel Broscheit | Éva Mújdricza-Maydt | Michael Strube
Incorporating Rule-based and Statistic-based Techniques for Coreference Resolution
Ruifeng Xu | Jun Xu | Jie Liu | Chengxiang Liu | Chengtian Zou | Lin Gui | Yanzhen Zheng | Peng Qu
Ruifeng Xu | Jun Xu | Jie Liu | Chengxiang Liu | Chengtian Zou | Lin Gui | Yanzhen Zheng | Peng Qu
Illinois-Coref: The UI System in the CoNLL-2012 Shared Task
Kai-Wei Chang | Rajhans Samdani | Alla Rozovskaya | Mark Sammons | Dan Roth
Kai-Wei Chang | Rajhans Samdani | Alla Rozovskaya | Mark Sammons | Dan Roth
System paper for CoNLL-2012 shared task: Hybrid Rule-based Algorithm for Coreference Resolution.
Heming Shou | Hai Zhao
Heming Shou | Hai Zhao
up
INLG 2012 Proceedings of the Seventh International Natural Language Generation Conference
INLG 2012 Proceedings of the Seventh International Natural Language Generation Conference
Barbara Di Eugenio | Susan McRoy
Barbara Di Eugenio | Susan McRoy
Expressive NLG for Next-Generation Learning Environments: Language, Affect, and Narrative
James Lester
James Lester
Learning Preferences for Referring Expression Generation: Effects of Domain, Language and Algorithm
Ruud Koolen | Emiel Krahmer | Mariët Theune
Ruud Koolen | Emiel Krahmer | Mariët Theune
Referring in Installments: A Corpus Study of Spoken Object References in an Interactive Virtual Environment
Kristina Striegnitz | Hendrik Buschmeier | Stefan Kopp
Kristina Striegnitz | Hendrik Buschmeier | Stefan Kopp
MinkApp: Generating Spatio-temporal Summaries for Nature Conservation Volunteers
Nava Tintarev | Yolanda Melero | Somayajulu Sripada | Elizabeth Tait | Rene Van Der Wal | Chris Mellish
Nava Tintarev | Yolanda Melero | Somayajulu Sripada | Elizabeth Tait | Rene Van Der Wal | Chris Mellish
Perceptions of Alignment and Personality in Generated Dialogue
Alastair Gill | Carsten Brockmann | Jon Oberlander
Alastair Gill | Carsten Brockmann | Jon Oberlander
Optimising Incremental Generation for Spoken Dialogue Systems: Reducing the Need for Fillers
Nina Dethlefs | Helen Hastie | Verena Rieser | Oliver Lemon
Nina Dethlefs | Helen Hastie | Verena Rieser | Oliver Lemon
Linguist’s Assistant: A Multi-Lingual Natural Language Generator based on Linguistic Universals, Typologies, and Primitives
Tod Allman | Stephen Beale | Richard Denton
Tod Allman | Stephen Beale | Richard Denton
On generating coherent multilingual descriptions of museum objects from Semantic Web ontologies
Dana Dannélls
Dana Dannélls
Reformulating student contributions in tutorial dialogue
Pamela Jordan | Sandra Katz | Patricia Albacete | Michael Ford | Christine Wilson
Pamela Jordan | Sandra Katz | Patricia Albacete | Michael Ford | Christine Wilson
Planning Accessible Explanations for Entailments in OWL Ontologies
Tu Anh T. Nguyen | Richard Power | Paul Piwek | Sandra Williams
Tu Anh T. Nguyen | Richard Power | Paul Piwek | Sandra Williams
Interactive Natural Language Query Construction for Report Generation
Fred Popowich | Milan Mosny | David Lindberg
Fred Popowich | Milan Mosny | David Lindberg
Blogging birds: Generating narratives about reintroduced species to promote public engagement
Advaith Siddharthan | Matthew Green | Kees van Deemter | Chris Mellish | René van der Wal
Advaith Siddharthan | Matthew Green | Kees van Deemter | Chris Mellish | René van der Wal
Natural Language Generation for a Smart Biology Textbook
Eva Banik | Eric Kow | Nikhil Dinesh | Vinay Chaudhri | Umangi Oza
Eva Banik | Eric Kow | Nikhil Dinesh | Vinay Chaudhri | Umangi Oza
Generating Natural Language Summaries for Multimedia
Duo Ding | Florian Metze | Shourabh Rawat | Peter Schulam | Susanne Burger
Duo Ding | Florian Metze | Shourabh Rawat | Peter Schulam | Susanne Burger
The Surface Realisation Task: Recent Developments and Future Plans
Anja Belz | Bernd Bohnet | Simon Mille | Leo Wanner | Michael White
Anja Belz | Bernd Bohnet | Simon Mille | Leo Wanner | Michael White
KBGen – Text Generation from Knowledge Bases as a New Shared Task
Eva Banik | Claire Gardent | Donia Scott | Nikhil Dinesh | Fennie Liang
Eva Banik | Claire Gardent | Donia Scott | Nikhil Dinesh | Fennie Liang
up
Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Gary Geunbae Lee | Jonathan Ginzburg | Claire Gardent | Amanda Stent
Gary Geunbae Lee | Jonathan Ginzburg | Claire Gardent | Amanda Stent
An End-to-End Evaluation of Two Situated Dialog Systems
Lina M. Rojas-Barahona | Alejandra Lorenzo | Claire Gardent
Lina M. Rojas-Barahona | Alejandra Lorenzo | Claire Gardent
“Love ya, jerkface”: Using Sparse Log-Linear Models to Build Positive and Impolite Relationships with Teens
William Yang Wang | Samantha Finkelstein | Amy Ogan | Alan W Black | Justine Cassell
William Yang Wang | Samantha Finkelstein | Amy Ogan | Alan W Black | Justine Cassell
Enhancing Referential Success by Tracking Hearer Gaze
Alexander Koller | Konstantina Garoufi | Maria Staudte | Matthew Crocker
Alexander Koller | Konstantina Garoufi | Maria Staudte | Matthew Crocker
Unsupervised Topic Modeling Approaches to Decision Summarization in Spoken Meetings
Lu Wang | Claire Cardie
Lu Wang | Claire Cardie
An Unsupervised Approach to User Simulation: Toward Self-Improving Dialog Systems
Sungjin Lee | Maxine Eskenazi
Sungjin Lee | Maxine Eskenazi
Hierarchical Conversation Structure Prediction in Multi-Party Chat
Elijah Mayfield | David Adamson | Carolyn Penstein Rosé
Elijah Mayfield | David Adamson | Carolyn Penstein Rosé
Rapid Development Process of Spoken Dialogue Systems using Collaboratively Constructed Semantic Resources
Masahiro Araki
Masahiro Araki
The Effect of Cognitive Load on a Statistical Dialogue System
Milica Gašić | Pirros Tsiakoulis | Matthew Henderson | Blaise Thomson | Kai Yu | Eli Tzirkel | Steve Young
Milica Gašić | Pirros Tsiakoulis | Matthew Henderson | Blaise Thomson | Kai Yu | Eli Tzirkel | Steve Young
Predicting Adherence to Treatment for Schizophrenia from Dialogue Transcripts
Christine Howes | Matthew Purver | Rose McCabe | Patrick G. T. Healey | Mary Lavelle
Christine Howes | Matthew Purver | Rose McCabe | Patrick G. T. Healey | Mary Lavelle
Reinforcement Learning of Question-Answering Dialogue Policies for Virtual Museum Guides
Teruhisa Misu | Kallirroi Georgila | Anton Leuski | David Traum
Teruhisa Misu | Kallirroi Georgila | Anton Leuski | David Traum
From Strangers to Partners: Examining Convergence within a Longitudinal Study of Task-Oriented Dialogue
Christopher M. Mitchell | Kristy Elizabeth Boyer | James C. Lester
Christopher M. Mitchell | Kristy Elizabeth Boyer | James C. Lester
Improving Implicit Discourse Relation Recognition Through Feature Set Optimization
Joonsuk Park | Claire Cardie
Joonsuk Park | Claire Cardie
A Temporal Simulator for Developing Turn-Taking Methods for Spoken Dialogue Systems
Ethan O. Selfridge | Peter A. Heeman
Ethan O. Selfridge | Peter A. Heeman
Estimating Adaptation of Dialogue Partners with Different Verbal Intelligence
Kseniya Zablotskaya | Fernando Fernández-Martínez | Wolfgang Minker
Kseniya Zablotskaya | Fernando Fernández-Martínez | Wolfgang Minker
A Demonstration of Incremental Speech Understanding and Confidence Estimation in a Virtual Human Dialogue System
David DeVault | David Traum
David DeVault | David Traum
Integrating Location, Visibility, and Question-Answering in a Spoken Dialogue System for Pedestrian City Exploration
Srinivasan Janarthanam | Oliver Lemon | Xingkun Liu | Phil Bartie | William Mackaness | Tiphaine Dalmas | Jana Goetze
Srinivasan Janarthanam | Oliver Lemon | Xingkun Liu | Phil Bartie | William Mackaness | Tiphaine Dalmas | Jana Goetze
A Mixed-Initiative Conversational Dialogue System for Healthcare
Fabrizio Morbini | Eric Forbell | David DeVault | Kenji Sagae | David Traum | Albert Rizzo
Fabrizio Morbini | Eric Forbell | David DeVault | Kenji Sagae | David Traum | Albert Rizzo
A Reranking Model for Discourse Segmentation using Subtree Features
Ngo Xuan Bach | Nguyen Le Minh | Akira Shimazu
Ngo Xuan Bach | Nguyen Le Minh | Akira Shimazu
Landmark-Based Location Belief Tracking in a Spoken Dialog System
Yi Ma | Antoine Raux | Deepak Ramachandran | Rakesh Gupta
Yi Ma | Antoine Raux | Deepak Ramachandran | Rakesh Gupta
Exploiting Machine-Transcribed Dialog Corpus to Improve Multiple Dialog States Tracking Methods
Sungjin Lee | Maxine Eskenazi
Sungjin Lee | Maxine Eskenazi
A Bottom-Up Exploration of the Dimensions of Dialog State in Spoken Interaction
Nigel G. Ward | Alejandro Vega
Nigel G. Ward | Alejandro Vega
Using Group History to Identify Character-Directed Utterances in Multi-Child Interactions
Hannaneh Hajishirzi | Jill F. Lehman | Jessica K. Hodgins
Hannaneh Hajishirzi | Jill F. Lehman | Jessica K. Hodgins
Dialog System Using Real-Time Crowdsourcing and Twitter Large-Scale Corpus
Fumihiro Bessho | Tatsuya Harada | Yasuo Kuniyoshi
Fumihiro Bessho | Tatsuya Harada | Yasuo Kuniyoshi
Automatically Acquiring Fine-Grained Information Status Distinctions in German
Aoife Cahill | Arndt Riester
Aoife Cahill | Arndt Riester
A Unified Probabilistic Approach to Referring Expressions
Kotaro Funakoshi | Mikio Nakano | Takenobu Tokunaga | Ryu Iida
Kotaro Funakoshi | Mikio Nakano | Takenobu Tokunaga | Ryu Iida
Combining Verbal and Nonverbal Features to Overcome the “Information Gap” in Task-Oriented Dialogue
Eun Young Ha | Joseph F. Grafsgaard | Christopher Mitchell | Kristy Elizabeth Boyer | James C. Lester
Eun Young Ha | Joseph F. Grafsgaard | Christopher Mitchell | Kristy Elizabeth Boyer | James C. Lester
Semantic Specificity in Spoken Dialogue Requests
Ben Hixon | Rebecca J. Passonneau | Susan L. Epstein
Ben Hixon | Rebecca J. Passonneau | Susan L. Epstein
Contingency and Comparison Relation Labeling and Structure Prediction in Chinese Sentences
Hen-Hsen Huang | Hsin-Hsi Chen
Hen-Hsen Huang | Hsin-Hsi Chen
A Study in How NLU Performance Can Affect the Choice of Dialogue System Architecture
Anton Leuski | David DeVault
Anton Leuski | David DeVault
Integrating Incremental Speech Recognition and POMDP-Based Dialogue Systems
Ethan O. Selfridge | Iker Arizmendi | Peter A. Heeman | Jason D. Williams
Ethan O. Selfridge | Iker Arizmendi | Peter A. Heeman | Jason D. Williams
Improving Sentence Completion in Dialogues with Multi-Modal Features
Anruo Wang | Barbara Di Eugenio | Lin Chen
Anruo Wang | Barbara Di Eugenio | Lin Chen