Other Workshops and Events (2005)
Volumes
- Proceedings of the Workshop on Frontiers in Corpus Annotations II: Pie in the Sky 13 papers
- Proceedings of the ACL Workshop on Feature Engineering for Machine Learning in Natural Language Processing 10 papers
- Proceedings of the Workshop on Psychocomputational Models of Human Language Acquisition 12 papers
- Proceedings of the ACL Workshop on Building and Using Parallel Texts 37 papers
- Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization 10 papers
- Proceedings of the ACL-SIGLEX Workshop on Deep Lexical Acquisition 12 papers
- Proceedings of Workshop on Software 10 papers
- Proceedings of the ACL Workshop on Empirical Modeling of Semantic Equivalence and Entailment 11 papers
- Proceedings of the ACL-ISMB Workshop on Linking Biological Literature, Ontologies and Databases: Mining Biological Semantics 9 papers
- Workshop on open-source machine translation 4 papers
- Workshop on example-based machine translation 17 papers
- Workshop on patent translation 9 papers
- Workshop on Semantic Web technologies for machine translation 5 papers
- Proceedings of the Fourth SIGHAN Workshop on Chinese Language Processing 37 papers
- Proceedings of the Fifth Workshop on Asian Language Resources (ALR-05) and First Symposium on Asian Language Resources Network (ALRN) 12 papers
- Proceedings of the Third International Workshop on Paraphrasing (IWP2005) 14 papers
- Proceedings of the Sixth International Workshop on Linguistically Interpreted Corpora (LINC-2005) 13 papers
- Proceedings of the Second ACL Workshop on Effective Tools and Methodologies for Teaching NLP and CL TeachingNLP 12 papers
- Proceedings of the Second Workshop on Building Educational Applications Using NLP BEA 14 papers
- Proceedings of the ACL Workshop on Computational Approaches to Semitic Languages SEMITIC 13 papers
- Proceedings of the Ninth International Workshop on Parsing Technology IWPT 30 papers
- Proceedings of the Tenth European Workshop on Natural Language Generation (ENLG-05) ENLG 29 papers
- Proceedings of the Second International Workshop on Spoken Language Translation IWSLT 27 papers
- Proceedings of the 6th SIGdial Workshop on Discourse and Dialogue SIGDIAL 28 papers
- Proceedings of the Australasian Language Technology Workshop 2005 ALTA 34 papers
- Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005) CoNLL 40 papers
- Proceedings of the 15th Nordic Conference of Computational Linguistics (NODALIDA 2005) NoDaLiDa 30 papers
up
Proceedings of the Workshop on Frontiers in Corpus Annotations II: Pie in the Sky
Merging PropBank, NomBank, TimeBank, Penn Discourse Treebank and Coreference
James Pustejovsky | Adam Meyers | Martha Palmer | Massimo Poesio
James Pustejovsky | Adam Meyers | Martha Palmer | Massimo Poesio
A Unified Representation for Morphological, Syntactic, Semantic, and Referential Annotations
Erhard W. Hinrichs | Sandra Kübler | Karin Naumann
Erhard W. Hinrichs | Sandra Kübler | Karin Naumann
Attribution and the (Non-)Alignment of Syntactic and Discourse Arguments of Connectives
Nikhil Dinesh | Alan Lee | Eleni Miltsakaki | Rashmi Prasad | Aravind Joshi | Bonnie Webber
Nikhil Dinesh | Alan Lee | Eleni Miltsakaki | Rashmi Prasad | Aravind Joshi | Bonnie Webber
A Framework for Annotating Information Structure in Discourse
Sasha Calhoun | Malvina Nissim | Mark Steedman | Jason Brenier
Sasha Calhoun | Malvina Nissim | Mark Steedman | Jason Brenier
A Parallel Proposition Bank II for Chinese and English
Martha Palmer | Nianwen Xue | Olga Babko-Malaya | Jinying Chen | Benjamin Snyder
Martha Palmer | Nianwen Xue | Olga Babko-Malaya | Jinying Chen | Benjamin Snyder
Semantically Rich Human-Aided Machine Annotation
Marjorie McShane | Sergei Nirenburg | Stephen Beale | Thomas O’Hara
Marjorie McShane | Sergei Nirenburg | Stephen Beale | Thomas O’Hara
up
Proceedings of the ACL Workshop on Feature Engineering for Machine Learning in Natural Language Processing
Proceedings of the ACL Workshop on Feature Engineering for Machine Learning in Natural Language Processing
Eric Ringger
Eric Ringger
A Novel Machine Learning Approach for the Identification of Named Entity Relations
Tianfang Yao | Hans Uszkoreit
Tianfang Yao | Hans Uszkoreit
Feature Engineering and Post-Processing for Temporal Expression Recognition Using Conditional Random Fields
Sisay Fissaha Adafre | Maarten de Rijke
Sisay Fissaha Adafre | Maarten de Rijke
Using Semantic and Syntactic Graphs for Call Classification
Dilek Hakkani-Tür | Gokhan Tur | Ananlada Chotimongkol
Dilek Hakkani-Tür | Gokhan Tur | Ananlada Chotimongkol
Identifying Non-Referential it: A Machine Learning Approach Incorporating Linguistically Motivated Patterns
Adriane Boyd | Whitney Gegg-Harrison | Donna Byron
Adriane Boyd | Whitney Gegg-Harrison | Donna Byron
Engineering of Syntactic Features for Shallow Semantic Parsing
Alessandro Moschitti | Bonaventura Coppola | Daniele Pighin | Roberto Basili
Alessandro Moschitti | Bonaventura Coppola | Daniele Pighin | Roberto Basili
up
Proceedings of the Workshop on Psychocomputational Models of Human Language Acquisition
Proceedings of the Workshop on Psychocomputational Models of Human Language Acquisition
William Gregory Sakas | Alexander Clark | Royal Holloway | James Cussens | Aris Xanthos
William Gregory Sakas | Alexander Clark | Royal Holloway | James Cussens | Aris Xanthos
Using Morphology and Syntax Together in Unsupervised Learning
Yu Hu | Irina Matveeva | John Goldsmith | Colin Sprague
Yu Hu | Irina Matveeva | John Goldsmith | Colin Sprague
Refining the SED Heuristic for Morpheme Discovery: Another Look at Swahili
Yu Hu | Irina Matveeva | John Goldsmith | Colin Sprague
Yu Hu | Irina Matveeva | John Goldsmith | Colin Sprague
A Connectionist Model of Language-Scene Interaction
Marshall R. Mayberry | Matthew W. Crocker | Pia Knoeferle
Marshall R. Mayberry | Matthew W. Crocker | Pia Knoeferle
A Second Language Acquisition Model Using Example Generalization and Concept Categories
Ari Rappoport | Vera Sheinman
Ari Rappoport | Vera Sheinman
Statistics vs. UG in Language Acquisition: Does a Bigram Analysis Predict Auxiliary Inversion?
Xuân-Nga Cao Kam | Iglika Stoyneshka | Lidiya Tornyova | William Gregory Sakas | Janet Dean Fodor
Xuân-Nga Cao Kam | Iglika Stoyneshka | Lidiya Tornyova | William Gregory Sakas | Janet Dean Fodor
Climbing the Path to Grammar: A Maximum Entropy Model of Subject/Object Learning
Felice Dell’Orletta | Alessandro Lenci | Simonetta Montemagni | Vito Pirrelli
Felice Dell’Orletta | Alessandro Lenci | Simonetta Montemagni | Vito Pirrelli
up
Proceedings of the ACL Workshop on Building and Using Parallel Texts
Proceedings of the ACL Workshop on Building and Using Parallel Texts
Philipp Koehn | Joel Martin | Rada Mihalcea | Christof Monz | Ted Pedersen
Philipp Koehn | Joel Martin | Rada Mihalcea | Christof Monz | Ted Pedersen
Cross Language Text Categorization by Acquiring Multilingual Domain Models from Comparable Corpora
Alfio Gliozzo | Carlo Strapparava
Alfio Gliozzo | Carlo Strapparava
Bilingual Word Spectral Clustering for Statistical Machine Translation
Bing Zhao | Eric P. Xing | Alex Waibel
Bing Zhao | Eric P. Xing | Alex Waibel
Revealing Phonological Similarities between Related Languages from Automatically Generated Parallel Corpora
Karin Müller
Karin Müller
Augmenting a Small Parallel Text with Morpho-Syntactic Language
Maja Popović | David Vilar | Hermann Ney | Slobodan Jovičić | Zoran Šarić
Maja Popović | David Vilar | Hermann Ney | Slobodan Jovičić | Zoran Šarić
Induction of Fine-Grained Part-of-Speech Taggers via Classifier Combination and Crosslingual Projection
Elliott Drábek | David Yarowsky
Elliott Drábek | David Yarowsky
A Hybrid Approach to Align Sentences and Words in English-Hindi Parallel Corpora
Niraj Aswani | Robert Gaizauskas
Niraj Aswani | Robert Gaizauskas
NUKTI: English-Inuktitut Word Alignment System Description
Philippe Langlais | Fabrizio Gotti | Guihong Cao
Philippe Langlais | Fabrizio Gotti | Guihong Cao
Symmetric Probabilistic Alignment
Ralf D. Brown | Jae Dong Kim | Peter J. Jansen | Jaime G. Carbonell
Ralf D. Brown | Jae Dong Kim | Peter J. Jansen | Jaime G. Carbonell
Comparison, Selection and Use of Sentence Alignment Algorithms for New Language Pairs
Anil Kumar Singh | Samar Husain
Anil Kumar Singh | Samar Husain
Shared Task: Statistical Machine Translation between European Languages
Philipp Koehn | Christof Monz
Philipp Koehn | Christof Monz
PORTAGE: A Phrase-Based Machine Translation System
Fatiha Sadat | Howard Johnson | Akakpo Agbago | George Foster | Roland Kuhn | Joel Martin | Aaron Tikuisis
Fatiha Sadat | Howard Johnson | Akakpo Agbago | George Foster | Roland Kuhn | Joel Martin | Aaron Tikuisis
Statistical Machine Translation of Euparl Data by using Bilingual N-grams
Rafael E. Banchs | Josep M. Crego | Adrià de Gispert | Patrik Lambert | José B. Mariño
Rafael E. Banchs | Josep M. Crego | Adrià de Gispert | Patrik Lambert | José B. Mariño
Improving Phrase-Based Statistical Translation by Modifying Phrase Extraction and Including Several Features
Marta Ruiz Costa-jussà | José A. R. Fonollosa
Marta Ruiz Costa-jussà | José A. R. Fonollosa
Competitive Grouping in Integrated Phrase Segmentation and Alignment Model
Ying Zhang | Stephan Vogel
Ying Zhang | Stephan Vogel
Deploying Part-of-Speech Patterns to Enhance Statistical Phrase-Based Machine Translation Resources
Christina Lioma | Iadh Ounis
Christina Lioma | Iadh Ounis
Novel Reordering Approaches in Phrase-Based Statistical Machine Translation
Stephan Kanthak | David Vilar | Evgeny Matusov | Richard Zens | Hermann Ney
Stephan Kanthak | David Vilar | Evgeny Matusov | Richard Zens | Hermann Ney
up
Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization
Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization
Jade Goldstein | Alon Lavie | Chin-Yew Lin | Clare Voss
Jade Goldstein | Alon Lavie | Chin-Yew Lin | Clare Voss
A Methodology for Extrinsic Evaluation of Text Summarization: Does ROUGE Correlate?
Bonnie Dorr | Christof Monz | Stacy President | Richard Schwartz | David Zajic
Bonnie Dorr | Christof Monz | Stacy President | Richard Schwartz | David Zajic
Preprocessing and Normalization for Automatic Evaluation of Machine Translation
Gregor Leusch | Nicola Ueffing | David Vilar | Hermann Ney
Gregor Leusch | Nicola Ueffing | David Vilar | Hermann Ney
Evaluating Automatic Summaries of Meeting Recordings
Gabriel Murray | Steve Renals | Jean Carletta | Johanna Moore
Gabriel Murray | Steve Renals | Jean Carletta | Johanna Moore
Evaluating DUC 2004 Tasks with the QARLA Framework
Enrique Amigó | Julio Gonzalo | Anselmo Peñas | Felisa Verdejo
Enrique Amigó | Julio Gonzalo | Anselmo Peñas | Felisa Verdejo
up
Proceedings of the ACL-SIGLEX Workshop on Deep Lexical Acquisition
Proceedings of the ACL-SIGLEX Workshop on Deep Lexical Acquisition
Timothy Baldwin | Anna Korhonen | Aline Villavicencio
Timothy Baldwin | Anna Korhonen | Aline Villavicencio
Automatically Distinguishing Literal and Figurative Usages of Highly Polysemous Verbs
Afsaneh Fazly | Ryan North | Suzanne Stevenson
Afsaneh Fazly | Ryan North | Suzanne Stevenson
Automatic Extraction of Idioms using Graph Analysis and Asymmetric Lexicosyntactic Patterns
Dominic Widdows | Beate Dorow
Dominic Widdows | Beate Dorow
Morphology vs. Syntax in Adjective Class Acquisition
Gemma Boleda | Toni Badia | Sabine Schulte im Walde
Gemma Boleda | Toni Badia | Sabine Schulte im Walde
up
Proceedings of Workshop on Software
Evaluating and Integrating Treebank Parsers on a Biomedical Corpus
Andrew B. Clegg | Adrian J. Shepherd
Andrew B. Clegg | Adrian J. Shepherd
up
Proceedings of the ACL Workshop on Empirical Modeling of Semantic Equivalence and Entailment
Proceedings of the ACL Workshop on Empirical Modeling of Semantic Equivalence and Entailment
Bill Dolan | Ido Dagan
Bill Dolan | Ido Dagan
Local Textual Inference: Can it be Defined or Circumscribed?
Annie Zaenen | Lauri Karttunen | Richard Crouch
Annie Zaenen | Lauri Karttunen | Richard Crouch
Discovering Entailment Relations Using “Textual Entailment Patterns”
Fabio Massimo Zanzotto | Maria Teresa Pazienza | Marco Pennacchiotti
Fabio Massimo Zanzotto | Maria Teresa Pazienza | Marco Pennacchiotti
up
Proceedings of the ACL-ISMB Workshop on Linking Biological Literature, Ontologies and Databases: Mining Biological Semantics
Proceedings of the ACL-ISMB Workshop on Linking Biological Literature, Ontologies and Databases: Mining Biological Semantics
K. Bretonnel Cohen | Lynette Hirschman | Hagit Shatkay | Christian Blaschke
K. Bretonnel Cohen | Lynette Hirschman | Hagit Shatkay | Christian Blaschke
Weakly Supervised Learning Methods for Improving the Quality of Gene Name Normalization Data
Ben Wellner
Ben Wellner
Adaptive String Similarity Metrics for Biomedical Reference Resolution
Ben Wellner | José Castaño | James Pustejovsky
Ben Wellner | José Castaño | James Pustejovsky
Unsupervised Gene/Protein Named Entity Normalization Using Automatically Extracted Dictionaries
Aaron Cohen
Aaron Cohen
A Machine Learning Approach to Acronym Generation
Yoshimasa Tsuruoka | Sophia Ananiadou | Jun’ichi Tsujii
Yoshimasa Tsuruoka | Sophia Ananiadou | Jun’ichi Tsujii
MedTag: A Collection of Biomedical Annotations
Lawrence H. Smith | Lorraine Tanabe | Thomas Rindflesch | W. John Wilbur
Lawrence H. Smith | Lorraine Tanabe | Thomas Rindflesch | W. John Wilbur
Corpus Design for Biomedical Natural Language Processing
K. Bretonnel Cohen | Lynne Fox | Philip V. Ogren | Lawrence Hunter
K. Bretonnel Cohen | Lynne Fox | Philip V. Ogren | Lawrence Hunter
up
Workshop on open-source machine translation
The Open A.I. Kit: General Machine Learning Modules from Statistical Machine Translation
Daniel J. Walker
Daniel J. Walker
The Open A.I. Kit implements the major components of Statistical Machine Translation as an accessible, extendable Software Development Kit with broad applicability beyond the field of Machine Translation. The high-level system design policies of the kit embrace the Open Source development model to provide a modular architecture and interface, which may serve as a basis for collaborative research and development for endeavors in Artificial Intelligence.
An Open Architecture for Transfer-based Machine Translation between Spanish and Basque
Iñaki Alegria | Arantza Diaz de Ilarraza | Gorka Labaka | Mikel Lersundi | Aingeru Mayor | Kepa Sarasola | Mikel L. Forcada | Sergio Ortiz-Rojas | Lluís Padró
Iñaki Alegria | Arantza Diaz de Ilarraza | Gorka Labaka | Mikel Lersundi | Aingeru Mayor | Kepa Sarasola | Mikel L. Forcada | Sergio Ortiz-Rojas | Lluís Padró
We present the current status of development of an open architecture for the translation from Spanish into Basque. The machine translation architecture uses an open source analyser for Spanish and new modules mainly based on finite-state transducers. The project is integrated in the OpenTrad initiative, a larger government funded project shared among different universities and small companies, which will also include MT engines for translation among the main languages in Spain. The main objective is the construction of an open, reusable and interoperable framework. This paper describes the design of the engine, the formats it uses for the communication among the modules, the modules reused from other project named Matxin and the new modules we are building.
Open Source Machine Translation with DELPH-IN
Francis Bond | Stephan Oepen | Melanie Siegel | Ann Copestake | Dan Flickinger
Francis Bond | Stephan Oepen | Melanie Siegel | Ann Copestake | Dan Flickinger
An Open-Source Shallow-Transfer Machine Translation Toolbox: Consequences of Its Release and Availability
Carme Armentano-Oller | Antonio M. Corbí-Bellot | Mikel L. Forcada | Mireia Ginestí-Rosell | Boyan Bonev | Sergio Ortiz-Rojas | Juan Antonio Pérez-Ortiz | Gema Ramírez-Sánchez | Felipe Sánchez-Martínez
Carme Armentano-Oller | Antonio M. Corbí-Bellot | Mikel L. Forcada | Mireia Ginestí-Rosell | Boyan Bonev | Sergio Ortiz-Rojas | Juan Antonio Pérez-Ortiz | Gema Ramírez-Sánchez | Felipe Sánchez-Martínez
By the time Machine Translation Summit X is held in September 2005, our group will have released an open-source machine translation toolbox as part of a large government-funded project involving four universities and three linguistic technology companies from Spain. The machine translation toolbox, which will most likely be released under a GPL-like license includes (a) the open-source engine itself, a modular shallow-transfer machine translation engine suitable for related languages and largely based upon that of systems we have already developed, such as interNOSTRUM for Spanish—Catalan and Traductor Universia for Spanish—Portuguese, (b) extensive documentation (including document type declarations) specifying the XML format of all linguistic (dictionaries, rules) and document format management files, (c) compilers converting these data into the high-speed (tens of thousands of words a second) format used by the engine, and (d) pilot linguistic data for Spanish—Catalan and Spanish—Galician and format management specifications for the HTML, RTF and plain text formats. After describing very briefly this toolbox, this paper aims at exploring possible consequences of the availability of this architecture, including the community-driven development of machine translation systems for languages lacking this kind of linguistic technology.
up
Workshop on example-based machine translation
An n-gram Approach to Exploiting a Monolingual Corpus for Machine Translation
Toni Badia | Gemma Boleda | Maite Melero | Antoni Oliver
Toni Badia | Gemma Boleda | Maite Melero | Antoni Oliver
Example-Based Machine Translation (EBMT) systems have typically operated on individual sentences without taking into account prior context. By adding a simple reweighting of retrieved fragments of training examples on the basis of whether the previous translation retrieved any fragments from examples within a small window of the current instance, translation performance is improved. A further improvement is seen by performing a similar reweighting when another fragment of the current input sentence was retrieved from the same training example. Together, a simple, straightforward implementation of these two factors results in an improvement on the order of 1.0–1.6% in the BLEU metric across multiple data sets in multiple languages.
Corpus-based MT systems that analyse and generalise texts beyond the surface forms of words require generation tools to re-generate the various internal representations into valid target language (TL) sentences. While the generation of word-forms from lemmas is probably the last step in every text generation process at its very bottom end, token-generation cannot be accomplished without structural and morpho-syntactic knowledge of the sentence to be generated. As in many other MT models, this knowledge is composed of a target language model and a bag of information transferred from the source language. In this paper we establish an abstracted, linguistically informed, target language model. We use a tagger, a lemmatiser and a parser to infer a template grammar from the TL corpus. Given a linguistically informed TL model, the aim is to see what need be provided from the transfer module for generation. During computation of the template grammar, we simultaneously build up for each TL sentence the content of the bag such that the sentence can be deterministically reproduced. In this way we control the completeness of the approach and will have an idea of what pieces of information we need to code in the TL bag.
This paper presents a generalization technique that induces translation templates from given translation examples by replacing differing parts in these examples with typed variables. Since the type of each variable is also inferred during the learning process, each induced template is associated with a set of type constraints. The type constraints that are associated with a translation template restrict the usage of that translation template in certain contexts in order to avoid some of wrong translations. The types of variables are induced using the type lattices designed for both source language and target language. The proposed generalization technique has been implemented as a part of an EBMT system.
METISII: Example-based Machine Translation Using Monolingual CorporaSystem Description
Peter Dirix | Ineke Schuurman | Vincent Vandeghinste
Peter Dirix | Ineke Schuurman | Vincent Vandeghinste
The METIS-II project is an example-based machine translation system, making use of minimal resources and tools for both source and target language, making use of a target-language (TL) corpus, but not of any parallel corpora. In the current paper, we discuss the view of our team on the general philosophy and outline of the METIS-II system.
Graph-based Retrieval for Example-based Machine Translation Using Edit-distance
Takao Doi | Hirofumi Yamamoto | Eiichiro Sumita
Takao Doi | Hirofumi Yamamoto | Eiichiro Sumita
We describe our use of RSS news feeds to quickly assemble a parallel English-Japanese corpus. Our method is simpler than other web mining approaches, and it produces a parallel corpus whose quality, quantity, and rate of growth are stable and predictable.
The example-based approach to MT is becoming increasingly popular. However, such is the variety of techniques and methods used that it is difficult to discern the overall conception of what example-based machine translation (EBMT) is and/or what its practitioners conceive it to be. Although definitions of MT systems are notoriously complex, an attempt is made to define EBMT in contrast to other MT architectures (RBMT and SMT).
EBMT by Tree-Phrasing: a Pilot Study
Philippe Langlais | Fabrizio Gotti | Didier Bourigault | Claude Coulombe
Philippe Langlais | Fabrizio Gotti | Didier Bourigault | Claude Coulombe
We present a study we conducted to build a repository storing associations between simple dependency treelets in a source language and their corresponding phrases in a target language. To assess the impact of this resource in EBMT, we used the repository to compute coverage statistics on a test bitext and on a n-best list of translation candidates produced by a standard phrase-based decoder.
The ‘purest’ EBMT System Ever Built: No Variables, No Templates, No Training, Examples, Just Examples, Only Examples
Yves Lepage | Etienne Denoual
Yves Lepage | Etienne Denoual
We designed, implemented and assessed an EBMT system that can be dubbed the “purest ever built”: it strictly does not make any use of variables, templates or training, does not have any explicit transfer component, and does not require any preprocessing of the aligned examples. It uses a specific operation, namely proportional analogy, that implicitly neutralises divergences between languages and captures lexical and syntactical variations along the paradigmatic and syntagmatic axes without explicitly decomposing sentences into fragments. In an experiment with a test set of 510 input sentences and an unprocessed corpus of almost 160,000 aligned sentences in Japanese and English, we obtained BLEU, NIST and mWER scores of 0.53, 8.53 and 0.39 respectively, well above a baseline simulating a translation memory.
Monolingual Corpus-based MT Using Chunks
Stella Markantonatou | Sokratis Sofianopoulos | Vassiliki Spilioti | Yiorgos Tambouratzis | Marina Vassiliou | Olga Yannoutsou | Nikos Ioannou
Stella Markantonatou | Sokratis Sofianopoulos | Vassiliki Spilioti | Yiorgos Tambouratzis | Marina Vassiliou | Olga Yannoutsou | Nikos Ioannou
In the present article, a hybrid approach is proposed for implementing a machine translation system using a large monolingual corpus coupled with a bilingual lexicon and basic NLP tools. In the first phase of the METIS system, a source language (SL) sentence, after being tagged, lemmatised and translated by a flat lemma-to-lemma lexicon, was matched against a tagged and lemmatised target language (TL) corpus using a pattern matching algorithm. In the second phase, translations are generated by combining sub-sentential structures. In this paper, the main features of the second phase are discussed while the system architecture and the corresponding translation approach are presented. The proposed methodology is illustrated with examples of the translation process.
Dependency Treelet Translation: The Convergence of Statistical and Example-based Machine-translation?
Arul Menezes | Chris Quirk
Arul Menezes | Chris Quirk
We describe a novel approach to machine translation that combines the strengths of the two leading corpus-based approaches: Phrasal SMT and EBMT. We use a syntactically informed decoder and reordering model based on the source dependency tree, in combination with conventional SMT models to incorporate the power of phrasal SMT with the linguistic generality available in a parser. We show that this approach significantly outperforms a leading string-based Phrasal SMT decoder and an EBMT system. We present results from two radically different language pairs, and investigate the sensitivity of this approach to parse quality by using two distinct parsers and oracle experiments. We also validate our automated BLEU scores with a small human evaluation.
Users of sign languages are often forced to use a language in which they have reduced competence simply because documentation in their preferred format is not available. While some research exists on translating between natural and sign languages, we present here what we believe to be the first attempt to tackle this problem using an example-based (EBMT) approach. Having obtained a set of English–Dutch Sign Language examples, we employ an approach to EBMT using the ‘Marker Hypothesis’ (Green, 1979), analogous to the successful system of (Way & Gough, 2003), (Gough & Way, 2004a) and (Gough & Way, 2004b). In a set of experiments, we show that encouragingly good translation quality may be obtained using such an approach.
A Machine Learning Approach to Hypotheses Selection of Greedy Decoding for SMT
Michael Paul | Eiichiro Sumita | Seiichi Yamamoto
Michael Paul | Eiichiro Sumita | Seiichi Yamamoto
This paper proposes a method for integrating example-based and rule-based machine translation systems with statistical methods. It extends a greedy decoder for statistical machine translation (SMT), which searches for an optimal translation by using SMT models starting from a decoder seed, i.e., the source language input paired with an initial translation hypothesis. In order to reduce local optima problems inherent in the search, the outputs generated by multiple translation engines, such as rule-based (RBMT) and example-based (EBMT) systems, are utilized as the initial translation hypotheses. This method outperforms conventional greedy decoding approaches using initial translation hypotheses based on translation examples retrieved from a parallel text corpus. However, the decoding of multiple initial translation hypotheses is computationally expensive. This paper proposes a method to select a single initial translation hypothesis before decoding based on a machine learning approach that judges the appropriateness of multiple initial translation hypotheses and selects the most confident one for decoding. Our approach is evaluated for the translation of dialogues in the travel domain, and the results show that it drastically reduces computational costs without a loss in translation quality.
A Semantics-based English-Bengali EBMT System for Translating News Headlines
Diganta Saha | Sivaji Bandyopadhyay
Diganta Saha | Sivaji Bandyopadhyay
The paper reports an Example based Machine Translation System for translating News Headlines from English to Bengali. The input headline is initially searched in the Direct Example Base. If it cannot be found, the input headline is tagged and the tagged headline is searched in the Generalized Tagged Example Base. If a match is obtained, the tagged headline in Bengali is retrieved from the example base, the output Bengali headline is generated after retrieving the Bengali equivalents of the English words from appropriate dictionaries and then applying relevant synthesis rules for generating the Bengali surface level words. If some named entities and acronyms are not present in the dictionary, transliteration scheme is applied for obtaining the Bengali equivalent. If a match is not found, the tagged input headline is analysed to identify the constituent phrase(s). The target translation is generated using English-Bengali phrasal example base, appropriate dictionaries and a set of heuristics for Bengali phrase reordering. If the headline still cannot be translated using example base strategy, a heuristic translation strategy will be applied. Any new input tagged headline along with its translation by the user will be inserted in the tagged Example base after generalization.
Example-based Translation Without Parallel Corpora: First Experiments on a Prototype
Vincent Vandeghinste | Peter Dirix | Ineke Schuurman
Vincent Vandeghinste | Peter Dirix | Ineke Schuurman
For the METIS-II project (IST, start: 10-2004 – end: 09-2007) we are working on an example-based machine translation system, making use of minimal resources and tools for both source and target language, i.e. making use of a target language corpus, but not of any parallel corpora. In the current paper, we present the results of the first experiments with our approach (CCL) within the METIS consortium : the translation of noun phrases from Dutch to English, using the British National Corpus as a target language corpus. Future research is planned along similar lines for the sentence as is presented here for the noun phrase.
up
Workshop on patent translation
The approach presented here enables Japanese users with no knowledge of English or legal English to generate patent claims in English from a Japanese-only interface. It exploits the highly determined structure of patent claims and merges Natural Language Generation (NLG) and Machine Translation (MT) techniques and resources as realized in the AutoPat and PC-Transfer applications. Due to its tuned MT engine, the approach can be seen as a human-aided machine translation (HAMT) system circumventing major obstacles in full-scale Japanese-English MT. The approach is fully implemented on a large scale and will be commercially released in autumn 2005.
Embedding MT for Generating Patent Claims in English from a Multilingual Interface
Svetlana Sheremetyeva
Svetlana Sheremetyeva
In this paper, we present a methodology for the development of interactive domain-tuned patent tools for generating patent claims in English from non-English interfaces. The methodology is based on a merger of an interactive English-to-English patent claim generator, AutoPat1 and any external MT engine that might be appropriate for a certain language. The translation procedure is reduced to translation words and phrases rather than a complex claim sentence. The approach has been successfully used in The J-E patent system 2 , a patent claim generator in English from a Japanese-only interface, and in Dan-Pat3, a similar tool for the Danish-English pair of languages. The two systems use different MT engines but feature similar overall architecture. The methodology is portable to other languages and MT engines.
It is well known that sentences in Japanese patents have long and complicated structures, especially necessary conditions and details. Here, patent sentences are analyzed and classified by pattern of modified relationships. Morphemes were first extracted using the famous morpheme analysis tool Chasen, and then the modified relations were extracted using the software Cabocha. Many modification mistakes were caused by long complicated structures, which required correction by humans. In the process of correction, the modification structure patterns were classified using about 200 sentences. This clarified the characteristics of Japanese patent sentences, and it is useful in machine translation of patent sentences.
A multilingual sense code may chart “constant-sense connection paths” across languages. A writer, not versed in any target language, may nonetheless proofread the sense for translation and edit it, to ensure that his meaning is conveyed as he wishes it, to other languages. A translation-ready format may be thus produced, to serve as a printing-press plate, for precise and automatic translation to any language, or to a plurality of languages. The translation-ready format may describe each word and the full document with a comprehensive code, which specifies the multilingual sense code and other relevant information about the word, in a standardized fashion, digitally, forming a unified, language-independent tagging system and a unified, language-independent lexicon.
Quality Analysis of Patent Parallel Corpus by the Scale
Isamu Okada | Shinichiro Miyazawa | Kazunari Ishida | Nobuhiko Shimizu | Toshizumi Ohta
Isamu Okada | Shinichiro Miyazawa | Kazunari Ishida | Nobuhiko Shimizu | Toshizumi Ohta
Large-scale parallel corpus is extremely important for translation memory, example-based machine translation, and the support system to create English sentences. Organized collection or establishment of large-scale corpus is currently ongoing; however it is a difficult project in terms of copyrights as well as economic efficiency. To investigate general tendency of large-scale corpus helps to improve economical efficiency of parallel corpus collection as well as system establishment. In this study, therefore, the relationship between the scale of parallel corpus and the degree of correspondence is clarified, using parallel corpus for patents.
The paper describes some ways to save on knowledge acquisition when developing MT systems for patents by reducing the size of resources to be acquired, and creating intelligent software for knowledge handling and access speed. The approach is illustrated by knowledge acquisition and maintenance in the APTrans system for translating patent claims. Domain tuned resources are based on contrastive studies of multilingual patent documents and are handled by an electronic dictionary with a powerful user-friendly environment for acquisition, editing, browsing, defaulting and coherence proofing.
The domain dependence of translations of nouns in English-to-Japanese patent translation is examined using an automatic method for identifying major translations from a pair of language corpora in the same domain. The method calculates the ratio of the number of associated words of a target word that suggest each translation of the target word to the total number of associated words. This ratio indicates how major a translation is in a domain. Application of the method to a bilingual patent-abstract corpus indicates the necessity and effectiveness of dividing the patent domain into subdomains and adapting a bilingual dictionary to subdomains.
This paper describes a method for retrieving technical terms and finding their translation candidates from patent corpora. The method improves the reliability of bilingual seed words that measure similarity between a target word and its translation candidates. We conducted an experiment with PAJ (Patent Abstracts of Japan), which is a collection of bilingual patent abstracts written in Japanese and English. The experiment result shows that our method achieves a precision of 53.5% and a recall of 75.4%.
Terminology Construction Workflow for Korean-English Patent MT
Young-Gil Kim | Seong-Il Yang | Munpyo Hong | Chang-Hyun Kim | Young-Ae Seo | Cheol Ryu | Sang-Kyu Park | Se-Young Park
Young-Gil Kim | Seong-Il Yang | Munpyo Hong | Chang-Hyun Kim | Young-Ae Seo | Cheol Ryu | Sang-Kyu Park | Se-Young Park
This paper addresses the workflow for terminology construction for Korean-English patent MT system. The workflow consists of the stage for setting lexical goals and the semi- automatic terminology construction stage. As there is no comparable system, it is difficult to determine how many terms are needed. To estimate the number of the needed terms, we analyzed 45,000 patent documents. Given the limited time and budget, we resorted to the semi-automatic methods to create the bilingual term dictionary in electronics domain. We will show that parenthesis information in Korean patent documents and bilingual title corpus can be successfully used to build a bilingual term dictionary.
up
Workshop on Semantic Web technologies for machine translation
Human translation is based on linguistic and extralinguistic knowledge. Despite promising pioneering advances, knowledge-based machine translation has remained a tempting vision. The bottleneck has been the engineering of sufficiently comprehensive bodies of relevant knowledge The Semantic Web offers opportunities for the gradual evolution of a global heterogeneous knowledge base. The immediate target has been the modelling of certain knowledge domains by practical ontologies. In the talk we will demonstrate the utilization of ontological knowledge indifferent crosslingual applications reaching from crosslingual document retrieval via crosslingual question answering to complex information services involving several crosslingual functionalities, including machine translation. We will then discuss the ramifications of this development and of the evolution of the World Wide Web for future directions in both statistical and rule-based machine translation.
Natural Language is considered the friendliest way of man-machine communication. However the implementation of natural language interfaces faces often the problem of lack of linguistic and world-knowledge, especially when the application domain is not very specific. This is exactly the case of Web-based applications, which aim to serve for retrieval of information in every-day areas of work. The recent Semantic Web activities had as consequence the development of large ontologies for a broad spectrum of domains, as well as of mechanisms for annotating the resources with semantic information. In this paper we present a new architecture aiming to bring together the advantages of natural language querying and the power of semantic W eb. W e will show also how described application can be easily adapted for other domains.
In this paper we give an overview of Semantic Web technologies and the impact of these ones for multilingual Web. We present a possible solution for improving the quality of on-line translation systems, using mechanisms and standards from Semantic Web. We focus on Example based machine translation and the automatization of the translation examples extraction by means of RDF-repositories.
The extraction of lexical sets from a corpus in Digital Signal Processing (DSP) has been detailed before on general sets, with direct ELT applications. In this contribution, a more specialized set is investigated to illustrate the possibility of actually using the results in more “intelligent” Text-Processing.
In this paper we present the actions we made to prepare an EBMT system to be integrated into the Semantic Web. We also described briefly the developed EBMT tool for translators.
up
Proceedings of the Fourth SIGHAN Workshop on Chinese Language Processing
Detecting Segmentation Errors in Chinese Annotated Corpus
Chengjie Sun | Chang-Ning Huang | Xiaolong Wang | Mu Li
Chengjie Sun | Chang-Ning Huang | Xiaolong Wang | Mu Li
Chinese Deterministic Dependency Analyzer: Examining Effects of Global Features and Root Node Finder
Yuchang Cheng | Masayuki Asahara | Yuji Matsumoto
Yuchang Cheng | Masayuki Asahara | Yuji Matsumoto
Morphological features help POS tagging of unknown words across language varieties
Huihsin Tseng | Daniel Jurafsky | Christopher Manning
Huihsin Tseng | Daniel Jurafsky | Christopher Manning
Product Named Entity Recognition Based on Hierarchical Hidden Markov Model
Feifan Liu | Jun Zhao | Bibo Lv | Bo Xu | Hao Yu
Feifan Liu | Jun Zhao | Bibo Lv | Bo Xu | Hao Yu
Chinese Sketch Engine and the Extraction of Grammatical Collocations
Chu-Ren Huang | Adam Kilgarriff | Yiching Wu | Chih-Ming Chiu | Simon Smith | Pavel Rychly | Ming-Hong Bai | Keh-Jiann Chen
Chu-Ren Huang | Adam Kilgarriff | Yiching Wu | Chih-Ming Chiu | Simon Smith | Pavel Rychly | Ming-Hong Bai | Keh-Jiann Chen
Word Meaning Inducing via Character Ontology: A Survey on the Semantic Prediction of Chinese Two-Character Words
Shu-Kai Hsieh
Shu-Kai Hsieh
Domain Specific Word Extraction from Hierarchical Web Documents: A First Step Toward Building Lexicon Trees from Web Corpora
Jing-Shin Chang
Jing-Shin Chang
Learning a Log-Linear Model with Bilingual Phrase-Pair Features for Statistical Machine Translation
Bing Zhao | Alex Waibel
Bing Zhao | Alex Waibel
NIL Is Not Nothing: Recognition of Chinese Network Informal Language Expressions
Yunqing Xia | Kam-Fai Wong | Wei Gao
Yunqing Xia | Kam-Fai Wong | Wei Gao
The Robustness of Domain Lexico-Taxonomy: Expanding Domain Lexicon with CiLin
Chu-Ren Huang | Xiang-Bing Li | Jia-Fei Hong
Chu-Ren Huang | Xiang-Bing Li | Jia-Fei Hong
Some Studies on Chinese Domain Knowledge Dictionary and Its Application to Text Classification
Jingbo Zhu | Wenliang Chen
Jingbo Zhu | Wenliang Chen
Combination of Machine Learning Methods for Optimum Chinese Word Segmentation
Masayuki Asahara | Kenta Fukuoka | Ai Azuma | Chooi-Ling Goh | Yotaro Watanabe | Yuji Matsumoto | Takashi Tsuzuki
Masayuki Asahara | Kenta Fukuoka | Ai Azuma | Chooi-Ling Goh | Yotaro Watanabe | Yuji Matsumoto | Takashi Tsuzuki
Unigram Language Model for Chinese Word Segmentation
Aitao Chen | Yiping Zhou | Anne Zhang | Gordon Sun
Aitao Chen | Yiping Zhou | Anne Zhang | Gordon Sun
Report to BMM-based Chinese Word Segmentor with Context-based Unknown Word Identifier for the Second International Chinese Word Segmentation Bakeoff
Jia-Lin Tsai
Jia-Lin Tsai
Perceptron Learning for Chinese Word Segmentation
Yaoyong Li | Chuanjiang Miao | Kalina Bontcheva | Hamish Cunningham
Yaoyong Li | Chuanjiang Miao | Kalina Bontcheva | Hamish Cunningham
Data-driven Language Independent Word Segmentation Using Character-Level Information
Dong-Hee Lim | Seung-Shik Kang
Dong-Hee Lim | Seung-Shik Kang
Description of the HKU Chinese Word Segmentation System for Sighan Bakeoff 2005
Guohong Fu | Kang-Kwong Luke | Percy Ping-Wai Wong
Guohong Fu | Kang-Kwong Luke | Percy Ping-Wai Wong
A Conditional Random Field Word Segmenter for Sighan Bakeoff 2005
Huihsin Tseng | Pichuan Chang | Galen Andrew | Daniel Jurafsky | Christopher Manning
Huihsin Tseng | Pichuan Chang | Galen Andrew | Daniel Jurafsky | Christopher Manning
Chinese Word Segmentation with Multiple Postprocessors in HIT-IRLab
Huipeng Zhang | Ting Liu | Jinshan Ma | Xiantao Liao
Huipeng Zhang | Ting Liu | Jinshan Ma | Xiantao Liao
Maximal Match Chinese Segmentation Augmented by Resources Generated from a Very Large Dictionary for Post-Processing
Ka-Po Chow | Andy C. Chin | Wing Fu Tsoi
Ka-Po Chow | Andy C. Chin | Wing Fu Tsoi
up
Proceedings of the Fifth Workshop on Asian Language Resources (ALR-05) and First Symposium on Asian Language Resources Network (ALRN)
Domain Knowledge Engineering Based on Encyclopedias and the Web Text
Zhifang Sui | Gaoying Cui | Wansong Ding | Qinlong Zhang
Zhifang Sui | Gaoying Cui | Wansong Ding | Qinlong Zhang
Evaluation of a Japanese CFG Derived from a Syntactically Annotated Corpus with Respect to Dependency Measures
Tomoya Noro | Chimato Koike | Taiichi Hashimoto | Takenobu Tokunaga | Hozumi Tanaka
Tomoya Noro | Chimato Koike | Taiichi Hashimoto | Takenobu Tokunaga | Hozumi Tanaka
An Integrated Framework for Archiving, Processing and Developing Learning Materials for an Endangered Aboriginal Language in Taiwan
Meng-Chien Yang | D. Victoria Rau
Meng-Chien Yang | D. Victoria Rau
Construction of Structurally Annotated Spoken Dialogue Corpus
Shingo Kato | Shigeki Matsubara | Yukiko Yamaguchi | Nobuo Kawaguchi
Shingo Kato | Shigeki Matsubara | Yukiko Yamaguchi | Nobuo Kawaguchi
Cross-lingual Conversion of Lexical Semantic Relations: Building Parallel Wordnets
Chu-Ren Huang | I-Li Su | Jia-Fei Hong | Xiang-Bing Li
Chu-Ren Huang | I-Li Su | Jia-Fei Hong | Xiang-Bing Li
up
Proceedings of the Third International Workshop on Paraphrasing (IWP2005)
Support Vector Machines for Paraphrase Identification and Corpus Construction
Chris Brockett | William B. Dolan
Chris Brockett | William B. Dolan
Using Machine Translation Evaluation Techniques to Determine Sentence-level Semantic Equivalence
Andrew Finch | Young-Sook Hwang | Eiichiro Sumita
Andrew Finch | Young-Sook Hwang | Eiichiro Sumita
Automated Generalization of Phrasal Paraphrases from the Web
Weigang Li | Ting Liu | Yu Zhang | Sheng Li | Wei He
Weigang Li | Ting Liu | Yu Zhang | Sheng Li | Wei He
Automatic generation of paraphrases to be used as translation references in objective evaluation measures of machine translation
Yves Lepage | Etienne Denoual
Yves Lepage | Etienne Denoual
up
Proceedings of the Sixth International Workshop on Linguistically Interpreted Corpora (LINC-2005)
Obtaining Japanese Lexical Units for Semantic Frames from Berkeley FrameNet Using a Bilingual Corpus
Toshiyuki Kanamaru | Masaki Murata | Kow Kuroda | Hitoshi Isahara
Toshiyuki Kanamaru | Masaki Murata | Kow Kuroda | Hitoshi Isahara
Integration of a Lexical Type Database with a Linguistically Interpreted Corpus
Chikara Hashimoto | Francis Bond | Takaaki Tanaka | Melanie Siegel
Chikara Hashimoto | Francis Bond | Takaaki Tanaka | Melanie Siegel
Building Dialogue Corpora for Nursing Activity Analysis
Hiromi itoh Ozaku | Akinori Abe | Noriaki Kuwahara | Futoshi Naya | Kiyoshi Kogure | Kaoru Sagara
Hiromi itoh Ozaku | Akinori Abe | Noriaki Kuwahara | Futoshi Naya | Kiyoshi Kogure | Kaoru Sagara
The Syntactically Annotated ICE Corpus and the Automatic Induction of a Formal Grammar
Alex Chengyu Fang
Alex Chengyu Fang
Linguistically enriched corpora for establishing variation in support verb constructions
Begoña Villada Moirón
Begoña Villada Moirón
Error Annotation for Corpus of Japanese Learner English
Emi Izumi | Kiyotaka Uchimoto | Hitoshi Isahara
Emi Izumi | Kiyotaka Uchimoto | Hitoshi Isahara
up
Proceedings of the Second ACL Workshop on Effective Tools and Methodologies for Teaching NLP and CL
Proceedings of the Second ACL Workshop on Effective Tools and Methodologies for Teaching NLP and CL
Chris Brew | Dragomir Radev
Chris Brew | Dragomir Radev
“Language and Computers”: Creating an Introduction for a General Undergraduate Audience
Chris Brew | Markus Dickinson | W. Detmar Meurers
Chris Brew | Markus Dickinson | W. Detmar Meurers
Language Technology from a European Perspective
Hans Uszkoreit | Valia Kordoni | Vladislav Kubon | Michael Rosner | Sabine Kirchmeier-Andersen
Hans Uszkoreit | Valia Kordoni | Vladislav Kubon | Michael Rosner | Sabine Kirchmeier-Andersen
Natural Language Processing at the School of Information Studies for Africa
Björn Gambäck | Gunnar Eriksson | Athanassia Fourla
Björn Gambäck | Gunnar Eriksson | Athanassia Fourla
up
Proceedings of the Second Workshop on Building Educational Applications Using NLP
Proceedings of the Second Workshop on Building Educational Applications Using NLP
Jill Burstein | Claudia Leacock
Jill Burstein | Claudia Leacock
Applications of Lexical Information for Algorithmically Composing Multiple-Choice Cloze Items
Chao-Lin Liu | Chun-Hung Wang | Zhao-Ming Gao | Shang-Ming Huang
Chao-Lin Liu | Chun-Hung Wang | Zhao-Ming Gao | Shang-Ming Huang
A Real-Time Multiple-Choice Question Generation For Language Testing: A Preliminary Study
Ayako Hoshino | Hiroshi Nakagawa
Ayako Hoshino | Hiroshi Nakagawa
Automatic Essay Grading with Probabilistic Latent Semantic Analysis
Tuomo Kakkonen | Niko Myller | Jari Timonen | Erkki Sutinen
Tuomo Kakkonen | Niko Myller | Jari Timonen | Erkki Sutinen
Towards a Prototyping Tool for Behavior Oriented Authoring of Conversational Agents for Educational Applications
Gahgene Gweon | Jaime Arguello | Carol Pai | Regan Carey | Zachary Zaiss | Carolyn Rosé
Gahgene Gweon | Jaime Arguello | Carol Pai | Regan Carey | Zachary Zaiss | Carolyn Rosé
Direkt Profil: A System for Evaluating Texts of Second Language Learners of French Based on Developmental Sequences
Jonas Granfeldt | Pierre Nugues | Emil Persson | Lisa Persson | Fabian Kostadinov | Malin Ågren | Suzanne Schlyter
Jonas Granfeldt | Pierre Nugues | Emil Persson | Lisa Persson | Fabian Kostadinov | Malin Ågren | Suzanne Schlyter
Measuring Non-native Speakers’ Proficiency of English by Using a Test with Automatically-Generated Fill-in-the-Blank Questions
Eiichiro Sumita | Fumiaki Sugaya | Seiichi Yamamoto
Eiichiro Sumita | Fumiaki Sugaya | Seiichi Yamamoto
Evaluating State-of-the-Art Treebank-style Parsers for Coh-Metrix and Other Learning Technology Environments
Christian F. Hempelmann | Vasile Rus | Arthur C. Graesser | Danielle S. McNamara
Christian F. Hempelmann | Vasile Rus | Arthur C. Graesser | Danielle S. McNamara
up
Proceedings of the ACL Workshop on Computational Approaches to Semitic Languages
Proceedings of the ACL Workshop on Computational Approaches to Semitic Languages
Kareem Darwish | Mona Diab | Nizar Habash
Kareem Darwish | Mona Diab | Nizar Habash
Memory-Based Morphological Analysis Generation and Part-of-Speech Tagging of Arabic
Erwin Marsi | Antal van den Bosch | Abdelhadi Soudi
Erwin Marsi | Antal van den Bosch | Abdelhadi Soudi
Examining the Effect of Improved Context Sensitive Morphology on Arabic Information Retrieval
Kareem Darwish | Hany Hassan | Ossama Emam
Kareem Darwish | Hany Hassan | Ossama Emam
Modifying a Natural Language Processing System for European Languages to Treat Arabic in Information Processing and Information Retrieval Applications
Gregory Grefenstette | Nasredine Semmar | Faïza Elkateb-Gara
Gregory Grefenstette | Nasredine Semmar | Faïza Elkateb-Gara
Choosing an Optimal Architecture for Segmentation and POS-Tagging of Modern Hebrew
Roy Bar-Haim | Khalil Sima’an | Yoad Winter
Roy Bar-Haim | Khalil Sima’an | Yoad Winter
up
Proceedings of the Ninth International Workshop on Parsing Technology
Probabilistic Models for Disambiguation of an HPSG-Based Chart Generator
Hiroko Nakanishi | Yusuke Miyao | Jun’ichi Tsujii
Hiroko Nakanishi | Yusuke Miyao | Jun’ichi Tsujii
Efficacy of Beam Thresholding, Unification Filtering and Hybrid Parsing in Probabilistic HPSG Parsing
Takashi Ninomiya | Yoshimasa Tsuruoka | Yusuke Miyao | Jun’ichi Tsujii
Takashi Ninomiya | Yoshimasa Tsuruoka | Yusuke Miyao | Jun’ichi Tsujii
Statistical Shallow Semantic Parsing despite Little Training Data
Rahul Bhagat | Anton Leuski | Eduard Hovy
Rahul Bhagat | Anton Leuski | Eduard Hovy
Generic Parsing for Multi-Domain Semantic Interpretation
Myroslava Dzikovska | Mary Swift | James Allen | William de Beaumont
Myroslava Dzikovska | Mary Swift | James Allen | William de Beaumont
Online Statistics for a Unification-Based Dialogue Parser
Micha Elsner | Mary Swift | James Allen | Daniel Gildea
Micha Elsner | Mary Swift | James Allen | Daniel Gildea
up
Proceedings of the Tenth European Workshop on Natural Language Generation (ENLG-05)
Proceedings of the Tenth European Workshop on Natural Language Generation (ENLG-05)
Graham Wilcock | Kristiina Jokinen | Chris Mellish | Ehud Reiter
Graham Wilcock | Kristiina Jokinen | Chris Mellish | Ehud Reiter
Interactive Authoring of Logical Forms for Multilingual Generation
Ofer Biller | Michael Elhadad | Yael Netzer
Ofer Biller | Michael Elhadad | Yael Netzer
A Context-dependent Algorithm for Generating Locative Expressions in Physically Situated Environments
John Kelleher | Geert-Jan Kruijff
John Kelleher | Geert-Jan Kruijff
Discrete Optimization as an Alternative to Sequential Processing in NLG
Tomasz Marciniak | Michael Strube
Tomasz Marciniak | Michael Strube
Evaluation of an NLG System using Post-Edit Data: Lessons Learnt
Somayajulu Sripada | Ehud Reiter | Lezan Hawizy
Somayajulu Sripada | Ehud Reiter | Lezan Hawizy
Exploiting OWL Ontologies in the Multilingual Generation of Object Descriptions
Ion Androutsopoulos | Spyros Kallonis | Vangelis Karkaletsis
Ion Androutsopoulos | Spyros Kallonis | Vangelis Karkaletsis
Towards Generating Procedural Texts: An Exploration of their Rhetorical and Argumentative Structure
Farida Aouladomar | Patrick Saint-Dizier
Farida Aouladomar | Patrick Saint-Dizier
The Types and Distributions of Errors in a Wide Coverage Surface Realizer Evaluation
Charles Callaway
Charles Callaway
An Evolutionary Approach to Referring Expression Generation and Aggregation
Raquel Hervás | Pablo Gervás
Raquel Hervás | Pablo Gervás
Using a Corpus of Sentence Orderings Defined by Many Experts to Evaluate Metrics of Coherence for Text Structuring
Nikiforos Karamanis | Chris Mellish
Nikiforos Karamanis | Chris Mellish
Reversibility and Re-usability of Resources in NLG and Natural Language Dialog Systems
Martin Klarner
Martin Klarner
up
Proceedings of the Second International Workshop on Spoken Language Translation
A decoding algorithm for word lattice translation in speech translation
Ruiqiang Zhang | Genichiro Kikui | Hirofumi Yamamoto | Wai-Kit Lo
Ruiqiang Zhang | Genichiro Kikui | Hirofumi Yamamoto | Wai-Kit Lo
Using multiple recognition hypotheses to improve speech translation
Ruiqiang Zhang | Genichiro Kikui | Hirofumi Yamamoto
Ruiqiang Zhang | Genichiro Kikui | Hirofumi Yamamoto
ALEPH: an EBMT system based on the preservation of proportional analogies between sentences across languages
Yves Lepage | Etienne Denoual
Yves Lepage | Etienne Denoual
Nobody is perfect: ATR’s hybrid approach to spoken language translation
Michael Paul | Takao Doi | Youngsook Hwang | Kenji Imamura | Hideo Okuma | Eiichiro Sumita
Michael Paul | Takao Doi | Youngsook Hwang | Kenji Imamura | Hideo Okuma | Eiichiro Sumita
The CMU Statistical Machine Translation System for IWSLT2005
Sanjika Hewavitharana | Bing Zhao | Hildebrand | Almut Silja | Matthias Eck | Chiori Hori | Stephan Vogel | Alex Waibel
Sanjika Hewavitharana | Bing Zhao | Hildebrand | Almut Silja | Matthias Eck | Chiori Hori | Stephan Vogel | Alex Waibel
Low Cost Portability for Statistical Machine Translation based on N-gram Frequency and TF-IDF
Matthias Eck | Stephan Vogel | Alex Waibel
Matthias Eck | Stephan Vogel | Alex Waibel
Edinburgh System Description for the 2005 IWSLT Speech Translation Evaluation
Philipp Koehn | Amittai Axelrod | Alexandra Birch Mayne | Chris Callison-Burch | Miles Osborne | David Talbot
Philipp Koehn | Amittai Axelrod | Alexandra Birch Mayne | Chris Callison-Burch | Miles Osborne | David Talbot
The ITC-irst SMT System for IWSLT-2005
Boxing Chen | Roldano Cattoni | Nicola Bertoldi | Mauro Cettolo | Marcello Federico
Boxing Chen | Roldano Cattoni | Nicola Bertoldi | Mauro Cettolo | Marcello Federico
The CASIA Phrase-Based Machine Translation System
Wei Pang | Zhendong Yang | Zhenbiao Chen | Wei Wei | Bo Xu | Chengqing Zong
Wei Pang | Zhendong Yang | Zhenbiao Chen | Wei Wei | Bo Xu | Chengqing Zong
The NTT Statistical Machine Translation System for IWSLT2005
Hajime Tsukada | Taro Watanabe | Jun Suzuki | Hideto Kazawa | Hideki Isozaki
Hajime Tsukada | Taro Watanabe | Jun Suzuki | Hideto Kazawa | Hideki Isozaki
NUT-NTT Statistical Machine Translation System for IWSLT 2005
Kazuteru Ohashi | Kazuhide Yamamoto | Kuniko Saito | Masaaki Nagata
Kazuteru Ohashi | Kazuhide Yamamoto | Kuniko Saito | Masaaki Nagata
Integrated Chinese Word Segmentation in Statistical Machine Translation
Jia Xu | Evgeny Matusov | Richard Zens | Hermann Ney
Jia Xu | Evgeny Matusov | Richard Zens | Hermann Ney
Evaluating Machine Translation Output with Automatic Sentence Segmentation
Evgeny Matusov | Gregor Leusch | Oliver Bender | Hermann Ney
Evgeny Matusov | Gregor Leusch | Oliver Bender | Hermann Ney
The RWTH Phrase-based Statistical Machine Translation System
Richard Zens | Oliver Bender | Sasa Hasan | Shahram Khadivi | Evgeny Matusov | Jia Xu | Yuqi Zhang | Hermann Ney
Richard Zens | Oliver Bender | Sasa Hasan | Shahram Khadivi | Evgeny Matusov | Jia Xu | Yuqi Zhang | Hermann Ney
Sehda S2MT: Incorporation of Syntax into Statistical Translation System
Yookyung Kim | Jun Huang | Youssef Billawala | Demitrios Master | Farzad Ehsani
Yookyung Kim | Jun Huang | Youssef Billawala | Demitrios Master | Farzad Ehsani
Rapid Development of an Afrikaans English Speech-to-Speech Translator
Herman A. Engelbrecht | Tanja Schultz
Herman A. Engelbrecht | Tanja Schultz
Ngram-based versus Phrase-based Statistical Machine Translation
Josep M. Crego | Marta R. Costa-Jussa | Jose B. Marino | Jose A. R. Fonollosa
Josep M. Crego | Marta R. Costa-Jussa | Jose B. Marino | Jose A. R. Fonollosa
up
Proceedings of the 6th SIGdial Workshop on Discourse and Dialogue
Where do we go from here? Research and Commercial Spoken Dialog Systems
Roberto Pieraccini | Juan Huerta
Roberto Pieraccini | Juan Huerta
Early Preparation of Experimentally Elicited Minimal Responses
Wieneke Wesseling | R. J. J. H. van Son
Wieneke Wesseling | R. J. J. H. van Son
Partially Observable Markov Decision Processes with Continuous Observations for Dialogue Management
Jason D. Williams | Pascal Poupart | Steve Young
Jason D. Williams | Pascal Poupart | Steve Young
Quantitative Evaluation of User Simulation Techniques for Spoken Dialogue Systems
Jost Schatzmann | Kallirroi Georgila | Steve Young
Jost Schatzmann | Kallirroi Georgila | Steve Young
Automatic Induction of Language Model Data for A Spoken Dialogue System
Grace Chung | Stephanie Seneff | Chao Wang
Grace Chung | Stephanie Seneff | Chao Wang
Does this Answer your Question? Towards Dialogue Management for Restricted Domain Question Answering Systems
Matthias Denecke | Norihito Yasuda
Matthias Denecke | Norihito Yasuda
Using Machine Learning for Non-Sentential Utterance Classification
Raquel Fernández | Jonathan Ginzburg | Shalom Lappin
Raquel Fernández | Jonathan Ginzburg | Shalom Lappin
Using Bigrams to Identify Relationships Between Student Certainness States and Tutor Responses in a Spoken Dialogue Corpus
Kate Forbes-Riley | Diane J. Litman
Kate Forbes-Riley | Diane J. Litman
A Corpus Collection and Annotation Framework for Learning Multimodal Clarification Strategies
Verena Rieser | Ivana Kruijff-Korbayová | Oliver Lemon
Verena Rieser | Ivana Kruijff-Korbayová | Oliver Lemon
A Corpus for Studying Addressing Behavior in Multi-Party Dialogues
Natasa Jovanovic | Rieks op den Akker | Anton Nijholt
Natasa Jovanovic | Rieks op den Akker | Anton Nijholt
Sorry and I Didn’t Catch That! - An Investigation of Non-understanding Errors and Recovery Strategies
Dan Bohus | Alexander I. Rudnicky
Dan Bohus | Alexander I. Rudnicky
Developing City Name Acquisition Strategies in Spoken Dialogue Systems Via User Simulation
Ed Filisko | Stephanie Seneff
Ed Filisko | Stephanie Seneff
GALATEA: A Discourse Modeller Supporting Concept-Level Error Handling in Spoken Dialogue Systems
Gabriel Skantze
Gabriel Skantze
Using Language Modelling to Integrate Speech Recognition with a Flat Semantic Analysis
Dirk Büler | Wolfgang Minker | Artha Elciyanti
Dirk Büler | Wolfgang Minker | Artha Elciyanti
Speech-Controlled Media File Selection on Embedded Systems
Yu-Fang H. Wang | Stefan W. Hamerich | Marcus E. Hennecke | Volker M. Schubert
Yu-Fang H. Wang | Stefan W. Hamerich | Marcus E. Hennecke | Volker M. Schubert
Dealing with Doctors: A Virtual Human for Non-team Interaction
David Traum | William Swartout | Jonathan Gratch | Stacy Marsella | Patrick Kenny | Eduard Hovy | Shri Narayanan | Ed Fast | Bilyana Martinovski | Rahul Baghat | Susan Robinson | Andrew Marshall | Dagen Wang | Sudeep Gandhe | Anton Leuski
David Traum | William Swartout | Jonathan Gratch | Stacy Marsella | Patrick Kenny | Eduard Hovy | Shri Narayanan | Ed Fast | Bilyana Martinovski | Rahul Baghat | Susan Robinson | Andrew Marshall | Dagen Wang | Sudeep Gandhe | Anton Leuski
up
Proceedings of the Australasian Language Technology Workshop 2005
Proceedings of the Australasian Language Technology Workshop 2005
Timothy Baldwin | James Curran | Menno van Zaanen
Timothy Baldwin | James Curran | Menno van Zaanen
A Statistical Approach towards Unknown Word Type Prediction for Deep Grammars
Yi Zhang | Valia Kordoni
Yi Zhang | Valia Kordoni
Word Prediction in a Running Text: A Statistical Language Modeling for the Persian Language
Masood Ghayoomi | Seyyed Mostafa Assi
Masood Ghayoomi | Seyyed Mostafa Assi
Faking it: Synthetic Text-to-speech Synthesis for Under-resourced Languages – Experimental Design
Harold Somers
Harold Somers
Dual-Type Automatic Speech Recogniser Designs for Spoken Dialogue Systems
Jason Littlefield | Michael Broughton
Jason Littlefield | Michael Broughton
Evaluating the Utility of Appraisal Hierarchies as a Method for Sentiment Classification
Jeremy Fletcher | Jon Patrick
Jeremy Fletcher | Jon Patrick
Automatic Induction of a POS Tagset for Italian
Raffaella Bernardi | Andrea Bolognesi | Corrado Seidenari | Fabio Tamburini
Raffaella Bernardi | Andrea Bolognesi | Corrado Seidenari | Fabio Tamburini
A Dual-Iterative Method for Concept-Word Acquisition from Large-Scale Chinese Corpora
Guogang Tian | Cungen Cao
Guogang Tian | Cungen Cao
A Distributed Architecture for Interactive Parse Annotation
Baden Hughes | James Haggerty | Joel Nothman | Saritha Manickam | James R. Curran
Baden Hughes | James Haggerty | Joel Nothman | Saritha Manickam | James R. Curran
Multi-document Summarisation and the PASCAL Textual Entailment Challenge
Nicola Stokes | Eamonn Newman
Nicola Stokes | Eamonn Newman
Design and Development of a Speech-driven Control for a In-car Personal Navigation System
Ying Su | Tao Bai | Catherine I. Watson
Ying Su | Tao Bai | Catherine I. Watson
up
Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005)
Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005)
Ido Dagan | Daniel Gildea
Ido Dagan | Daniel Gildea
Effective use of WordNet Semantics via Kernel-Based Learning
Roberto Basili | Marco Cammisa | Alessandro Moschitti
Roberto Basili | Marco Cammisa | Alessandro Moschitti
Search Engine Statistics Beyond the n-Gram: Application to Noun Compound Bracketing
Preslav Nakov | Marti Hearst
Preslav Nakov | Marti Hearst
New Experiments in Distributional Representations of Synonymy
Dayne Freitag | Matthias Blume | John Byrnes | Edmond Chow | Sadik Kapadia | Richard Rohwer | Zhiqiang Wang
Dayne Freitag | Matthias Blume | John Byrnes | Edmond Chow | Sadik Kapadia | Richard Rohwer | Zhiqiang Wang
Word Independent Context Pair Classification Model for Word Sense Disambiguation
Cheng Niu | Wei Li | Rohini K. Srihari | Huifeng Li
Cheng Niu | Wei Li | Rohini K. Srihari | Huifeng Li
Computing Word Similarity and Identifying Cognates with Pair Hidden Markov Models
Wesley Mackay | Grzegorz Kondrak
Wesley Mackay | Grzegorz Kondrak
A Bayesian Mixture Model for Term Re-occurrence and Burstiness
Avik Sarkar | Paul H Garthwaite | Anne De Roeck
Avik Sarkar | Paul H Garthwaite | Anne De Roeck
Discriminative Training of Clustering Functions: Theory and Experiments with Entity Identification
Xin Li | Dan Roth
Xin Li | Dan Roth
Using Uneven Margins SVM and Perceptron for Information Extraction
Yaoyong Li | Kalina Bontcheva | Hamish Cunningham
Yaoyong Li | Kalina Bontcheva | Hamish Cunningham
Improving Sequence Segmentation Learning by Predicting Trigrams
Antal van den Bosch | Walter Daelemans
Antal van den Bosch | Walter Daelemans
Investigating the Effects of Selective Sampling on the Annotation Task
Ben Hachey | Beatrice Alex | Markus Becker
Ben Hachey | Beatrice Alex | Markus Becker
Inferring Semantic Roles Using Sub-Categorization Frames and Maximum Entropy Model
Akshar Bharati | Sriram Venkatapathy | Prashanth Reddy
Akshar Bharati | Sriram Venkatapathy | Prashanth Reddy
Generalized Inference with Multiple Semantic Role Labeling Systems
Peter Koomen | Vasin Punyakanok | Dan Roth | Wen-tau Yih
Peter Koomen | Vasin Punyakanok | Dan Roth | Wen-tau Yih
Semantic Role Labeling System Using Maximum Entropy Classifier
Ting Liu | Wanxiang Che | Sheng Li | Yuxuan Hu | Huaijun Liu
Ting Liu | Wanxiang Che | Sheng Li | Yuxuan Hu | Huaijun Liu
Semantic Role Labeling as Sequential Tagging
Lluís Màrquez | Pere Comas | Jesús Giménez | Neus Català
Lluís Màrquez | Pere Comas | Jesús Giménez | Neus Català
Semantic Role Labeling Using Support Vector Machines
Tomohiro Mitsumori | Masaki Murata | Yasushi Fukuda | Kouichi Doi | Hirohumi Doi
Tomohiro Mitsumori | Masaki Murata | Yasushi Fukuda | Kouichi Doi | Hirohumi Doi
Hierarchical Semantic Role Labeling
Alessandro Moschitti | Ana-Maria Giuglea | Bonaventura Coppola | Roberto Basili
Alessandro Moschitti | Ana-Maria Giuglea | Bonaventura Coppola | Roberto Basili
Semantic Role Chunking Combining Complementary Syntactic Views
Sameer Pradhan | Kadri Hacioglu | Wayne Ward | James H. Martin | Daniel Jurafsky
Sameer Pradhan | Kadri Hacioglu | Wayne Ward | James H. Martin | Daniel Jurafsky
Applying Spelling Error Correction Techniques for Improving Semantic Role Labelling
Erik Tjong Kim Sang | Sander Canisius | Antal van den Bosch | Toine Bogers
Erik Tjong Kim Sang | Sander Canisius | Antal van den Bosch | Toine Bogers
up
Proceedings of the 15th Nordic Conference of Computational Linguistics (NODALIDA 2005)
Robust stochastic parsing: Comparing and combining two approaches for processing extra-grammatical sentences
Marita Ailomaa | Vladimír Kadlec | Martin Rajman | Jean-Cédric Chappelier
Marita Ailomaa | Vladimír Kadlec | Martin Rajman | Jean-Cédric Chappelier
A new semantic similarity measure evaluated in word sense disambiguation
Ergin Altintas | Elif Karsligil | Vedat Coskun
Ergin Altintas | Elif Karsligil | Vedat Coskun
Dictionary acquisition using parallel text and co-occurrence statistics
Chris Biemann | Uwe Quasthoff
Chris Biemann | Uwe Quasthoff
Creating bilingual lexica using reference wordlists for alignment of monolingual semantic vector spaces
Jon Holmlund | Magnus Sahlgren | Jussi Karlgren
Jon Holmlund | Magnus Sahlgren | Jussi Karlgren
Naive Bayes spam filtering using word-position-based attributes and length-sensitive classification thresholds
Johan Hovold
Johan Hovold
An analytical relation between analogical modeling and memory based learning
Christer Johansson | Lars G. Johnsen
Christer Johansson | Lars G. Johnsen
Towards modeling the semantics of calendar expressions as extended regular expressions
Jyrki Niemi | Lauri Carlson
Jyrki Niemi | Lauri Carlson
SUiS–cross-language ontology-driven information retrieval in a restricted domain
Kristina Nilsson | Hans Hjelm | Henrik Oxhammar
Kristina Nilsson | Hans Hjelm | Henrik Oxhammar
Towards automatic recognition of product names: an exploratory study of brand names in economic texts
Kristina Nilsson | Aisha Malmgren
Kristina Nilsson | Aisha Malmgren