Michał Stolarski
2006
UAM Text Tools - a flexible NLP architecture
Tomasz Obrębski
|
Michał Stolarski
Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06)
The paper presents a new language processing toolkit developed at Adam Mickiewicz University. Its functionality includes currently tokenization, sentence splitting, dictionary-based morphological analysis, heuristic morphological analysis of unknown words, spelling correction, pattern search, and generation of concordances. It is organized as a collection of command-line programs, each performing one operation. The components may be connected in various ways to provide various text processing services. Also new user-deoned components may be easily incorporated into the system. The toolkit is destined for processing raw (not annotated) text corpora. The system was originally intended for Polish, but its adaptation to other languages is possible.