A Topic-Aligned Multilingual Corpus of Wikipedia Articles for Studying Information Asymmetry in Low Resource Languages
Dwaipayan Roy | Sumit Bhatia | Prateek Jain
Proceedings of the Twelfth Language Resources and Evaluation Conference
Wikipedia is the largest web-based open encyclopedia covering more than three hundred languages. However, different language editions of Wikipedia differ significantly in terms of their information coverage. We present a systematic comparison of information coverage in English Wikipedia (most exhaustive) and Wikipedias in eight other widely spoken languages (Arabic, German, Hindi, Korean, Portuguese, Russian, Spanish and Turkish). We analyze the content present in the respective Wikipedias in terms of the coverage of topics as well as the depth of coverage of topics included in these Wikipedias. Our analysis quantifies and provides useful insights about the information gap that exists between different language editions of Wikipedia and offers a roadmap for the IR community to bridge this gap.
Anaphora Resolution in Multi-Person Dialogues
Prateek Jain | Manav Ratan Mital | Sumit Kumar | Amitabha Mukerjee | Achla M. Raina
Proceedings of the 5th SIGdial Workshop on Discourse and Dialogue at HLT-NAACL 2004
- Manav Ratan Mital 1
- Sumit Kumar 1
- Amitabha Mukerjee 1
- Achla M. Raina 1
- Dwaipayan Roy 1
- show all...