Girish Nath Jha
Also published as: Girish Jha
2026
POS Tagging in Low-Resource Maithili Language: Specific Challenges and Nuances
Shivani Priya | Shruti Jha | Urmila Jha | Girish Nath Jha | Deepali Tiwari | Jyoti Raj
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Shivani Priya | Shruti Jha | Urmila Jha | Girish Nath Jha | Deepali Tiwari | Jyoti Raj
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Abstract Part-of-Speech (POS) tagging is a key step in Natural Language Processing (NLP), laying the groundwork for more advanced syntactic and semantic tasks. Despite Maithili’s status as an Indo-Aryan language with a rich literary tradition and official recognition in India, computational resources for it are still very limited. In this paper, the creation of an annotated corpus of 25,000 sentences drawn from the fields of health, tourism, and administration is described with the hierarchical tagset currently used for Maithili. This paper also indicates that standard tagsets, typically adapted from English or Hindi, fail to capture the linguistic nuances of Maithili. This underestimates the need for a dedicated tagging framework that considers characteristics like vocative particles, verbal nuances, honorific complexities. Keywords: Parts of Speech, Natural Language Processing, Maithili, annotation
The Shabd Portal – Searchable Lexical Resources for Indian Languages by Government of India
Mercy Lalrohluo Hmar | Girish Nath Jha | Dhananjay Singh
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Mercy Lalrohluo Hmar | Girish Nath Jha | Dhananjay Singh
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
The shabd portal ( https://shabd.education.gov.in ) of Commission for Scientific and Technical Terminology (CSTT), a subordinate office under the Ministry of Education, Department of Higher Education, Government of India (GOI) is a data server designed and developed by Prof. Girish Nath Jha, former Chairman CSTT, featuring all the standardized scientific and technical glossaries of CSTT in digital searchable mode. The aim is to launch a central repository for the terminologies prepared in Indian Languages, thus enriching the language bank of India enabling user friendly and free access to standardized terminology. This website is available in 22 Indian languages. The data covers several domains of science, humanities, engineering, medical science and agriculture subjects. The data is dynamic with regular updates in various domains. Users can search the equivalents of terms in Indian languages and submit their feedback for those equivalents prepared by CSTT. The unique feature of the search platform is that users have various options for search, based on languages, subjects, dictionary type and language pairs. The user can also choose to search in a specific glossary or the entire collection which includes about 471 glossaries having about (29,56,125 headwords).
Development of Speech Corpus for Low-Resource Language- a Case of Sanskrit
Devendr Kumar | Girish Nath Jha | Khalid Choukri
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Devendr Kumar | Girish Nath Jha | Khalid Choukri
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
This paper presents a comprehensive framework for the development of a speech corpus for Sanskrit, designed to facilitate advances in Automatic Speech Recognition (ASR) and AI/ML research. The proposed corpus comprises over 107 hours of transcribed speech data, collected from diverse Sanskrit sources through a systematic and scalable pipeline. We detail the end-to-end methodology adopted for corpus creation, encompassing web crawling, data sanitization, audio downloading, and transcription alignment. Particular emphasis is placed on the methodological rigor applied at each stage, including source selection, preprocessing for quality assurance, transcription protocols, and forced alignment techniques. The paper further addresses the unique complexities inherent to Sanskrit, spanning its phonetic richness, intricate morphological structure, and distinctive syntactic patterns. By systematically addressing these dimensions, the resulting 107-hour corpus aims to serve as a foundational resource for speech technology research in Sanskrit.
Integrating Cultural Wisdom and Digital Technologies for Children’s Moral and Emotional Development
Ms Garima | Girish Nath Jha
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Ms Garima | Girish Nath Jha
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
The influence of technology in children’s education is increasing rapidly in the digital age, but with it the challenge of how to develop the cultural and moral development of children in a balanced manner in the technological environment. In traditional societies, moral and cultural teachings have often been imparted through religious and philosophical texts, memorization, interpretation, and oral tradition. This research presents an AI-based value-based learning framework, which aims to make cultural and ethical teachings more structured, simple, and technologically accessible to children. The study provides a brief analysis of memory-based teaching systems prevalent in various religious traditions and presents a model based on verses from the Bhagavad Gita as an example. The proposed system includes data generation and processing, simplified interpretation, semantic understanding, pronunciation analysis and interactive learning facilities based on selected material from cultural texts. The study indicates that through AI and modern technologies, traditional cultural knowledge can be delivered to children in a more effective and engaging form, developing new possibilities for reinforcing their moral and cultural development.
Preserving Civilisation Memory: A Digital Humanities Approach to the Ramayan
Shashank Tiwari | Girish Nath Jha
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Shashank Tiwari | Girish Nath Jha
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Abstract The Ramayana takes a leading role in the list of the most important texts in the world literature, with a multiplicity of textual traditions of unparalleled numbers of more than three hundred variants throughout South and Southeast. Stone manuscripts, bamboo manuscripts, and palm leaf manuscripts have been passed down in palm-leaf codices, inscriptions on temple walls, in highly illustrated folios and through generations of oral performance. But the corpus now faces serious challenges due to the destruction of the environment, material frailty, fragmentation, and the scope of modern script recognition methods. The current analysis examines how digitization projects are re-defining the preservation of Ramayana in a heritage system that is networked across the world. In this paper, the author critically assesses the work of large-scale projects thru the use of a qualitative research design that has been conducted between the years 2003 and 2026, including the National Mission for Manuscripts (NMM), the digital reunification of the Mewar Ramayana, and efforts by southeast Asian countries to document adaptations, like the Reamker by Cambodia). It predicts imaging standards, metadata formatting policies, integration of optical character recognition (OCR) and digital access structures, and struggles with the problem of multi-script complexity (Grantha, Devanagari, Kawi), partial coverage of variant texts, and infrastructural inequities. The results support that digitization has a significant positive impact on scholarly accessibility and comparative research but the advantages are unexpressed, especially relating to oral traditions. To make the endeavor sustainable preservation, interoperable standards must be adopted, the script recognition with the help of AI should be encouraged, the community should be involved, and cross-border collaboration institutionalized to protect the long-term cultural viability of the Ramayana. Keywords: Ramayana, Manuscript Preservation, Digitization of Cultural Heritage, Digital Humanities, Textual Transmission, Palm-Leaf Manuscripts, Grantha and Kawi Scripts, Metadata Standards (METS/XML), IIIF Interoperability, AI-Assisted Philology, Intangible, Cultural Heritage, Archival Sustainability, Cultural Heritage Informatics, Open Access Repositories, Civilizational Memory.
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Kalika Bali | Sobha L | Devendr Kumar
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Kalika Bali | Sobha L | Devendr Kumar
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
2024
Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Sobha L. | Kalika Bali | Atul Kr. Ojha
Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Sobha L. | Kalika Bali | Atul Kr. Ojha
Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation
2022
Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference
Girish Nath Jha | Sobha L. | Kalika Bali | Atul Kr. Ojha
Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference
Girish Nath Jha | Sobha L. | Kalika Bali | Atul Kr. Ojha
Proceedings of the WILDRE-6 Workshop within the 13th Language Resources and Evaluation Conference
2021
Prosody Labelled Dataset for Hindi
Esha Banerjee | Atul Kr. Ojha | Girish Jha
Proceedings of the Workshop on Speech and Music Processing 2021
Esha Banerjee | Atul Kr. Ojha | Girish Jha
Proceedings of the Workshop on Speech and Music Processing 2021
This study aims to develop an intonation labelled database for Hindi, for enhancing prosody in ASR and TTS systems, which is also helpful for building Speech to Speech Machine Translation systems. Although no single standard for prosody labelling exists in Hindi, researchers in the past have employed perceptual and statistical methods in literature to draw inferences about the behaviour of prosody patterns in Hindi. Based on such existing research and largely agreed upon intonational theories in Hindi, this study attempts to develop a manually annotated prosodic corpus of Hindi speech data, which can be used for training speech models for natural-sounding speech in the future. 500 sentences (2,550 words) for declarative and interrogative types have been labelled using Praat.
2020
Abstractive Text Summarization for Sanskrit Prose: A Study of Methods and Approaches
Shagun Sinha | Girish Jha
Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation
Shagun Sinha | Girish Jha
Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation
The authors present a work-in-progress in the field of Abstractive Text Summarization (ATS) for Sanskrit Prose – a first attempt at ATS for Sanskrit (SATS). We will evaluate recent approaches and methods used for ATS and argue for the ones to be adopted for Sanskrit prose considering the unique properties of the language. There are three goals of SATS - to make manuscript summaries, to enrich the semantic processing of Sanskrit, and to improve the information retrieval systems in the language. While Extractive Text Summarization (ETS) is an important method, the summaries it generates are not always coherent. For qualitative coherent summaries, ATS is considered a better option by scholars. This paper reviews various ATS/ETS approaches for Sanskrit and other Indian Languages done till date. In the preliminary overview, authors conclude that of the two available approaches - structure-based and semantic-based - the latter would be viable owing to the rich morphology of Sanskrit. Moreover, a graph-based method may also be suitable. The second suggested method is the supervised-learning method. The authors also suggest attempting cross-lingual summarization as an extension to this work in future.
Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Kalika Bali | Sobha L. | S. S. Agrawal | Atul Kr. Ojha
Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation
Girish Nath Jha | Kalika Bali | Sobha L. | S. S. Agrawal | Atul Kr. Ojha
Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Evaluation
Formal Sanskrit Syntax: A Specification for Programming Language
K. Kabi Khanganba | Girish Jha
Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop
K. Kabi Khanganba | Girish Jha
Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop
The paper discusses the syntax of the primary statements of the Sanskritam, a programming language specification based on natural Sanskrit under a doctoral thesis. By a statement, we mean a syntactic unit regardless of its computational operations of variable declarations, program executions or evaluations of Boolean expressions etc. We have selected six common primary statements of declaration, assignment, inline initialization, if-then-else, for loop and while loop. The specification partly overlaps the ideas of natural language programming, Controlled Natural Language (Kunh, 2013), and Natural Language subset. The practice and application of structured natural language set in a discourse are deeply rooted in the theoretical text tradition of Sanskrit, like the sūtra-based disciplines and Navya-Nyāya (NN) formal language, etc. The effort is a kind of continuation and application of such traditions and their techniques in the modern field of Sanskrit NLP.
2016
The IMAGACT4ALL Ontology of Animated Images: Implications for Theoretical and Machine Translation of Action Verbs from English-Indian Languages
Pitambar Behera | Sharmin Muzaffar | Atul Ku. Ojha | Girish Jha
Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016)
Pitambar Behera | Sharmin Muzaffar | Atul Ku. Ojha | Girish Jha
Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016)
Action verbs are one of the frequently occurring linguistic elements in any given natural language as the speakers use them during every linguistic intercourse. However, each language expresses action verbs in its own inherently unique manner by categorization. One verb can refer to several interpretations of actions and one action can be expressed by more than one verb. The inter-language and intra-language variations create ambiguity for the translation of languages from the source language to target language with respect to action verbs. IMAGACT is a corpus-based ontological platform of action verbs translated from prototypic animated images explained in English and Italian as meta-languages. In this paper, we are presenting the issues and challenges in translating action verbs of Indian languages as target and English as source language by observing the animated images. Among the ten Indian languages which have been annotated so far on the platform are Sanskrit, Hindi, Urdu, Odia (Oriya), Bengali, Manipuri, Tamil, Assamese, Magahi and Marathi. Out of them, Manipuri belongs to the Sino-Tibetan, Tamil comes off the Dravidian and the rest owe their genesis to the Indo-Aryan language family. One of the issues is that the one-word morphological English verbs are translated into most of the Indian languages as verbs having more than one-word form; for instance as in the case of conjunct, compound, serial verbs and so on. We are further presenting a cross-lingual comparison of action verbs among Indian languages. In addition, we are also dealing with the issues in disambiguating animated images by the L1 native speakers using competence-based judgements and the theoretical and machine translation implications they bear.
Issues and Challenges in Annotating Urdu Action Verbs on the IMAGACT4ALL Platform
Sharmin Muzaffar | Pitambar Behera | Girish Jha
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Sharmin Muzaffar | Pitambar Behera | Girish Jha
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
In South-Asian languages such as Hindi and Urdu, action verbs having compound constructions and serial verbs constructions pose serious problems for natural language processing and other linguistic tasks. Urdu is an Indo-Aryan language spoken by 51, 500, 0001 speakers in India. Action verbs that occur spontaneously in day-to-day communication are highly ambiguous in nature semantically and as a consequence cause disambiguation issues that are relevant and applicable to Language Technologies (LT) like Machine Translation (MT) and Natural Language Processing (NLP). IMAGACT4ALL is an ontology-driven web-based platform developed by the University of Florence for storing action verbs and their inter-relations. This group is currently collaborating with Jawaharlal Nehru University (JNU) in India to connect Indian languages on this platform. Action verbs are frequently used in both written and spoken discourses and refer to various meanings because of their polysemic nature. The IMAGACT4ALL platform stores each 3d animation image, each one of them referring to a variety of possible ontological types, which in turn makes the annotation task for the annotator quite challenging with regard to selecting verb argument structure having a range of probability distribution. The authors, in this paper, discuss the issues and challenges such as complex predicates (compound and conjunct verbs), ambiguously animated video illustrations, semantic discrepancies, and the factors of verb-selection preferences that have produced significant problems in annotating Urdu verbs on the IMAGACT ontology.
2010
The TDIL Program and the Indian Langauge Corpora Intitiative (ILCI)
Girish Nath Jha
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Girish Nath Jha
Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
India is considered a linguistic ocean with 4 language families and 22 scheduled national languages, and 100 un-scheduled languages reported by the 2001 census. This puts tremendous pressures on the Indian government to not only have comprehensive language policies, but also to create resources for their maintenance and development. In the age of information technology, there is a greater need to have a fine balance between allocation of resources to each language keeping in view the political compulsions, electoral potential of a linguistic community and other issues. In this connection, the government of India through various ministries and a think tank consisting of eminent linguistics and policy makers has done a commendable job despite the obvious roadblocks. This paper describes the Indian government’s policies towards language development and maintenance in the age of technology through the Ministry of HRD through its various agencies and the Ministry of Communications & Information Technology (MCIT) through its dedicated program called TDIL (Technology Development for Indian Languages). The paper also describes some of the recent activities of the TDIL in general and in particular, an innovative corpora project called ILCI - Indian Languages Corpora Initiative.
2008
A Common Parts-of-Speech Tagset Framework for Indian Languages
Baskaran Sankaran | Kalika Bali | Monojit Choudhury | Tanmoy Bhattacharya | Pushpak Bhattacharyya | Girish Nath Jha | S. Rajendran | K. Saravanan | L. Sobha | K.V. Subbarao
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
Baskaran Sankaran | Kalika Bali | Monojit Choudhury | Tanmoy Bhattacharya | Pushpak Bhattacharyya | Girish Nath Jha | S. Rajendran | K. Saravanan | L. Sobha | K.V. Subbarao
Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)
We present a universal Parts-of-Speech (POS) tagset framework covering most of the Indian languages (ILs) following the hierarchical and decomposable tagset schema. In spite of significant number of speakers, there is no workable POS tagset and tagger for most ILs, which serve as fundamental building blocks for NLP research. Existing IL POS tagsets are often designed for a specific language; the few that have been designed for multiple languages cover only shallow linguistic features ignoring linguistic richness and the idiosyncrasies. The new framework that is proposed here addresses these deficiencies in an efficient and principled manner. We follow a hierarchical schema similar to that of EAGLES and this enables the framework to be flexible enough to capture rich features of a language/ language family, even while capturing the shared linguistic structures in a methodical way. The proposed common framework further facilitates the sharing and reusability of scarce resources in these languages and ensures cross-linguistic compatibility.
Search
Fix author
Co-authors
- Kalika Bali 6
- Sobha L 6
- Atul Kr. Ojha 5
- Pitambar Behera 2
- Tanmoy Bhattacharya 2
- Pushpak Bhattacharyya 2
- Devendr Kumar 2
- Sharmin Muzaffar 2
- S. Rajendran 2
- Baskaran Sankaran 2
- K Saravanan 2
- Subbarao K. V 2
- S. S. Agrawal 1
- Esha Banerjee 1
- Monojit Choudhury 1
- Khalid Choukri 1
- Ms Garima 1
- Mercy Lalrohluo Hmar 1
- Shruti Jha 1
- Urmila Jha 1
- K. Kabi Khanganba 1
- Shivani Priya 1
- Jyoti Raj 1
- Dhananjay Singh 1
- Shagun Sinha 1
- Deepali Tiwari 1
- Shashank Tiwari 1