Challenge: a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization.
Approach: They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities.
Outcome: The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri.

Similar Papers

Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)

Copied to clipboard

Challenge: a library for low-level processing of brahmic scripts is available for free.
Approach: They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts.
Outcome: The proposed library supports low-level processing of ten major south Asian Brahmic scripts.
Criteria for Useful Automatic Romanization in South Asian Languages (2022.lrec-1)

Copied to clipboard

Challenge: a number of possible criteria for systems that transliterate South Asian languages are considered . romanization is the special case where the target script is the Latin script.
Approach: They propose a set of criteria for systems that transliterate South Asian languages . criteria include fidelity to human linguistic behavior, processing utility for people, invertibility . they then propose several algorithms that address different criteria .
Outcome: The proposed algorithms address linguistic considerations in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam.
Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities (2023.acl-long)

Copied to clipboard

Challenge: linguistically under-represented communities have an extraordinary opportunity to create content in their native languages.
Approach: They propose to solve the problem of script normalization for languages written in a Perso-Arabic script and use a transformer-based model to analyze the noise levels.
Outcome: The proposed model can normalize a language written in a Perso-Arabic script and improve machine translation and language identification tasks.
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)

Copied to clipboard

Challenge: a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet .
Approach: They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages.
Outcome: The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset .
Script-Agnosticism and its Impact on Language Identification for Dravidian Languages (2025.naacl-long)

Copied to clipboard

Challenge: a recent study shows that modern systems are script-dependent in language identification (langID) many languages are written in multiple writing systems, and script diversity is common in low-resource languages.
Approach: They propose to learn script-agnostic representations using different strategies . they use word-level script randomization and script exposure to a language written in multiple scripts .
Outcome: The proposed methods exploit script randomization and exposure to a language written in multiple scripts to improve language identification while maintaining competitive performance on naturally occurring text.
ScriptMix: Mixing Scripts for Low-resource Language Parsing (2024.naacl-long)

Copied to clipboard

Challenge: Existing work has considered transliteration and vocabulary augmentation, but the consideration of combining the two has been lacking.
Approach: They propose a multilingual pretrained language model that combines two strengths and overcomes the hurdle of combining them.
Outcome: The proposed model improves POS accuracy by 14% and improves DEP LAS score by 5.6%.
A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)

Copied to clipboard

Challenge: Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory.
Approach: They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction.
Outcome: Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions.
Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in India (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known .
Approach: They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions .
Outcome: The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script .
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)

Copied to clipboard

Challenge: ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones.
Approach: They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions.
Outcome: The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations