Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)
Copied to clipboard
| Challenge: | a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization. |
| Approach: | They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities. |
| Outcome: | The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri. |
Similar Papers
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)
Copied to clipboard
| Challenge: | a library for low-level processing of brahmic scripts is available for free. |
| Approach: | They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. |
| Outcome: | The proposed library supports low-level processing of ten major south Asian Brahmic scripts. |
Criteria for Useful Automatic Romanization in South Asian Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | a number of possible criteria for systems that transliterate South Asian languages are considered . romanization is the special case where the target script is the Latin script. |
| Approach: | They propose a set of criteria for systems that transliterate South Asian languages . criteria include fidelity to human linguistic behavior, processing utility for people, invertibility . they then propose several algorithms that address different criteria . |
| Outcome: | The proposed algorithms address linguistic considerations in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam. |
Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities (2023.acl-long)
Copied to clipboard
| Challenge: | linguistically under-represented communities have an extraordinary opportunity to create content in their native languages. |
| Approach: | They propose to solve the problem of script normalization for languages written in a Perso-Arabic script and use a transformer-based model to analyze the noise levels. |
| Outcome: | The proposed model can normalize a language written in a Perso-Arabic script and improve machine translation and language identification tasks. |
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)
Copied to clipboard
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
| Challenge: | a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet . |
| Approach: | They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. |
| Outcome: | The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset . |
Script-Agnosticism and its Impact on Language Identification for Dravidian Languages (2025.naacl-long)
Copied to clipboard
| Challenge: | a recent study shows that modern systems are script-dependent in language identification (langID) many languages are written in multiple writing systems, and script diversity is common in low-resource languages. |
| Approach: | They propose to learn script-agnostic representations using different strategies . they use word-level script randomization and script exposure to a language written in multiple scripts . |
| Outcome: | The proposed methods exploit script randomization and exposure to a language written in multiple scripts to improve language identification while maintaining competitive performance on naturally occurring text. |
ScriptMix: Mixing Scripts for Low-resource Language Parsing (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing work has considered transliteration and vocabulary augmentation, but the consideration of combining the two has been lacking. |
| Approach: | They propose a multilingual pretrained language model that combines two strengths and overcomes the hurdle of combining them. |
| Outcome: | The proposed model improves POS accuracy by 14% and improves DEP LAS score by 5.6%. |
A Lightweight Modeling Middleware for Corpus Processing (L18-1)
Copied to clipboard
| Challenge: | Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources. |
| Approach: | They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools. |
| Outcome: | The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface. |
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)
Copied to clipboard
| Challenge: | Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory. |
| Approach: | They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction. |
| Outcome: | Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions. |
Evaluating the Efficacy of Large Acoustic Model for Documenting Non-Orthographic Tribal Languages in India (2024.lrec-main)
Copied to clipboard
| Challenge: | Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known . |
| Approach: | They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions . |
| Outcome: | The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script . |
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)
Copied to clipboard
| Challenge: | ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones. |
| Approach: | They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions. |
| Outcome: | The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format. |