Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities (2023.acl-long)
Copied to clipboard
| Challenge: | linguistically under-represented communities have an extraordinary opportunity to create content in their native languages. |
| Approach: | They propose to solve the problem of script normalization for languages written in a Perso-Arabic script and use a transformer-based model to analyze the noise levels. |
| Outcome: | The proposed model can normalize a language written in a Perso-Arabic script and improve machine translation and language identification tasks. |
Similar Papers
Unknown Script: Impact of Script on Cross-Lingual Transfer (2024.naacl-srw)
Copied to clipboard
| Challenge: | Existing models for high-resource languages are not available for all languages, and the vast majority of the world's languages are excluded from these models. |
| Approach: | They propose to use pre-trained models to analyze the effect of the target language and its script on cross-lingual transfer. |
| Outcome: | The proposed model is based on six models pre-trained on NER and POS tasks in the original script and romanized version. |
Can Large Language Models Translate Unseen Languages in Underrepresented Scripts? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive performance in machine translation, but struggle with unseen low-resource languages. |
| Approach: | They propose a benchmark to evaluate translation for Mongolian and Yi using linguistic resources. |
| Outcome: | The proposed model can translate Mongolian (in traditional script) and Yi with the help of linguistic resources, but is limited in its ability to handle these languages effectively. |
Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-Switching (2024.eacl-long)
Copied to clipboard
| Challenge: | Multilingual models exhibit impressive cross-lingual transfer capabilities on unseen languages, but performance is impacted when there is a script disparity with the languages used in the model’s pre-training data. |
| Approach: | They propose a novel method to align a resource-rich language's script with a target language and train a classifier that can make informed decisions regarding the appropriate processing of each token. |
| Outcome: | The proposed model can be used to transfer a language's scripts across multiple languages, but it is suboptimal for mixed languages, where only a subset benefits while the rest is impeded. |
Normalising Non-standardised Orthography in Algerian Code-switched User-generated Data (D19-55)
Copied to clipboard
| Challenge: | a new corpus of unstructured data from social media is presenting challenges to NLP research . standardisation is neither natural nor universal, it is rather a human invention. |
| Approach: | They compile a parallel corpus of Arabic textual data matched with human annotations . they use a deep neural model designed to deal with context-dependent spelling correction . |
| Outcome: | The proposed model performs best with two CNN sub-network encoders and an LSTM decoder . pre-processing data token-by-token with edit-distance aligner significantly improves performance . |
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)
Copied to clipboard
| Challenge: | a library for low-level processing of brahmic scripts is available for free. |
| Approach: | They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. |
| Outcome: | The proposed library supports low-level processing of ten major south Asian Brahmic scripts. |
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)
Copied to clipboard
| Challenge: | a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin. |
| Approach: | They propose a fully automated, context-aware machine translation approach with fewer stages of processing. |
| Outcome: | The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods . |
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)
Copied to clipboard
| Challenge: | a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization. |
| Approach: | They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities. |
| Outcome: | The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri. |
Text Normalization Infrastructure that Scales to Hundreds of Language Varieties (L18-1)
Copied to clipboard
| Challenge: | a multi-language text normalization infrastructure is used to train language models for keyboards and speech recognition systems. |
| Approach: | They describe a multi-language text normalization infrastructure that prepares textual data to train language models used in Google's keyboards and speech recognition systems. |
| Outcome: | The proposed system can normalize training data across hundreds of languages . it can detect errors in training data and detect corruption issues . |
Low-resource Neural Machine Translation with Cross-modal Alignment (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing neural machine translation techniques rely on large monolingual corpus, which is costly for some low-resource languages. |
| Approach: | They propose a cross-modal contrastive learning method to learn a shared space for all languages by additional visual modality. |
| Outcome: | The proposed method can learn cross-modal and cross-lingual alignment with small amount of image-text pairs and achieves significant improvements over the text-only baseline. |
Local Languages, Third Spaces, and other High-Resource Scenarios (2022.acl-long)
Copied to clipboard
| Challenge: | In one view, languages exist on a resource continuum and the challenge is to scale existing solutions, bringing under-resourced languages into the high-resource world. |
| Approach: | They propose to scale existing solutions to bring under-resourced languages into the high-resource world by bringing standardised languages into high-level global information society. |
| Outcome: | The proposed language technology agendas address the diverse situations of the world's languages. |