Challenge: a library for low-level processing of brahmic scripts is available for free.
Approach: They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts.
Outcome: The proposed library supports low-level processing of ten major south Asian Brahmic scripts.

Similar Papers

Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)

Copied to clipboard

Challenge: a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization.
Approach: They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities.
Outcome: The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri.
Criteria for Useful Automatic Romanization in South Asian Languages (2022.lrec-1)

Copied to clipboard

Challenge: a number of possible criteria for systems that transliterate South Asian languages are considered . romanization is the special case where the target script is the Latin script.
Approach: They propose a set of criteria for systems that transliterate South Asian languages . criteria include fidelity to human linguistic behavior, processing utility for people, invertibility . they then propose several algorithms that address different criteria .
Outcome: The proposed algorithms address linguistic considerations in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam.
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)

Copied to clipboard

Challenge: a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks.
Approach: They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community .
Outcome: The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges.
Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities (2023.acl-long)

Copied to clipboard

Challenge: linguistically under-represented communities have an extraordinary opportunity to create content in their native languages.
Approach: They propose to solve the problem of script normalization for languages written in a Perso-Arabic script and use a transformer-based model to analyze the noise levels.
Outcome: The proposed model can normalize a language written in a Perso-Arabic script and improve machine translation and language identification tasks.
Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Application to Text-to-Speech (2020.lrec-1)

Copied to clipboard

Challenge: Using crowd-sourced speech corpus and finite-state transducer grammars, we build a text-to-speech system for Burmese, a tonal Southeast Asian language from the Sino-Tibetan family.
Approach: They propose an open-source crowd-sourced multi-speaker speech corpus and finite-state grammars for performing grapheme-to-phoneme conversion for Burmese.
Outcome: The proposed system performs well for Burmese in a low-resource setting.
A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.
High Performance Natural Language Processing (2020.emnlp-tutorials)

Copied to clipboard

Challenge: a tutorial on scaling natural language processing will recapitulate the state-of-the-art in the field .
Approach: This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective.
Outcome: This cutting-edge tutorial recapitulates the state-of-the-art in natural language processing with scale in perspective.
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)

Copied to clipboard

Challenge: a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin.
Approach: They propose a fully automated, context-aware machine translation approach with fewer stages of processing.
Outcome: The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods .
Efficient Methods for Natural Language Processing: A Survey (2023.tacl-1)

Copied to clipboard

Challenge: Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data, but using only scale to improve performance means resource consumption also grows.
Approach: They propose to use data, time, storage, or energy to improve model performance.
Outcome: The proposed methods and findings provide guidance for conducting NLP under limited resources and point towards promising research directions for developing more efficient methods.
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) (D19-61)

Copied to clipboard

Challenge: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China .
Approach: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . call for papers for this second workshop met with a strong response .
Outcome: the EMNLP-IJCNLP 2019 workshop on deep learning approaches for low-resource natural language processing takes place in Hong Kong, China.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations