Papers by Arya McCarthy

5 papers
Pre-Trained Multilingual Sequence-to-Sequence Models: A Hope for Low-Resource Language Translation? (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained multilingual sequence-to-sequence models like mBART and mT5 can be used to translate low-resource languages, but their practical application is unclear.
Approach: They conduct an empirical experiment in 10 languages to determine what can pre-trained multilingual sequence-to-sequence models like mBART do to translate low-resource languages?
Outcome: The proposed models are robust to domain differences, but translations for unseen and typologically distant languages remain below 3.0 BLEU.
UniMorph 2.0: Universal Morphology (L18-1)

Copied to clipboard

Challenge: The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages.
Approach: They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features.
Outcome: The project releases annotated morphological data using a universal tagset, the UniMorph schema.
Morphological Processing of Low-Resource Languages: Where We Are and What’s Next (2022.findings-acl)

Copied to clipboard

Challenge: Existing models for morphological processing are not suitable for low-resource languages, but they are still lacking in the field of computational morphology.
Approach: They propose to bridge two unsupervised models to understand a language’s morphology from raw text alone and propose to use them to improve their models.
Outcome: The proposed models perform reasonably, but there is room for improvement.
Unsupervised Morphological Paradigm Completion (2020.acl-main)

Copied to clipboard

Challenge: a task of generating morphological paradigms is a challenging unsupervised task for natural language processing systems . acuidados y acciones del idioma es a problem in linguistic annotators.
Approach: They propose a task of unsupervised morphological paradigm completion using raw text and a lemma list.
Outcome: The proposed system outperforms trivial baselines on 14 typologically diverse languages with ease and higher accuracy than minimally supervised systems.
Long-Form Speech Translation through Segmentation with Finite-State Decoding Constraints on Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: a challenge in speech translation is that plenty of spoken content is long-form, but short units are necessary for obtaining high-quality translations.
Approach: They propose a large language model to split long ASR transcripts into segments that can be independently translated to maximize translation quality.
Outcome: The proposed model improves the average BLEU by 2.9 points for English–German, English–Spanish, and English–Arabic TED talk translation in 9 sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations