Papers by Naveen Arivazhagan

7 papers
Simultaneous Translation (2020.emnlp-tutorials)

Copied to clipboard

Challenge: Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation.
Approach: This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation.
Outcome: This tutorial will examine the design and evaluation of policies for simultaneous translation .
Multilingual Document-Level Translation Enables Zero-Shot Transfer From Sentences to Documents (2022.acl-long)

Copied to clipboard

Challenge: Document-level neural machine translation (DocNMT) is a powerful tool for integrating cross-sentence context into translations.
Approach: They explore whether and how contextual modeling in DocNMT is transferable via multilingual modeling.
Outcome: The proposed model can be used to transfer from teacher languages to student languages with no documents but sentence level data.
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Existing multilingual NMT approaches do not utilize the abundance of monolingual data, especially in low-resource languages.
Approach: They propose to combine monolingual data with self-supervision to pre-train translation models and fine-tune on small amounts of supervised data.
Outcome: The proposed approach improves translation quality of low-resource languages and zero-shot translation quality.
Language-agnostic BERT Sentence Embedding (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for learning bilingual sentence embeddings are not well explored.
Approach: They propose to combine best methods for learning multilingual sentence embeddings with pre-trained models to achieve 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba.
Outcome: The proposed model achieves 83.7% bi-text retrieval accuracy over 112 languages on Tatoeba, above the 65.5% achieved by LASER.
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)

Copied to clipboard

Challenge: a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks.
Approach: They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community .
Outcome: The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges.
Monotonic Infinite Lookback Attention for Simultaneous Machine Translation (P19-1)

Copied to clipboard

Challenge: Simultaneous machine translation begins to translate each source sentence before the source speaker has finished speaking, with applications to live and streaming scenarios.
Approach: They propose a simultaneous translation system that learns an adaptive schedule with a neural machine translation model that attends over all source tokens read thus far.
Outcome: The proposed system can achieve latency-quality trade-offs favorable to a proposed wait-k strategy for many latency values.
Small and Practical BERT Models for Sequence Labeling (D19-1)

Copied to clipboard

Challenge: Existing models for morphosyntactic tagging have focused on building separate models for each language or for a small group of related languages.
Approach: They propose a scheme to train a single multilingual sequence labeling model that is small and fast enough to run on a CPU.
Outcome: The proposed model outperforms state-of-the-art models on low-resource languages and low-level models on codemixed inputs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations