Papers by Rajat Bhatnagar

2 papers
CHIA: CHoosing Instances to Annotate for Machine Translation (2022.findings-emnlp)

Copied to clipboard

Challenge: Neural machine translation systems perform poorly on low-resource language pairs, for which large-scale parallel data is unavailable.
Approach: They propose a method for selecting instances to annotate for machine translation using existing multi-way parallel datasets.
Outcome: The proposed method outperforms unsupervised methods on 20 languages and a multi-way parallel dataset on high-resource languages.
Don’t Rule Out Monolingual Speakers: A Method For Crowdsourcing Machine Translation Data (2021.acl-short)

Copied to clipboard

Challenge: High-performing machine translation systems require large amounts of training data in the form of parallel sentences, and translators are difficult to find and expensive.
Approach: They propose a data collection strategy which uses graphics interchange formats (GIFs) as a pivot to collect parallel sentences from monolingual annotators.
Outcome: The proposed method collects parallel sentences from monolingual annotators in Hindi, Tamil and English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations