Papers by Rajat Bhatnagar
CHIA: CHoosing Instances to Annotate for Machine Translation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Neural machine translation systems perform poorly on low-resource language pairs, for which large-scale parallel data is unavailable. |
| Approach: | They propose a method for selecting instances to annotate for machine translation using existing multi-way parallel datasets. |
| Outcome: | The proposed method outperforms unsupervised methods on 20 languages and a multi-way parallel dataset on high-resource languages. |
Don’t Rule Out Monolingual Speakers: A Method For Crowdsourcing Machine Translation Data (2021.acl-short)
Copied to clipboard
| Challenge: | High-performing machine translation systems require large amounts of training data in the form of parallel sentences, and translators are difficult to find and expensive. |
| Approach: | They propose a data collection strategy which uses graphics interchange formats (GIFs) as a pivot to collect parallel sentences from monolingual annotators. |
| Outcome: | The proposed method collects parallel sentences from monolingual annotators in Hindi, Tamil and English. |