Papers by Rachel Wicks
Recovering document annotations for sentence-level bitext (2024.findings-acl)
Copied to clipboard
| Challenge: | In machine translation, historical models were incapable of handling longer contexts, so the lack of document-level datasets was less noticeable. |
| Approach: | They propose a document-level filtering technique that discards document- level metadata. |
| Outcome: | The proposed method improves translation without degradation of sentence-level translation. |
The Effects of Language Token Prefixing for Multilingual Machine Translation (2022.aacl-short)
Copied to clipboard
| Challenge: | In recent years, the field has moved towards large neural models either translating from or into many languages. |
| Approach: | They propose to prefix language tokens onto a source or target sequence to improve translation performance. |
| Outcome: | The proposed methods improve translation performance and source side prefixes improve translation. |
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)
Copied to clipboard
Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
| Challenge: | Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families. |
| Approach: | They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations. |
| Outcome: | The results show that the Bible provides high coverage of core vocabulary. |
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc. |
| Approach: | They propose a new context-based modeling approach that can be trained on noisily-annotated data. |
| Outcome: | The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set. |