Papers by Harritxu Gete
Handle with Care: A Case Study in Comparable Corpora Exploitation for Neural Machine Translation (2020.lrec-1)
Copied to clipboard
| Challenge: | Comparable corpora are an important source of potential parallel data, suitable for training data-driven machine translation systems. |
| Approach: | They present a case study on the exploitation of comparable corpora for machine translation. |
| Outcome: | The results show that filtering in terms of alignment thresholds and length-difference outliers has a significant impact on translation quality. |
TANDO: A Corpus for Document-level Machine Translation (2022.lrec-1)
Copied to clipboard
Harritxu Gete, Thierry Etchegoyhen, David Ponce, Gorka Labaka, Nora Aranberri, Ander Corral, Xabier Saralegi, Igor Ellakuria, Maite Martin
| Challenge: | Document-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information. |
| Approach: | They propose to use document-level corpus for Basque-Spanish language pairs to take into account contextual information and perform fine-grained evaluations of gender and gender. |
| Outcome: | The proposed corpus is suitable for fine-grained evaluation of document-level machine translation systems. |
To Case or not to case: Evaluating Casing Methods for Neural Machine Translation (2020.lrec-1)
Copied to clipboard
| Challenge: | Comparative evaluation of casing methods for Neural Machine Translation . evaluators evaluated methods for tokenisation and word segmentation into subword units . |
| Approach: | They evaluate three main casing methods for Neural Machine Translation to determine optimal handling of capitalisation. |
| Outcome: | The proposed methods are used to handle capitalisation on English-German and English-Turkish datasets. |
Does Context Help Mitigate Gender Bias in Neural Machine Translation? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Neural machine translation models perpetuate gender bias in their training data distribution. |
| Approach: | They examine the gender bias in Neural Machine Translation by using context-aware models to enhance translation accuracy for feminine terms and translation with non-informative context in Basque to Spanish. |
| Outcome: | The proposed models can maintain or even amplify gender bias in translations of stereotypical professions in English and with non-informative context in Basque to Spanish. |
Split and Rephrase with Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Split and Rephrase (SPRP) tasks require modelling complex grammatical aspects to provide optimal splits and appropriate rephrasing. |
| Approach: | They evaluate large language models on the Split and Rephrase task . they show they can provide large improvements over the state of the art on main metrics . |
| Outcome: | The proposed model outperforms the state-of-the-art model on the Split and Rephrase task on the main metric, but still lacks in splitting compliance. |
Using Discourse Information for Education with a Spanish-Chinese Parallel Corpus (L18-1)
Copied to clipboard
| Challenge: | Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages. |
| Approach: | They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education. |
| Outcome: | The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source. |