Papers by Harritxu Gete

6 papers
Handle with Care: A Case Study in Comparable Corpora Exploitation for Neural Machine Translation (2020.lrec-1)

Copied to clipboard

Challenge: Comparable corpora are an important source of potential parallel data, suitable for training data-driven machine translation systems.
Approach: They present a case study on the exploitation of comparable corpora for machine translation.
Outcome: The results show that filtering in terms of alignment thresholds and length-difference outliers has a significant impact on translation quality.
TANDO: A Corpus for Document-level Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Document-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information.
Approach: They propose to use document-level corpus for Basque-Spanish language pairs to take into account contextual information and perform fine-grained evaluations of gender and gender.
Outcome: The proposed corpus is suitable for fine-grained evaluation of document-level machine translation systems.
To Case or not to case: Evaluating Casing Methods for Neural Machine Translation (2020.lrec-1)

Copied to clipboard

Challenge: Comparative evaluation of casing methods for Neural Machine Translation . evaluators evaluated methods for tokenisation and word segmentation into subword units .
Approach: They evaluate three main casing methods for Neural Machine Translation to determine optimal handling of capitalisation.
Outcome: The proposed methods are used to handle capitalisation on English-German and English-Turkish datasets.
Does Context Help Mitigate Gender Bias in Neural Machine Translation? (2024.findings-emnlp)

Copied to clipboard

Challenge: Neural machine translation models perpetuate gender bias in their training data distribution.
Approach: They examine the gender bias in Neural Machine Translation by using context-aware models to enhance translation accuracy for feminine terms and translation with non-informative context in Basque to Spanish.
Outcome: The proposed models can maintain or even amplify gender bias in translations of stereotypical professions in English and with non-informative context in Basque to Spanish.
Split and Rephrase with Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Split and Rephrase (SPRP) tasks require modelling complex grammatical aspects to provide optimal splits and appropriate rephrasing.
Approach: They evaluate large language models on the Split and Rephrase task . they show they can provide large improvements over the state of the art on main metrics .
Outcome: The proposed model outperforms the state-of-the-art model on the Split and Rephrase task on the main metric, but still lacks in splitting compliance.
Using Discourse Information for Education with a Spanish-Chinese Parallel Corpus (L18-1)

Copied to clipboard

Challenge: Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages.
Approach: They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education.
Outcome: The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations