Papers by Alessandro Raganato
MAMMOTH: Massively Multilingual Modular Open Translation @ Helsinki (2024.eacl-demo)
Copied to clipboard
Timothee Mickus, Stig-Arne Grönroos, Joseph Attieh, Michele Boggia, Ona De Gibert, Shaoxiong Ji, Niki Andreas Loppi, Alessandro Raganato, Raúl Vázquez, Jörg Tiedemann
| Challenge: | a growing trend towards modularization is limiting the size and information that can be handled in large language models. |
| Approach: | They propose a framework for training massively multilingual modular machine translation systems at scale. |
| Outcome: | The proposed framework is adapted to train multilingual models at scale on NVIDIA GPUs. |
An Evaluation Benchmark for Testing the Word Sense Disambiguation Capabilities of Machine Translation Systems (2020.lrec-1)
Copied to clipboard
| Challenge: | Lexical ambiguity is one of the many challenging linguistic phenomena involved in translation, i.e., translating an ambiguous word with its correct sense. |
| Approach: | They propose to use training data to measure the sense distributions of a machine translation system to measure lexical ambiguity. |
| Outcome: | The proposed benchmark builds upon the multilingual sense inventory of BabelNet, the multilinguistic neural parsing pipeline TurkuNLP, and the OPUS collection of translated texts from the web. |
XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for assessing distinct meanings of words are tied to sense inventories, restricting their usage to knowledge-based representation techniques. |
| Approach: | They propose a multilingual benchmark that models distinct meanings of words in English . they use a binary disambiguation task with gold standards in 12 new languages . |
| Outcome: | The proposed model can model distinct meanings of words in English even when no tagged instances are available for a target language. |
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have shown that attention heads learn simple positional patterns . |
| Approach: | They propose to replace all but one attention head of each encoder layer with simple fixed – non-learnable – attentive patterns that are solely based on position and do not require external knowledge. |
| Outcome: | The proposed model improves translation quality and improves BLEU scores by up to 3 points in low-resource scenarios. |
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
An Empirical Investigation of Word Alignment Supervision for Zero-Shot Multilingual Neural Machine Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent work has highlighted several flaws of MNMT models in zero-shot scenarios where language labels are ignored and the wrong language is generated. |
| Approach: | They propose to combine explicit alignment to language labels with word alignment supervision to improve zero-shot translations. |
| Outcome: | The proposed model improves on three multilingual MT benchmarks. |
Wikipedia Entities as Rendezvous across Languages: Grounding Multilingual Language Models by Predicting Wikipedia Hyperlinks (2021.naacl-main)
Copied to clipboard
| Challenge: | Masked language models have become the de facto standard when processing text . however, these models are evaluated in a monolingual setting only . |
| Approach: | They propose a language-independent entity prediction task as an intermediate training procedure to ground word representations on entity semantics and bridge the gap between different languages. |
| Outcome: | The proposed approach bridges the gap between word representations and knowledge graphs by using a shared vocabulary of entities. |
AdaKron: An Adapter-based Parameter Efficient Model Tuning with Kronecker Product (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Pretrained Language Models (PLMs) have billions of parameters, causing computational challenges to fine-tuning models. |
| Approach: | They propose an Adapter-based fine-tuning with the Kronecker product that combine the outputs of two small networks to form a final vector whose dimension is the product of the dimensions of the individual outputs. |
| Outcome: | The proposed method achieves the same performance levels as state-of-the-art methods on the GLUE benchmark . |