Papers by Marco Dinarelli
Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models (2022.acl-long)
Copied to clipboard
| Challenge: | Multi-encoder models aim to improve translation quality by encoding document-level contextual information alongside the current sentence. |
| Approach: | They propose to pre-train contextual parameters over split sentence pairs to improve contextual encoding . they propose four different splitting methods to improve learning of contextual parameters . |
| Outcome: | The proposed model improves learning of contextual parameters, both in low and high resource settings. |
DOLFIN - Document-Level Financial Test-Set for Machine Translation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing document-level machine translation test-sets cover general domain but fall short on specialised domains, such as legal and financial. |
| Approach: | They propose to use a document-level machine translation test-set to replace perfectly aligned sentences by presenting data in units of sections rather than sentences. |
| Outcome: | The proposed dataset is built from specialised financial documents and it shows that it can discriminate between context-sensitive and context-agnostic models and shows the weaknesses when models fail to accurately translate financial texts. |
TArC: Tunisian Arabish Corpus, First complete release (2022.lrec-1)
Copied to clipboard
| Challenge: | a project focused on Tunisian Arabic encoded in Arabizi is a hybrid approach to linguistics and linguistic research . Arabic dialects are notoriously under-resourced linguistic systems . |
| Approach: | They propose to use Arabic script as a linguistic corpus and a neural network architecture to annotate the latter with various levels of linguistic information. |
| Outcome: | The proposed approach is hybrid and combines linguistic and linguistic tools . the proposed approach produces in cascade different levels of annotation . |
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Arabish is a spontaneous coding of Arabic dialects in Latin characters and "arithmographs" this code-system was developed by Arabic-speaking users of social media . little research has been dedicated to Tunisian Arabish (TA) |
| Approach: | They describe the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus . they describe preliminary work on the TArC semi-automatic construction process . |
| Outcome: | The first morpho-syntactically annotated Tunisian Arabish corpus (TArC) was developed by arab-speaking users of social media . the code-system will be a useful support for different types of analyses, computational and linguistic, as well as for NLP tools training. |
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)
Copied to clipboard
| Challenge: | ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones. |
| Approach: | They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions. |
| Outcome: | The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format. |