Papers by Marco Dinarelli

6 papers
Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models (2022.acl-long)

Copied to clipboard

Challenge: Multi-encoder models aim to improve translation quality by encoding document-level contextual information alongside the current sentence.
Approach: They propose to pre-train contextual parameters over split sentence pairs to improve contextual encoding . they propose four different splitting methods to improve learning of contextual parameters .
Outcome: The proposed model improves learning of contextual parameters, both in low and high resource settings.
DOLFIN - Document-Level Financial Test-Set for Machine Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing document-level machine translation test-sets cover general domain but fall short on specialised domains, such as legal and financial.
Approach: They propose to use a document-level machine translation test-set to replace perfectly aligned sentences by presenting data in units of sections rather than sentences.
Outcome: The proposed dataset is built from specialised financial documents and it shows that it can discriminate between context-sensitive and context-agnostic models and shows the weaknesses when models fail to accurately translate financial texts.
TArC: Tunisian Arabish Corpus, First complete release (2022.lrec-1)

Copied to clipboard

Challenge: a project focused on Tunisian Arabic encoded in Arabizi is a hybrid approach to linguistics and linguistic research . Arabic dialects are notoriously under-resourced linguistic systems .
Approach: They propose to use Arabic script as a linguistic corpus and a neural network architecture to annotate the latter with various levels of linguistic information.
Outcome: The proposed approach is hybrid and combines linguistic and linguistic tools . the proposed approach produces in cascade different levels of annotation .
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models are the de facto backbone of most state-of-the-art NLP systems.
Approach: They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law.
Outcome: The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets.
TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Arabish is a spontaneous coding of Arabic dialects in Latin characters and "arithmographs" this code-system was developed by Arabic-speaking users of social media . little research has been dedicated to Tunisian Arabish (TA)
Approach: They describe the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus . they describe preliminary work on the TArC semi-automatic construction process .
Outcome: The first morpho-syntactically annotated Tunisian Arabish corpus (TArC) was developed by arab-speaking users of social media . the code-system will be a useful support for different types of analyses, computational and linguistic, as well as for NLP tools training.
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)

Copied to clipboard

Challenge: ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones.
Approach: They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions.
Outcome: The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations