Challenge: Existing datasets for evaluating MT systems in this domain are limited.
Approach: They propose to use a multi-parallel corpus from the European Central Bank to analyze the impact of domain-specific terminology on multilingual machine translation for finance.
Outcome: The proposed method compares open-source multilingual MT systems with large language models (LLMs) that possess multilingual capabilities.

Similar Papers

Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education (2025.emnlp-main)

Copied to clipboard

Challenge: Existing public terminology datasets for MT research are limited in language coverage or domain specificity, making it difficult to assess or improve MT systems in specialized settings.
Approach: They propose a multilingual terminology resource for tax and financial education covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.
Outcome: The proposed terminology resource covers seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole.
MultiFin: A Dataset for Multilingual Financial NLP (2023.findings-eacl)

Copied to clipboard

Challenge: Multilingual models are needed to process financial text, which is produced across the world and requires a large dataset.
Approach: They propose to annotate a publicly available financial dataset using a hierarchical label structure and an annotation schema based on a real-world application.
Outcome: The proposed model can be used in high-resource languages, but there is room for improvement in low-resourced languages.
SEDAR: a Large Scale French-English Financial Domain Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches for neural machine translation use small amount of data or monolingual data.
Approach: They describe acquisition, preprocessing and characteristics of a large English-French parallel corpus for the financial domain.
Outcome: The proposed corpus contains 8.6 million high quality sentence pairs . the first release of the corpus is available on github.
Machine Translation of Restaurant Reviews: New Corpus for Domain Adaptation and Robustness (D19-56)

Copied to clipboard

Challenge: BLEU: MT is a very robust and efficient way to translate user-generated content.
Approach: They propose a task to encourage research on MT robustness and domain adaptation . they ask professionals to translate 11.5k french 4SQ reviews to English .
Outcome: The proposed task improves on the existing MT systems in a real-world scenario . the proposed methods improve translation accuracy and sentiment analysis .
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks .
Approach: They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain.
Outcome: The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English.
Revisiting Machine Translation for Cross-lingual Classification (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work in cross-lingual learning has pivoted around multilingual models, which are typically pretrained on unlabeled corpora in multiple languages using some form of language modeling objective.
Approach: They propose to use a stronger machine translation system to mitigat mismatch between training on original text and running inference on machine translated text.
Outcome: The proposed approach is highly task dependent and calls into question the dominance of multilingual models for cross-lingual classification.
Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing financial PLMs are not pretrained on sufficiently diverse financial data, leading to subpar generalization performance.
Approach: They propose to pretrain financial PLMs on financial corpus and train financial models on financial data.
Outcome: The proposed financial language models outperform existing financial PLMs on financial tasks even for unseen corpus groups.
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains.
Approach: They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs.
Outcome: The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs.
m^4 Adapter: Multilingual Multi-Domain Adaptation for Machine Translation with a Meta-Adapter (2022.findings-emnlp)

Copied to clipboard

Challenge: Multilingual neural machine translation models (MNMT) are effective on transferring knowledge between high-resource languages to low-resourced languages.
Approach: They propose a multilingual multi-domain adapter which combines domain and language knowledge using meta-learning with adapters.
Outcome: The proposed model outperforms other adapter methods in a domain shift and language pair translation task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations