Papers with Multilinguality
Copied to clipboard
| Challenge: | Using a transformer architecture, we study coreference phenomena in three neural machine translation systems. |
| Approach: | They analyse coreference phenomena in three neural machine translation systems . they manually annotate (the possibly incorrect) coreference chains in the outputs . |
| Outcome: | The proposed model shows stronger translationese effects in machine translated outputs than in human translations. |
Copied to clipboard
| Challenge: | Existing work on improving cross-lingual transferability of NMT model is under-explored. |
| Approach: | They propose a model that leverages a multilingual pretrained encoder to improve cross-lingual transferability. |
| Outcome: | The proposed model outperforms mBART and m2m-100 on a zero-shot cross-lingual transfer task. |
Copied to clipboard
| Challenge: | Multilingual pretrained language models have shown impressive results for cross-lingual transfer, but due to the constant model capacity, multilingual pre-training usually lags behind the monolingual competitors. |
| Approach: | They propose to transfer the knowledge from monolingual pretrained models to multilingual ones to improve zero-shot cross-lingual classification by using machine translation systems. |
| Outcome: | The proposed methods outperform vanilla multilingual fine-tuning on two cross-lingual classification benchmarks. |
Copied to clipboard
| Challenge: | Mauritian Creole is a French-based creole and a lingua franca of the Republic of Mauritius. |
| Approach: | They describe a dataset for benchmarking machine translation quality of Mauritian Creole. |
| Outcome: | The proposed dataset compares KreolMorisienMT with existing models and human evaluation reveals the systems’ high translation quality. |
Copied to clipboard
| Challenge: | In this tutorial, we will cover the latest advances in NMT to enhance low-resource translation. |
| Approach: | They will cover the latest advances in NMT approaches that leverage multilingualism . they will focus on topics such as language divergence, transfer learning and pivoting . |
| Outcome: | This tutorial will cover the latest advances in NMT to enhance low-resource translation models. |
Copied to clipboard
| Challenge: | a new method for information extraction from Greek corpora is being developed for low-resource languages. |
| Approach: | They propose a methodology that aims at bridging the gap between high and low-resource languages in the context of Open Information Extraction. |
| Outcome: | The proposed method outperforms the current state-of-the-art for the Greek language on benchmark datasets. |
Copied to clipboard
| Challenge: | Quality estimation models are often opaque and computationally expensive, making them impractical to be part of large-scale pipelines. |
| Approach: | They propose an uncertainty-aware quality estimation model that matches previous approaches at a fraction of their costs. |
| Outcome: | The proposed method reduces evaluation costs by 50% and improves reranking performance. |
Copied to clipboard
| Challenge: | Existing approaches to improve model performance are finetuning on all acquired data after each round, which is computationally expensive in multilingual and low-resource settings. |
| Approach: | They evaluate continual finetuning (CF) against full finetuned (FA) across 28 African languages using MasakhaNEWS and SIB-200. |
| Outcome: | The proposed approach outperforms full finetuning (FA) in 28 African languages, achieving up to 35% reductions in GPU memory, FLOPs, and training time. |
Copied to clipboard
| Challenge: | Current MT evaluation measures pay the same attention to each sentence component . in real-world examinations, the questions vary in difficulty and weightings . |
| Approach: | They propose a difficulty-aware MT evaluation metric that takes translation difficulty into account . they propose to use this metric to evaluate machine translation (MT) results . |
| Outcome: | The proposed method outperforms most MT evaluation metrics in terms of human correlation. |
Copied to clipboard
| Challenge: | Neural Chat Translation (NCT) models that use dialogue characteristics of chat are often incoherent and speakerirrelevant. |
| Approach: | They propose to introduce the modeling of dialogue characteristics into the NCT model by capturing the inherent dialogue characteristics. |
| Outcome: | The proposed model can translate conversational text between speakers of different languages. |
Copied to clipboard
| Challenge: | Existing linguistic knowledge bases such as URIEL+ lack a principled method for aggregating these signals into a single, comprehensive score. |
| Approach: | They propose a framework for type-matched language distances that unifies these signals into a robust, task-agnostic composite distance. |
| Outcome: | The proposed representations improve transfer performance when the distance type is relevant to the task, while yielding gains in most tasks. |
Copied to clipboard
| Challenge: | MT-Telescope is an open source, written in Python, and is built around a user friendly and dynamic web interface. |
| Approach: | They propose a platform to facilitate comparative analysis of the output quality of two Machine Translation (MT) systems. |
| Outcome: | The proposed platform supports fine-grained segment-level analysis and interactive visualisations that expose the fundamental differences in the performance of the compared systems. |
Copied to clipboard
| Challenge: | Existing studies suggest that accuracy and fluency should trade off against each other, and that capturing every detail of the source is difficult for human raters to distinguish. |
| Approach: | They propose to evaluate the relationship between accuracy and fluency at the segment level and to use probabilities to estimate probabilities. |
| Outcome: | The proposed model relies on human judgments of accuracy and fluency collected in prior work on translation quality estimation. |
Copied to clipboard
| Challenge: | Existing studies use pretrained motion detection models as verb sense ambiguity representations to solve the verb sense problem. |
| Approach: | They propose to use video contents as auxiliary information to address the word sense ambiguity problem in machine translation. |
| Outcome: | Experiments on the VATEX dataset show that the proposed system achieves 35.86 BLEU-4 score, which is 0.51 score higher than the single model of the SOTA method. |
Copied to clipboard
| Challenge: | SR is one of the main tasks involved in Natural Language Generation. |
| Approach: | They propose a system which divides the SR task into two independent subtasks, namely word order prediction and morphology inflection prediction. |
| Outcome: | The proposed system is a direct successor to the architecture presented at SR'19. |
Copied to clipboard
| Challenge: | We submitted two systems for scientific paper subtask and timely disclosure subtask . we evaluated the usefulness of incorporating external data from a wide variety of web pages to improve the translation quality. |
| Approach: | They describe two different translation tasks submitted to WAT 2019 . they submitted scientific paper subtasks and timely disclosure subtask . |
| Outcome: | The proposed system performed better on scientific paper and timely disclosure subtasks. |
Copied to clipboard
| Challenge: | a new method of "humanizing" automatic translations has been developed for the translation industry . a demonstration of an online learning system for machine translation in a production environment . |
| Approach: | They present a system which implements online learning for neural machine translation in a production environment. |
| Outcome: | The proposed system saves post-editing effort and adapts to a specific domain or user style. |
Copied to clipboard
| Challenge: | a growing trend towards modularization is limiting the size and information that can be handled in large language models. |
| Approach: | They propose a framework for training massively multilingual modular machine translation systems at scale. |
| Outcome: | The proposed framework is adapted to train multilingual models at scale on NVIDIA GPUs. |
Copied to clipboard
| Challenge: | Existing code-switching-based cross-lingual spoken language understanding frameworks are limited to low-resource languages. |
| Approach: | They propose a cross-lingual spoken language understanding framework that leverages both code-switched and original sentences to achieve multi-level alignment. |
| Outcome: | The proposed framework can achieve multi-level alignment on two benchmarks across ten languages. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a promising approach for low resource languages. |
| Approach: | They propose to use both Recurrent Neural Networks & Transformer architectures to train NMT models. |
| Outcome: | The proposed model outperforms Statistical Machine Translation (SMT) techniques on a low resource Hindi-English language pair. |
Copied to clipboard
| Challenge: | a small-scale human evaluation confirms that the segments are highly parallel, making the dataset suitable for NLP applications. |
| Approach: | They present a first parallel corpus of Romansh idioms from 291 schoolbooks . they use automatic alignment methods to extract 207k multi-parallel segments from the books . |
| Outcome: | The proposed corpus is based on 291 schoolbook volumes, which are comparable in content for the five idioms. |
Copied to clipboard
| Challenge: | Identifying offensive spans in texts is the goal of the SemEval-2021 Task 5: Toxic Spans Detection . previous work focused on post level annotations, but identifying offensive span is useful in many ways. |
| Approach: | They propose a Python-based system to detect offensive spans in texts with pre-trained models and a user-friendly web-based interface. |
| Outcome: | The proposed system is based on a Python-based framework and a user-friendly web-based interface. |
Copied to clipboard
| Challenge: | Language Technologies can help in promoting and facilitating multilingualism in the Social Sciences and Humanities domain. |
| Approach: | They propose to use Natural Language Processing and Machine Translation to provide tools to foster multilingual access and discovery to SSH content across different languages. |
| Outcome: | The proposed tools prove to be a valid asset to translation tasks . validation of results by domain experts proficient in the language is an unavoidable phase of the whole workflow. |
Copied to clipboard
| Challenge: | E-commerce stores increasingly use Large Language Models to improve catalog data quality . a critical challenge is accurately predicting missing structured attribute values . |
| Approach: | They propose a retrieval-augmented system that leverages existing product catalog entries to guide LLM predictions for missing attributes. |
| Outcome: | The proposed system improves catalog data quality by 34% and accuracy by 0.8% . the proposed model can predict missing attributes in multilingual product catalogs . |
Copied to clipboard
| Challenge: | evaluating machine translation (MT) with cross-lingual information retrieval is relatively time-consuming and subjective. |
| Approach: | They propose a toolkit that evaluates machine translation with a proxy task of cross-lingual information retrieval. |
| Outcome: | The proposed toolkit is based on the "metrics shared task" of WMT2019. |
Copied to clipboard
| Challenge: | TokLens is an open-source toolkit for evaluating tokenizer quality across languages . authors evaluated 24 tokenizers from major LLM families across 15 typologically diverse languages - a gap that is stark in Japanese . |
| Approach: | They evaluate 24 tokenizers from major LLM families across 15 typologically diverse languages and correlate these metrics with downstream performance. |
| Outcome: | The proposed tokenizers produce 56x more tokens per word in Japanese than in English . the newer tokenizer Qwen2.5 and Gemma-2 reduce this gap to under 4x . |
Copied to clipboard
| Challenge: | Appraise is an open-source framework for crowd-based annotation tasks . it is used for shared tasks at the conference on machine translation and at IWSLT 2017 . |
| Approach: | They present an open-source framework for crowd-based annotation tasks . they describe the entire lifecycle of an Appraise evaluation campaign . |
| Outcome: | The proposed framework is used to run evaluation campaigns at the WMT Conference on Machine Translation and at IWSLT 2017 . it has been adopted by the translator team at Microsoft Translator for internal quality monitoring . |
Copied to clipboard
| Challenge: | Current multilingual agreement (MA) methods require parallel data between multiple language pairs, which is not always realistic and optimize the agreement in an ambiguous direction, which hampers the translation performance. |
| Approach: | They propose a novel multilingual agreement framework that optimizes agreement bidirectionally with the Kullback-Leibler Divergence loss. |
| Outcome: | The proposed method improves strong baselines on the task of multilingual neural machine translation with three benchmarks: TED Talks, News, and Europarl. |
Copied to clipboard
| Challenge: | Pretrained multilingual translation models with massive coverage are becoming of the backbone of many translation systems. |
| Approach: | They propose to use a gradient-based inference-time controller to control a pretrained multilingual model by using a model with attribute annotations. |
| Outcome: | The proposed model performs well on pretrained multilingual models and is attribute- rather than language-specific. |
Copied to clipboard
| Challenge: | Using cloud-based translation providers carries privacy risks, as users lose control of their data once it enters the web. |
| Approach: | They propose a desktop translation application that runs locally on a user's desktop or laptop CPU. translateLocally delivers cloud-like translation speed and quality even on 10 year old hardware. |
| Outcome: | The open-source translation system runs on Linux, Windows and macOS on desktops and laptops. |
Copied to clipboard
| Challenge: | Current approaches resort to suboptimal compromises and computational methods remain inadequate for translation. |
| Approach: | They propose a Constant-Variable Optimization (CVO) model for translation strategy and an Ovl metric for translation quality assessment that adapts to Chinese and English. |
| Outcome: | The proposed model improves performance on textual and visual puns while maintaining linguistic mechanisms and humorous effects. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models suffer from performance degradation when learning multiple languages. |
| Approach: | They propose to use LaSS to jointly train a single unified multilingual MT model. |
| Outcome: | The proposed model gains on 36 language pairs by up to 1.2 BLEU and zero-shot translation with 8.3 BLUE on 30 language pairs. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) are highly effective on mathematical, scientific, and other question-answering tasks. |
| Approach: | They compare an LRM's reasoning in English to that of a multilingual question . they find that English reasoning traces exhibit a substantially higher presence of cognitive behaviors . |
| Outcome: | The LRMs generate reasoning sequences in English, but the language of the question is not. |
Copied to clipboard
| Challenge: | Existing systems that bridge language barriers can introduce errors leading to misunderstandings and conversation breakdown. |
| Approach: | They propose a framework to integrate contextual information into automatic translation systems . they validate the framework on customer chat and user-assistant interaction . |
| Outcome: | The proposed framework consistently produces better translations than state-of-the-art systems on two task-oriented domains. |
Copied to clipboard
| Challenge: | a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence. |
| Approach: | They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence. |
| Outcome: | The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text. |
Copied to clipboard
| Challenge: | a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. |
| Approach: | They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages. |
| Outcome: | The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages. |
Copied to clipboard
| Challenge: | Existing approaches to identify complex semantic structures are difficult to train from under-annotated sources. |
| Approach: | They exploit relation- and event-relevant language-universal features to train relation or event extractors from source annotations and apply them to target languages. |
| Outcome: | The proposed approach achieves comparable performance to state-of-the-art models trained on 3,000 manually annotated mentions. |
Copied to clipboard
| Challenge: | Graph-based semantic parsing is one of the most promising general-purpose meaning representations . owing to this heterogeneity, most research focused on solutions specific to a given formalism . |
| Approach: | They propose a multilingual neural machine translation framework for Graph-based semantic parsing . they propose Graph2seq architecture that trains with an MNMT objective . |
| Outcome: | The proposed framework outperforms all competitors on cross-lingual parsing tasks. |
Copied to clipboard
| Challenge: | Large language models respond well in high-resource languages but struggle in low-resourced languages. |
| Approach: | They propose a method to construct cross-lingual instruction following samples with instruction in English and response in low-resource languages. |
| Outcome: | The proposed method builds a large-scale cross-lingual instruction tuning dataset on 10 languages. |
Copied to clipboard
| Challenge: | Existing models that only use auxiliary languages to encourage multilingual agreement ignore the relationships between different language pairs. |
| Approach: | They propose a multilingual agreement-based method which explicitly models the agreement between different translation directions by randomly substituting some fragments of the source language with their counterpart translations of auxiliary languages. |
| Outcome: | The proposed method improves on the multilingual translation task of 10 language pairs. |
Copied to clipboard
| Challenge: | Using a prototype, we present an automatic speech translation system for live subtitling of conference speech . the system is routinely tested in recognizing English, Czech, and German speech - and presenting it simultaneously into 42 target languages. |
| Approach: | They propose an automatic speech translation system aimed at live subtitling of conference presentations. |
| Outcome: | The proposed system is a working prototype that is routinely tested in recognizing English, Czech, and German speech and presenting it translated simultaneously into 42 target languages. |
Copied to clipboard
| Challenge: | X-STA is a new approach for cross-lingual machine reading comprehension . the variation of answer span positions in different languages makes it difficult to transfer knowledge across languages. |
| Approach: | They propose a method that leverages an attentive teacher to subtly transfer the answer spans of the source language to the answer output space of the target. |
| Outcome: | The proposed method outperforms state-of-the-art approaches on three multi-lingual datasets. |
Copied to clipboard
| Challenge: | Query Translator is a cross-lingual messaging app for the travel domain that automatically translates conversations . the application addresses common cross-linguistic communication issues such as translation accuracy, speed, privacy and personalization. |
| Approach: | They present a cross-lingual messaging app that automatically translates conversations while supporting keyword-to-sentence matching. |
| Outcome: | The proposed app translates conversations while supporting keyword-to-sentence matching. |
Copied to clipboard
| Challenge: | Large pretrained multilingual models have delivered promising results due to cross-lingual learning capabilities on a variety of language tasks. |
| Approach: | They propose to use language phylogenetic information to improve cross-lingual transfer by leveraging closely related languages in a structured, linguistically-informed manner. |
| Outcome: | The proposed model significantly improves on the baseline model on languages unseen during training. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) suffers from the scarcity of annotated training data, especially for low-resource languages without labeled data. |
| Approach: | They propose a cross-lingual entity projection framework to enable zero-shot cross-linguistic NER with the help of a multilingual labeled sequence translation model. |
| Outcome: | The proposed method outperforms the baseline method on two benchmarks by a large margin of +3 7 F1 scores and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing machine learning models for code comment generation are poorly suited for Russian . existing datasets that contain simple comments and docstrings in English are not suitable for function-level documentation generation. |
| Approach: | They propose a dataset specifically designed for Russian code documentation. |
| Outcome: | The first large-scale dataset specifically designed for Russian code documentation is based on human-written comments from GitHub repositories with synthetically generated ones. |
Copied to clipboard
| Challenge: | Pre-trained language models have achieved remarkable results on several NLP tasks. |
| Approach: | They propose three new masking strategies for cross-lingual visual pre-training that focus on learning different linguistic patterns. |
| Outcome: | The proposed methods outperform the baseline model and achieve state-of-the-art accuracy on the Portuguese-English MMT task. |
Copied to clipboard
| Challenge: | Recent studies in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way. |
| Approach: | They develop a multilingual discourse-aware benchmark to evaluate model performance on discourse phenomena in a given dataset. |
| Outcome: | The proposed model improves on previously studied phenomena while uncovering others which were not addressed. |
Copied to clipboard
| Challenge: | Existing trend prediction methods only make predictions within a language, but this is not enough to predict cross-lingual trends. |
| Approach: | They propose a method to predict which microblog trends will cross linguistic boundaries to become popular in other languages and when. |
| Outcome: | The proposed model outperforms existing trend prediction methods and LLM-based approaches by 4% in F1-score . |
Copied to clipboard
| Challenge: | Commercial translation systems support only one hundred languages or fewer . commercial translation systems do not make these models available for transfer to low resource languages . |
| Approach: | They propose a multilingual neural machine translation model that can translate from 500 source languages to English. |
| Outcome: | The proposed model can translate from 500 source languages to English, or be used as a parent model for low-resource languages. |
Copied to clipboard
| Challenge: | Using a cosine distance in a joint multilingual sentence embedding, we filter out noisy parallel data and mine for bitexts in large news collections. |
| Approach: | They propose to learn a joint multilingual sentence embedding and use the distance between sentences in different languages to filter noisy parallel data and to mine for parallel data in large monolingual texts. |
| Outcome: | The proposed approach improves a competitive baseline on the WMT'14 task by 0.3 BLEU by filtering out 25% of the training data. |
Copied to clipboard
| Challenge: | In-image machine translation is a sub-task of Image-Based Machine Translation that aims to substitute text embedded in images with its translation into another language. |
| Approach: | They propose a simple task that renders parallel text over a plain background and a pipeline that obtains the transcript of the original image, translates it, and generates a new image similar to the original one. |
| Outcome: | The proposed approach outperforms existing models including an end-to-end approach and is competitive with other similar approaches. |
Copied to clipboard
| Challenge: | Unsupervised BWE methods are evaluated on word translation or word similarity tasks. |
| Approach: | They propose a method that learns sentiment-specific word representations for two languages in a common space without cross-lingual supervision. |
| Outcome: | The proposed method outperforms previous unsupervised BWE methods and even supervised Bwe methods on three language pairs for cross-lingual sentiment analysis. |
Copied to clipboard
| Challenge: | Neural machine translation models with deeper neural networks are difficult to train. |
| Approach: | They propose a MultiScale Collaborative framework to boost gradient back-propagation . they let each encoder block learn a fine-grained representation and enhance it . |
| Outcome: | The proposed framework outperforms baseline models on translation tasks with three translation directions and achieves a BLEU score of 30.56 on the English-to-German task. |
Copied to clipboard
| Challenge: | Existing methods for extracting complete (binary) parses from pre-trained language models are expensive and time-consuming. |
| Approach: | They propose a chart-based method and an effective top-K ensemble technique to extractbinary parses from PLMs. |
| Outcome: | The proposed method can induce non-trivial parses for sentences from nine languages in an integrated and language-agnostic manner, and is robust to cross-lingual transfer. |
Copied to clipboard
| Challenge: | Existing methods for learning foreign languages are to use a spaced repetition system to learn new vocabulary. |
| Approach: | They use large language models to generate personalized stories using only the vocabulary they know. |
| Outcome: | The generated stories are more grammatical, coherent, and provide better examples of word usage than the standard beam search approach. |
Copied to clipboard
| Challenge: | Current multimodal machine translation systems rely on fully supervised data, which is costly to collect and prevents extension of MMT to language pairs with no such data. |
| Approach: | They propose a method to bypass the need for fully supervised data to train MMT systems . they adapt a strong text-only machine translation model to a visually conditioned language model and a divergence test set to evaluate how well models use images to disambiguate translations. |
| Outcome: | The proposed method can generalize to languages with no fully supervised training data. |
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
Copied to clipboard
| Challenge: | Specific-domain bilingual lexicons are composed of MultiWord Expressions (MWEs) the manual construction of MWEs bilingual dictionaries is costly and time-consuming. |
| Approach: | They propose to use word alignment approaches to automatically construct bilingual lexicons of MWEs from parallel corpora by formalizing the alignment process as an integer linear programming problem. |
| Outcome: | The proposed approach extracts and aligns multiword expressions from parallel corpora and then filters them using linguistic patterns to build bilingual lexicons. |
Copied to clipboard
| Challenge: | Realignment techniques are often employed to enhance cross-lingual transfer in multilingual language models, but can degrade performance in languages that differ significantly from the fine-tuned source language. |
| Approach: | They propose a method that freezes either the lower half or upper half of the layers during realignment to prevent performance degradation. |
| Outcome: | The proposed method improves Part-of-Speech (PoS) tagging performance in languages where realignment fails. |
Copied to clipboard
| Challenge: | End-to-end neural machine translation (NMT) has attracted increasing attention in recent years. |
| Approach: | They propose an adaptive multi-pass decoder which introduces a flexible multi- pass polishing mechanism to extend the capacity of NMT via reinforcement learning. |
| Outcome: | The proposed architecture improves Chinese-English translation with 1.55 BLEU . the proposed architecture adopts a flexible multi-pass polishing mechanism . |
Copied to clipboard
| Challenge: | Existing methods waiting-and-translating for a fixed duration break speech acoustic units . Existing models waiting-for a set duration and generating partial sentences are not effective . |
| Approach: | They propose a monotonic segmentation module inside an encoder-decoder model to detect proper speech unit boundaries for a streaming speech input. |
| Outcome: | The proposed method outperforms existing methods on a speech translation dataset and achieves the best trade-off between translation quality and latency. |
Copied to clipboard
| Challenge: | Existing studies have shown that NMT models trained to generate target syntax exhibit improved sentence structure relative to those trained on plain-text. |
| Approach: | They propose an approach to decoding ensembles of models generating different representations, focusing on models generating syntax. |
| Outcome: | The proposed approach gives state-of-the-art performance on a difficult Japanese-English task. |
Copied to clipboard
| Challenge: | Recent studies have shown that multilingual NMT models can handle more than one translation direction with a single system. |
| Approach: | They propose a multilingual neural machine translation model that can handle more than one translation direction with a single system. |
| Outcome: | The proposed model performs well in low-resource settings against bilingual systems. |
Copied to clipboard
| Challenge: | mEdIT is a multi-lingual extension to CoEdit for writing assistance. |
| Approach: | They propose to train multi-lingual large language models (LLMs) by fine-tuning them via instruction tuning. |
| Outcome: | The proposed model performs well on multilingual text editing benchmarks and generalizes well to new languages. |
Copied to clipboard
| Challenge: | Unlike annotation projection techniques, our model does not need parallel data during inference time. |
| Approach: | They propose a cross-lingual Encoder-Decoder model that simultaneously translates and generates sentences with semantic role annotations in a resource-poor target language. |
| Outcome: | The proposed model can be applied in monolingual, multilingual and cross-lingual settings and produces dependency-based and span-based annotations. |
Copied to clipboard
| Challenge: | Recent research in cross-lingual learning has found that combining large-scale pretrained multilingual language models with machine translation can yield good performance. |
| Approach: | They propose a model architecture that jointly encodes a source language input sentence with its translation to the target language during training and takes a target language sentence with it as input during evaluation. |
| Outcome: | The proposed model architecture can integrate machine translation to improve event extraction while adding machine-translated data yields unstable performance due to representational gap. |
Copied to clipboard
| Challenge: | Existing studies have shown that token overlap is a strong predictor of multilinguality and cross-lingual knowledge transfer between languages with different scripts. |
| Approach: | They propose a subword token alignability metric to understand the impact and quality of multilingual tokenisation. |
| Outcome: | The proposed metric predicts multilinguality much better when scripts are disparate and the overlap of literal tokens is low. |
Copied to clipboard
| Challenge: | Recent research in multilingual coreference and automatic pronoun translation has led to important insights into the problem and some promising results. |
| Approach: | They propose a corpus annotated with full coreference chains that addresses a problem that machine translation and other multilingual natural language processing (NLP) technologies face: translation of coreference across languages. |
| Outcome: | The proposed corpus contains parallel texts for the language pair English-German, two major European languages. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual text classification leverage text classifiers trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning. |
| Approach: | They propose to combine a neural machine translator and a text classifier trained in a high-resource language to perform text classification in other languages with no or minimal fine-tuning. |
| Outcome: | The proposed approach significantly improves over a baseline approach. |
Copied to clipboard
| Challenge: | This paper concerns the use of religious texts in natural language processing (NLP) religious texts are expressions of culturally important values, and machine learning models reproduce cultural values encoded in training data. |
| Approach: | They argue that NLP's use of religious texts raises considerations beyond model biases . authors argue that religious texts are culturally important and are often used by researchers . |
| Outcome: | The proposed method repurposes translations from their original uses and motivations, and raises considerations beyond model biases. |
Copied to clipboard
| Challenge: | Multilingual models are needed to process financial text, which is produced across the world and requires a large dataset. |
| Approach: | They propose to annotate a publicly available financial dataset using a hierarchical label structure and an annotation schema based on a real-world application. |
| Outcome: | The proposed model can be used in high-resource languages, but there is room for improvement in low-resourced languages. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains. |
| Approach: | They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs. |
| Outcome: | The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs. |
Copied to clipboard
| Challenge: | Current LLMs are primarily trained on English data but also include data from other languages. |
| Approach: | They propose to use a pre-translation strategy to translate a task prompt into English before inference . they use 'a modular entity' that could be translated into four different languages . |
| Outcome: | The proposed strategies are based on a set of pre-trained data across 35 languages covering both low and high-resource languages. |
Copied to clipboard
| Challenge: | Despite the success of low-resource neural machine translation, there is a data scarcity problem in many languages . large-scale, high-quality, and widecoverage bilingual corpora do not exist for most language pairs . |
| Approach: | They propose to quantify confidence of NMT models based on model uncertainty . they propose to use uncertainty-based confidence measures to improve back-translation . |
| Outcome: | The proposed model outperforms conventional statistical machine translation (SMT) on Chinese-English and English-German translation tasks. |
Copied to clipboard
| Challenge: | Existing textless speech-to-speech translation models have two main challenges: 1) learning cross-modal features and 2) learning alignment of difference languages in long sequences. |
| Approach: | They propose a unit language to overcome two main modeling challenges . they propose task prompt modeling to utilize the unit language in guiding the modeling process. |
| Outcome: | The proposed language improves over a strong baseline and achieves comparable performance to models trained with text. |
Copied to clipboard
| Challenge: | a large corpus of online content has been developed via large-scale crowdsourcing. |
| Approach: | They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform. |
| Outcome: | The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines. |
Copied to clipboard
| Challenge: | Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender. |
| Approach: | They propose to use probing methods to assess gender encoding across ST models. |
| Outcome: | The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed method achieves consistent improvements with faster convergence speed. |
| Approach: | They propose a curriculum learning method to gradually utilize pseudo bi-texts based on their quality from multiple granularities. |
| Outcome: | The proposed method achieves consistent improvements with faster convergence speed on WMT 14 En-Fr, WMT14 En-De, and LDC En-Zh translation tasks. |
Copied to clipboard
| Challenge: | Existing models for bilingual language modeling are limited due to lack of training data and syntactic structure. |
| Approach: | They propose a bilingual attention language model that performs language modeling objective with a quasi-translation objective to model the monolingual and cross-lingual sequential dependency. |
| Outcome: | The proposed model reduces the perplexity of 20.5% over the best-reported model. |
Copied to clipboard
| Challenge: | Existing studies have demonstrated the effectiveness of iterative back-translation, but its reason has not been sufficiently elucidated. |
| Approach: | They propose a method for machine translation known as iterative back-translation . they use two monolingual data to create a pseudo-bilingual data and update translation models . |
| Outcome: | The proposed method improves translation quality and improves BLEU. |
Copied to clipboard
| Challenge: | a recent study shows that tonal languages like Chinese have a higher classification performance than non-tonal languages like English. |
| Approach: | a new study examines the differences between tonal and non-tonal language classifications . they hypothesize that the difference is rooted in language typology . early cognitive decline is notoriously difficult to detect . |
| Outcome: | The proposed method compared to TAUKADIAL audio shows that Chinese and English perform better on Chinese . the findings suggest that language typology should inform the design of audio-based cognitive screening tools . |
Copied to clipboard
| Challenge: | Existing corpus ParCorFull contains parallel texts for English-German, French and Portuguese . translation of coreference across languages is challenging for MT and other NLP applications . |
| Approach: | They describe a parallel corpus annotated with full coreference chains for multiple languages . they use the existing corpus ParCorFull to study translation of coreference across languages - a challenge for machine translation and NLP . |
| Outcome: | The proposed corpus addresses translation of coreference across languages, a problem still challenging for machine translation and other multilingual natural language processing applications. |
Copied to clipboard
| Challenge: | Africa has over 2000 indigenous languages but they are under-represented in NLP research due to lack of datasets. |
| Approach: | They propose to use a dataset to classify sentiments for cross-domain adaptation for Nigerian and other African languages. |
| Outcome: | The proposed dataset compares the performance of cross-domain adaptation from Twitter domain and cross-lingual adaptation from English domain. |
Copied to clipboard
| Challenge: | Nigeria is a multilingual country with 500+ languages. |
| Approach: | They propose to use a pidgin and a creole to analyze the pidgins of Nigeria . they also use machine translation to analyze their results . |
| Outcome: | The results show that the two pidgins do not represent each other and are hard to teach . the results show the pidgin varieties are underrepresented in Generative AI . |
Copied to clipboard
| Challenge: | Despite recent advances in multilingual information retrieval, a significant gap remains between research efforts and real-world deployment. |
| Approach: | They propose to use Quranic multilingual corpus to develop an ad-hoc IR system that can satisfy users’ information needs in multiple languages. |
| Outcome: | The proposed model achieves promising results across diverse retrieval scenarios. |
Copied to clipboard
| Challenge: | FREME framework bridges Language Technologies (LT) and Linked Data (LD) core attributes of FREMe are usability, reusability and interoperability. |
| Approach: | They define user types and user levels and describe how they influence design decisions in a LT and Linked Data processing framework. |
| Outcome: | The proposed framework bridges Language Technologies (LT) and Linked Data (LD) it addresses common challenges that researchers and industry face when integrating LT and LD: interoperability, "silo" solutions and the lack of adequate tooling. |
Copied to clipboard
| Challenge: | MT metrics trained on segment-level human judgments are inherently non-transparent and reflect undesirable biases. |
| Approach: | They propose to use a type-based classifier metric to evaluate machine translation and compare it with a supervised and unsupervised one. |
| Outcome: | The proposed model outperforms other models in indicating cross-lingual information retrieval task performance and shows that it can be used to compare supervised and unsupervised neural machine translation. |
Copied to clipboard
| Challenge: | Existing literature on populism has only limited agreement on its exact properties . |
| Approach: | They propose a cross-lingual dataset to identify populist rhetoric in text . they propose 'hierarchical' annotation procedure to annotate populist references . |
| Outcome: | The proposed dataset can be used to investigate how political actors talk about The Elite and The People and to study how populist rhetoric is used as a strategic device. |
Copied to clipboard
| Challenge: | Hate speech and abusive language are global phenomena that need sociocultural background knowledge to be understood, identified, and moderated. |
| Approach: | They propose to use a multilingual dataset to collect hate speech and abusive language in 15 African languages to help improve model performance. |
| Outcome: | The proposed datasets are based on tweets annotated by native speakers familiar with the regional culture and show that they perform well in low-resource settings. |
Copied to clipboard
| Challenge: | Literature in Natural Language Processing (NLP) typically labels whole language with strict type of morphology, e.g. fusional or agglutinative. |
| Approach: | They propose to quantify morphological typology at the word and segment level by using two indices: synthesis (e.g. analytic to polysynthetic) and fusion (agglutinative to fusional). |
| Outcome: | The proposed method reduces the rigidity of NLP classification claims by measuring morphological diversity at the word and segment level. |
Copied to clipboard
| Challenge: | Pretrained language models learn cross-lingual knowledge and perform well on diverse tasks when finetuned. |
| Approach: | They propose a zero-shot prompting approach that captures cross-lingual word sense with a contextual prompt. |
| Outcome: | The proposed approach outperforms baselines on recall in many evaluation languages without additional training or finetuning. |
Copied to clipboard
| Challenge: | Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing. |
| Approach: | They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets. |
| Outcome: | The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter. |
Copied to clipboard
| Challenge: | a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data is presented. |
| Approach: | They propose a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual language models. |
| Outcome: | The proposed framework can be used to evaluate cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual models. |
Copied to clipboard
| Challenge: | Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages. |
| Approach: | They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties. |
| Outcome: | The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines. |
Copied to clipboard
| Challenge: | Complex Word Identification (CWI) is the task of identifying which words or phrases in a sentence are difficult to understand by a specific type of reader. |
| Approach: | They propose to use monolingual and cross-lingual CWI models to make predictions for languages not seen during training. |
| Outcome: | The proposed models perform as well as (or better than) most models submitted to the latest CWI Shared Task. |
Copied to clipboard
| Challenge: | a new method to identify code-switched data from the web is needed . code-witching is defined as the tendency of bilinguals to switch between languages . |
| Approach: | They propose a method that automatically collects code-switched tweets from the web . they use crowd-sourcing to obtain language identifiers for a subset of 8,000 tweets . |
| Outcome: | The proposed method identifies tweets as code-switched in languages L1 and L2 . it is compared to a Spanish-English corpus of code-witched tweets . |
Copied to clipboard
| Challenge: | Existing methods for identifying nuanced sociological concepts fail to capture domain-specific subtleties or require extensive parallel data. |
| Approach: | a new approach to aligning nuanced sociological concepts is proposed . a dual-branch LoRA approach captures core semantics and counteracts specific language perturbations. |
| Outcome: | a new approach outperforms baselines on cross-lingual sociological concept retrieval across 10 languages. |
Copied to clipboard
| Challenge: | Existing work on quantifying the prevalence of syntactic divergences across languages has not been done. |
| Approach: | They propose a framework for extracting divergence patterns for any language pair from a parallel corpus building on Universal Dependencies. |
| Outcome: | The proposed framework provides a detailed picture of cross-language divergences, generalizes previous approaches, and lends itself to full automation. |
Copied to clipboard
| Challenge: | Multilingual Machine Translation (MMT) benefits from knowledge transfer across different language pairs, but performance differences between one-to-many and many-to-1 translation are negligible. |
| Approach: | They conduct a large-scale study that varies the auxiliary target-side languages along two dimensions to show the dynamic impact of knowledge transfer on the main language pairs. |
| Outcome: | The proposed model can translate between multiple languages with minimal positive transfer ability. |
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Copied to clipboard
| Challenge: | End-to-end speech translation models have limited training data and are often inefficient due to the inconsistency of length and representation between speech and text. |
| Approach: | They find that the "modality gap" between speech and text data is not a major problem in E2E ST . they decouple the encoder to speech encoder and text encoder, and they find that there is a 'capacity gap' |
| Outcome: | The proposed model achieves 29.0 for en-de and 40.3 for fr on the MuST-C dataset. |
Copied to clipboard
| Challenge: | Existing models can reproduce existing social inequalities but cannot be reduced. |
| Approach: | They propose that models should maintain uncertainty when input is ambiguous to avoid reinforcing biases. |
| Outcome: | The proposed model can detect gender bias when translated to ambiguous and unambiguous sources and shows that it does not correlate with high translation accuracy and debiases the two cases differently. |
Copied to clipboard
| Challenge: | Existing paradigms for multilingual neural machine translation do not make full use of language commonality and parameter sharing. |
| Approach: | They propose a multilingual neural machine translation paradigm with one encoder-decoder model that makes full use of language commonality and parameter sharing. |
| Outcome: | The proposed method outperforms strong standard multilingual translation systems on WMT and IWSLT datasets. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a common linguistic phenomenon wherein speakers fluidly transition between languages in conversation. |
| Approach: | They propose to use a part-of-speech (POS)-based analysis of Spanish-English and Mandarin-English corpora to examine the propensity of bilinguals to engage in CS. |
| Outcome: | The findings confirm the existence of a statistically significant connection between POS and the likelihood of CS across language pairs, but show that it diminishes as tokens distance themselves from CS instances. |
Copied to clipboard
| Challenge: | Parallel sentence extraction is a task addressing the data sparsity problem found in multilingual natural language processing applications. |
| Approach: | They propose a bidirectional recurrent neural network based approach to extract parallel sentences from multilingual corpora. |
| Outcome: | The proposed approach outperforms existing approaches on noisy parallel corpora and shows significant improvements in translation performance. |
Copied to clipboard
| Challenge: | Parallel corpora play a vital role in advanced multilingual natural language processing tasks, notably in machine translation (MT). |
| Approach: | They manually and automatically evaluated four well-known publicly available parallel corpora across eleven language pairs. |
| Outcome: | The results show that the four well-known parallel corpora have a substantial amount of noisy sentence pairs, while CCMatrix and CCAligned have low quality sentences. |
Copied to clipboard
| Challenge: | Pre-trained multilingual language models are often better on English than other languages . however, they are trained on varying amounts of data for each language . |
| Approach: | They apply the MORALDIRECTION framework to multilingual models and analyse their results . they find that PMLMs encode differing moral biases, but these do not correspond to cultural differences or commonalities in human opinions. |
| Outcome: | The proposed model captures moral norms from English and imposes them on other languages. |
Copied to clipboard
| Challenge: | Existing models for multilingual RS are limited by capacity and data distribution skew . we propose Conditional Generative Matching models (CGM) to overcome these challenges . |
| Approach: | They propose Conditional Generative Matching models to address multilingual RS challenges . they use expressive message conditional priors, mixture densities and latent alignment . results exceed ROUGE scores by 10% on average, and 16% for low resource languages . |
| Outcome: | The proposed model exceeds baselines in relevance by 10% on average and 16% for low resource languages. |
Copied to clipboard
| Challenge: | Existing multilingual word translation methods focus on learning mappings from each language to a shared space. |
| Approach: | They propose a multilingual translation procedure that uses all the learned mappings to translate a word from one language to another. |
| Outcome: | Experiments on a standard multilingual word translation benchmark show that the proposed translation procedure outperforms state-of-the-art translation methods. |
Copied to clipboard
| Challenge: | Existing sentence alignment systems focus on auxiliary information such as document metadata and hyperparameter-sensitive techniques, and neglect the crucial role that context plays in the alignment process. |
| Approach: | They propose a context-aware, end-to-end and fully-neural architecture for sentence alignment that maps source and target sentences in long documents by contextualizing their sentence embeddings with respect to the other sentences in the document. |
| Outcome: | The proposed system maps source and target sentences in long documents by contextualizing their sentence embeddings with respect to the other sentences in the document. |
Copied to clipboard
| Challenge: | Using frameworks such as Universal Dependencies (UD) to transfer knowledge between languages can be challenging because of variation in syntactic structures. |
| Approach: | They propose a typologically driven method which reduces anisomorphism in UD treebanks by considering both morphological and structural properties. |
| Outcome: | The proposed method is effective for machine translation and cross-lingual sentence similarity. |
Copied to clipboard
| Challenge: | Existing parallel datasets limit multilingual evaluations due to the nature of linguistic annotation, which is tedious, subjective, and costly. |
| Approach: | They extend the Cross-lingual Natural Language Inference corpus with Croatian and use Facebook's 1.2B parameter m2m_100 model to analyze the train set and compare its quality with the existing machine-translated German set. |
| Outcome: | The proposed model is consistent with other XNLI dubs and is compared with the existing machine-translated German train set. |
Copied to clipboard
| Challenge: | a monolingual speaker can learn to translate by looking up a bilingual dictionary . a novel task of machine translation (MT) is based on no parallel sentences but can refer to a ground-truth bilingual dictionary and large-scale monolingual corpora. |
| Approach: | They propose a task of machine translation that uses a bilingual dictionary and large-scale monolingual corpora to translate a monolingual speaker. |
| Outcome: | The proposed task is based on a bilingual dictionary and large scale monolingual corpora, while being independent on parallel sentences. |
Copied to clipboard
| Challenge: | End-to-end speech-totext translation (ST) models require large amounts of data to train, but their size is considerably smaller than text-based MT data. |
| Approach: | They propose a method to convert MT data to ST data via text-to-speech systems. |
| Outcome: | The proposed method improves translation quality by an average of 1.83 BLEU score while performing equally well as TTS-generated speech in improving translation quality. |
Copied to clipboard
| Challenge: | a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs. |
| Approach: | They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus. |
| Outcome: | The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it . |
Copied to clipboard
| Challenge: | Temporal Expression Extraction (TEE) is essential for understanding time in natural language. |
| Approach: | They propose a framework for multilingual Temporal Expression Extraction that leverages pre-trained language models to prompt cross-language knowledge transfer from English to non-English languages. |
| Outcome: | The proposed framework outperforms the existing SOTA methods on French, Spanish, Portuguese, and Basque by large margins. |
Copied to clipboard
| Challenge: | Using the existing English dataset, we can use the subjectivity classification to test the ability of pre-trained multilingual models to transfer knowledge between languages. |
| Approach: | They propose to use a Czech subjectivity dataset of 10k manually annotated subjective and objective sentences as a cross-lingual benchmark. |
| Outcome: | The proposed dataset is the first subjectivity dataset for the Czech language and also includes 200k automatically labeled sentences. |
Copied to clipboard
| Challenge: | Existing studies of entrainment in code-switched domains have been limited to human-machine textual interactions. |
| Approach: | They propose to use acoustic-prosodic features to identify multiple dimensions and feature sets of entrainment in code-switched speech. |
| Outcome: | The findings give rise to important implications for the potentially “universal” nature of entrainment as a communication phenomenon and potential applications in inclusive and interactive speech technology. |
Copied to clipboard
| Challenge: | a theoretical analysis of crosslingual transfer in probabilistic topic models is presented . we use Gibbs sampling to quantify the loss of knowledge across languages . |
| Approach: | They propose a method to quantify the loss of knowledge across languages during crosslingual transfer in probabilistic topic models. |
| Outcome: | The proposed model quantifies the loss of knowledge across languages during this process . it is validated on a diverse set of five languages and discusses best practices for data collection and model design . |
Copied to clipboard
| Challenge: | Existing approaches to improve neural machine translation use token-level adaptive training . however, standard models make predictions on condition of previous contexts . |
| Approach: | They propose a target-context-aware metric which can be supplemented by statistical metrics . they propose an adaptive training approach based on token- and sentence-level CBMI . |
| Outcome: | The proposed model outperforms the Transformer baseline and other similar approaches on English-German and Chinese-English tasks. |
Copied to clipboard
| Challenge: | Existing cross-lingual topic models depend on sparse bilingual resources and often yield incoherent or weakly aligned topics. |
| Approach: | They propose a framework that integrates LLM-guided topic refinement with self-consistency uncertainty quantification to enable black-box, stable, and scalable enhancement of cross-lingual topic models. |
| Outcome: | Experiments on multilingual corpora show that the proposed framework achieves superior topic coherence and alignment while reducing reliance on bilingual dictionaries and expensive LLM calls. |
Copied to clipboard
| Challenge: | Existing models require associated image with input sentence, which is difficult to satisfy at inference. |
| Approach: | They propose to use synthetic and authentic images to generate translations using text-to-image generation models. |
| Outcome: | The proposed model achieves state-of-the-art performance on En-De and En-Fr datasets while remaining independent of authentic images during inference. |
Copied to clipboard
| Challenge: | a new CLTL model is proposed to facilitate cross-linguistic transfer learning between distant languages . a key to CLTL is to learn a shared representation space for the given source-target language pair. |
| Approach: | They propose a new CLTL model that integrates machine translation with MT . they use an unannotated data technique to make use of the model's pre-training and fine-tuning . |
| Outcome: | The proposed model achieves better CLTL performance than the baseline model without more annotated data. |
Copied to clipboard
| Challenge: | a neural machine translation system generates a translation t in the target language, but for any sentence of non-trivial complexity, the translation s is not unique. |
| Approach: | They propose a method to quantify the amount of information missing in a machine translation system. |
| Outcome: | The proposed model captures extra information from a single float representation of the target sentence and reproduces it with two 32-bit floats per target token. |
Copied to clipboard
| Challenge: | appositives are phrases that appear next to a noun phrase and serve an explicative function. |
| Approach: | They propose a more realistic end-to-end definition of appositive generation with a dataset that spans four languages and two entity types. |
| Outcome: | The proposed model is non-trivial and leaves plenty of room for improvement. |
Copied to clipboard
| Challenge: | XCOPA dataset provides a typologically diverse dataset for commonsense reasoning in 11 languages . current methods for evaluating commonsensible reasoning in resource-poor languages are weak compared to translation-based transfer. |
| Approach: | They propose a typologically diverse multilingual dataset for causal commonsense reasoning in 11 languages. |
| Outcome: | The proposed model performs better than current methods on a resource-poor dataset compared to translation-based transfer in the 11 languages studied . |
Copied to clipboard
| Challenge: | Zero pronouns (ZPs) are often omitted in pro-drop languages, but should be recalled in non-pro-drop language. |
| Approach: | They propose to analyze the literature on zero pronoun translation after the neural revolution . they uncover that data limitation causes learning bias in languages and domains . |
| Outcome: | The proposed method and methods are compared to other models and evaluation metrics on different benchmarks. |
Copied to clipboard
| Challenge: | Existing studies have shown that existing models amplify biases observed in training data. |
| Approach: | They propose to use MT and NLP to amplify biases observed in training data to investigate how bias amplification might affect language in a broader sense. |
| Outcome: | The proposed model amplifys biases observed in training data and could lead to an artificially impoverished language, the authors show. |
Copied to clipboard
| Challenge: | Using zero-shot rankers, cross-lingual IR models are limited by their language coverage. |
| Approach: | They propose to train ranking models on artificially code-switched data instead of using a dictionary. |
| Outcome: | The proposed approach is robust towards the ratio of code-switched tokens and extends to unseen languages. |
Copied to clipboard
| Challenge: | despite the growing need for advanced signing technologies, signed language resources remain scarce. |
| Approach: | They propose a linguistically informed alignment algorithm that matches instances between signed languages . they compare similarities and differences across three signed languages to develop a model . |
| Outcome: | The proposed algorithm performs well on automatic metrics for sign-to-sign translation and generation. |
Copied to clipboard
| Challenge: | We present a novel system for cross-lingual summarization that can be applied to low-resource languages. |
| Approach: | They propose a neural abstractive summarization system that can be applied to low-resource languages . they use machine translation and the New York Times summarizing corpus to create a corpus . |
| Outcome: | The proposed system achieves higher fluency than standard summarizers on translated documents . the proposed system can be easily applied to new low-resource languages . |
Copied to clipboard
| Challenge: | a large number of scientific journals are published exclusively in English . this creates barriers for non-native English speakers to access scientific knowledge . |
| Approach: | They propose a way to translate scientific articles while preserving native JATS XML formatting. |
| Outcome: | The proposed approach shows that the key scientific details are accurately conveyed. |
Copied to clipboard
| Challenge: | Neural syntactic distance (NSD) is used to represent constituent trees using a sequence whose length is identical to the number of words in the sentence. |
| Approach: | They propose five strategies to improve NMT with explicit use of syntactic information . et al., 2014) propose a set of five strategies that incorporate syntastic information into the encoder and/or decoder of the baseline model. |
| Outcome: | The proposed strategies improve translation performance of the baseline model (+2.1 (En–Ja), +1.3 (Ja–En), +1.2 (En-Ch), and +1.0 (Ch–En) BLEU. |
Copied to clipboard
| Challenge: | Historically, metrics for evaluating the quality of machine translation (MT) have relied on basic, lexical-level features such as counting the number of matching n-grams between the MT hypothesis and the reference translation. |
| Approach: | They propose a neural framework for training multilingual machine translation evaluation models which exploits human judgements to obtain new state-of-the-art levels of correlation with MT quality. |
| Outcome: | The proposed framework achieves state-of-the-art performance on the WMT 2019 Metrics shared task and demonstrate robustness to high-performing systems. |
Copied to clipboard
| Challenge: | a growing demand for translations and multilingual content is surpassing the supply of professional translation services. |
| Approach: | They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality. |
| Outcome: | The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities. |
Copied to clipboard
| Challenge: | a new benchmarking suite for natural language processing (NLP) is proposed to measure progress in the area of natural language understanding. |
| Approach: | They propose to use machine translation to translate a superGLUE benchmark into Slovene . they propose to combine monolingual, cross-lingual, and multilingual models . |
| Outcome: | The proposed model is superior to multilingual models but lags behind the best English models. |
Copied to clipboard
| Challenge: | Experimental results show that cross-language data expansion results in performance degradation. |
| Approach: | They leverage cross-language data expansion and retraining to enhance neural Event Detection on English ACE corpus. |
| Outcome: | The proposed method improves ED performance by 1.6% over the straight data combination. |
Copied to clipboard
| Challenge: | Experiments show that ShifCon significantly enhances the performance of non-dominant languages due to the imbalance in training data across languages. |
| Approach: | They propose a Shift-based multilingual Contrastive framework that aligns the internal forward process of other languages toward that of the dominant one. |
| Outcome: | The proposed framework significantly improves performance of non-dominant languages, particularly for low-resource ones. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings. |
| Approach: | They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size . |
| Outcome: | The proposed approach performs best in MT tasks but under-performs in other languages. |
Copied to clipboard
| Challenge: | Existing studies on unlearning in multilingual large language models focus on monolingual settings, typically English. |
| Approach: | They propose to use a multilingual data and concept unlearning model to investigate the problem . they extend benchmarks for factual knowledge and stereotypes into ten languages . |
| Outcome: | The proposed model is able to unlearning in 10 languages across five languages and resource levels. |
Copied to clipboard
| Challenge: | Experimental results show document-level neural machine translation improves lexical consistency . inconsistent translations tend to confuse readers in some cases . |
| Approach: | They propose to use a word link to obtain a document word link and an auxiliary loss function to constrain that their translation should be consistent. |
| Outcome: | The proposed approach improves translation consistency on ChineseEnglish and EnglishFrench translation tasks. |
Copied to clipboard
| Challenge: | Experimental results show that bidirectional training pushes the SOTA neural machine translation performance significantly higher. |
| Approach: | They propose a bidirectional training strategy that updates model parameters at the early stage and tunes it normally. |
| Outcome: | The proposed approach pushes the SOTA neural machine translation performance significantly higher on 15 translation tasks on 8 language pairs. |
Copied to clipboard
| Challenge: | Prior approaches for predicting code-switching only consider shallow linguistic context. |
| Approach: | They hypothesize that enriching models with speaker information can guide them to pick up on relevant inductive biases. |
| Outcome: | The proposed model improves on a speaker-driven task in English–Spanish bilingual dialogues by adding sociolinguistically-grounded speaker features as prepended prompts. |
Copied to clipboard
| Challenge: | Word Sense Disambiguation is a crucial task in Natural Language Processing . supervised systems need to be trained on word-by-word basis, a problem that is beyond reach for resource-rich languages like English. |
| Approach: | They release six large-scale sense-annotated datasets in multiple languages to pave the way for supervised multilingual Word Sense Disambiguation. |
| Outcome: | The results show that large-scale sense annotations can be used as training sets for supervised systems. |
Copied to clipboard
| Challenge: | Vision-and-language models with separate encoders for each modality are limited in availability. |
| Approach: | They propose a multilingual benchmark that offers (partial) translations of ImageNet labels to 100 languages, built without machine translation or manual annotation. |
| Outcome: | The proposed model outperforms models on English and low-resource languages. |
Copied to clipboard
| Challenge: | Document machine translation typically suffers from a lack of document-level bilingual data. |
| Approach: | They propose a document machine translation model that incorporates contextual information into the training signals by capturing cross-sentence dependency within the target document and cross sentence translation to make better use of contextual information. |
| Outcome: | The proposed model outperforms baselines on three benchmark datasets and significantly outperformed previous approaches. |
Copied to clipboard
| Challenge: | a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations. |
| Approach: | They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation. |
| Outcome: | The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation. |
Copied to clipboard
| Challenge: | Intuitively, Hindi and English corpora should aid improve task performance on code-switched Hindi-English. |
| Approach: | They propose a meta-learning framework that utilizes the labelled resources of the downstream tasks in the constituent languages to improve task performance. |
| Outcome: | The proposed framework improves the performance on downstream tasks on code-switched Hindi-English. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) generates translations in isolation, resulting in translation inconsistency and ambiguity. |
| Approach: | They propose to incorporate referring process into translation decoding of NMT by using local coordinates coding to obtain global context vectors containing monolingual and bilingual contextual information. |
| Outcome: | The proposed model improves translation quality with lightweight computation cost on Chinese-English and English-German translation tasks. |
Copied to clipboard
| Challenge: | Modern artificial intelligence is characterized by large pretrained language models with strong language capabilities to be adapted to various downstream tasks. |
| Approach: | They propose to use the task of speech translation (ST) to pretrain speech models for end-to-end SLU on intra- and cross-lingual scenarios. |
| Outcome: | The proposed model achieves higher performance over baselines on monolingual and multilingual intent classification as well as spoken question answering using SLURP, MINDS-14, and NMSQA benchmarks. |
Copied to clipboard
| Challenge: | Recent efforts to train cross-lingual models on source language fail to take advantage of data transfer . current methods focus on learning task-specific information from syntactical features or word-label relations in target language. |
| Approach: | They propose a hybrid knowledge-transfer approach that leverages a teacher-student framework . the model is evaluated on a distinct target language for which there is no labeled data . |
| Outcome: | The proposed model achieves state-of-the-art results on 9 morphologically-diverse target languages across 3 distinct datasets. |
Copied to clipboard
| Challenge: | Empirical results show that a sentence-level agreement module can significantly improve the performance of neural machine translation (NMT) |
| Approach: | They propose a sentence-level agreement module to minimize the difference between the representation of source and target sentences. |
| Outcome: | Empirical results show the proposed agreement module significantly improves translation performance. |
Copied to clipboard
| Challenge: | Existing approaches to train multiple languages with a shared encoder and multiple decoders are based on denoising autoencoding of each language and back-translating between English and multiple non-English languages. |
| Approach: | They propose a multilingual unsupervised NMT scheme which trains multiple languages with a shared encoder and multiple decoders. |
| Outcome: | The proposed model performs better than the separately trained bilingual models on monolingual corpora and improves by 1.48 BLEU points on WMT test sets. |
Copied to clipboard
| Challenge: | Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language . however, many languages lack the labeled training data necessary for training deep neural nets for varying NLP tasks. |
| Approach: | They propose a cross-lingual transfer learning method that leverages annotated data from other languages to build NLP models for a target language. |
| Outcome: | The proposed model achieves significant performance gains over prior art over multiple text classification and sequence tagging tasks including a large-scale industry dataset. |
Copied to clipboard
| Challenge: | Existing models for translation of ambiguous text use context to disambiguate meaning . current models for MTs consistently translate English idioms literally, whereas LMs are context-aware . |
| Approach: | They use a dataset of 512 pairs of English sentences to study semantic ambiguities . they use literal and figurative idioms to disambiguate intended meaning . |
| Outcome: | The results show that current models translate English idioms literally, even when the context suggests a figurative interpretation. |
Copied to clipboard
| Challenge: | We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script. |
| Approach: | They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic. |
| Outcome: | The proposed training recipe improves multilingual transliteration for Indic languages. |
Copied to clipboard
| Challenge: | Recent studies have employed machine translation systems for cross-lingual VQA tasks . however, translated texts contain unique characteristics distinct from human-written ones, referred to as translation artifacts. |
| Approach: | They propose a machine translation system that can train models in multiple languages . they propose augmentation strategies that reduce translation artifacts in translated texts . |
| Outcome: | The proposed approach reduces translation artifacts in models across languages and languages. |
Copied to clipboard
| Challenge: | a method for automatic extraction of bilingual multiword units (BMWUs) from a parallel corpus has been shown to be useful for estimating human translation quality. |
| Approach: | They applied a method for automatic extraction of bilingual multiword units from a parallel corpus in order to investigate their contribution to translation quality in terms of adequacy and fluency. |
| Outcome: | The method is based on generalized additive modelling and it shows that normalized BMWU ratios can be useful for estimating human translation quality. |
Copied to clipboard
| Challenge: | Existing efforts on text synthesis for code-switching require training on code-witched texts in the target language pairs. |
| Approach: | They propose a model that synthesizes code-switched texts for language pairs absent from training data by adding an additional code-sharing module to a pre-trained machine translation model. |
| Outcome: | The proposed model synthesizes code-switched texts for language pairs lacking from training data. |
Copied to clipboard
| Challenge: | a corpus of Guarani sentences with sentence-level alignment is presented . the corpus contains 228,000 Guaran tokens along with 336,000 Spanish tokens . |
| Approach: | They propose to develop a Guarani - Spanish parallel corpus with sentence-level alignment . the corpus contains 228,000 Guaran tokens along with 336,000 Spanish tokens . |
| Outcome: | The proposed corpus contains 22,800 Guarani tokens along with 336,000 Spanish tokens extracted from web sources. |
Copied to clipboard
| Challenge: | Prior work favors simplified label translation or relying on word-level alignments for label projection. |
| Approach: | They propose a novel approach CLaP which translates text to target language and performs *contextual translation* on the labels using the translated text as the context. |
| Outcome: | The proposed approach improves translation accuracy on two prediction tasks and shows 2.4 F1 improvement for EAE and 1.4 F1 for named entity recognition. |
Copied to clipboard
| Challenge: | Synthetic translations have been used for a wide range of NLP tasks, but it remains unclear how they differ from naturally occurring data. |
| Approach: | They propose to use a semantic equivalence classifier to improve bitext quality without additional bilingual supervision to replace the originals. |
| Outcome: | The proposed samples improve bitext quality without additional bilingual supervision and are validated intrinsically and extrinsically through bilingual induction and MT tasks. |
Copied to clipboard
| Challenge: | Existing methods for machine translation evaluation use source sentences as pseudo references instead of word symbols. |
| Approach: | They propose an automatic machine translation evaluation method that uses source sentences as pseudo references instead of source sentences. |
| Outcome: | The proposed method achieves higher correlation with human judgments than baseline evaluation method that uses only hypothesis and reference sentences. |
Copied to clipboard
| Challenge: | Existing approaches to learn orthogonal matrix aligning bilingual lexicons are suboptimal . resulting models suffer from "hubness problem" because word vectors tend to be nearest neighbors of abnormally high number of other words. |
| Approach: | They propose a unified formulation that directly optimizes a retrieval criterion in an end-to-end fashion. |
| Outcome: | The proposed approach outperforms the state-of-the-art on word translation on standard benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to extract bilingual terminologies from corpora are limited . MWTs pose serious challenges for alignment and machine translation systems . |
| Approach: | They propose an approach to build comparable corpora and bilingual term dictionaries that evaluate bilingual term alignment in comparable corpus. |
| Outcome: | The proposed method is validated on an existing dataset and manually annotated data. |
Copied to clipboard
| Challenge: | Pronouns are often omitted in pro-drop languages, such as Chinese . this leads to various translation problems in terms of completeness, syntax and semantics . |
| Approach: | They propose a reconstruction-based approach to alleviate dropped pronoun (DP) translation problems for neural machine translation models by employing a shared reconstructor and a joint learning approach. |
| Outcome: | The proposed approach improves translation performance and accuracy of DP predictions. |
Copied to clipboard
| Challenge: | Existing mRAG systems suffer from a language bias during reranking, systematically favoring English and the query’s native language. |
| Approach: | They propose a language-agnostic utility-driven reranker alignment technique to mitigate language bias during re-ranking. |
| Outcome: | The proposed approach mitigates language bias and consistently improves mRAG performance across languages. |
Copied to clipboard
| Challenge: | Tokenization is the first step of most NLP pipelines. |
| Approach: | They propose a parity-aware byte pair encoder that maximizes the compression gain of the currently worst-compressed language for cross-lingual parity. |
| Outcome: | a new algorithm reduces tokenization inequality by 89% compared to classical BPE . the proposed algorithm is based on a fair-max rule that maximizes the compression gain of the currently worst-compressed language . |
Copied to clipboard
| Challenge: | End-to-end speech translation requires a powerful encoder to transcribe, understand and learn cross-lingual semantics simultaneously. |
| Approach: | They propose a curriculum pre-training method that includes an elementary course for transcription learning and two advanced courses for understanding the utterance and mapping words in two languages. |
| Outcome: | The proposed method improves on En-De and En-Fr speech translation benchmarks. |
Copied to clipboard
| Challenge: | Using recurrent neural networks to build language models for code-switched text is an important problem with implications to downstream applications such as speech recognition and machine translation. |
| Approach: | They propose a novel recurrent neural network unit with dual components that focus on each language in the code-switched text separately and a generative model estimated using the training data. |
| Outcome: | The proposed techniques yield significant reductions in perplexity on Mandarin-English task and improve on baseline models. |
Copied to clipboard
| Challenge: | Multilingual machine translation (MMT) is a key tool for improving translation in low-resource languages. |
| Approach: | They examine how denoising autoencoding and backtranslation impact multilingual machine translation under different data conditions and model scales. |
| Outcome: | The proposed method improves translation efficiency in low-resource languages by using denoising autoencoding (DAE) and backtranslation (BT) . |
Copied to clipboard
| Challenge: | despite advances in English-Thai MT, common MT approaches often underperform in the medical field due to their inability to precisely translate medical terminologies. |
| Approach: | They propose to maintain medical terminology in English within translated text through code-switched translation. |
| Outcome: | The proposed method shows that medical professionals prefer CS translations that maintain critical English terms accurately, even if it slightly compromises fluency. |
Copied to clipboard
| Challenge: | Word embedding is central to neural machine translation, but indirectly interfaces with other layers, making them comparatively isolated. |
| Approach: | They propose a shared-private bilingual word embedding which gives a closer relationship between the source and target embedders and reduces the number of model parameters. |
| Outcome: | The proposed model improves on 5 language pairs belonging to 6 different language families and written in 5 different alphabets and significantly reduces model parameters. |
Copied to clipboard
| Challenge: | Pre-trained multilingual language models represent multiple languages in a single vector space, a feature hypothesized to enable impressive crosslingual transfer capabilities. |
| Approach: | They propose to use a multilingual representation space that sorts axes based on their language-separability to determine whether geometric distances between languages correlate with crosslingual transfer performance. |
| Outcome: | The proposed measures do not generalize well across models, layers, and tasks. |
Copied to clipboard
| Challenge: | Recent work in multilingual machine translation (MMT) has focused on the potential of positive transfer between languages. |
| Approach: | They propose to augment training data with alternative signals that unify different writing systems, such as phonetic, romanized, and transliterated input. |
| Outcome: | The proposed model outperforms strong ensemble baselines on Indic and Turkic languages by 1.3 BLEU points on both languages. |
Copied to clipboard
| Challenge: | Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research . |
| Approach: | They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora. |
| Outcome: | The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs . |
Copied to clipboard
| Challenge: | Existing multilingual benchmarks show severe drawbacks, such as overly translated content, the absence of difficulty control, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. |
| Approach: | They propose a multilingual benchmark that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening to provide a comprehensive, difficult multilingual assessment. |
| Outcome: | The proposed benchmark features 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. |
Copied to clipboard
| Challenge: | Using labeled NLI datasets for learning sentence embeddings leads to improved performance for natural language understanding tasks. |
| Approach: | They compare two data augmentation techniques for learning better sentence embeddings . they use a cross-lingual transfer technique that exploits English resources as training data to yield non-English sentence embeds as zero-shot inference . |
| Outcome: | The proposed techniques yield better performance on Japanese and Korean sentences. |
Copied to clipboard
| Challenge: | rumors with multimedia content are becoming more and more common on social networks . a new feature set is proposed to verify rumors pivoting on multimedia content . |
| Approach: | They propose to use multimedia content to find external information on social media platforms . they propose to leverage semantic similarity between rumors and external information . |
| Outcome: | The proposed approach achieves state-of-the-art results on social networks . it leverages semantic similarity between rumors and external information . |
Copied to clipboard
| Challenge: | a method for completing multilingual translation dictionaries is proposed . a 27% relative improvement in whole-word accuracy is achieved when multilingual data is unavailable . |
| Approach: | They propose a method for completing multilingual translation dictionaries using multilingual inputs and multilingual decoding objective. |
| Outcome: | The proposed method can synthesize new word forms in multilingual translation dictionaries . it can perform in settings where correct translations have not been observed in text . |
Copied to clipboard
| Challenge: | Multilingual pre-trained language models have demonstrated impressive (zero-shot) cross-lingual transfer abilities, however, their performance is hindered when the target language has distant typology from the source language or when pre-training data is limited in size. |
| Approach: | They propose a method that contextually retrieves prompts as flexible guidance for encoding instances conditionally. |
| Outcome: | The proposed method improves on the XTREME task and also for low-resource languages in unsupervised sentence retrieval. |
Copied to clipboard
| Challenge: | Existing multimodal machine translation methods require paired input of source sentence and image, which makes them suffer from shortage of sentence-image pairs. |
| Approach: | They propose a phrase-level retrieval-based method to get visual information from existing sentence-image data sets. |
| Outcome: | The proposed method significantly outperforms strong baselines on multiple MMT datasets, especially when the textual context is limited. |
Copied to clipboard
| Challenge: | Recent advances in multilingual pretrained models have proven effective at zero-shot transfer to a wide variety of languages, but this transfer is not universal, with many languages not currently understood by multilingual approaches. |
| Approach: | They propose a general approach that requires only unlabelled text to detect which languages are not well understood by a cross-lingual model. |
| Outcome: | The proposed model can detect which languages are not well understood by a multilingual model on 350 low-resource languages. |
Copied to clipboard
| Challenge: | In this paper, we propose to extend the recently introduced model-agnostic meta-learning algorithm for low-resource neural machine translation (NMT). |
| Approach: | They propose to extend the recently introduced meta-learning algorithm for low-resource neural machine translation (NMT) they frame low-Resource translation as a meta- learning problem where we learn to adapt to low-REsource languages based on multilingual high-resourced language tasks. |
| Outcome: | The proposed meta-learning algorithm outperforms the multilingual, transfer learning based approach and can train a competitive NMT system with only a fraction of training examples. |
Copied to clipboard
| Challenge: | Existing methods for bilingual Lexicon Induction use nonparallel corpora, but hubness often degrades accuracy. |
| Approach: | They propose a method to create a lexicon of translation equivalents from non-parallel corpora by aligning two word embedding spaces and retrieving the nearest neighbor (NN) this method reduces hubness, which is necessary for retrieval tasks. |
| Outcome: | The proposed method outperforms NN, Inverted SoFtmax and other state-of-the-art methods. |
Copied to clipboard
| Challenge: | Language Technologies (LTs) are a powerful means to break down language barriers impacting business, cross-lingual and cross-cultural communication in Europe. |
| Approach: | They present an overview of the European LT landscape and the current state of play in industry and the LT market. |
| Outcome: | The present study outlines funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. |
Copied to clipboard
| Challenge: | We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations. |
| Approach: | They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data . |
| Outcome: | The proposed method improves on a standard attention baseline by using copy attention. |
Copied to clipboard
| Challenge: | Word sense disambiguation is a widely studied NLP task of identifying the meaning of a word in context. |
| Approach: | They propose a method to create parallel sense-annotated datasets in English . they use machine translation, word alignment, sense projection, and sense filtering to produce silver annotations . |
| Outcome: | The proposed method produces parallel sense-annotated datasets on Farsi, Chinese, and Bengali . the results are higher than those obtained with recent multilingual systems, the authors say . |
Copied to clipboard
| Challenge: | Communication practices vary across cultures. Inherent differences in how people think and behave influence cultural norms. |
| Approach: | They propose a framework to extract stylistic differences from multilingual language models (LMs) they use a multilingual lexica to consolidate feature importances into comparable lexical categories . |
| Outcome: | The proposed framework generates comprehensive style lexica in any language and consolidates feature importances from LMs into comparable lexical categories. |
Copied to clipboard
| Challenge: | State-of-the-art unsupervised multilingual models generalize in zero-shot cross-lingual setting . generalization ability attributed to shared subword vocabulary and joint training across multiple languages . |
| Approach: | They propose an approach that transfers a monolingual model to new languages at the lexical level. |
| Outcome: | The proposed approach is competitive with multilingual BERT on cross-lingual classification benchmarks and on a new cross-linguistic question answering dataset. |
Copied to clipboard
| Challenge: | Medical visual question answering (VQA) and federated learning (FL) are important tools for privacy-preserving collaborative learning. |
| Approach: | They propose a cross-modal FL framework that uses modality-expert low-rank adaptation for medical visual question answering (VQA) X-FLoRA enables the synthesis of images from one modality to another without requiring data sharing . |
| Outcome: | Experiments show that X-FLoRA outperforms existing FL methods in terms of performance . XFLorage enables synthesis of images from one modality to another without data sharing . |
Copied to clipboard
| Challenge: | a new study analyzes the nature of twitter data and compares it with other social networking websites. |
| Approach: | They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool. |
| Outcome: | The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets. |
Copied to clipboard
| Challenge: | Several recent papers claim to have achieved human parity at sentence-level machine translation. |
| Approach: | They propose to use a dataset with rich discourse annotations to evaluate MT performance . they find that MT outputs differ fundamentally from human translations in terms of latent discourse structures. |
| Outcome: | The proposed dataset builds upon the large-scale parallel corpus BWB . it covers 15,095 entity mentions in both languages and compares them to human translations . |
Copied to clipboard
| Challenge: | Existing datasets for question answering (QA) tasks mostly support only English . however, existing resources for these tasks are labor intensive . |
| Approach: | They propose to combine Korean QA datasets with machine-translated English resources to build seed resources. |
| Outcome: | The proposed approach leads to 71.50 F1 on Korean QA (comparable to 77.3 F1) |
Copied to clipboard
| Challenge: | Existing approaches for neural machine translation use small amount of data or monolingual data. |
| Approach: | They describe acquisition, preprocessing and characteristics of a large English-French parallel corpus for the financial domain. |
| Outcome: | The proposed corpus contains 8.6 million high quality sentence pairs . the first release of the corpus is available on github. |
Copied to clipboard
| Challenge: | In natural language, we often omit some words that are easily understandable from the context. |
| Approach: | They propose to use a dataset to evaluate whether translation models can resolve zero pronoun problems in Japanese to English translations. |
| Outcome: | The proposed model can resolve the zero pronoun problem in Japanese to English translations. |
Copied to clipboard
| Challenge: | a societal movement towards using gender-fair language exists, but gender-free German is barely supported in machine translation. |
| Approach: | They propose to use a community-created gender-fair language dictionary to study gender-neutral German . they also use encyclopedic text and parliamentary speeches to translate the words in isolation . |
| Outcome: | The proposed study shows that most systems produce mainly masculine forms and rarely gender-neutral variants. |
Copied to clipboard
| Challenge: | In the context of under-resourced neural machine translation, transfer learning from an NMT model trained on a high resource language pair, or from a multilingual NMT (M-NMT) model, has been shown to boost performance to a large extent. |
| Approach: | They propose to use a multilingual NMT model to train on an under-resourced child and to use large sub-word vocabularies to improve performance. |
| Outcome: | The proposed approach involving dynamic vocabularies is both practical and effective on two under-resourced language pairs, i.e. Icelandic-English and Irish-English. |
Copied to clipboard
| Challenge: | a corpus of code-switched speech from soap operas is compiled from soaps . the corpus contains 14.3 hours of annotated and segmented speech . |
| Approach: | They propose a speech corpus containing multilingual code-switching from soap operas . the corpus contains English, isiZulu, isisXhosa, Setswana and Sesotho speech . |
| Outcome: | The corpus contains 14.3 hours of annotated and segmented speech from soap operas . the speech rate is 1.22 to 1.83 times higher than prompted speech in the same languages . |
Copied to clipboard
| Challenge: | Lexical ambiguity is one of the many challenging linguistic phenomena involved in translation, i.e., translating an ambiguous word with its correct sense. |
| Approach: | They propose to use training data to measure the sense distributions of a machine translation system to measure lexical ambiguity. |
| Outcome: | The proposed benchmark builds upon the multilingual sense inventory of BabelNet, the multilinguistic neural parsing pipeline TurkuNLP, and the OPUS collection of translated texts from the web. |
Copied to clipboard
| Challenge: | Existing clustering methods cannot handle asymmetric problem in multilingual NMT . existing models cannot handle the asymmetry problem since there are thousands of languages involved . |
| Approach: | They propose a fuzzy task clustering method to address the asymmetric problem in multilingual NMT by using task affinity as the clustering criterion. |
| Outcome: | The proposed method outperforms baselines for a multilingual model and the existing models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive language capabilities, but most of them have very unbalanced performance across different languages. |
| Approach: | They propose to use question translation data to enhance LLMs' multilingual capabilities by using mechanistic interpretability methods. |
| Outcome: | The proposed method improves multilingual alignment even with unannotated answers in English and a wide range of languages even with instruction-tuned LLMs. |
Copied to clipboard
| Challenge: | Text Image Machine Translation (TIMT) is a critical subfield of machine translation . it requires accurate optical character recognition, robust visual-text reasoning, and high-quality translation a challenge . |
| Approach: | They propose a multi-task optimization framework to specialize MLLMs into expert TIMT models. |
| Outcome: | The proposed model outperforms baselines on the latest in-domain MIT-10M benchmark. |
Copied to clipboard
| Challenge: | Existing detectors for translating texts fail to detect a text from a strange translator . Existing methods for detection of translated texts use text structure and complex words to detect translations . |
| Approach: | They propose a detector using text similarity with round-trip translation (TSRT) TSRT achieves 86.9% accuracy in detecting a translated text from a strange translator . Existing detectors have been built around a specific translator but fail to detect a translation from skeptics . |
| Outcome: | Existing detectors fail to detect translated texts from a strange translator . a detector achieves 86.9% accuracy in detecting a translated text from skeptic translators . |
Copied to clipboard
| Challenge: | Comparative evaluation of casing methods for Neural Machine Translation . evaluators evaluated methods for tokenisation and word segmentation into subword units . |
| Approach: | They evaluate three main casing methods for Neural Machine Translation to determine optimal handling of capitalisation. |
| Outcome: | The proposed methods are used to handle capitalisation on English-German and English-Turkish datasets. |
Copied to clipboard
| Challenge: | a large corpus covering 22 Turkic languages is included in this paper . low-resource MT evaluation has traditionally focused on European languages due to limitations of available technology and resources. |
| Approach: | They present a case study of the practical application of MT in the Turkic language family . they propose to realize the gains of NMT for Turkic languages under high-resource to extremely low-resourced scenarios. |
| Outcome: | The proposed study shows that the new methods can be used in the Turkic language family . the results highlight bottlenecks in building competitive systems . |
Copied to clipboard
| Challenge: | Existing methods to compress Transformer are limited to sub-components, e.g., selfattention networks or embedding layer. |
| Approach: | They propose a Hybrid Tensor-Train decomposition which retains full rank and meanwhile reduces operations and parameters. |
| Outcome: | The proposed model outperforms light-weight SOTA methods on three translation tasks and achieves 7.1 points absolute improvement in BLEU and 1.27 X speedup on IWSLT’14 De-En task. |
Copied to clipboard
| Challenge: | Streaming MT is an extension of simultaneous MT to the incremental translation of a continuous input text stream. |
| Approach: | They propose to extend simultaneous machine translation to streaming setups by leveraging streaming history. |
| Outcome: | The proposed system compares favorably to the best performing systems on IWSLT Translation Tasks. |
Copied to clipboard
| Challenge: | Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. |
| Approach: | They exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. |
| Outcome: | The proposed method can label documents at 94.5% across languages with high precision . the proposed method is useful for low-resource languages with limited resources . |
Copied to clipboard
| Challenge: | a novel method for clustering news across languages is proposed . a key challenge in handling news streams is that they must be generated on the fly . |
| Approach: | They propose a method for clustering news across languages into monolingual and crosslingual clusters . they use real news datasets in multiple languages to find an ever growing number of cluster labels . |
| Outcome: | The proposed method produces state-of-the-art results on real news datasets in German, English and Spanish. |
Copied to clipboard
| Challenge: | Existing studies show that a sentence has less ambiguity than a single word . if the word semantics is changed in translation, then a better translation is possible. |
| Approach: | They propose a linear cross-lingual mapping to improve multilingual embeddings . they also consider deviation from orthogonality conditions as a measure of deficiency . |
| Outcome: | The proposed method improves the multilingual embeddings by allowing for a linear cross-lingual mapping. |
Copied to clipboard
| Challenge: | Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet lack a robust methodology to dissect these phenomena comprehensively. |
| Approach: | They propose a multilingual dataset centered on food-related cultural facts and variations in food practices. |
| Outcome: | The proposed model incorporates cultural context significantly and improves its ability to access cultural knowledge. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have the potential to generate harmful content, posing risks to users. |
| Approach: | They propose a dataset specifically designed for safety evaluation in Kazakh and Russian . they use a bilingual context in Kazakhstan where both Kazakh (a low-resource language) and Russian (a high-resourced language) |
| Outcome: | The proposed dataset is designed for safety evaluation in Kazakh and Russian . it shows that both multilingual and language-specific LLMs perform better than others . |
Copied to clipboard
| Challenge: | Using monolingual tools, code-switching is a problem in the natural language processing community. |
| Approach: | They introduce the Canberra Vietnamese-English Code-switching corpus (CanVEC) which is an original corpus of mixed speech annotated with language information, part of speech tags and Vietnamese translations. |
| Outcome: | The proposed corpus was annotated with language information, part of speech tags and Vietnamese translations using pipelining and monolingual toolkits. |
Copied to clipboard
| Challenge: | Existing work on adding syntactic information to NMT systems is limited to linguistically-inspired tree structures. |
| Approach: | They propose an NMT model that can naturally generate the topology of an arbitrary tree structure on the target side. |
| Outcome: | The proposed model outperforms standard seq2seq models by 2.1 BLEU points and other methods for incorporating target-side syntax by 0.7 BLUE points. |
Copied to clipboard
| Challenge: | Swiss-AL is a multilingual web corpus for Applied Linguistics that supports data-based and data-driven research on societal and political discourses in Switzerland. |
| Approach: | They propose a multilingual Swiss web corpus for Applied Linguistics that supports data-based research on societal and political discourses in Switzerland. |
| Outcome: | The Swiss Web Corpus for Applied Linguistics (SWS) is a multilingual collection of texts from selected web sources. |
Copied to clipboard
| Challenge: | International NPO/NGOs are struggling with the design and development of tools and systems for multi-language communication in the real world. |
| Approach: | They propose a framework for service design with the Language Grid by bridging the gap between language service infrastructures and multi-language systems. |
| Outcome: | The proposed framework bridges the gap between language service infrastructures and multi-language systems by allowing users to design and develop multilingual communication services and tools in the real world. |
Copied to clipboard
| Challenge: | Using captioned images, we can quantify language function and semantics using a grounded typology approach . linguistic typology is the study of patterns and variation across the world's languages . |
| Approach: | They propose a grounded typology approach that uses images captioned across languages to quantify meaning and semantics. |
| Outcome: | The proposed approach can quantify language function and semantics using images captioned across languages. |
Copied to clipboard
| Challenge: | WikiPron is an open-source command-line tool for extracting pronunciation data from Wiktionary . the tool generates a database of 1.7 million pronunciations from 165 languages . |
| Approach: | They propose a command-line tool for extracting pronunciation data from Wiktionary . they use it to generate a database of 1.7 million pronunciations from 165 languages . |
| Outcome: | The proposed software generates a database of pronunciations for 165 languages . the proposed model is then validated by a grapheme-to-phoneme model . |
Copied to clipboard
| Challenge: | Non-literal translations are difficult to produce even for human translators, especially for foreign language learners, and machine translations have not yet been developed to simulate human translations. |
| Approach: | They propose to fine-tune generic sentence representations produced by a pre-trained cross-lingual language model to detect non-literal translations. |
| Outcome: | The proposed model can predict human translations and distinguish literal and non-literal translations at phrase level with a moderate positive correlation. |
Copied to clipboard
| Challenge: | Homophone normalization is a pre-processing step used in Amharic natural language processing (NLP) but it also results in models that are unable to process different forms of writing in a single language. |
| Approach: | They propose a method where normalization is applied to model predictions instead of training data and a scheme where normalized data is preserved in training. |
| Outcome: | The proposed model achieves an increase in BLEU score of up to 1.03 while preserving language features in training. |
Copied to clipboard
| Challenge: | Existing datasets for machine translation quality estimation and post-editing have several shortcomings. |
| Approach: | They propose a dataset for machine translation quality estimation and automatic post-editing . they report the performance of baseline systems trained on the MLQE-PE dataset . |
| Outcome: | The proposed dataset contains human labels for up to 10,000 translations per language pair. |
Copied to clipboard
| Challenge: | Recent advances in neural language modeling and multilingual training have prompted widespread adoption of machine translation (MT) technologies across an unprecedented range of world languages. |
| Approach: | They propose to use a dataset to assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity. |
| Outcome: | The proposed model is faster than translation from scratch, but the magnitude of productivity gains varies widely across systems and languages. |
Copied to clipboard
| Challenge: | a corpus of urban varieties of Portuguese is being studied in Angola, Mozambique and So Tomé and Prncipe . the corpora are transcribed spoken data, complemented by metadata describing the setting of the audio recordings and sociolinguistic information about the speakers. |
| Approach: | They present three new corpora of urban varieties of Portuguese spoken in Angola, Mozambique and So Tomé and Prncipe . they provide new, contemporary data for the study of each variety and for comparative research on African, Brazilian and European varieties . |
| Outcome: | The corpora are transcribed spoken data and annotated with POS and lemma information . they are already being used for comparative research on possession and location . |
Copied to clipboard
| Challenge: | Cross-lingual open-ended generation is an important yet understudied problem. |
| Approach: | They propose XL-Instruct, a novel technique for generating high-quality synthetic data, and introduce Xl-AlpacaEval, evaluating cross-lingual generation capabilities of large language models. |
| Outcome: | The proposed technique improves model performance by fine tuning with just 8K instructions generated using XL-Instruct, and also by improving on several fine-grained quality metrics. |
Copied to clipboard
| Challenge: | MULTI-EURLEX is a dataset for topic classification of EU legal documents . fine-tuning a multilingually pretrained model in a single source language leads to catastrophic forgetting of multilingual knowledge and poor zero-shot transfer to other languages. |
| Approach: | They propose to use the dataset as a testbed for zero-shot cross-lingual transfer to exploit annotated training documents in one language to classify documents in another language. |
| Outcome: | The proposed model can be used to classify EU legal documents in other languages without a single source language and retain multilingual knowledge. |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations in machine translation are limited for low-resource languages. |
| Approach: | They evaluate sentence-level hallucination detection approaches using Large Language Models (LLMs) they find that the choice of model is essential for performance. |
| Outcome: | The proposed models outperform the existing models in HRLs and LRLs on average by 0.16 MCC. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown they can match or surpass finetuned models on many natural language processing tasks. |
| Approach: | They propose to use in-context learning and pivot translation to improve code-switching translation. |
| Outcome: | The proposed models show strong ability for cross-lingual understanding in a code-switching setting. |
Copied to clipboard
| Challenge: | Developing NLP methods for historical corpora is difficult, as only domain experts can label them . off-the-shelf models are trained on modern language texts, rendering them weaker for historical documents . |
| Approach: | They propose to use an annotated newspaper dataset to extract historical data from a novel domain of texts. |
| Outcome: | The proposed method performs well on a multilingual dataset in English, French, and Dutch . it is possible to extract surprisingly good results even with scarce annotated data using existing models and datasets for modern languages . |
Copied to clipboard
| Challenge: | Hate speech detection models are evaluated on a held-out test data, but they are incapable of identifying weaknesses. |
| Approach: | They propose to use multilingual hate speech detection models to evaluate their performance on social media conversation. |
| Outcome: | The proposed model can detect hate speech in multiple languages using a real-world conversation on social media. |
Copied to clipboard
| Challenge: | Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages. |
| Approach: | They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages. |
| Outcome: | The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures. |
Copied to clipboard
| Challenge: | Large text corpora are increasingly important for a wide variety of NLP tasks. |
| Approach: | They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages. |
| Outcome: | The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation. |
Copied to clipboard
| Challenge: | MultiUAT dynamically adjusts training data usage based on model’s uncertainty on a small set of trusted clean data for multi-corpus machine translation. |
| Approach: | They propose an approach that dynamically adjusts the training data usage based on the model’s uncertainty on a small set of trusted clean data for multi-corpus machine translation. |
| Outcome: | The proposed approach outperforms baselines on 16 languages and 2 domains on English-German translation. |
Copied to clipboard
| Challenge: | European "Tenders Electronic Daily" is a valuable source of semi-structured and multilingual data . collecting and managing such kind of data is incredibly burdensome and takes time and resources . |
| Approach: | They describe two documented and easy-to-use multilingual corpora extracted from the TED web site . they propose to make the extracted dataset available to the scientific community . |
| Outcome: | The proposed dataset is based on the European tenders electronic daily (TED) web site . it is easy to use and can be used for text mining and natural language processing tasks. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for assessing distinct meanings of words are tied to sense inventories, restricting their usage to knowledge-based representation techniques. |
| Approach: | They propose a multilingual benchmark that models distinct meanings of words in English . they use a binary disambiguation task with gold standards in 12 new languages . |
| Outcome: | The proposed model can model distinct meanings of words in English even when no tagged instances are available for a target language. |
Copied to clipboard
| Challenge: | LLMs are a popular evaluation strategy, but their reliability in multilingual evaluation remains uncertain. |
| Approach: | They evaluate five models from different model families across five diverse tasks involving 25 languages. |
| Outcome: | The models perform poorly across languages and average Fleiss’ Kappa is 0.3 . |
Copied to clipboard
| Challenge: | Zero copulas are the phenomenon that nominal predicates lack an explicit verbal copule in default present tense 3rd person indicative cases. |
| Approach: | They propose a tool that can identify and mark the location of zero copulas in Hungarian clauses that contain nominal predicates at the right position. |
| Outcome: | The proposed tool can identify and mark the location of zero copulas, i.e. where an overt copulan would appear in the non-default cases. |
Copied to clipboard
| Challenge: | Existing methods for truthfulness enhancement in English are limited to multilingual scenarios. |
| Approach: | They propose a method for cross-lingual truthfulness transfer that uses language bias and transfer contributions to select an optimal subset of all tested languages and employ translation instruction tuning for cross language truthfulness transfers. |
| Outcome: | The proposed method reduces multilingual representation disparity and boosts cross-lingual truthfulness transfer of LLMs. |
Copied to clipboard
| Challenge: | Terminology standardization plays an important role in the management of terminological resources. |
| Approach: | They propose to re-model an existing multilingual terminological database for the medical domain, TriMED, and propose a method to make it compliant to the latest ISO/TC 37 standards. |
| Outcome: | The proposed model should be compliant with the three most recent ISO/TC 37 standards and has a new data category repository and a Web application that can be used to access the multilingual terminological records. |
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
Copied to clipboard
| Challenge: | MULTITuDE benchmarks lack authentic and machine-generated text in languages other than English . defining characteristic of new generation of LLMs is increased quality of text . |
| Approach: | They propose a benchmarking dataset for multilingual machine-generated text detection that compares detectors with authentic and machine-generated texts in 11 languages. |
| Outcome: | The proposed dataset compares detectors with zero-shot and fine-tuned detectors in 11 languages. |
Copied to clipboard
| Challenge: | Effective projection-based cross-lingual word embedding induction relies on the iterative self-learning procedure. |
| Approach: | They propose a classification-based approach to self-learning that allows for integration of diverse features into the iterative process. |
| Outcome: | The proposed method improves bilingual lexicon induction on a weakly supervised setup with 28 language pairs. |
Copied to clipboard
| Challenge: | Existing approaches to multilingual entity linking are cross-lingual, with a focus on zero-shot evaluation. |
| Approach: | They propose a new formulation for multilingual entity linking where language-specific mentions resolve to a language-agnostic Knowledge Base. |
| Outcome: | The proposed model outperforms state-of-the-art models on a large multilingual dataset and shows that frequency-based analysis provided key insights for the model and training enhancements. |
Copied to clipboard
| Challenge: | a new study examines the performance of code-switching IR in monolingual contexts . code-witching is a pervasive linguistic phenomenon in global communication . |
| Approach: | They propose a benchmark to evaluate code-switching IR in monolingual contexts . they propose CS-MTEB, which measures performance declines of up to 27% . |
| Outcome: | The proposed benchmark shows that code-switching performance is degraded by 27% . the proposed benchmark is based on a dataset of mixed-language queries . |
Copied to clipboard
| Challenge: | Existing approaches to domain-specific neural machine translation (NMT) are lexically constrained and draw from domain- specific dictionaries. |
| Approach: | They propose a lexically constrained neural machine translation system that disambiguates between multiple dictionary candidates. |
| Outcome: | The proposed system disambiguates between multiple candidate translations derived from dictionaries on English-Hindi, English-German, and English-French datasets. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) are inherently multilingual agents . concerns regarding their safety have emerged . |
| Approach: | They propose a framework to synthesize red-teaming queries and investigate their safety . they demonstrate that the framework outperforms existing red- teaming techniques . |
| Outcome: | The proposed framework outperforms existing red-teaming techniques in the safety domain . it generates code-switching attack prompts in monolingual data . |
Copied to clipboard
| Challenge: | Existing methods for developing broad-coverage semantic dependency parsers for languages without semantically annotated data are limited to English, Czech and Chinese. |
| Approach: | They propose a multitask learning framework coupled with annotation projection to build broad-coverage semantic dependency parsers for languages without annotated resources. |
| Outcome: | The proposed model improves labeled F1 score on multitask tasks from English to Czech compared to baseline models . |
Copied to clipboard
| Challenge: | Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries. |
| Approach: | They propose a large and highly multilingual dataset for sign language translation: JWSign. |
| Outcome: | The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. |
Copied to clipboard
| Challenge: | Existing multilingual models such as XLM-R support only approximately 100-200 languages, leaving nearly 7,000 low-resource languages untapped. |
| Approach: | They construct and open-source a dataset of four-language corpora obtained through machine translation into Chinese, Uyghur and Tibetan. |
| Outcome: | The proposed dataset includes two resource-rich languages and two low-resource languages. |
Copied to clipboard
| Challenge: | Documents as short as a single sentence may reveal sensitive information about authors . style transfer is effective but a number of current methods cause a drop in down-stream utility . |
| Approach: | They propose a method to remove sensitive information from documents by multilingual back-translation using off-the-shelf translation models. |
| Outcome: | The proposed method lowers adversarial gender and race prediction by 22% while retaining 95% of original utility on downstream tasks. |
Copied to clipboard
| Challenge: | Existing data selection methods do not work well for multiple domains . multiple aspects need to be considered for training a multi-domain model . |
| Approach: | They propose a dynamic data selection method to multi-domain NMT that incorporates instance-level domain-relevance features and a curriculum to gradually focus on multi- domain relevant data batches. |
| Outcome: | The proposed model outperforms no-curriculum training on multiple domains and reaches or outperformed individual performance. |
Copied to clipboard
| Challenge: | Evaluating cross-lingual knowledge transfer in large language models is challenging, as correct answers in a target language may arise either from genuine transfer or from prior exposure during pre-training. |
| Approach: | They propose a pipeline to isolate and measure cross-lingual knowledge transfer by identifying self-contained, time-sensitive knowledge entities from real-world domains and generating factual questions. |
| Outcome: | The proposed pipeline analyzes multiple LLMs across five languages and shows that cross-lingual transfer is strongly influenced by linguistic distance and often asymmetric across language directions. |
Copied to clipboard
| Challenge: | Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets. |
| Approach: | They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies. |
| Outcome: | The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models are often prone to parameter interference . a common problem is that the model compromises with the language diversity to find a solution . |
| Approach: | They propose a method that allocates parameters based on consistency between the gradients of the individual language and the average gradient. |
| Outcome: | The proposed method reduces parameter interference and improves translation quality. |
Copied to clipboard
| Challenge: | Existing grammar error correction systems have been trained on monolingual data and not developed for CSW text. |
| Approach: | They propose a method of generating synthetic CSW GEC datasets by translating different spans of text within existing GEC corpora and investigate different methods of selecting these spans based on CSW ratio, switch-point factor and linguistic constraints. |
| Outcome: | The proposed model achieves an average increase of 1.57 F0.5 across 3 CSW test sets (English-Chinese, English-Korean and English-Japanese) without affecting the model’s performance on a monolingual dataset. |
Copied to clipboard
| Challenge: | Prior studies focused on English posts to provide early warnings for epidemic prediction, but these work focused on non-English posts. |
| Approach: | They propose a multilingual event extraction framework for extracting epidemic event information for any disease and language using 5.1K tweets in four languages. |
| Outcome: | The proposed framework can provide epidemic warnings for COVID-19 in its earliest stages in Dec 2019 (3 weeks before global discussions) and aggregate community epidemic discussions like symptoms and cure measures, aiding misinformation detection and public attention monitoring. |
Copied to clipboard
| Challenge: | Existing multilingual vision-language pretrained models are biased towards English due to the lack of sufficient non-English image-text pairs. |
| Approach: | They propose to train a retrieval-efficient dual-stream multilingual VLP model by aligning CLIP model and a multilingual text encoder through a novel Triangle Cross-modal Knowledge Distillation method. |
| Outcome: | Empirical results show that mCLIP achieves new state-of-the-art performance for both zero-shot and finetuned multilingual image-text retrieval tasks. |
Copied to clipboard
| Challenge: | Existing methods for enhancing sign language text data are insufficient . fewer studies have been performed on text data augmentation compared to video data . |
| Approach: | They propose three methods to augment sign language text data using Korean sign language gloss dictionary. |
| Outcome: | The proposed method improves translation performance by 0.204 and 0.170 compared to the original data. |
Copied to clipboard
| Challenge: | Multilingual models aim for language-invariant representations but still encode language identity. |
| Approach: | They propose a multi-task learning framework that induces language invariance in multilingual retrieval by reducing language-specific signals in the embedding space. |
| Outcome: | The proposed learning framework improves language-invariant dense retrieval over baselines on English retrieval data and general multilingual corpora. |
Copied to clipboard
| Challenge: | Translation difficulty is a problem when translators are required to resolve translation ambiguity from multiple possible translations. |
| Approach: | They use word alignments computed over large scale bilingual corpora to develop predictors of lexical translation difficulty. |
| Outcome: | The proposed method improves on a previous embedding-based approach and can contribute to a deeper understanding of cross-lingual differences and of causes of translation difficulty. |
Copied to clipboard
| Challenge: | Multilingual large language models (LLMs) exhibit factual inconsistencies across languages . authors identify two primary sources of error: insufficient engagement of reliable English-centric mechanism for factual recall, and incorrect translation from English back into the target language for the final answer. |
| Approach: | They propose two vector interventions to redirect the model toward better internal paths for higher factual consistency. |
| Outcome: | The proposed interventions increase the recall accuracy by over 35 percent for the lowest-performing language. |
Copied to clipboard
| Challenge: | a large amount of troll accounts have emerged with efforts to manipulate public opinion on social network sites . a recent study found that trolled tweets spread misinformation, fake news, and propaganda . we use supervised classification to detect trol tweets in both English and Russian . |
| Approach: | They propose to detect troll tweets in English and Russian using machine learning algorithms . they use monolingual, cross-lingual, and bilingual training scenarios . |
| Outcome: | The proposed method uses monolingual, cross-lingual, and bilingual training scenarios. |
Copied to clipboard
| Challenge: | a lexical resource associates words with concepts in multiple languages, which makes it difficult to combine information from multiple resources. |
| Approach: | They propose a translation-based approach to mapping lexical resources . they use word-concept pairs to align WordNet/BabelNet to CLICS and OmegaWiki . |
| Outcome: | The proposed method achieves state-of-the-art accuracy without other sources of knowledge . it can be framed as word sense disambiguation, and it can improve on existing methods . |
Copied to clipboard
| Challenge: | Existing high-quality Vietnamese-English parallel datasets are inadequate for translation training. |
| Approach: | They introduce a high-quality Vietnamese-English parallel dataset for medical translation . they compare Google Translate, ChatGPT, and pre-trained bilingual/multilingual models . |
| Outcome: | The proposed dataset is compared with translation models from Google Translate and ChatGPT. |
Copied to clipboard
| Challenge: | Detoxifying multilingual Large Language Models (LLMs) has become crucial due to their increasing global use. |
| Approach: | They propose to use English preference tuning to study cross-lingual detoxification of LLMs. |
| Outcome: | The proposed method reduces toxicity in multilingual LLMs by reducing the probability of mGPT-1.3B generating toxic continuations across 17 languages. |
Copied to clipboard
| Challenge: | a systematic review of 300 publications reveals a language gap in LLM safety research . even high-resource non-English languages receive little attention, authors note . |
| Approach: | They propose to focus on safety evaluation, training data generation, and crosslingual safety generalization based on their findings. |
| Outcome: | The authors suggest that the field can develop more robust, inclusive safety practices for diverse global populations. |
Copied to clipboard
| Challenge: | KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage-level relevance judgements. |
| Approach: | They introduce a Kazakh open-domain question answering dataset that can be used in reading comprehension and full ODQA settings. |
| Outcome: | The proposed dataset can be used in reading comprehension and full ODQA settings, as well as for information retrieval experiments. |
Copied to clipboard
| Challenge: | Despite recent advances in Reasoning Language Models, most research focuses solely on English, even though many models are pretrained on multilingual data. |
| Approach: | They evaluate three open-source RLMs: DeepSeek R1, Qwen 2.5, and Qwend 3 across four math datasets and seven typologically diverse languages. |
| Outcome: | The proposed model reduces token usage and preserves accuracy even after translation into English. |
Copied to clipboard
| Challenge: | a dataset of 10122 hours in 87394 recordings is presented in a new journal . 43% of recordings have human-written subtitles, covering a total of 137 languages. |
| Approach: | They present a Khan Academy corpus with 10122 hours in 87394 recordings . 43% of recordings have human-written subtitles, and 137 languages are included . |
| Outcome: | The dataset can be used to train multilingual speech recognition and translation models. |
Copied to clipboard
| Challenge: | Existing studies show that the ability of large language models to generate contextual understanding of the sentence can degrade translation quality. |
| Approach: | They propose a method that generates contextual understanding for both source and target languages separately. |
| Outcome: | The proposed method outperforms strong comparison methods in multiple domains. |
Copied to clipboard
| Challenge: | linguistic affiliation of languages to a common language family is traditionally carried out manually . large-scale standardized collections of multilingual wordlists and grammatical language structures could improve this . |
| Approach: | They propose to use lexical and grammatical data to classify languages into families using neural network models. |
| Outcome: | The proposed models outperform models trained on lexical and grammatical data while combining both types of data yields even better performance. |
Copied to clipboard
| Challenge: | Currently, existing systems cannot accurately identify most of the world's 7000 languages due to lack of data and computational challenges. |
| Approach: | They propose a misprediction-resolution hierarchical model, LIMIT, that reduces error by 55% on a children's stories dataset and by 40% on 'fLORES-200' benchmark. |
| Outcome: | The proposed model reduces error by 55% on the MCS-350 and 40% on the FLORES-200 benchmarks. |
Copied to clipboard
| Challenge: | XC-Translate is a large-scale, manually-created benchmark for machine translation . current systems struggle to translate texts containing entity names, but KG-MT outperforms state-of-the-art approaches . |
| Approach: | They propose a method to integrate multilingual knowledge into a neural machine translation model . XC-Translate is the first large-scale, manually-created benchmark for machine translation . they propose KG-MT to integrate cultural-related references into MT models . |
| Outcome: | The proposed method outperforms state-of-the-art approaches by a large margin compared to NLLB-200 and GPT-4 . the proposed method is based on a multilingual knowledge graph and dense retrieval mechanism . |
Copied to clipboard
| Challenge: | Existing work on discourse understanding is constrained by framework-dependent discourse representations. |
| Approach: | They examine whether large language models capture discourse knowledge that generalizes across languages and frameworks. |
| Outcome: | The proposed model can generalize discourse information across languages and frameworks. |
Copied to clipboard
| Challenge: | Existing systems that use a left-to-right completion paradigm are inefficient and expensive. |
| Approach: | They propose an open-source end-to-end interactive machine translation system platform . they propose to use a prefix-constrained decoding approach to achieve end- to-end evaluation . |
| Outcome: | The proposed system can guarantee high-quality, error-free translations . it uses prefix-constrained decoding and improves on previous systems . |
Copied to clipboard
| Challenge: | Synthetic data generation relies on a single oracle teacher model, which can lead to model collapse and bias propagation. |
| Approach: | They propose a multilingual arbitration approach that exploits performance variations among multiple models for each language. |
| Outcome: | The proposed approach surpasses single-teacher distillation with 80% win rates over proprietary and open-weight models with the largest improvements in low-resource languages. |
Copied to clipboard
| Challenge: | Existing approaches to sign language translation use gloss annotations as an intermediary . a new approach to use large language models and word embeddings to improve Gloss2Text translation is needed. |
| Approach: | They propose to leverage large language models pre-trained on expansive and diverse corpora to improve Gloss2Text translation stage by using data augmentation and label-smoothing loss function. |
| Outcome: | The proposed approach surpasses state-of-the-art methods on the PHOENIX Weather 2014T dataset . it shows that gloss annotations can be used to guide the translation process . |
Copied to clipboard
| Challenge: | Existing approaches to multilingual neural machine translation (MNMT) are limited in their ability to handle large amounts of data. |
| Approach: | They propose a framework which only requires target-side monolingual data and a bilingual dictionary to improve the performance of the MNMT model. |
| Outcome: | The proposed framework is more effective than baselines in long-tail and high-resource languages. |
Copied to clipboard
| Challenge: | Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks . |
| Approach: | They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. |
| Outcome: | The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English. |
Copied to clipboard
| Challenge: | Existing code translation benchmarks focus on individual functions, overlooking repository-level challenges like intermodule coherence and dependency management. |
| Approach: | They propose a framework for benchmarking Java-to-C# translation at the repository level . it uses a translation framework guided by skeletons and fine-grained quality evaluation . |
| Outcome: | The proposed framework improves Java-to-C# translation quality at the repository level. |
Copied to clipboard
| Challenge: | Existing models for multi-domain translation tasks only use monolingual data, whereas bilingual data is indispensable for improving the models. |
| Approach: | They propose a modular strategy that facilitates the cooperation of monolingual and bilingual knowledge in translation tasks by avoiding catastrophic forgetting. |
| Outcome: | The proposed model exhibits superior generalization and robustness over the conventional approach. |
Copied to clipboard
| Challenge: | Text sanitization is the task of detecting and removing personal information from the text. |
| Approach: | They propose a dataset for multilingual named entities that can be used for text sanitization. |
| Outcome: | The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language. |
Copied to clipboard
| Challenge: | Existing measures of code-switching (CS) complexity are word-based, meaning any word is equally likely to switch between any two words. |
| Approach: | They adapt two NLP metrics, multilinguality and CS probability, and put forward Intonation Units (IUs) as basic tokens for transcribed bilingual speech. |
| Outcome: | The proposed measures account for prosodic and prosodic constraints on CS in bilingual speech. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on English or use translated data, which fails to capture cultural nuances. |
| Approach: | They propose to use a multilingual end-to-end Meta-Evaluation RAG benchmark MEMERAG to assess accuracy and faithfulness of RAG systems. |
| Outcome: | The proposed benchmark can identify improvements offered by advanced prompting techniques and LLMs. |
Copied to clipboard
| Challenge: | ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages . |
| Approach: | They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation . |
| Outcome: | The proposed resource is the largest of its type to date and performs well on 10 languages. |
Copied to clipboard
| Challenge: | Maltese is a Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. |
| Approach: | They investigate whether Arabic-language resources can support Maltese natural language processing . they introduce transliteration schemes and machine translation approaches to align Arabic text with Maltesen . |
| Outcome: | The proposed techniques can significantly improve Maltese natural language processing tasks. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) demonstrate multilingual abilities, yet they are English-centric due to dominance of English in training corpora. |
| Approach: | They propose to use a synthetic English-korean CS question-answering dataset to investigate this potential. |
| Outcome: | The proposed model can activate, identify and leverage knowledge for reasoning in low-resource languages. |
Copied to clipboard
| Challenge: | The Schema Learning Corpus is a linguistic resource designed to support research into the structure of complex events in multilingual data. |
| Approach: | The Schema Learning Corpus is a linguistic resource that includes large volumes of background data in English, Spanish and Russian. |
| Outcome: | The SLC defines 100 complex events (CEs) across 12 domains and multiple documents labeled for each . multiple documents contain evidence for each step, plus labeles events and relations along with their arguments across a large tag set. |
Copied to clipboard
| Challenge: | Existing benchmarking datasets for Bangla LLMs are not available for all languages. |
| Approach: | They present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. |
| Outcome: | The proposed model outperforms existing models in Bangla, but not always in the first place. |
Copied to clipboard
| Challenge: | a number of languages are used in online conversations, resulting in code-mixing . the problem is largely unexplored due to the lack of annotated data and noise . |
| Approach: | They propose a robust perturbation-based joint-training model that learns to handle noise in code-mixed text by parameter sharing across clean and noisy words. |
| Outcome: | The proposed model learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words. |
Copied to clipboard
| Challenge: | a recent study addresses the challenge of adapting loanwords during the translation process in low-resource languages. |
| Approach: | They propose a method that augments source sentences with loanword constraints . they then integrate loanwords as external linguistic knowledge into machine translation systems . |
| Outcome: | The proposed approach improves translation quality and handling loanword adaptation correctly in target languages. |
Copied to clipboard
| Challenge: | a new dataset focuses on gender-neutral terms that necessitate gendered translations in Catalan. |
| Approach: | They propose to use a new dataset to evaluate gender bias in machine translation . they train four MT systems using different tokenization techniques . |
| Outcome: | The proposed dataset focuses on gender-neutral terms necessitating gendered translations in Catalan. |
Copied to clipboard
| Challenge: | Recent reasoning language models (RLMs) achieve strong performance on complex reasoning tasks, yet they still exhibit a multilingual reasoning gap. |
| Approach: | They propose a strategy that incorporates an English translation into the initial reasoning trace when an understanding failure is detected. |
| Outcome: | The proposed strategy incorporates an English translation into the initial reasoning trace when an understanding failure is detected. |
Copied to clipboard
| Challenge: | a growing number of researchers are examining whether large language models can learn to translate a "new" language using grammar books. |
| Approach: | They examine an LLM's ability to learn new languages using grammar books . authors suggest alternative fine-tuning strategies to improve explicit learning . |
| Outcome: | The proposed model can learn low-resource languages described in grammar books but lacking extensive corpora. |
Copied to clipboard
| Challenge: | AfroCS-xs is a low-quality dataset for code-switching in multilingual communities . code-witching is prevalent in multicultural societies but lacks high-quality data for model development . |
| Approach: | They propose to use human-validated synthetic code-switched datasets to generate code-witched sentences for four African languages and English within a specific domain—agriculture. |
| Outcome: | The proposed model improves translation accuracy on the high-quality dataset for four African languages and English within a specific domain—agriculture. |
Copied to clipboard
| Challenge: | Existing models generate morpheme-level glosses but assign them to whole words without predicting the actual morphological boundaries, making them less interpretable and therefore untrustworthy to human annotators. |
| Approach: | They propose to use neural networks to predict interlinear glosses and morphological segmentation from raw text. |
| Outcome: | The proposed model outperforms GlossLM on glossing and beats open-source models on segmentation, glossing, and alignment. |
Copied to clipboard
| Challenge: | Existing public terminology datasets for MT research are limited in language coverage or domain specificity, making it difficult to assess or improve MT systems in specialized settings. |
| Approach: | They propose a multilingual terminology resource for tax and financial education covering seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole. |
| Outcome: | The proposed terminology resource covers seven typologically diverse languages: English, Spanish, Russian, Vietnamese, Korean, Chinese (traditional and simplified) and Haitian Creole. |
Copied to clipboard
| Challenge: | Existing methods to improve performance of large language models rely on additional training objectives or language-specific parameters. |
| Approach: | They propose a bidirectional language projection framework that enables efficient multilingual alignment and language shift using the intrinsic parameters. |
| Outcome: | The proposed framework improves performance of non-dominant languages and improves internal representations. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. |
| Approach: | They propose a multilingual chart question answering benchmark that enables efficient multilingual generation via data translation and code reuse. |
| Outcome: | The proposed benchmark systematically evaluates multilingual chart understanding on state-of-the-art LVLMs and shows a significant performance gap between English and other languages. |