Papers by Graham Neubig
Copied to clipboard
| Challenge: | Existing approaches to morphological tagging are limited by the assumption that tag sets overlap . a limited amount of data is available for most languages to learn these morphology taggers. |
| Approach: | They propose a method for cross-lingual morphological tagging that relaxes this assumption . they use factorial conditional random fields with neural network potentials to smooth over superficial differences in the surface forms . |
| Outcome: | The proposed model can smooth over superficial differences in the surface forms and generate unseen or rare tag sets. |
Copied to clipboard
| Challenge: | Recent studies in context-aware MT attempt to target a small set of discourse phenomena during evaluation, however not in a fully systematic way. |
| Approach: | They develop a multilingual discourse-aware benchmark to evaluate model performance on discourse phenomena in a given dataset. |
| Outcome: | The proposed model improves on previously studied phenomena while uncovering others which were not addressed. |
Copied to clipboard
| Challenge: | 'synthetic data' is a data generated with the assistance of large language models to make dataset construction faster and cheaper. |
| Approach: | This tutorial seeks to build a shared understanding of recent progress in synthetic data generation from NLP and related fields by grouping and describing major methods, applications, and open problems. |
| Outcome: | This tutorial will describe methods, applications, and open problems that have been developed and are being used to improve the quality and efficiency of synthetic data generation. |
Copied to clipboard
| Challenge: | Language model performance is largely dependent on pretraining decisions, but scaling laws based on only these two aspects do not always explain downstream task performance. |
| Approach: | They meta-analyze 92 open-source pretrained models to quantify their impact on performance. |
| Outcome: | The framework lays a foundation for more systematic investigation of how model development choices shape final capabilities. |
Copied to clipboard
| Challenge: | Natural language processing (NLP) is a vast field, with a wide variety of tasks, languages, and domains. |
| Approach: | They build regression models to predict evaluation score of an NLP experiment . they find that their models can produce meaningful predictions over unseen languages . |
| Outcome: | The proposed model outperforms baseline models and human experts on 9 different tasks. |
Copied to clipboard
| Challenge: | Using manual content to learn languages is expensive and time consuming. |
| Approach: | They propose a method for automatically identifying fine-grained lexical distinctions and extracting rules explaining them in a human- and machine-readable format. |
| Outcome: | The proposed method is able to identify fine-grained distinctions and explain them in a human- and machine-readable format. |
Copied to clipboard
| Challenge: | Noisy input text can cause disastrous mistranslations in most modern machine translation systems. |
| Approach: | They propose a benchmark dataset for Machine Translation of Noisy Text (MTNT) they use reddit comments and professionally sourced translations to examine noise types. |
| Outcome: | The proposed dataset can provide an attractive testbed for noise-robust machine translation systems. |
Copied to clipboard
| Challenge: | Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains. |
| Approach: | They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains. |
| Outcome: | The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation. |
Copied to clipboard
| Challenge: | a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations . |
| Approach: | They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search . |
| Outcome: | The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX. |
Copied to clipboard
| Challenge: | Unsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters. |
| Approach: | They propose a generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a structured generative prior. |
| Outcome: | The proposed model outperforms state-of-the-art models on part-of speech (POS) induction and unsupervised dependency parsing without gold POS annotation. |
Copied to clipboard
| Challenge: | Existing work on scientific information extraction (SciIE) considers extraction solely based on the content of an individual paper, without considering the paper’s place in the broader literature. |
| Approach: | They propose to automate the extraction of key information from scientific documents by leveraging a complementary source: the citation graph of referential links between citing and cited papers. |
| Outcome: | The proposed model improves on a set of English-language scientific documents. |
Copied to clipboard
| Challenge: | Existing approaches to neural machine translation (NMT) are dependent on limited parallel data, and can be difficult to use for many language pairs. |
| Approach: | They propose a method where target-language sentences are re-ordered to match the order of the source and used as an additional source of training-time supervision. |
| Outcome: | The proposed method improves on simulated low-resource Japanese-to-English and real low-demand Uyghur-to English scenarios. |
Copied to clipboard
| Challenge: | Existing methods for assessing the robustness of sequence-to-sequence models have been ignored by the literature. |
| Approach: | They propose an evaluation framework for adversarial attacks on seq2seq models that takes the semantic equivalence of the pre- and post-perturbation input into account. |
| Outcome: | The proposed framework breaks the assumption that source perturbations should not result in changes in the expected output, but allows for meaning-preserving perturbations that change the output sequence. |
Copied to clipboard
| Challenge: | Existing work on word alignment has focused on unsupervised learning on parallel text. |
| Approach: | They propose to combine pre-trained contextualized word embeddings with multilingually trained language models to achieve competitive results on word alignment tasks. |
| Outcome: | The proposed model outperforms state-of-the-art models on five language pairs and can train multilingual word aligners that can obtain robust performance on different language pairs. |
Copied to clipboard
| Challenge: | Existing approaches to generalization to resource-rich languages are difficult . a recent study shows that word representations can be useful in low resource languages . |
| Approach: | They propose two approaches for improving generalization to low-resource languages by adapting continuous word representations using linguistically motivated subword units. |
| Outcome: | The proposed method improves generalization to low resource languages . it requires neither parallel corpora nor bilingual dictionaries and requires no parallel training . |
Copied to clipboard
| Challenge: | mSimCSE can learn high-quality universal cross-lingual sentence embeddings without any parallel data. |
| Approach: | They propose a new language-based sentence embedding system that extends SimCSE to multilingual settings. |
| Outcome: | The proposed method improves existing methods on retrieval and multilingual STS tasks. |
Copied to clipboard
| Challenge: | Recent studies have shown that language models capture different types of knowledge regarding facts or commonsense knowledge. |
| Approach: | They examine how language models can be calibrated to make their confidence scores correlate better with the likelihood of correctness. |
| Outcome: | The proposed calibration methods improve confidence scores on QA tasks and improve accuracy. |
Copied to clipboard
| Challenge: | Existing work on adding syntactic information to NMT systems is limited to linguistically-inspired tree structures. |
| Approach: | They propose an NMT model that can naturally generate the topology of an arbitrary tree structure on the target side. |
| Outcome: | The proposed model outperforms standard seq2seq models by 2.1 BLEU points and other methods for incorporating target-side syntax by 0.7 BLUE points. |
Copied to clipboard
| Challenge: | Large Language Model (LLM) agents produce rich, multi-step trajectories that interleave observations, internal reasoning, and tool actions. |
| Approach: | They propose an open-source framework for diagnosing agent trajectories that quantifies five core agentic competencies and a visualization module that highlights trajectory semantics. |
| Outcome: | The proposed framework is extensible and compatible with most agent trajectories. |
Copied to clipboard
| Challenge: | Using leaderboards, researchers can track the performance of various systems on various NLP tasks. |
| Approach: | They propose a new conceptualization and implementation of NLP evaluation using a leaderboard. |
| Outcome: | The ExplainaBoard is an evaluation tool for natural language processing (NLP) it covers more than 400 systems, 50 datasets, 40 languages, and 12 tasks. |
Copied to clipboard
| Challenge: | a recent study examined how large language models handle interactions in meaning across words and larger syntactic forms. |
| Approach: | They propose to use a dataset to examine the linguistic properties of optionally transitive English verbs to examine their agentivity. |
| Outcome: | The proposed model outperforms all other models in the evaluation dataset . the results are better correlated with human judgements than syntactic and semantic corpus statistics . |
Copied to clipboard
| Challenge: | Experimental results on a newly-annotated version of the NAIST Simultaneous Translation Corpus indicate the promise of our proposed method. |
| Approach: | They propose a task of predicting which terminology simultaneous interpreters will leave untranslated using supervised sequence taggers. |
| Outcome: | The proposed method predicts which terminology interpreters leave untranslated . it is based on an annotated version of the NAIST Simultaneous Translation Corpus . |
Copied to clipboard
| Challenge: | Existing methods to generate program source code from natural language are not able to generate complex code due to a lack of ability to memorize large and complex structures. |
| Approach: | They propose a method that uses subtree retrieval to explicitly reference existing code examples within a neural code generation model. |
| Outcome: | The proposed method improves performance on two code generation tasks by up to +2.6 BLEU. |
Copied to clipboard
| Challenge: | Existing methods for evaluating image transcreation have relied on human evaluation. |
| Approach: | They propose a suite of automatic evaluation metrics inspired by machine translation metrics . they identify cultural relevance, semantic equivalence and visual similarity as critical dimensions of image transcreation . |
| Outcome: | The proposed evaluation metrics agree with human ratings across 7 countries. |
Copied to clipboard
| Challenge: | Prior studies have focused on developing effective data generation methods, but lack systematic comparison of different LMs as data generators in a unified setting. |
| Approach: | They propose to use a benchmark to compare language models' data generation abilities against a set of standardized settings and metrics. |
| Outcome: | The proposed benchmark provides standardized settings and metrics to evaluate LMs’ data generation abilities. |
Copied to clipboard
| Challenge: | Semantic parsing is the task of transducing natural language (NL) utterances into formal meaning representations (MRs), commonly represented as tree structures. |
| Approach: | They propose a variational auto-encoding model for semi-supervised semantic parsing which learns from limited amounts of parallel data and readily-available unlabeled NL utterances. |
| Outcome: | Experiments on ATIS domain and Python show that with extra unlabeled data, StructVAE outperforms strong supervised models. |
Copied to clipboard
| Challenge: | Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages. |
| Approach: | They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties. |
| Outcome: | The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines. |
Copied to clipboard
| Challenge: | Modern machine learning relies on datasets to develop and validate research ideas. |
| Approach: | They propose a dataset recommendation system that uses a training set and an evaluation set to help people find relevant datasets. |
| Outcome: | The proposed model finds more relevant search results than existing third-party search engines. |
Copied to clipboard
| Challenge: | Existing studies on building language agents have not addressed this social learning gap. |
| Approach: | They propose an interactive learning method that improves the social intelligence of language agents by using behavior cloning and self-reinforcement based training on filtered social interaction data. |
| Outcome: | The proposed method allows a 7B LLM to reach the social goal completion ability of an expert model (GPT-4-based agent) without the loss of more generic abilities, such as the ability to answer knowledge-based questions. |
Copied to clipboard
| Challenge: | Variational Autoencoders are powerful language models and effective representation learning frameworks. |
| Approach: | They propose a fix for posterior collapse which improves held-out likelihood, reconstruction and latent representation learning . |
| Outcome: | The proposed fix significantly improves held-out likelihood, reconstruction, and latent representation learning compared with previous state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing models generate morpheme-level glosses but assign them to whole words without predicting the actual morphological boundaries, making them less interpretable and therefore untrustworthy to human annotators. |
| Approach: | They propose to use neural networks to predict interlinear glosses and morphological segmentation from raw text. |
| Outcome: | The proposed model outperforms GlossLM on glossing and beats open-source models on segmentation, glossing, and alignment. |
Copied to clipboard
| Challenge: | Existing neural semantic parsers only focus on a small subset of tasks, such as SQL queries, robotic commands, and even general-purpose programming languages like Java. |
| Approach: | They propose a transition-based neural semantic parser that maps natural language utterances into formal meaning representations (MRs) they use an abstract syntax description language to constrain the output space and model the information flow. |
| Outcome: | Experiments on four different semantic parsing and code generation tasks show that the proposed system is generalizable, extensible, and effective. |
Copied to clipboard
| Challenge: | (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results. |
| Approach: | They propose to create a dataset for named entity recognition (NER) in ten African languages. |
| Outcome: | The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP. |
Copied to clipboard
| Challenge: | Language model (LM) evaluators that generate chain-of-thought reasoning are widely used for the assessment of LM responses. |
| Approach: | They investigate whether increasing LMs' "thinking" time through scaling test-time compute can improve an LM's evaluation capability. |
| Outcome: | The proposed reasoning models improve evaluation performance monotonically with the number of reasoning tokens generated, mirroring trends seen in LM reasoning. |
Copied to clipboard
| Challenge: | Using a large-scale chess commentary dataset, we generate a set of comments for individual moves in a game. |
| Approach: | They propose a large-scale chess commentary dataset and a method to generate commentary for individual moves in a chessian game. |
| Outcome: | The proposed method is rated similar to ground truth commentary texts in terms of correctness and fluency. |
Copied to clipboard
| Challenge: | Language models (LMs) capture factual knowledge by filling in the blanks of cloze-style prompts. |
| Approach: | They propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge and verify its effectiveness on several benchmark languages. |
| Outcome: | The proposed method improves the ability of multilingual LMs to access knowledge and verify its effectiveness on several benchmark languages. |
Copied to clipboard
| Challenge: | Pre-trained cross-lingual encoders have proven impressively effective at enabling transfer-learning of NLP systems from high-resource languages to low-resourced languages. |
| Approach: | They propose a method to align multilingual encoders using two explicit alignment objectives that align the multilingual representations at different granularities. |
| Outcome: | The proposed method achieves gains of up to 1.1 average F1 score on sequence tagging and 27.3 average accuracy on retrieval over the XLM-R-large model. |
Copied to clipboard
| Challenge: | Existing methods to train multilingual machine translation models are imbalanced and heterogeneous data is wildly varying. |
| Approach: | They propose a method that automatically learns how to weight training data through a data scorer that is optimized to maximize performance on all test languages. |
| Outcome: | The proposed method outperforms baselines on two sets of languages under one-to-many and many-to-1 MT settings and offers flexible control over which languages are optimized. |
Copied to clipboard
| Challenge: | Interlinear Glossed Text (IGT) is a form of linguistic annotation that can support documentation and resource creation for endangered languages. |
| Approach: | They propose a task in which these four annotation components are extracted automatically from speech and introduce a dataset to lay the groundwork for future research on IGT generation from speech. |
| Outcome: | The proposed dataset provides the first dataset to lay the groundwork for future research on IGT generation from speech, including end-to-end versus cascaded, monolingual versus multilingual, and single-task versus multiple-task approaches. |
Copied to clipboard
| Challenge: | Automated evaluation metrics are an essential part of the development of text-generation tasks such as summarization. |
| Approach: | They propose to use top-scoring system outputs to assess the reliability of automatic evaluation metrics for text summarization. |
| Outcome: | The proposed evaluation method is based on human judgments from 25 top-scoring neural summarization systems. |
Copied to clipboard
| Challenge: | Existing open-source evaluation paradigms lack flexibility and performance . language model-based evaluation is cheap and scalable, but it is difficult to evaluate . |
| Approach: | They propose a language model-based evaluation paradigm that uses a scalar indicator of quality to assess LM outputs. |
| Outcome: | The proposed language model-based evaluation model is more powerful than its predecessor. |
Copied to clipboard
| Challenge: | Multimodal Retrieval Augmented Generation (MMRAG) is a powerful approach to question-answering over multimodal documents. |
| Approach: | They propose a synthetic data generation framework that leverages interplay between a retriever, large language model and large multimodal model to generate question and answer pairs directly from multimodal documents. |
| Outcome: | The proposed framework generates question and answer pairs from 1024 questions over Wikipedia documents and evaluates state-of-the-art models using it. |
Copied to clipboard
| Challenge: | We show that when people use large language models to generate recommendations, the LLMs produce responses that reflect both what the user wants and who the user is. |
| Approach: | They propose that chatbots should transparently indicate when user’s revealed identity influences model recommendations but fail to do so . |
| Outcome: | The proposed model generates racially stereotypical recommendations regardless of whether the user revealed their identity intentionally or unintentionally through implicit cues. |
Copied to clipboard
| Challenge: | Existing approaches to compositional generalization in semantic parsers focus on word-level alignments, but they focus on spans. |
| Approach: | They propose a span-level supervised attention loss that improves compositional generalization in semantic parsers by focusing on spans. |
| Outcome: | The proposed method improves on three benchmarks of compositional generalization. |
Copied to clipboard
| Challenge: | Existing methods for data augmentation for text-based tasks such as machine translation are limited due to noise and noise. |
| Approach: | They propose a data augmentation policy with desirable properties as an optimization problem and propose 'SwitchOut' switchout randomly replaces words in both the source and target sentences with other random words from their corresponding vocabularies. |
| Outcome: | The proposed method outperforms strong alternatives such as word dropout on three translation datasets. |
Copied to clipboard
| Challenge: | Pretrained multilingual models can perform cross-lingual transfer in a zero-shot setting, even for unseen languages. |
| Approach: | They propose to extend XNLI to 10 indigenous languages of the Americas and test multiple zero-shot and translation-based approaches. |
| Outcome: | The proposed model can perform cross-lingual transfer in a zero-shot setting even for languages unseen during pretraining. |
Copied to clipboard
| Challenge: | Performance prediction is a task of estimating a system’s performance without performing experiments. |
| Approach: | They propose to understand reliability of performance prediction models from two angles: confidence intervals and calibration. |
| Outcome: | The proposed methods demonstrate the feasibility of fine-grained performance prediction and the necessity to perform reliability analysis for performance prediction methods in the future. |
Copied to clipboard
| Challenge: | Existing subword regularization methods for multilingual pretrained representations are suboptimal for multi-lingual transfer. |
| Approach: | They propose a method that enforces consistency between standard and probabilistic segmentations. |
| Outcome: | The proposed method improves the effectiveness of cross-lingual transfer by 2.5 points over standard methods. |
Copied to clipboard
| Challenge: | Specialized language and task adapters have been proposed to facilitate cross-lingual transfer of multilingual pretrained models. |
| Approach: | They propose a method that optimizes the ensemble weights of pretrained adapters for each test sentence by minimizing the entropy of its predictions. |
| Outcome: | The proposed method improves robustness to uncovered languages without training new adapters. |
Copied to clipboard
| Challenge: | specialized AI agents with task-specific tools or architectures fail to generalize beyond their intended scope. |
| Approach: | They propose a single-agent system with a modest number of general tools . they propose to generalize across software engineering, deep research and web browsing . |
| Outcome: | The proposed system achieves superior or competitive performance over specialized agents on three benchmarks. |
Copied to clipboard
| Challenge: | idioms are common in everyday language, but often pose a challenge to translators because their meanings do not follow from the meanings of their parts. |
| Approach: | They propose to use retrieval-augmented models to increase the accuracy of a strong pretrained machine translation model on idiomatic sentences by up to 13%. |
| Outcome: | The proposed techniques improve the accuracy of a strong pretrained model on idiomatic sentences by up to 13% in absolute accuracy, and holds potential benefits for non-idiomatic phrases. |
Copied to clipboard
| Challenge: | Existing algorithms for annotating parts of speech are not optimal for all languages. |
| Approach: | They propose to use a data selection algorithm to select useful training samples to minimize annotation cost. |
| Outcome: | The proposed strategy outperforms existing strategies on six typologically diverse languages. |
Copied to clipboard
| Challenge: | Recent studies have found that the performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text in a target language. |
| Approach: | They propose to use bilingual lexicons to synthesize textual or labeled data and combine it with monolingual or parallel text when available. |
| Outcome: | The proposed methods improve performance for 19 under-represented languages with and without extra monolingual text. |
Copied to clipboard
| Challenge: | Many-shot in-context learning shifts computational burden from training-time to inference-time, making deployment of many-shot ICL challenging to justify in-practice. |
| Approach: | They propose a method for retrieval-based many-shot in-context learning that uses blocks-sparse attention and retrieval of cached demonstrations to achieve comparable per-example latency to finetuning. |
| Outcome: | The proposed method achieves comparable per-example latency to finetuning while maintaining on average >95% of the best method’s accuracy across strong ICL and finetuned baselines. |
Copied to clipboard
| Challenge: | Existing named entity recognition models use gazetteers to improve performance, but they are limited in coverage and do not exist in low-resource languages. |
| Approach: | They propose a method that integrates Wikipedia information into named entity models by cross-lingual entity linking. |
| Outcome: | The proposed method improves on four low-resource languages with Wikipedia . it incorporates available information from english knowledge bases into neural models . |
Copied to clipboard
| Challenge: | Creating a descriptive grammar is an indispensable step for language documentation but it is tedious and time-consuming. |
| Approach: | They propose a framework for extracting a first-pass grammatical specification from raw text in a concise, human- and machine-readable format. |
| Outcome: | The proposed framework extracts a grammatical specification that is nearly equivalent to those created with large amounts of gold-standard annotated data. |
Copied to clipboard
| Challenge: | Recent studies have performed zero-shot learning by synthesizing training examples of canonical utterances and programs from a grammar, and further paraphrasing these utterrances to improve linguistic diversity. |
| Approach: | They propose to bridge gaps between canonical and real-world user-issued examples by using stronger paraphrasers and improved grammars. |
| Outcome: | The proposed model achieves strong performance on two semantic parsing benchmarks with zero labeled data. |
Copied to clipboard
| Challenge: | Semantic parsing is the task of transducing natural language utterances into machine executable meaning representations (e.g., Python code). |
| Approach: | They propose to rerank an n-best list of predicted MRs and use features to fix observed problems with baseline models to improve parser performance. |
| Outcome: | The proposed method outperforms the best published neural parser on four datasets and improves the baseline parsing performance by 5.7% and 2.9%. |
Copied to clipboard
| Challenge: | In-context learning-based evaluators are competitive with learned evaluation frameworks for text summarization tasks. |
| Approach: | They propose to use large language models as multi-dimensional evaluators using in-context learning to evaluate text summarization tasks. |
| Outcome: | The proposed frameworks are competitive with existing frameworks on relevance and factual consistency, the authors show . |
Copied to clipboard
| Challenge: | Existing methods to combine evidence annotations with document labels are limited to a minority of training examples. |
| Approach: | They propose to combine evidence annotations with abundant document labels for evidence extraction task. |
| Outcome: | The proposed method outperforms baselines on two classification tasks with evidence annotations. |
Copied to clipboard
| Challenge: | Prior work on argumentation in the NLP community has focused mainly on the first goal and has missed more nuanced and complex details of viewpoints. |
| Approach: | They propose a neural architecture that explicitly models the interplay between an Opinion Holder's (OH's) reasoning and a challenger's argument to predict if the argument succeeded in altering the OH' s view. |
| Outcome: | The proposed model outperforms several baselines on discussions on the Change My View forum on Reddit. |
Copied to clipboard
| Challenge: | a new spelling correction toolkit is available for free. |
| Approach: | They propose an open-source toolkit for spelling correction in English . they train neural models using spelling errors in context and using richer contextual representations. |
| Outcome: | The proposed spell-checker improves accuracy on synthetic examples and richer representations of the context. |
Copied to clipboard
| Challenge: | Recent years have witnessed the burgeoning of pretrained language models (LMs) for text-based natural language understanding tasks. |
| Approach: | They propose a pretrained language model that jointly learns representations for NL sentences and (semi-)structured tables. |
| Outcome: | The proposed model performs best on the weakly-supervised semantic parsing benchmark WikiTableQuestions while performing competitively on the text-to-SQL dataset Spider. |
Copied to clipboard
| Challenge: | Existing methods for learning paraphrastic sentence embeddings on bitext are expensive and require manual annotation. |
| Approach: | They propose a method that trains paraphrastic sentence embeddings directly from bitext, eliminating the time-consuming step of creating paraphrase corpora. |
| Outcome: | The proposed model outperforms and is faster than state-of-the-art models on cross-lingual tasks. |
Copied to clipboard
| Challenge: | Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges. |
| Approach: | They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources. |
| Outcome: | The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. |
Copied to clipboard
| Challenge: | a new system allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. |
| Approach: | They propose a system that allows users to train their own paraphrastic sentence representations in a variety of languages. |
| Outcome: | The proposed models outperform previous models on monolingual and cross-lingual tasks and can be used on CPUs with little difference in inference speed. |
Copied to clipboard
| Challenge: | Current instruction-tuning datasets focus on simplistic visual question answering tasks, and provide phrase-level answers without any intermediate rationales. |
| Approach: | They propose to use open-source multimodal large language models to train MLLMs on a dataset with 12M instruction-response pairs to elicit CoT reasoning. |
| Outcome: | The proposed model achieves state-of-the-art performance on benchmarks such as MathVerse, MMMU-Pro, and MuirBench, and gains improvements of up to 4% on non-reasoning-based benchmarks. |
Copied to clipboard
| Challenge: | a large number of natural language processing tasks are generated with specially designed architectures. |
| Approach: | They propose to represent a wide variety of tasks in a single unified format . they perform extensive experiments to demonstrate benefits of multi-task learning . |
| Outcome: | The proposed model performs comparable to state-of-the-art models on 10 tasks . it also shows that it can analyze differences and similarities in how the model handles different tasks compared to other models . |
Copied to clipboard
| Challenge: | lexicon induction evaluation dictionaries are mostly between English and another language, and the English hub is selected by default as the hub . lexiconic embeddings are often learned with a two-step process, whether under bilingual or multilingual settings. |
| Approach: | They propose to use English as the hub language for lexicon induction evaluation . they also expand a standard English-centered evaluation dictionary collection to include all language pairs . |
| Outcome: | The proposed method can significantly improve lexicon induction performance over multiple languages. |
Copied to clipboard
| Challenge: | Figures permeate human communication, but are understudied in NLP. |
| Approach: | They create a figurative language inference dataset for seven languages associated with a variety of cultures, using cultural and regional concepts for figurativ expressions. |
| Outcome: | The results show that the most common figurative expressions are found in Hindi, Indonesian, Javanese, Kannada, Sundanese, Swahili and Yoruba. |
Copied to clipboard
| Challenge: | Existing models that capture speaker-related variations do not include explicit information about the speaker. |
| Approach: | They propose a method that adapts the bias of the output softmax to each particular user . they propose to model speaker-related variations as an additional bias vector in the softmax layer . |
| Outcome: | The proposed technique improves translation accuracy and better reflection of speaker traits in target text. |
Copied to clipboard
| Challenge: | Existing methods for text style transfer lack parallel corpora, which makes it impossible to train supervised models. |
| Approach: | They propose to use semantic similarity metrics to explicitly assess the preservation of content between system outputs and inputs. |
| Outcome: | The proposed methods provide significant gains in automatic and human evaluation over strong baselines. |
Copied to clipboard
| Challenge: | ODEX is the first open-domain EXecution-based natural language (NL) to Python code generation dataset. |
| Approach: | They propose to use a dataset to extend the scope of coding queries to more realistic settings by using open-domain EXecution-based natural language (NL) to Python. |
| Outcome: | The proposed dataset has 945 NL-Code pairs and 1,707 human-written test cases. |
Copied to clipboard
| Challenge: | Existing work on figurative language has not been done on literal language models. |
| Approach: | They propose a Winograd-style task to evaluate figurative phrases with divergent meanings by interpreting paired figurativ phrases with a human input. |
| Outcome: | The proposed task outperforms state-of-the-art models on a nonliteral language understanding task in zero-shot settings. |
Copied to clipboard
| Challenge: | Abstractive summarization models are flexible, but they can be difficult to control. |
| Approach: | They propose a general and extensible guided summarization framework that takes different kinds of guidance as input and perform experiments across different varieties. |
| Outcome: | The proposed framework can generate more faithful summaries and different types of guidance generate qualitatively different summary. |
Copied to clipboard
| Challenge: | Existing methods to predict interpreter confidence and the adequacy of the interpreted message are lacking. |
| Approach: | They propose to extend a QE pipeline to estimate interpreter performance by using five settings in three language pairs. |
| Outcome: | The proposed method can predict interpreter confidence and adequacy over five settings in three language pairs and improves interpretation strategy and evaluation measures. |
Copied to clipboard
| Challenge: | Existing methods for MT have problems with translating homographs, as it is difficult to select the correct translation based on the context. |
| Approach: | They propose to model the context of the input word with context-aware word embeddings that help to differentiate the word sense before feeding it into the encoder. |
| Outcome: | The proposed models improve translation accuracy and BLEU score on three language pairs. |
Copied to clipboard
| Challenge: | Existing resources for standardized, easily accessible IGT data limit their applicability to linguistic research. |
| Approach: | They compile the largest existing corpus of interlinear glossed text data from a variety of sources and use it to generate annotated text. |
| Outcome: | The proposed model outperforms SOTA models on monolingual corpora by 6.6%. |
Copied to clipboard
| Challenge: | Language documentation involves recording the speech of native speakers. |
| Approach: | They propose to use a neural network architecture to model phonemes and tones versus modelling them separately. |
| Outcome: | The proposed method improves efficiency, minimizes typographical errors and maintains transcription faithfulness to acoustic signal while highlighting phonetic and phonemic facts for linguistic consideration. |
Copied to clipboard
| Challenge: | Existing sequence generation models produce outputs in one pass, usually left-to-right . current models model only a single edit step, and do not fully model editing . |
| Approach: | They propose to model editing processes, modeling the whole process of iteratively generating sequences. |
| Outcome: | The proposed model improves performance on a variety of axes compared to previous models . iterative refinement and editing are central parts of human creative workflow . |
Copied to clipboard
| Challenge: | Generative question answering (QA) models generate answers to complex questions, but their mechanism for doing so is still poorly understood. |
| Approach: | They decompose multi-hop questions into multiple corresponding single-hop question chains and find marked inconsistency in QA models’ answers on these pairs of ostensibly identical question chains. |
| Outcome: | The proposed models lack zero-shot multi-hop reasoning ability when trained on single-hop questions and on logical forms. |
Copied to clipboard
| Challenge: | Currently, there is little to no data available to build natural language processing models for endangered languages. |
| Approach: | They propose a benchmark dataset of transcriptions for scanned books in three critically endangered languages and a method to improve OCR in these data-scarce settings. |
| Outcome: | The proposed method reduces the recognition error rate by 34% across the three endangered languages. |
Copied to clipboard
| Challenge: | despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages. |
| Approach: | They propose to use global benchmarks to track progress on all NLP datasets in all languages. |
| Outcome: | a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench. |
Copied to clipboard
| Challenge: | Recent advances in morphological inflection generation have limited resources . antonisa and colleagues present a battery of improvements to improve performance under low-resource conditions . |
| Approach: | They propose a two-step attention architecture for the inflection decoder that uses two-segments attention and a multi-single-syllabic attention architecture. |
| Outcome: | The proposed model outperforms the state-of-the-art in low-resource languages by 15 percentage points . the proposed model also shows that it can be used to model monolingual data hallucinations . |
Copied to clipboard
| Challenge: | Existing work on style transfer has focused on controlling formality, authorial style, and sentiment of text. |
| Approach: | They propose a style transfer task that reframes a dialogue from informal first person to formal third person rephrasing . they use a dataset to annotate dialogues from a text summarization corpus . |
| Outcome: | The proposed task improves the performance of extractive models on a dialogue summarization dataset. |
Copied to clipboard
| Challenge: | Pre-trained multilingual models have enabled deployment of NLP technologies for multiple languages, but their performance under an annotation budget remains an open question. |
| Approach: | They propose a framework that prescribes the exact data-points to label from vast amounts of unlabelled multilingual data, having unknown degrees of overlap with the target set. |
| Outcome: | The proposed framework outperforms strong baselines in 84% of the test cases in the zero-shot setting of disjoint source and target language sets. |
Copied to clipboard
| Challenge: | Prior work on LM and acceptability judgments treat these effects uniformly across models, making a strong assumption that models require the same degree of adjustment to control for length and unigram frequency effects. |
| Approach: | They propose a linking theory where the optimal level of adjustment is estimated from data via learned parameters for length and unigram frequency. |
| Outcome: | The proposed theory outperforms a commonly used linking theory for acceptability—SLOR—across two families of transformer LMs. |
Copied to clipboard
| Challenge: | Language teachers need to be accessible and have the necessary resources to create effective content for their students. |
| Approach: | They propose to extract grammar descriptions from a natural text corpus that answer questions about morphosyntax and semantics from lexical corpus. |
| Outcome: | The proposed method is applied to two Indian languages, Kannada and Marathi, which, unlike English, do not have well-developed resources for second language learning. |
Copied to clipboard
| Challenge: | Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding . |
| Approach: | They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another. |
| Outcome: | The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers . |
Copied to clipboard
| Challenge: | Context-aware machine translation models fail to leverage contextual information to resolve ambiguous words and pronouns. |
| Approach: | They propose a new dataset that includes supporting context words for 14K translations that professional translators found useful for pronoun disambiguation. |
| Outcome: | The proposed model can automatically disambiguate pronouns and polysemous words when they are not in the same context. |
Copied to clipboard
| Challenge: | Non-parametric neural language models (NLMs) learn text distributions by memorizing training data points. |
| Approach: | They propose to use an external datastore to learn from a non-parametric language model. |
| Outcome: | The proposed methods achieve up to a 6x speed-up in inference speed while retaining comparable performance. |
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models have led to progress in tackling complex reasoning tasks that combine textual and visual information. |
| Approach: | They introduce a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. |
| Outcome: | The proposed model performs lower on MMMU-Pro than on the previous benchmark, ranging from 16.8% to 26.9%. |
Copied to clipboard
| Challenge: | Existing web agents use browsers to facilitate human activities such as online shopping, online planning, and other work-related tasks. |
| Approach: | They propose to use web browsers as an interface to interact with online content through application programming interfaces (APIs) they propose to call APIs and use Hybrid Agents to perform online tasks. |
| Outcome: | The proposed agents outperform web Browsing Agents on a widely-used and realistic benchmark for web navigation tasks. |
Copied to clipboard
| Challenge: | Existing NMT systems require specialized heuristics and large batch sizes. |
| Approach: | They propose a curriculum learning framework for NMT that reduces training time and costs . framework consists of a principled way of deciding which training samples are shown to the model . |
| Outcome: | The proposed framework can reduce training time and improve performance of recurrent neural network models and Transformers. |
Copied to clipboard
| Challenge: | NLP models today strive for supporting multiple languages and modalities, improving accessibility for diverse users. |
| Approach: | They propose a translation-test approach to tackle multilinguality, visual programming approach to break down complex reasoning, and a method that leverages image captioning to address multimodality. |
| Outcome: | The proposed interventions boost open models LLaVA-v1.5-13B by 13.4%, LLva-v1.6-34B by 20.3%, and Qwen-VL by 16.7% while minorly improving GPT-4V’s performance. |
Copied to clipboard
| Challenge: | Neural networks are data hungry and domain sensitive, so it is difficult to obtain labeled data for every domain. |
| Approach: | They propose a framework for domain adaptation where we model the difference between domains instead of smoothing over them. |
| Outcome: | The proposed framework improves on domain adaptation in multiple experimental settings. |
Copied to clipboard
| Challenge: | Contrastive learning is the dominant paradigm for learning text representations from parallel text, but finding negative examples can be expensive in terms of compute or manual effort. |
| Approach: | They propose a generative model for learning multilingual text embeddings which encourages source separation in multilingual contexts by an approximation. |
| Outcome: | The proposed model outperforms both a strong contrastive and generative baseline on a suite of tasks including semantic similarity, bitext mining, and cross-lingual question retrieval. |
Copied to clipboard
| Challenge: | Existing approaches to multilingual neural machine translation lack language-specific parameterization. |
| Approach: | They propose a modification to existing neural machine translation models that allows for language specific parameterization and domain adaptation. |
| Outcome: | The proposed model surpasses state-of-the-art for both the IWSLT-15 and IWSTL-17 datasets and can perform zero-shot translation. |
Copied to clipboard
| Challenge: | Recent work on tokenizer-free models shows promising results in cross-lingual transfer . previous work focused on reporting accuracy on a limited set of tasks and data settings . |
| Approach: | They compare tokenizer-free and subword-based models using various dimensions . they find subword models are still the most practical choice in many settings . |
| Outcome: | The proposed model improves cross-lingual transfer and reduces engineering overhead. |
Copied to clipboard
| Challenge: | Existing text-to-video diffusion models rely on text-only encoders for their pretraining, restricting their versatility and application in multimodal integration. |
| Approach: | They propose a multimodal conditional video generation framework for pretraining on augmented text prompts and then utilize a two-stage training strategy to enable diverse video generation tasks within a model. |
| Outcome: | The proposed model can synthesize consistent and temporally coherent videos with large motion while retaining the semantic control. |
Copied to clipboard
| Challenge: | Attention mechanisms are ubiquitous components in neural network architectures and are often claimed to confer interpretability. |
| Approach: | They propose a method for training models to produce deceptive attention masks by combining weights assigned to designated impermissible tokens with a weighted sum. |
| Outcome: | The proposed method reduces the weight assigned to designated impermissible tokens while still using them across multiple models and tasks. |
Copied to clipboard
| Challenge: | Unlike machine translation, natural language outputs are nuanced and there are no clear yes/no distinctions about whether they are correct or not. |
| Approach: | They describe compare-mt, a tool for holistic analysis and comparison of the results of systems for language generation tasks such as machine translation. |
| Outcome: | The compare-mt tool is an open-source pure-python package that has already proven useful to generate analyses that have been used in our papers. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are a step backward from traditional special-purpose NLP models . they require extensive computational resources for deployment and can be gated behind APIs . |
| Approach: | They propose a general-purpose method that takes a natural language task description and uses it to train a special-purpose model. |
| Outcome: | The proposed method outperforms a strong LLM by 20% while being 700 times smaller. |
Copied to clipboard
| Challenge: | a tutorial will explore the intersection of programming and natural language to make this goal a reality . |
| Approach: | This tutorial will focus on machine learning models of programs and natural language . it will discuss similarities and differences between programming and natural languages . |
| Outcome: | This tutorial will discuss the intersection of programming and natural language . it will cover automatic explanation of programs in natural language and automatic generation of programs from natural language specifications . |
Copied to clipboard
| Challenge: | Semantic sentence embedding models encode natural language sentences into vectors, such that closeness in embeddable space indicates closeness of semantics between the sentences. |
| Approach: | They propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors. |
| Outcome: | The proposed model outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations. |
Copied to clipboard
| Challenge: | In-context learning is limited by context length, but it can be used for many tasks. |
| Approach: | They study the behavior of in-context learning at an extreme context length . example retrieval shows excellent performance at low context lengths but has diminished gains . |
| Outcome: | The proposed model can perform many tasks with reasonable accuracy when a few examples are provided in-context. |
Copied to clipboard
| Challenge: | Text generation systems are ubiquitous in natural language processing applications, but evaluation of these systems remains a challenge, especially in multilingual settings. |
| Approach: | They propose a metric to evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
| Outcome: | The proposed metric can evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve. |
| Approach: | They propose a benchmark to assess the performance of multimodal web agents . they use visual and textual inputs to process and interpret natural language instructions . |
| Outcome: | a new benchmark assesses the performance of multimodal agents on visually grounded tasks . the benchmark identifies limitations of text-only agents and offers insights towards building stronger agents for the web . |
Copied to clipboard
| Challenge: | Connectionist Temporal Classification (CTC) is widely used for automatic speech recognition (ASR) but lags behind attentional decoder approaches in terms of translation quality. |
| Approach: | They propose to use a CTC/attention framework to validate this hypothesis by modifying the Hybrid CTC-Attention model proposed for automatic speech recognition to support text-to-text translation (MT) and speech-totext translation. |
| Outcome: | The proposed model outperforms pure-attention baselines across six translation tasks. |
Copied to clipboard
| Challenge: | Experimental results demonstrate that our methods achieve improvements of up to 1.8 BLEU points over competitive baselines. |
| Approach: | They propose a data selection and weighting strategy to iterate back-translation models and apply it to it . they use a target language to back-transcribe monolingual data, which is of high quality and reflect the target domain. |
| Outcome: | The proposed approach achieves 1.8 BLEU points over baselines on domain adaptation, low-resource, and high-resourced MT settings and on two language pairs. |
Copied to clipboard
| Challenge: | Open information extraction (IE) is the task of extracting open-domain assertions from natural language sentences. |
| Approach: | They propose an additional binary classification loss to calibrate the extraction likelihood . they propose an iterative learning process where extractions generated by the open IE model are incrementally included as training samples to help the model learn from trial and error. |
| Outcome: | Experiments on open information extraction (IE) show that the extraction likelihood is not well calibrated when comparing quality of extracted assertions. |
Copied to clipboard
| Challenge: | Existing models that generate words from a fixed vocabulary are linguistically nave . authors present an open-vocabulary language model that incorporates morphological knowledge into a neural framework . |
| Approach: | They propose a model that incorporates morphological knowledge into a neural model by generating words as a sequence of characters, generating full word forms and combining them with a hand-written morphology analyzer. |
| Outcome: | The proposed model outperforms character-based models on Finnish, Turkish, and Russian on three languages. |
Copied to clipboard
| Challenge: | Current systems for syntactic analysis tasks rely heavily on large scale annotated data. |
| Approach: | They propose to learn a generative model with a structured prior that uses labeled source and unlabeled target data jointly. |
| Outcome: | The proposed model improves on part-of-speech tagging and dependency parsing tasks on English as the only source corpus and on a wide range of target languages. |
Copied to clipboard
| Challenge: | Despite advances in machine translation quality estimation and evaluation, decoding is mostly oblivious to this. |
| Approach: | They propose to use a decoding framework that is quality-aware for neural machine translation . they compare various methods like N-best reranking and minimum Bayes risk decoding . |
| Outcome: | The proposed quality-aware decoding outperforms MAP-based decoding on four datasets and two model classes. |
Copied to clipboard
| Challenge: | eschewing separate architecture and training for knowledge-intensive tasks is cumbersome . end-to-end training only based on supervision from the end task is awkward . |
| Approach: | They propose a single Transformer that performs retrieval as attention and end-to-end training solely based on supervision from the end QA task. |
| Outcome: | The proposed model outperforms state-of-the-art retrievers and readers on in-domain datasets. |
Copied to clipboard
| Challenge: | Large language models store factual knowledge in parameters, but it can become outdated as the work evolves . pre-instruction-tuning improves ability of LLMs to absorb knowledge from new documents . |
| Approach: | They propose a method that instruction-tunes on questions prior to training on documents . they propose to use QA pairs to update factual knowledge of large language models . |
| Outcome: | The proposed method outperforms instruction-tuning on documents by 17.8%. |
Copied to clipboard
| Challenge: | Recent approaches to cross-lingual word embeddings have been based on linear transformations between the embeddable vectors in the two languages. |
| Approach: | They propose a method that expresses two monolingual embedding spaces as probability densities and matches them using a Gaussian mixture model. |
| Outcome: | The proposed method can achieve competitive or superior performance on bilingual lexicon induction and cross-lingual word similarity data. |
Copied to clipboard
| Challenge: | Prior work shows that program-aided reasoning improves accuracy but also requires reasoners to "know what they know". |
| Approach: | They compare the calibration of program-aided language models (PAL) and text-based Chain-of-thought (COT) prompting techniques over 5 datasets and 2 model types . |
| Outcome: | The proposed methods improve accuracy and calibrate the models over 5 datasets and 2 model types. |
Copied to clipboard
| Challenge: | despite advances in large language models, task-specific data is not available for many use cases . a new method to improve automated dataset generation uses publicly available datasets . |
| Approach: | They propose a method to make better use of existing datasets to improve automatic dataset generation. |
| Outcome: | The proposed method outperforms existing methods on language-based tasks . it significantly increases diversity and difficulty of generated data on many tasks compared to other methods . |
Copied to clipboard
| Challenge: | Recent advances in multilingual natural language processing have improved performance on benchmarks such as XTREME and XGLUE by 13 points . however, improvements have been easier to achieve in some tasks than others . |
| Approach: | They extend XTREME to XTRAME-R, which includes ten natural language understanding tasks and covers 50 typologically diverse languages. |
| Outcome: | The proposed framework improves the performance on the XTREME multilingual benchmark by 13 points compared to human-level performance on English transfer learning. |
Copied to clipboard
| Challenge: | Existing methods for evaluating language models are brittle, corpus-level perplexities are vague, and the choice of benchmarks is endless. |
| Approach: | They propose a method that uses contextual embeddings to find fine-grained features of text where one model outperforms another. |
| Outcome: | The proposed method extracts features that demonstrate differences with respect to ease of generation between two language models. |
Copied to clipboard
| Challenge: | Recent studies have focused on domain adaptation for neural machine translation systems where in-domain data is scarce or nonexistent. |
| Approach: | They propose an approach that adapts models with domain-aware feature embeddings, which are learned via an auxiliary language modeling task. |
| Outcome: | The proposed model performs better in multiple experimental settings and with back translation. |
Copied to clipboard
| Challenge: | Continual learning methods tackle the problem of a changing world by incrementally training on new information. |
| Approach: | They propose a retrieval-augmented generation approach that incrementally deletes or rewrites other entries in the knowledge base each time a document is added. |
| Outcome: | The proposed model improves accuracy relative to conventional retrieval-augmented generation by 7-13% and 6-10% absolute. |
Copied to clipboard
| Challenge: | Existing evaluation methods for named entity recognition tasks are difficult to interpret . authors present a general methodology for interpretable evaluation for named entities . |
| Approach: | They propose a general methodology for interpretable evaluation for named entity recognition task. |
| Outcome: | The proposed evaluation method enables researchers to interpret differences in models and datasets . it makes it easy for future researchers to run similar analyses and drive progress in this area . |
Copied to clipboard
| Challenge: | Existing work has treated procedures as shallow structures without modeling the parent-child relation. |
| Approach: | They propose to construct an open-domain hierarchical knowledge-base (KB) of procedures based on wikiHow . they link steps in an article to other articles with similar goals, recursively building the KB . |
| Outcome: | The proposed method significantly outperforms baselines according to automatic evaluation, human judgment, and application to downstream tasks such as instructional video retrieval. |
Copied to clipboard
| Challenge: | Compositionality is a hallmark of human language, but many phrases are non-compositional . a study by a team of researchers shows that LMs may not be able to distinguish between compositional and non-composable phrases. |
| Approach: | They propose to predict LM-internal representations of longer phrases given their constituents . they find that the representation of a parent phrase can be predicted with some accuracy . |
| Outcome: | The proposed model can predict a parent phrase with some accuracy given its children's transformations, but this is not the case. |
Copied to clipboard
| Challenge: | Language models excel at generating code, but many programs are difficult to generate using only parametric knowledge. |
| Approach: | They propose a retrieval-augmented code generation benchmark that provides reproducible evaluations on retrieval and end-to-end code generation performance. |
| Outcome: | The proposed benchmark covers programming, open-domain, and repository-level tasks and provides reproducible evaluations on retrieval and end-to-end code generation performance. |
Copied to clipboard
| Challenge: | Abstractive summarization models are often trained with maximum likelihood estimation (MLE) . mLE assumes a deterministic (one-point) target distribution, but can cause performance degradation . |
| Approach: | They propose a new training paradigm which assumes a non-deterministic distribution so that different candidate summaries are assigned probability mass according to their quality. |
| Outcome: | The proposed model can estimate probabilities of candidate summaries that are more correlated with their level of quality. |
Copied to clipboard
| Challenge: | Low-resource language pairs with a lack of parallel data pose challenges for machine translation . data augmentation using monolingual data is an effective way to alleviate the problem . |
| Approach: | They propose a general framework for data augmentation for low-resource machine translation using monolingual data and a related high-resourced language. |
| Outcome: | The proposed method improves translation quality by 1.5 to 8 BLEU points under extreme low-resource settings compared to baselines. |
Copied to clipboard
| Challenge: | Existing work has extended recurrent neural networks to model lattice inputs but these models suffer from slow computation speeds. |
| Approach: | They propose to extend the paradigm of self-attention to handle lattice inputs by adding probabilistic reachability masks that incorporate latticae structure into the model and support lattics if available. |
| Outcome: | The proposed model outperforms baseline models while being much faster to compute than previous models. |
Copied to clipboard
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) is sensitive to domain shift, resulting in failure for sentences with large numbers of unknown words and lack of supervision for domain-specific words. |
| Approach: | They propose an unsupervised method which fine-tunes a pre-trained out-of-domain NMT model using a pseudo-in-domain corpus. |
| Outcome: | The proposed method improves in five domains without using in-domain parallel sentences and up to 2 BLEU over strong back-translation baselines. |
Copied to clipboard
| Challenge: | Multilingual data is more beneficial for NMT models that translate from the LRL to a target language than those that translate into the LLLs. |
| Approach: | They propose a decoder that embeds character n-grams into NMT models that translate from an LRL to a target language. |
| Outcome: | The proposed decoder improves the performance of NMT models that translate from an LRL to a target language. |
Copied to clipboard
| Challenge: | Recent work in neural machine translation has demonstrated the necessity and feasibility of using inter-sentential context, but it is often not clear how much they actually utilize it at translation time. |
| Approach: | They propose a conditional cross-mutual information metric to quantify usage of context by model architectures that can use it at translation time. |
| Outcome: | The proposed method increases context usage and improves translation quality according to BLEU and COMET metrics. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual entity linking (XEL) do not extend well to low-resource languages with few Wikipedia pages. |
| Approach: | They propose to improve the model by combining Wikipedia references with a list of plausible candidate entities. |
| Outcome: | The proposed method yields 16.9% in Top-30 gold candidate recall compared with state-of-the-art models. |
Copied to clipboard
| Challenge: | Recent work on bilingual lexicon induction (BLI) relies on an assumption about the isometry of two embedding spaces. |
| Approach: | They propose a semi-supervised approach that relaxes the isometric assumption while leveraging limited aligned bilingual lexicons and a larger set of unaligned word embeddings. |
| Outcome: | The proposed method obtains state-of-the-art results on 15 of 18 language pairs on the MUSE dataset and does particularly well when the embedding spaces don’t appear isometric. |
Copied to clipboard
| Challenge: | Recent work examines knowledge contained in language models by having the LM fill in the blanks of prompts such as “Obama is a __ by profession”. |
| Approach: | They propose mining-based and paraphrasing-based methods to automatically generate high-quality and diverse prompts, as well as ensemble methods to combine answers from different prompts. |
| Outcome: | The proposed methods improve accuracy from 31.1% to 39.6% on the LAMA benchmark for extracting relational knowledge from LMs. |
Copied to clipboard
| Challenge: | Existing approaches to adapt neural machine translation systems to low-resource languages are difficult to implement and require large amounts of training data. |
| Approach: | They propose a method to train neural machine translation systems to new low-resource languages . they propose to start with massively multilingual "seed models" and continue training on data related to the LRL . |
| Outcome: | The proposed method achieves BLEU scores of up to 15.5 with no data from the LRL and improves over other adaptation methods by 1.7 BLUE points average over 4 LRL settings. |
Copied to clipboard
| Challenge: | a table-based question answering system requires complex reasoning and alignment between questions and tables. |
| Approach: | They propose a table-based QA model that consumes both natural and synthetic data . they combine retrieval with masking to pair natural sentences with QA . |
| Outcome: | The proposed model outperforms existing models in few-shot and full settings and on WikiTableQuestions. |
Copied to clipboard
| Challenge: | Existing approaches to grammar induction focus on discovering constituents or dependencies. |
| Approach: | They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies . |
| Outcome: | The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods. |
Copied to clipboard
| Challenge: | Recent work has shown that pre-trained language models can perform zero-shot generalization to new tasks without annotated examples. |
| Approach: | They propose to regularize prompt consistency to encourage consistent predictions over a diverse set of prompts. |
| Outcome: | The proposed approach outperforms the state-of-the-art zero-shot learner, T0, on 9 out of 11 datasets across 4 NLP tasks by 10.6 absolute points in terms of accuracy. |
Copied to clipboard
| Challenge: | Earlier named entity translation methods focus on phonetic transliteration, which ignores the sentence context for translation. |
| Approach: | They propose a DEnoising Entity Pre-training method that leverages monolingual data and a knowledge base to improve named entity translation accuracy within sentences. |
| Outcome: | The proposed method improves on three language pairs and denoising auto-encoding baselines. |
Copied to clipboard
| Challenge: | Existing systems that make predictions and ask questions are unable to have a mutual exchange of opinions. |
| Approach: | They propose to use a dataset and computational framework to allow systems to have beneficial discussions with humans, improving the accuracy by 25 points on a natural language inference task. |
| Outcome: | The proposed system improves accuracy by 25 points on a natural language inference task. |
Copied to clipboard
| Challenge: | Named-entity recognition (NER) models are highly dependent on large amounts of labeled data. |
| Approach: | They propose a method that finds translations based on bilingual word embeddings . they also propose 'self-attention' which allows for a degree of flexibility with respect to word order . |
| Outcome: | The proposed method achieves state-of-the-art or competitive performance on common languages with lower resource requirements than previous approaches. |
Copied to clipboard
| Challenge: | XEL is challenging for most languages because of limited availability of requisite resources . simulated environments that use significant resources are not available in truly low-resource languages . |
| Approach: | They propose improvements to entity candidate generation and disambiguation to make better use of the limited resources available in low-resource languages. |
| Outcome: | The proposed model gains 6-20% end-to-end linking accuracy on four low-resource languages. |
Copied to clipboard
| Challenge: | Current text-to-image models produce homogeneous outputs given under-specified prompts and their outputs are disproportionately biased toward Western cultures. |
| Approach: | They propose a framework that assesses the degree of cultural relevance of an image, given a user-defined set of labels. |
| Outcome: | The proposed evaluation metric surpasses baselines on a manually curated dataset of culturally salient but rare items built using language models by 22% F1 points. |
Copied to clipboard
| Challenge: | Using culture-agnostic subsets, performance drops in many LMMs when evaluated in Japanese. |
| Approach: | They introduce a Japanese benchmark to evaluate large multimodal models on expert-level tasks based on the Japanese cultural context. |
| Outcome: | The proposed benchmark enables comparisons with other benchmarks in other languages based on cultural contexts. |
Copied to clipboard
| Challenge: | Named entity recognition models rely on large amounts of labeled data, making them challenging to extend to new, lower-resource languages. |
| Approach: | They propose a method for bootstrapping named entity recognition models in under-resourced languages . they use cross-lingual transfer learning and targeted annotation of only uncertain entities . |
| Outcome: | The proposed method achieves competitive accuracy with just one-tenth of training data. |
Copied to clipboard
| Challenge: | Existing methods to explain predictions by highlighting salient features are often unstated. |
| Approach: | They propose a framework to quantify the value of explanations via the accuracy gains that they confer on a student model trained to simulate a teacher model. |
| Outcome: | The proposed framework allows principled, automatic, model-agnostic evaluation of attributions. |
Copied to clipboard
| Challenge: | Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text. |
| Approach: | They propose to use automatic stress tests to evaluate models' ability to make inferential decisions. |
| Outcome: | The proposed model performs well across genres of text, but lacks the ability to make inferential decisions. |
Copied to clipboard
| Challenge: | Multilingual training is an essential ingredient in machine translation systems . but it has different effects in different multilingual settings, such as many-to-one, one-tomany and many- to-many learning . |
| Approach: | They compare multilingual training settings with encoders and decoders initialized by multilingual learning . they find important attention heads for each language pair and compare their correlations during inference . |
| Outcome: | The proposed models outperform the best models for high-resource languages and one-to-many models for low-resourced languages. |
Copied to clipboard
| Challenge: | Prior work on text style transfer has not focused on politeness as a style transfer task and we argue that defining it is cumbersome. |
| Approach: | They propose a task of politeness transfer which involves converting non-polite sentences to polite sentences while preserving the meaning. |
| Outcome: | The proposed model outperforms state-of-the-art methods on content preservation and style transfer accuracy. |
Copied to clipboard
| Challenge: | Existing tools and research focus on how to interpret and manipulate data, despite its crucial role in machine learning, . existing tools and researchers focus on systems on top of existing data, rather than how to use it. |
| Approach: | They propose a unified data-oriented platform that allows users to interactively analyze the characteristics of data and provides a standard interface for many data processing operations. |
| Outcome: | The proposed platform allows users to analyze the characteristics of data and provides a standardized interface so that many data processing operations can be provided within a single interface. |
Copied to clipboard
| Challenge: | Existing approaches to generate graphs using pre-trained language models hinder their ability to generate them correctly. |
| Approach: | They propose to frame structured commonsense reasoning tasks as code generation tasks instead of serializing the output graph as a flat list of nodes and edges. |
| Outcome: | The proposed approach outperforms natural-language LMs in three natural language tasks even when the downstream task does not involve source code at all. |
Copied to clipboard
| Challenge: | Existing methods for contextual guessing and definition generation do not take clues from local contexts. |
| Approach: | They propose a neural description model that takes clues from local and global contexts . they assume that the target phrase is newly emerged and there is no global context . |
| Outcome: | The proposed model takes clues from local and global contexts over existing methods . it is more effective than existing methods for non-standard English explanation . |
Copied to clipboard
| Challenge: | Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages. |
| Approach: | They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
| Outcome: | The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP. |
Copied to clipboard
| Challenge: | Existing studies show that training on a single related language is more effective than using all data. |
| Approach: | They propose an efficient algorithm that first samples a target sentence, and then conditionally samples its source sentence. |
| Outcome: | The proposed algorithm brings significant gains on three of four languages with minimal training overhead. |
Copied to clipboard
| Challenge: | Existing studies on named entity recognition methods for African languages focus on English as the source language, but there is evidence that it is not the best for low-resource languages. |
| Approach: | They propose to use human-annotated datasets to analyze named entity recognition tasks in 20 African languages to test whether they are effective. |
| Outcome: | The proposed method improves zero-shot F1 scores by 14% over 20 languages compared to using English . |
Copied to clipboard
| Challenge: | Existing methods for digitizing text in endangered languages rely on manual data curated by the user. |
| Approach: | They propose a semi-supervised learning method that utilizes raw images to improve performance. |
| Outcome: | The proposed method reduces errors by 15%–29% on four endangered languages. |
Copied to clipboard
| Challenge: | Neural sequence-to-sequence models are autoregressive, meaning they factor the joint probability of the output sequence into the product of probabilities over the next to-ken. |
| Approach: | They propose a non-autoregressive sequence generation model using latent variables . they use generative flow to model complex distributions using neural networks . |
| Outcome: | The proposed model performs comparable to state-of-the-art models and has constant decoding time w.r.t the sequence length. |
Copied to clipboard
| Challenge: | Existing methods for multilingual entity linking are limited by textual contexts and limited resources. |
| Approach: | They propose a testbed system for multilingual multimodal entity linking using BBC news articles paired with corresponding images in five languages. |
| Outcome: | The proposed system improves accuracy for entities with ambiguous textual contexts and models with weak multilingual abilities. |
Copied to clipboard
| Challenge: | Neural sequence models can generate fluent sentences, but they can also hallucinate additional content not supported by the input. |
| Approach: | They propose a task to predict whether each token in the output sequence is hallucinated and collect manually annotated evaluation sets for this task. |
| Outcome: | The proposed method outperforms baseline methods on machine translation and abstractive summarization datasets and achieves significant improvements in both supervised and unsupervised settings. |
Copied to clipboard
| Challenge: | Knowledge Graphs (KGs) store information in the form of (head, predicate, tail)-triples. |
| Approach: | They propose a framework for performing fine-grained evaluation on meaningful subsets of data. |
| Outcome: | The proposed framework tests models on meaningful subsets of the data, which would have been impossible to detect with standard averaged single-score metrics. |
Copied to clipboard
| Challenge: | FAIL-TaLMs contains 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage. |
| Approach: | They introduce a benchmark to examine the shortcomings of tool-augmented language models (TaLMs) that assume 'perfect' information access and tool availability. |
| Outcome: | The proposed benchmark systematically evaluates 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage. |
Copied to clipboard
| Challenge: | Documentation is not a cure-all for language loss, but it is an important part of language preservation. |
| Approach: | They propose to use multi-source neural models to create automatic glossing models . they also explore cross-lingual transfer and a simple output length control mechanism . |
| Outcome: | The proposed model outperforms state-of-the-art models on low-resource scenarios. |
Copied to clipboard
| Challenge: | Current reasoning large language models (RLMs) are trained on data that is primarily in English, resulting in lower performance when asked the same question in a non-English language. |
| Approach: | They propose a framework for enhancing multilingual reasoning without any data in the target language(s). |
| Outcome: | The proposed framework outperforms Qwen2.5-7B-Instruct on 4 math and non-math tasks with less than 1/8 of the training data (125). |
Copied to clipboard
| Challenge: | Existing approaches to dependency parsing are local and greedy transitionbased . StackPtr parsers use the information of whole sentences and previously derived subtree structures . |
| Approach: | They propose a stack-pointer network-based dependency parser that reads whole sentence and builds dependency tree top-down in a depth-first fashion. |
| Outcome: | The proposed model reads and encodes whole sentence, then builds dependency tree top-down (from root-to-leaf) in a depth-first fashion. |
Copied to clipboard
| Challenge: | This tutorial will focus on NLP for endangered languages documentation and revitalization. |
| Approach: | This tutorial will focus on NLP for endangered languages documentation and revitalization . the goal is to motivate more NLP practitioners to work towards this important direction . |
| Outcome: | This tutorial will acquaint attendees with the process and the challenges of language documentation and revitalization. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, but the data imbalance hinders it from performing uniformly across language pairs. |
| Approach: | They propose a distributionally robust optimization objective which minimizes the worst-case expected loss over the set of language pairs. |
| Outcome: | The proposed learning objective outperforms baseline methods on three sets of languages and shows that it is cost-effective and efficient. |
Copied to clipboard
| Challenge: | a new data-centric approach could address cultural gaps in multimodal large language models . despite being trained on billions of image-text pairs, today's models are biased towards English and Western data. |
| Approach: | They propose a data-centric approach that directly grounds MLLMs in cultural knowledge. |
| Outcome: | The proposed approach outperforms open-source models on cultural-focused benchmarks without degrading results on mainstream vision–language tasks. |
Copied to clipboard
| Challenge: | The 3rd Workshop on Neural Machine Translation and Generation (WNGT) was held in concert with the annual conference of the Empirical Methods in Natural Language Processing (EMNLP 2019). |
| Approach: | They describe the results of the third workshop on Neural Generation and Translation held in concert with the annual conference of the Empirical Methods in Natural Language Processing (EMNLP 2019). |
| Outcome: | The results of the 3rd Workshop on Neural Machine Translation and Generation (WNGT) were summarized in Sections 3 and 4. |
Copied to clipboard
| Challenge: | a new task is to translate images to make them culturally relevant . currently, translation systems focus on translating words and images . |
| Approach: | They propose a task of translating images to make them culturally relevant . they build pipelines comprising state-of-the-art generative models to do the task . |
| Outcome: | The proposed pipelines can translate only 5% of translated images for some countries and no translation is successful for others. |
Copied to clipboard
| Challenge: | Cross-lingual transfer is a useful tool for improving performance of natural language processing (NLP) on low-resource languages. |
| Approach: | They propose to use cross-lingual transfer to improve accuracy of low-resource languages . they build models that consider features to perform prediction on such languages based on ranking problem . |
| Outcome: | The proposed model predicts good transfer languages much better than baselines considering single features in isolation. |
Copied to clipboard
| Challenge: | Phonemes are contrastive phonological units, and allophones are their various concrete realizations. |
| Approach: | They propose a resource that maps allophones to phonemes for 14 languages . they propose phonological representations that are much closer to a universal transcription . |
| Outcome: | The proposed resource maps from 218 allophones to phonemes for 14 languages. |
Copied to clipboard
| Challenge: | Existing work on open-domain code generation focuses on limited domains or domain-specific languages with limited set of operators. |
| Approach: | They incorporate external knowledge into NL-to-code generation by combining StackOverflow and programming language API documentation with data augmentation and retrieval-based data re-sampling. |
| Outcome: | The proposed approach improves the current state-of-the-art by up to 2.2% absolute BLEU score on the code generation testbed CoNaLa. |
Copied to clipboard
| Challenge: | Recent work shows that multilingual representations are disjointed across languages, bringing additional challenges for transfer onto extremely low-resource languages. |
| Approach: | They propose a meta-learning based framework that learns to transform representations judiciously from auxiliary languages to a target one and brings their representation spaces closer for effective transfer. |
| Outcome: | The proposed framework learns to transform representations from auxiliary languages to a target language and brings their representation spaces closer for effective transfer. |
Copied to clipboard
| Challenge: | Recent work on MT robustness has demonstrated the need to build or adapt systems that are resilient to such noise. |
| Approach: | They propose to synthesize natural noise in social media data to enhance robustness of MT systems by leveraging natural noise. |
| Outcome: | The proposed method can make a vanilla MT system more resilient to noise, partially mitigating loss in accuracy resulting therefrom. |
Copied to clipboard
| Challenge: | Recent work has shown that optimizing neural machine translation systems to directly improve evaluation metrics such as BLEU can improve final translation accuracy. |
| Approach: | They propose a reward function that assigns partial credit to BLEU and provides more diversity in scores than BLUE. |
| Outcome: | The proposed reward function improves translation accuracy, semantic similarity, and human evaluation on four languages trans-lated to English and the optimization procedure converges faster. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) has trouble with lowfrequency words or phrases and generalizing across domains. |
| Approach: | They propose a method for recalling low-frequency words and phrases into neural machine translation by retrieving n-grams from a search engine and incorporating them into the decoding process. |
| Outcome: | The proposed method improves translation results up to 6 BLEU points on three narrow domain translation tasks where repetitiveness of the target sentences is particularly salient. |
Copied to clipboard
| Challenge: | Recent trends in NLP use of pre-trained weights raise security questions . authors show that pre-training weights can be injected with vulnerabilities . |
| Approach: | They propose to build "weight poisoning" attacks where pre-trained weights are injected with vulnerabilities that expose "backdoors" they outline practical defenses against such attacks. |
| Outcome: | The proposed attacks expose "backdoors" after fine-tuning models . the proposed attacks are widely applicable and pose a serious threat . |
Copied to clipboard
| Challenge: | Existing methods for evaluating text quality are discriminative and generative . current methods use manual annotation of human judgements to train them . |
| Approach: | They propose a framework that combines the best of both worlds by using supervised and unsupervised signals from whatever data we have available. |
| Outcome: | The proposed method outperforms existing metrics on 5 datasets, 19 languages and 280 systems. |
Copied to clipboard
| Challenge: | Pre-trained word embeddings have proven to be invaluable for improving performance in natural language analysis tasks where large-scale parallel corpora cannot be obtained. |
| Approach: | They perform five sets of experiments to analyze when pre-trained word embeddings can be useful in NMT tasks. |
| Outcome: | The embeddings provide gains of up to 20 BLEU points in the most favorable setting. |
Copied to clipboard
| Challenge: | Generative language models (LMs) have a tendency to hallucinate and create inaccurate output. |
| Approach: | They propose a method which iteratively uses a prediction of the upcoming sentence to anticipate future content. |
| Outcome: | The proposed method achieves superior or competitive performance on all tasks . iteratively uses a prediction of the upcoming sentence to anticipate future content . |
Copied to clipboard
| Challenge: | NLCode generates long expressions and statements rather than a single next-token . evaluating and comparing different models has remained a challenge . |
| Approach: | They propose a code-generating evaluation metric built on BERTScore . they use five language-specific pretrained models to evaluate their code . |
| Outcome: | The proposed evaluation metric achieves higher correlation with human preference and functional correctness than existing metrics across four programming languages. |
Copied to clipboard
| Challenge: | MCoNaLa benchmarks natural language code generation in languages that are not native to English. |
| Approach: | They propose to benchmark natural language code generation from natural language commands extending beyond English by using a multilingual dataset. |
| Outcome: | The proposed dataset compares natural language commands with code generation systems in three languages. |