Papers by John Wieting

22 papers
RankGen: Improving Text Generation with Large Ranking Models (2022.emnlp-main)

Copied to clipboard

Challenge: Modern language models assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix.
Approach: They propose a 1.2B parameter encoder model for English that scores model generations given a prefix.
Outcome: The proposed model outperforms decoding algorithms on automatic metrics and human evaluations with English writers.
Evaluating and Modeling Attribution for Cross-Lingual Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Open-retrieval question answering systems are lacking in attribution for cross-lingual question answering . open-research questions are available in 20 languages, but their raw generation often falls short in factuality .
Approach: They are the first to study attribution for cross-lingual question answering . they collect data in 5 languages to assess the attribution level of a state-of-the-art QA system .
Outcome: The proposed approach improves the attribution level of a state-of-the-art cross-lingual QA system.
Augmenting Pre-trained Language Models with QA-Memory for Open-Domain Question Answering (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for open-domain question-answering use an open book approach . a recent alternative is to retrieve from a collection of previously-generated question-annwer pairs .
Approach: They propose a new QA system that augments a text-to-text model with a large memory of question-answer pairs and a task for the latent step of question retrieval.
Outcome: The proposed system outperforms closed-book QA and can answer multi-hop questions.
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators . a dataset of non-English language novels is used to study literary MT .
Approach: They use a dataset of non-English language novels aligned to human and automatic English translations to study literary MT.
Outcome: The proposed model prefers human translations over machine translations at a rate of 84% . state-of-the-art MT metrics do not correlate with preferences, the study finds .
On The Ingredients of an Effective Zero-shot Semantic Parser (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have performed zero-shot learning by synthesizing training examples of canonical utterances and programs from a grammar, and further paraphrasing these utterrances to improve linguistic diversity.
Approach: They propose to bridge gaps between canonical and real-world user-issued examples by using stronger paraphrasers and improved grammars.
Outcome: The proposed model achieves strong performance on two semantic parsing benchmarks with zero labeled data.
Evaluating Large Language Models on Controlled Generation Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have looked into the ability of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc. However, few studies investigate the controllability of large languages.
Approach: They propose to compare large language models with state-of-the-start finetuned smaller models to find that large language model controls are comparable to smaller models.
Outcome: The proposed model can meet hard constraints and perform better than state-of-the-art models.
Canine: Pre-training an Efficient Tokenization-Free Encoder for Language Representation (2022.tacl-1)

Copied to clipboard

Challenge: End-to-end neural models have replaced the traditional pipeline and require an explicit tokenization step.
Approach: They propose a neural encoder that operates directly on character sequences without explicit tokenization or vocabulary and a pre-training strategy that optionally uses subwords as a soft inductive bias.
Outcome: The proposed model outperforms a comparable mBert model on a multilingual benchmark by 5.7 F1 on the TyDi QA benchmark.
Simple and Effective Paraphrastic Similarity from Parallel Translations (P19-1)

Copied to clipboard

Challenge: Existing methods for learning paraphrastic sentence embeddings on bitext are expensive and require manual annotation.
Approach: They propose a method that trains paraphrastic sentence embeddings directly from bitext, eliminating the time-consuming step of creating paraphrase corpora.
Outcome: The proposed model outperforms and is faster than state-of-the-art models on cross-lingual tasks.
Paraphrastic Representations at Scale (2022.emnlp-demos)

Copied to clipboard

Challenge: a new system allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages.
Approach: They propose a system that allows users to train their own paraphrastic sentence representations in a variety of languages.
Outcome: The proposed models outperform previous models on monolingual and cross-lingual tasks and can be used on CPUs with little difference in inference speed.
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval (2024.naacl-long)

Copied to clipboard

Challenge: et al., 2020: performance of dense retrieval models in multilingual retrieval is limited due to uneven and scarce training data available across multiple languages.
Approach: They propose a synthetic retrieval training dataset containing 33 languages for fine-tuning multilingual retrievers without human supervision.
Outcome: The proposed model outperforms human-supervised retrieval models on three retrieval benchmarks.
On Learning Text Style Transfer with Direct Rewards (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for text style transfer lack parallel corpora, which makes it impossible to train supervised models.
Approach: They propose to use semantic similarity metrics to explicitly assess the preservation of content between system outputs and inputs.
Outcome: The proposed methods provide significant gains in automatic and human evaluation over strong baselines.
Faithful to the Document or to the World? Mitigating Hallucinations via Entity-Linked Knowledge in Abstractive Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing abstractive summarization systems are hampered by content hallucinations in which models generate text that is not directly inferable from the source alone.
Approach: They propose to use external knowledge to latently connect entities and concepts to latences to lend provenance to many of these unfaithful yet factual entities.
Outcome: The proposed model can be used to improve the factuality of summarizations without simply making them more extractive.
Beyond Contrastive Learning: A Variational Generative Model for Multilingual Retrieval (2023.acl-long)

Copied to clipboard

Challenge: Contrastive learning is the dominant paradigm for learning text representations from parallel text, but finding negative examples can be expensive in terms of compute or manual effort.
Approach: They propose a generative model for learning multilingual text embeddings which encourages source separation in multilingual contexts by an approximation.
Outcome: The proposed model outperforms both a strong contrastive and generative baseline on a suite of tasks including semantic similarity, bitext mining, and cross-lingual question retrieval.
A Bilingual Generative Transformer for Semantic Sentence Embedding (2020.emnlp-main)

Copied to clipboard

Challenge: Semantic sentence embedding models encode natural language sentences into vectors, such that closeness in embeddable space indicates closeness of semantics between the sentences.
Approach: They propose a deep latent variable model that attempts to perform source separation on parallel sentences, isolating what they have in common in a latent semantic vector, and explaining what is left over with language-specific latent vectors.
Outcome: The proposed model outperforms the state-of-the-art on a standard suite of unsupervised semantic similarity evaluations.
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets are often informed by established research directions in the NLP community.
Approach: They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Outcome: The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Reformulating Unsupervised Style Transfer as Paraphrase Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing systems for style transfer warp the input’s meaning through attribute transfer, which changes semantic properties such as sentiment.
Approach: They propose a method for fine-tuning pretrained language models on automatically generated paraphrase data to improve the efficiency of style transfer.
Outcome: The proposed method outperforms state-of-the-art style transfer systems on human and automatic evaluations and proposes fixed variants.
PostMark: A Robust Blackbox Watermark for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to detect LLM-generated text require access to the underlying LLM’s logits, which LLM providers are loath to share due to fears of model distillation.
Approach: They develop a post-hoc watermarking procedure that inserts an input-dependent set of words into the text after the decoding process has completed.
Outcome: The proposed method is more robust to paraphrasing attacks than existing methods.
ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations (P18-1)

Copied to clipboard

Challenge: Using neural machine translation, we generate more than 50 million sentential paraphrase pairs from a large parallel corpus.
Approach: They use a dataset of more than 50 million English-English sentential paraphrase pairs to generate them automatically using neural machine translation.
Outcome: The proposed dataset outperforms all supervised systems on every SemEval semantic textual similarity competition and shows how it can be used for paraphrase generation.
Adversarial Example Generation with Syntactically Controlled Paraphrase Networks (N18-1)

Copied to clipboard

Challenge: Existing approaches to learn to do syntactically controlled paraphrase generation are limited . lexical, pragmatic, and syntaktic variation can hurt generalization of models trained on them .
Approach: They propose a new approach for learning to do syntactically controlled paraphrase generation using a parser.
Outcome: The proposed model generates paraphrases that follow their target specifications without decreasing paraphrase quality compared to baseline models . it improves the robustness of the models to syntactic variation when used to augment training data.
Improving Candidate Generation for Low-resource Cross-lingual Entity Linking (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to cross-lingual entity linking (XEL) do not extend well to low-resource languages with few Wikipedia pages.
Approach: They propose to improve the model by combining Wikipedia references with a list of plausible candidate entities.
Outcome: The proposed method yields 16.9% in Top-30 gold candidate recall compared with state-of-the-art models.
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)

Copied to clipboard

Challenge: a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks.
Approach: They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community .
Outcome: The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges.
Beyond BLEU:Training Neural Machine Translation with Semantic Similarity (P19-1)

Copied to clipboard

Challenge: Recent work has shown that optimizing neural machine translation systems to directly improve evaluation metrics such as BLEU can improve final translation accuracy.
Approach: They propose a reward function that assigns partial credit to BLEU and provides more diversity in scores than BLUE.
Outcome: The proposed reward function improves translation accuracy, semantic similarity, and human evaluation on four languages trans-lated to English and the optimization procedure converges faster.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations