Challenge: We propose a scalable decoding strategy for conversation infilling . large pretrained language models are effective solutions to many popular natural language generation tasks such as machine translation and conversational dialogue.
Approach: They propose a heuristic guided lookahead decoding strategy for conversation infilling which leverages a greedy lookalike phase before committing to any token.
Outcome: The proposed strategy outperforms baselines when evaluated with automatic and human evaluation metrics, which, we argue, are appropriate for the task.

Similar Papers

ConvFiT: Conversational Fine-Tuning of Pretrained Language Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing Transformer-based language models (LMs) are not effective as sentence encoders when used off-the-shelf.
Approach: They propose a method which turns a pretrained LM into a universal conversational encoder and task-specialised sentence encoder.
Outcome: The proposed framework achieves state-of-the-art ID performance across the board with particular gains in the most challenging, few-shot setups.
Enabling Language Models to Fill in the Blanks (2020.acl-main)

Copied to clipboard

Challenge: Infilling is the task of predicting missing spans of text at any position in a document.
Approach: They propose a framework which can be used to infill entire sentences . they train off-the-shelf LMs on sequences containing concatenation of masked text .
Outcome: The proposed approach can infill entire sentences on short stories, scientific abstracts, and lyrics.
ChatUIE: Exploring Chat-based Unified Information Extraction Using Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have shown impressive performance in general chat, but their domain-specific capabilities have certain limitations.
Approach: They propose a unified information extraction framework built upon ChatGLM that incorporates domain-specific modeling to extract structured information from natural language.
Outcome: The proposed framework significantly improves the performance of information extraction tasks with a slight decrease in chatting ability.
Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation (2024.findings-acl)

Copied to clipboard

Challenge: Stochastic sampling strategies are not widely used in open-domain dialogue systems.
Approach: They propose a dynamic decoding strategy which can adjust the decoding space w.r.t. different contexts.
Outcome: The proposed decoding strategy can improve the performance of pre-trained models when coupled with four well-used stochastic decoding algorithms.
Better Conversations by Modeling, Filtering, and Optimizing for Coherence and Diversity (D18-1)

Copied to clipboard

Challenge: Existing encoder-decoder models for open domain dialogue generate generic, uninformative, and non-coherent responses.
Approach: They propose to introduce a measure of coherence as the GloVe embedding similarity between dialogue context and generated response to improve output diversity.
Outcome: The proposed model improves on the OpenSubtitles corpus in terms of BLEU score and diversity metrics.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation (2022.acl-long)

Copied to clipboard

Challenge: Despite recent improvements in open-domain dialogue models, state of the art models are trained and evaluated on short conversations with little context.
Approach: They propose to use retrieval-augmented methods to summarize and recall past conversations to improve their models.
Outcome: The proposed models outperform the current state-of-the-art models on human-human chat sessions in both automatic and human evaluations.
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection (2025.acl-long)

Copied to clipboard

Challenge: Existing work has been using decoding-free candidate selection methods to obtain candidate probability from initial output logits over vocabulary.
Approach: They propose to evaluate a set of tasks using decoding-free candidate selection methods on a comprehensive set of questions.
Outcome: The proposed methods are evaluated on a set of tasks including five multiple-choice QA tasks with a small candidate pool and four clinical decision tasks with 10k+ options.
Thoughts to Target: Enhance Planning for Target-driven Conversation (2024.emnlp-main)

Copied to clipboard

Challenge: Empirical results demonstrate that our method significantly improves the planning ability of LLMs, especially in target-driven conversations.
Approach: They propose a two-stage framework to improve the LLMs’ capability in planning conversations towards designated targets by distilling natural language plans from a target-driven conversation corpus and generating new plans with demonstration-guided in-context learning.
Outcome: The proposed framework improves the ability of conversational models to plan towards designated targets and can be used to build extensive conversational AI.
Online Semantic Parsing for Latency Reduction in Task-Oriented Dialogue (2022.acl-long)

Copied to clipboard

Challenge: Standard conversational semantic parsing maps a user's intent into an executable program, but execution is slow when expensive function calls are included.
Approach: They propose a task of online semantic parsing to predict and execute function calls while the user is still speaking.
Outcome: The proposed approach reduces latency with good parsing quality and execution cost.
HiURE: Hierarchical Exemplar Contrastive Learning for Unsupervised Relation Extraction (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods to extract relational feature signals from natural language sentences use self-supervised clustering and classification that cause gradual drift problems.
Approach: They propose a framework that derives hierarchical signals from relational feature space using cross hierarchy attention and effectively optimizes relation representation of sentences under exemplar-wise contrastive learning.
Outcome: The proposed framework can extract the relationship between entities from natural language sentences without prior knowledge on relation scope or distribution.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations