Papers by Yonatan Bisk

24 papers
VISREAS: Complex Visual Reasoning with Unanswerable Questions (2024.findings-acl)

Copied to clipboard

Challenge: Logic2Vision is a visual question-answering dataset that validates question authenticity with the corresponding image and then reasoning over it.
Approach: They propose a compositional visual question-answering dataset, VisReas, that consists of answerable and unanswerable visual queries . they use visual genome scene graphs to generate the query and the reasoning steps to generate it.
Outcome: The proposed model outperforms generative models and the existing classification models and outperformed existing models.
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)

Copied to clipboard

Challenge: Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning.
Approach: They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks.
Outcome: The proposed model types do not consistently improve self-rationalization in multimodal tasks.
Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games (2020.coling-main)

Copied to clipboard

Challenge: Existing models fail to learn multi-modal representations, relying on category labels at inference time.
Approach: They propose a "imagination" module that learns context-aware and category-awful latent embeddings without relying on category labels at inference time.
Outcome: The imagination module outperforms state-of-the-art competitors by 8.26% gameplay accuracy in the CompGuessWhat?! benchmark.
EvEntS ReaLM: Event Reasoning of Entity States via Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model event implications fail to reason about the world, despite their knowledge of physical attributes.
Approach: They propose to use a model prompting technique to prompt models of event implications by targeting their understanding of physical attributes.
Outcome: The proposed model prompting technique is especially useful for unseen attributes or when only limited data is available.
SOTOPIA-π: Interactive Learning of Socially Intelligent Language Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on building language agents have not addressed this social learning gap.
Approach: They propose an interactive learning method that improves the social intelligence of language agents by using behavior cloning and self-reinforcement based training on filtered social interaction data.
Outcome: The proposed method allows a 7B LLM to reach the social goal completion ability of an expert model (GPT-4-based agent) without the loss of more generic abilities, such as the ability to answer knowledge-based questions.
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a growing influx of misinformation across news and social media is hampered by outdated foundation model training data.
Approach: They propose to use large language models to scale up online policing mechanisms . they evaluate foundation model performance without continual updating .
Outcome: The proposed model can improve performance without continual updating . the proposed model improves on two widely used benchmarks .
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference (D18-1)

Copied to clipboard

Challenge: a new dataset presents a task of grounded commonsense inference, unifying natural language inference and commonsensical reasoning.
Approach: They propose a procedure that constructs a de-biased dataset by iteratively training stylistic classifiers and using them to filter the data.
Outcome: The proposed procedure oversamples a de-biased dataset using state-of-the-art language models . human models struggle on the proposed procedure, indicating significant opportunities for future research.
HellaSwag: Can a Machine Really Finish Your Sentence? (P19-1)

Copied to clipboard

Challenge: Existing commonsense models struggle to perform inferences that are trivial for humans, but are often misclassified by state-of-the-art models.
Approach: They propose a dataset that is adversarial to state-of-the-art commonsense reasoning and use it to build a model that is surprisingly robust.
Outcome: The proposed dataset is compared with existing models and scaled up towards a critical 'Goldilocks zone' wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models.
The Framework Tax: Disparities Between Inference Efficiency in NLP Research and Deployment (2023.emnlp-main)

Copied to clipboard

Challenge: Inference is estimated to make up 80 to 90% of ML cloud computing demand .
Approach: They propose to identify bottlenecks in deep learning frameworks that are causing the disparity in model latency as hardware speed increases over time.
Outcome: The proposed models show that the framework tax is increasing as the hardware speed increases over time.
Experience Grounds Language (2020.emnlp-main)

Copied to clipboard

Challenge: aaron carroll: language understanding research is held back by a failure to relate language to the physical world it describes and to social interactions it facilitates. carroll says successful linguistic communication relies on a shared experience of the world.
Approach: They propose to use a broader physical and social context to address communication problems . they argue that the current success of representation learning approaches is limited .
Outcome: a new study suggests that the current success of representation learning requires a parallel tradition of research on the broader physical and social context of language to address the deeper questions of communication.
Tools Fail: Detecting Silent Errors in Faulty Tools (2024.emnlp-main)

Copied to clipboard

Challenge: a failure in one tool can trigger a cascade of errors, leading to complete task failure.
Approach: They propose a framework for tools more broadly which explores a model’s ability to detect “silent” tool errors and reflect on how to plan.
Outcome: The proposed approach shows that the model can detect "silent" tool errors and plan.
Don’t Copy the Teacher: Data and Model Challenges in Embodied Dialogue (2022.emnlp-main)

Copied to clipboard

Challenge: Embodied dialogue instruction following requires an agent to complete a complex sequence of tasks from a natural language exchange.
Approach: They argue that imitation learning and low-level metrics are misleading . they compare existing models with IL and argue evaluation should focus on higher-level semantic goals .
Outcome: The proposed model evaluations are based on three models and compare them with benchmarks . they show that existing models fail to ground query utterances, which are essential for task completion .
MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Correction (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown growing potential in molecular sciences, but they often produce chemically inaccurate descriptions and struggle to recognize or justify potential errors.
Approach: They propose a benchmark to assess LLMs on error detection and correction in molecular descriptions.
Outcome: The proposed benchmark targets LLMs on error detection and correction in molecular descriptions.
RMM: A Recursive Mental Model for Dialogue Navigation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing work on language-guided robots focuses on the latter, but little attention is paid to the guiding agent.
Approach: They propose a two-agent task where one agent navigates and asks questions that a second, guiding agent answers.
Outcome: The proposed model can be generalized to novel environments.
Grounding ‘Grounding’ in NLP (2021.findings-acl)

Copied to clipboard

Challenge: Cognitive Science defines "grounding" as the process of establishing mutual information between two interlocutors.
Approach: They examine the gaps between NLP and Cognitive Science definitions of "grounding" they propose ways to create new tasks or repurpose existing ones to achieve a more complete sense of grounding .
Outcome: The authors examine the gaps between definitions of grounding and cognitive science . they show that there are ways to improve existing tasks or repurpose existing ones .
Shifting the Baseline: Single Modality Performance on Visual Navigation & QA (N19-1)

Copied to clipboard

Challenge: Existing work on unimodal approaches often lacks dataset biases . we present unimod ablations on three recent datasets in visual navigation and QA .
Approach: They propose unimodal ablations for visual navigation and QA using egocentric vision . they argue that unimodulated models better capture and reflect dataset biases .
Outcome: The proposed models outperform full models on visual navigation and QA tasks with language only on three recent datasets.
KAT: A Knowledge Augmented Transformer for Vision-and-Language (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for knowledge retrieval and answer prediction have left open questions about the quality and relevance of the retrieved knowledge and how the reasoning processes over implicit and explicit knowledge should be integrated.
Approach: They propose a Knowledge Augmented Transformer which integrates both implicit and explicit knowledge in an encoder-decoder architecture while simultaneously reasoning over both knowledge sources during answer generation.
Outcome: The proposed model achieves a strong state-of-the-art (+6% absolute) on the open-domain multimodal task of OK-VQA.
Gradient Localization Improves Lifelong Pretraining of Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for continual learning do not account for locality of knowledge . however, in practice language models are deployed in dynamic real-world settings and their learned knowledge becomes stale over time.
Approach: They examine two types of knowledge relating to temporally sensitive entities . they hypothesize that lack of consideration of locality contributes to failed uptake of new information .
Outcome: The proposed model can be improved by updating parameters to relevant layers . the proposed model is based on a large static web-scale dataset .
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies have demonstrated that direct preference optimization (DPO) can be effective in generalizing large language models, but its effectiveness in video domain remains limited.
Approach: They propose a framework that utilizes detailed video captions as a proxy of video content to enable language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions.
Outcome: The proposed framework shows that it can be used to align language models with video content and improves performance on open-ended video QA tasks.
The Return of Lexical Dependencies: Neural Lexicalized PCFGs (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to grammar induction focus on discovering constituents or dependencies.
Approach: They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies .
Outcome: The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods.
Benchmarking Hierarchical Script Knowledge (N19-1)

Copied to clipboard

Challenge: Understanding procedural language requires reasoning about hierarchical and temporal relations between events.
Approach: They propose a hierarchical script learning dataset and a cloze task to match video captions with missing procedural details.
Outcome: The proposed model matches video captions with missing procedural details to find out if they can understand the language.
Energy Considerations of Large Language Model Inference and Efficiency Optimizations (2025.acl-long)

Copied to clipboard

Challenge: Prior benchmarking efforts focused on latency reduction in idealized settings, often overlooking real-world inference workloads that shape energy use.
Approach: They propose a modeling approach that approximates real-world LLM workflows . they show that the effectiveness of inference optimizations is sensitive to workload geometry .
Outcome: The proposed approach reduces energy use by 73% from unoptimized baselines.
Robust Navigation with Language Pretraining and Stochastic Sampling (D19-1)

Copied to clipboard

Challenge: Existing methods to learn visual representations and action decoding schemes are limited to previously unseen instructions and environments.
Approach: They propose a stochastic sampling scheme to reduce the gap between the expert actions in training and sampled actions in test to correct its own mistakes.
Outcome: The proposed methods achieve 6% absolute gain over the previous best results on the Room-to-Room benchmark.
An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games (2021.eacl-main)

Copied to clipboard

Challenge: Using guessing games, an artificial agent can learn to perform on novel downstream tasks such as Visual Question Answering (VQA).
Approach: They propose a supervised learning scenario in which an agent learns to mimic successful guessing games and a novel way for an agent to play by itself, called Self-play via Iterated Experience Learning.
Outcome: The proposed model can be applied to a VQA dataset using a supervised learning scenario and a novel way for an agent to play by itself.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations