Papers by Reza Haf

14 papers
Systematic Assessment of Factual Knowledge in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing question-answering benchmarks for large language models have limitations regarding factual knowledge coverage, as they focus on generic domains and overlap with pretraining data.
Approach: They propose a framework to assess the factual knowledge of large language models by leveraging knowledge graphs.
Outcome: The proposed framework generates questions and expected answers from the facts stored in a given knowledge graph and evaluates them with KGs in generic and specific domains.
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach (2024.findings-eacl)

Copied to clipboard

Challenge: Large pre-trained language models such as GPT-3.5 and GPT-4 have gained significant attention in natural language research due to limited computational resources or inaccessible parameters.
Approach: They propose a neural programmer-interpreter approach that preserves the domain generalization ability of LLMs while editing their output.
Outcome: The proposed framework significantly improves GPT-3.5’s performance in logical form-to-text conversion and low-resource machine translation, surpassing other state-of-the-art (SOTA) LLM post-editing methods in cross-domain settings.
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights (2024.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models have demonstrated a strong capability to understand multimodal information and interact with human users.
Approach: They propose a speech-specific risk taxonomy to assess LMMs' ability to detect high-risk interactions in multimodal settings.
Outcome: The proposed model is based on a speech-specific risk taxonomy covering 8 risk categories . it shows that the models are ineffective in detecting paralinguistic-specific risks in speech .
DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions.
Approach: They propose a methodology to study the soundness of Social-IQ by applying simple perturbations to a dataset of multiple choice questions on videos of complex social interactions.
Outcome: The proposed method reduces biases in the original dataset and improves performance.
Fire Burns, Sword Cuts: Commonsense Inductive Bias for Exploration in Text-based Games (2022.acl-short)

Copied to clipboard

Challenge: Existing RL agents are far away from solving text-based games due to their combinatorially large action spaces that hinders efficient exploration.
Approach: They propose an exploration technique that injects external commonsense knowledge, via a pretrained language model, into the agent during training when the agent is the most uncertain about its next action.
Outcome: The proposed method exhibits improvement on the collected game scores during the training in four out of nine games from Jericho.
MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations (2024.acl-short)

Copied to clipboard

Challenge: a new problem setting is designed to detect critical moments in conversations . a human-annotated multi-modal dataset is used to classify and detect turning points .
Approach: They propose a problem setting focusing on turning points in conversations as TPs . they propose MTPC, MTPD, & MTPR tasks to classify and detect turning points .
Outcome: The proposed model achieves an F1-score of 0.88 in classification and 0.61 in detection . it uses state-of-the-art vision-language models to construct a narrative from the videos .
RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations (2024.findings-naacl)

Copied to clipboard

Challenge: Norm violations occur when individuals fail to conform to culturally accepted behaviors, which may lead to potential conflicts.
Approach: They propose to use a large corpus of 9,258 multi-turn dialogues annotated with social norms to equip AI systems with a remediation ability.
Outcome: The proposed system can understand and remediate norm violations step by step.
IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models (2024.acl-long)

Copied to clipboard

Challenge: IMO is a machine learning model that learns invariant features from unseen domains.
Approach: They propose IMO: Invariant features Masks for Out-of-Distribution text classification to achieve OOD generalization by learning invariant feature masks.
Outcome: The proposed model outperforms baseline models in various evaluation metrics and settings.
Causal Discovery Inspired Unsupervised Domain Adaptation for Emotion-Cause Pair Extraction (2024.findings-emnlp)

Copied to clipboard

Challenge: Emotion-cause pair extraction is a task that aims to extract emotions and the events causing such emotions.
Approach: They propose a deep latent model which captures the underlying latent structures of data and utilizes the easily transferable knowledge of emotions as the bridge to link the distributions of events in different domains.
Outcome: The proposed model outperforms the strongest baseline by approximately 11.05% on a Chinese benchmark and 2.45% on an English benchmark in terms of weighted-average F1 score.
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models are fine-tuned on diverse datasets to develop a range of skills . each skill has unique characteristics, and datasets are heterogeneous and imbalanced . a general, model-agnostic, reinforcement learning framework is proposed to optimize data usage .
Approach: They propose a general, model-agnostic, reinforcement learning framework that optimizes data usage automatically during the fine-tuning process.
Outcome: The proposed framework optimizes data usage automatically during the fine-tuning process.
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific domains remain challenging.
Approach: They propose a task that focuses on transcribing scientific conference videos by leveraging visual information from slides to enhance the accuracy of technical terminologies.
Outcome: The proposed framework improves transcript quality through post-editing and improves performance over speech-only baselines.
Let’s Negotiate! A Survey of Negotiation Dialogue Systems (2024.findings-eacl)

Copied to clipboard

Challenge: Recent research has focused on negotiation dialogue systems, but no systematic review of this task has been conducted.
Approach: They propose to provide a systematic review of negotiation dialogue systems and to provide an overview of current research.
Outcome: The proposed systems are based on the literature and are compared against existing systems.
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that virtual agents can help humans achieve task and social goals.
Approach: They propose a tuning-free and label-free method to identify high-quality ICL exemplars for the remediator agent and propose measurable criteria to measure the quality of the negotiation outcomes.
Outcome: The proposed model is able to improve negotiation outcomes across three negotiation topics.
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) have shown impressive generalization ability on vision and language tasks, but their spatial understanding is under-explored.
Approach: They construct a VQA dataset to analyze LMMs' spatial reasoning capabilities.
Outcome: The proposed model is stronger at basic object detection than complex spatial reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations