Papers by Dinesh Agarwal

5 papers
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses.
Approach: They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos.
Outcome: EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos.
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing research explores the use of Large Language Models (LLMs) as the backbone of agentic systems.
Approach: They propose a model trained using a multi-task training approach on seven fundamental tasks encompassed in function calling that has better generalizability on multiple tasks across seven evaluation benchmarks.
Outcome: The proposed model outperforms more than 15 other models on out-of-domain datasets and ranks among the top on the Berkeley Function Calling Leaderboard (BFCL).
MedCLIP: Contrastive Learning from Unpaired Medical Images and Text (2022.emnlp-main)

Copied to clipboard

Challenge: Existing vision-text contrastive learning methods encounter many false negatives, i.e., images and reports from separate patients probably carry the same semantics but are wrongly treated as negatives.
Approach: They propose to decouple medical image-text contrastive learning and replace it with semantic matching loss based on medical knowledge to eliminate false negatives in contrastive training.
Outcome: The proposed framework outperforms state-of-the-art methods on zero-shot prediction, supervised classification, and image-text retrieval with only 20K pre-training data.
Saliency-Aware Interpolative Augmentation for Multimodal Financial Prediction (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in the Financial AI realm have expanded the scope of data and methods they use, such as textual and audio cues from financial earnings calls, but limitations exist.
Approach: They propose a Saliency-guided Hierarchical Mixup augmentation technique for multimodal financial prediction tasks.
Outcome: The proposed technique outperforms state-of-the-art methods by 3-7% on financial earnings and conference call datasets.
End-to-End Learning of Flowchart Grounded Task-Oriented Dialogs (2021.emnlp-main)

Copied to clipboard

Challenge: Existing systems that use human-to-human dialogs to help users with specific tasks are still unexplored.
Approach: They propose a problem in which a dialog system mimics a troubleshooting agent . they use a dataset grounded on 12 different troubleshooking flowcharts to train the agent a neural model .
Outcome: The proposed model can do zero-shot transfer to unseen flowcharts and sets a strong baseline for future research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations