Papers by Gaurav Verma

11 papers
MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms (2024.findings-acl)

Copied to clipboard

Challenge: Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces.
Approach: They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance.
Outcome: The proposed model performs better in a zero-shot setting, suggesting potential improvements.
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations (2025.acl-long)

Copied to clipboard

Challenge: State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs).
Approach: They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability.
Outcome: The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations.
Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to model multimodal data do not leverage cross-modal information . augmenting input text using cross-module attribute insertions results in poor performance .
Approach: They propose a multimodal deep learning approach that adds visual attributes to inputs to enhance model robustness.
Outcome: The proposed approach is modular, controllable, and task-agnostic.
Learning the Visualness of Text Using Large Vision-Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Visual text evokes an image in a person’s mind, while non-visual text fails to do so.
Approach: They propose a method to automatically detect visualness in text to enable text-to-image retrieval and generation models to augment text with relevant images.
Outcome: The proposed method performs better than several baseline models and heuristics for the task.
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use (2025.naacl-long)

Copied to clipboard

Challenge: Adverse Drug Reactions (ADRs) from psychiatric medications are the leading cause of hospitalizations among mental health patients.
Approach: They propose a benchmark and a framework to evaluate LLMs' ability to detect ADRs . they find that LLM responses are more complex and harder to read than experts .
Outcome: The proposed framework evaluates LLMs' ability to detect and deliver expert-aligned mitigation strategies.
Cross-Modal Projection in Multimodal LLMs Doesn’t Really Project Visual Attributes to Textual Space (2024.acl-short)

Copied to clipboard

Challenge: Existing multimodal large language models are limited to general-purpose multimodal tasks like question-answering on natural images.
Approach: They propose to use cross-modal projection networks and a large language model to model domain-specific visual attributes of MLLMs.
Outcome: The proposed models gain domain-specific visual capabilities when the projection is fine-tuned, but the updates do not extract relevant domain-specific visual attributes.
Incorporating Stylistic Lexical Preferences in Generative Language Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in language modeling have resulted in powerful generation models, but their style is implicitly dependent on the training data and cannot emulate a specific target style.
Approach: They propose an approach to induce certain target-author attributes by incorporating continuous multi-dimensional lexical preferences of an author into generative language models.
Outcome: The proposed model generates text that aligns with a given target author’s lexical style and is competitive with baselines.
Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Recent few-shot learning methods focus on improving downstream task performance, but there is limited understanding of the adversarial robustness of such methods.
Approach: They evaluate prompt-based FSL methods against fully fine-tuned models to better understand the impact of various factors towards robustness.
Outcome: The proposed methods show that they are less robust in the face of adversarial perturbations than fully fine-tuned models.
DRAG: Director-Generator Language Modelling Framework for Non-Parallel Author Stylized Rewriting (2021.eacl-main)

Copied to clipboard

Challenge: Recent work in this area has focused on author stylized rewriting but is limited by the lack of explicit control of target attributes and being data-driven.
Approach: They propose a Director-Generator framework to rewrite input text in the target author’s style, specifically focusing on certain target attributes.
Outcome: The proposed framework has better meaning retention and results in more fluent generations on a small corpus of text authored by three distinct authors.
A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech (2024.acl-long)

Copied to clipboard

Challenge: Using data from 420k Twitter posts, we characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Approach: They develop a codebook to characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Outcome: The proposed codebook analyzes 420k tweets over 3 years and compares classifiers with hateful speech classifier classifier to detect hateful content.
Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks.
Approach: They develop a model that generates additional dilution text that maintains relevance and topical coherence with the image and existing text, and when added to the original text, leads to misclassification of the multimodal input.
Outcome: The proposed model outperforms fusion-based classifiers on Crisis Humanitarianism and Sentiment Detection tasks by 23.3% and 22.5% in presence of dilutions generated by the model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations