Challenge: Mobile Agents are a key component of the “Agentic Economy” where they perform high-stakes financial transactions.
Approach: They propose a systemic vulnerability termed Visual Dominance Hallucination (VDH) VDH exploits the modality gap in CLIP-based encoders via a novel Semantic-Decoupling Loss.
Outcome: The proposed framework exploits the modality gap in CLIP-based encoders by preserving fidelity.

Similar Papers

Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization (2026.findings-eacl)

Copied to clipboard

Challenge: Existing white-box jailbreak methods require full model accessibility and require computational costs.
Approach: They propose a black-box jailbreak attack using Zeroth-Order optimization using ZO-SPSA.
Outcome: The proposed method achieves highest jailbreak success rate on three LVLMs, including InstructBLIP, LLaVA and MiniGPT-4.
Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) are increasingly being deployed as content moderators . however, they exploit the Human-AI capability gap and create adversarial environments . smuggling attacks exploit the human-AI gap and exploit the vulnerability .
Approach: They construct a benchmark to evaluate the vulnerability of MLLMs as content moderators . they identify three root causes: limited capabilities of vision encoders, robustness gap in OCR .
Outcome: The proposed model exploits the Human-AI capability gap and is vulnerable to smuggling attacks.
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects (2026.findings-acl)

Copied to clipboard

Challenge: Large Vision-Language Models are hindered by a systemic efficiency barrier known as visual token dominance.
Approach: They propose a systematic taxonomy of efficiency techniques structured around the inference lifecycle . they examine visual encoding, prefilling, and decoding to understand bottlenecks .
Outcome: The proposed techniques reveal how upstream decisions dictate downstream bottlenecks . the proposed techniques include hybrid compression and modality-aware decoding .
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation (2026.acl-long)

Copied to clipboard

Challenge: Search agents are a pivotal paradigm for solving open-ended, knowledge-intensive reasoning tasks.
Approach: They propose a search agent simulation environment that bootstraps robust search agents using Reinforcement Learning.
Outcome: The proposed model outperforms the web-enhanced ASearcher model by 10.6%.
MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs (2026.acl-long)

Copied to clipboard

Challenge: MLLMs assume linguistic context invariably enhances visual understanding . a diagnostic benchmark is used to evaluate ML models under hierarchical linguistic interference .
Approach: They propose a diagnostic benchmark to evaluate MLLMs under hierarchical linguistic interference.
Outcome: The proposed benchmark compared 402 videos with a physical constraint set to evaluate MLLMs under hierarchical linguistic interference.
Adaptive Weighted Proxy Tuning: Efficient Gray-Box Steering for Image Captioning. (2026.acl-industry)

Copied to clipboard

Challenge: Proxy tuning is a decoding-time approach that fails to account for instance-specific variations in model certainty and domain shift.
Approach: They propose a gray-box steering framework that dynamically modulates the logit contributions of a large base model, a fine-tuned expert, and an untune .
Outcome: Adaptive Weighted Proxy Tuning achieves performance parity with fine-tuned models while remaining parameter-free.
Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning (2026.acl-long)

Copied to clipboard

Challenge: Existing research focuses on prompt injection or immediate adversarial misclassification of user-uploaded images.
Approach: They propose a dual-process defense framework inspired by human cognition to mitigate this vulnerability by injecting triggers into user-uploaded images that act as "sleeper agents"
Outcome: The proposed framework achieves about 85% Goal-Hit Rate (GHR) while reducing the risk to 10% with configurable latency trade-offs.
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to jailbreak Large Vision Language Models do not consider interaction between images and text.
Approach: They propose a prior-guided bimodal interactive black-box jailbreak attack for toxicity maximization that exploits the interaction of images and text.
Outcome: The proposed method outperforms state-of-the-art jailbreak methods in black box scenarios and in closed-source LVLMs.
GOBench: Stage-Wise Diagnostics and the Visual Paradox in Multimodal Graph Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks fail to represent multimodal problem specifications, score outcomes only and cannot localize where failures occur along the modeling pipeline.
Approach: They propose a Graph Optimization benchmark that aligns multiple modalities with solver-derived oracles and a diagnostic protocol that evaluates intermediate artifacts as well as end results.
Outcome: Graph Optimization benchmark (GOBench) evaluates intermediate artifacts as well as end results . vision reliably increases inference cost, while reliability impact is regime-dependent . current benchmarks fail to represent multimodal problem specifications, fail to localize failures .
TokenPenalty: Alleviating Attention Sinks and Positional Decay in LVLMs (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) often hallucinate due to two relevant phenomena: massive activation phenomenon and positional information decay.
Approach: They propose a token-level intervention strategy that dynamically suppresses irrelevant visual tokens while preserving key contextual signals.
Outcome: Experiments show that TokenTruth significantly improves factual consistency across MLLMs on standard image understanding benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations