Papers with LVLMs
Copied to clipboard
| Challenge: | Recent research has focused on addressing multimodal hallucinations in Large Vision-Language Models (LVLMs) however, these methods lack fine-grained visual contrast mechanisms and rely on single-margin optimization. |
| Approach: | They propose a framework that integrates text-conditioned preference loss with visual ranking-based objective. |
| Outcome: | The proposed framework improves cross-modal alignment and fine-grained visual grounding. |
Copied to clipboard
| Challenge: | Existing methods for visual and language alignment depend on external models or data, leading to uncontrollable and unstable results. |
| Approach: | They propose a framework that enhances visual and language alignment without external dependencies by incorporating an in-context self-critic mechanism that constructs preference pairs for tuning. |
| Outcome: | The proposed framework outperforms existing methods and improves performance on 14 hallucination and comprehensive benchmarks. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) suffer from object hallucinations, i.e., they tend to generate objects inconsistent with the target images in the descriptions. |
| Approach: | They propose to integrate powerful large vision-language models (LVLMs) they propose a polling-based query method to evaluate object hallucination . |
| Outcome: | The proposed model can evaluate object hallucination in a more stable and flexible way. |
Copied to clipboard
| Challenge: | Large-scale scientific research on medieval Arabic manuscripts remains challenging due to the need for advanced paleographic and linguistic training and the lack of assisting software. |
| Approach: | They propose an end-to-end Arabic manuscript analysis tool for manuscript-based analytics and research hypothesis testing. |
| Outcome: | The proposed tool overcomes the limitations of existing tools and can be used in large-scale scientific research. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have transformed image captioning . existing evaluations lack standardized criteria and a standardized evaluation framework . |
| Approach: | They propose a leaderboard for evaluating detailed captions that addresses three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations. |
| Outcome: | The proposed model evaluates caption quality, descriptiveness, risks, and societal biases while tailoring criteria to user preferences. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are highly sensitive to end-of-input artifacts in fine-tuning and inference data, e.g., whether input sequences end with punctuation or newline characters. |
| Approach: | They propose to convert generative LVLMs into vision-language encoders via contrastive learning objectives and use supervised contrastive objectives to train them. |
| Outcome: | The proposed approach improves visual and text representations and improves retrieval and (semantic) similarity tasks. |
Copied to clipboard
| Challenge: | MM-RAG is a promising approach for enhancing the reliability and factuality of large vision-language models . current methods focus on component-level optimizations and necessitate extensive component-specific training datasets . |
| Approach: | They propose a new paradigm that backpropagates global rewards to each component . this backpropage transforms local losses into specific local losses . |
| Outcome: | The proposed paradigm achieves high training efficiency on knowledge-intensive multimodal benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to measure image common sense inconsistentness are difficult to implement because of their complexity. |
| Approach: | They propose a visual commonsense model that leverages large vision-language models to extract atomic facts from images and a compact attention-pooling classifier to fine-tune it over encoded atomic fact. |
| Outcome: | The proposed method outperforms existing methods on the WHOOPS! and WEIRD datasets while maintaining a compact attention-pooling classifier over encoded atomic facts. |
Copied to clipboard
| Challenge: | PatentVision integrates textual and visual inputs to generate patent specifications . existing systems fail to capture the nuanced interplay between textual, visual components . |
| Approach: | They propose a multimodal framework that integrates textual and visual inputs to generate patent specifications. |
| Outcome: | The proposed framework surpasses text-only methods in patent writing, the authors show . it integrates visual data to better represent intricate design features and functional connections . |
Copied to clipboard
| Challenge: | Lack of perceptual grounding limits vision-language models' ability to interpret visual data . prior work on visualized data understanding focused on adapting VLMs to instruction tuning and chain-of-thought supervision . |
| Approach: | They propose a framework that enhances visual reasoning through human-like interpretation grounding. |
| Outcome: | The proposed framework improves on ChartQA and ChartQAPro benchmarks by +11.2%. |
Copied to clipboard
| Challenge: | Existing methods to jailbreak Large Vision Language Models do not consider interaction between images and text. |
| Approach: | They propose a prior-guided bimodal interactive black-box jailbreak attack for toxicity maximization that exploits the interaction of images and text. |
| Outcome: | The proposed method outperforms state-of-the-art jailbreak methods in black box scenarios and in closed-source LVLMs. |
Copied to clipboard
| Challenge: | Recent studies have revealed significant deficiencies of LVLMs in understanding visual contents, leaving the gap between current embodied intelligence and large vision-language models (LVLM) . |
| Approach: | They propose to use a benchmark to evaluate LVLMs' spatial understanding of embodied environments to evaluate their ability to understand visual contents. |
| Outcome: | The proposed benchmark is derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Large Vision Language Model (LVLMs) exhibit advanced proficiency in language reasoning and comprehension across a wide array of languages. |
| Approach: | They propose to use a multitask language understanding benchmark specifically designed for the Malay language to assess their proficiency. |
| Outcome: | The proposed model performs well in well-resourced languages, but in low-resource languages such as Bahasa Melayu, they are less studied due to a lack of studies and benchmarks. |
Copied to clipboard
| Challenge: | Multimodal in-context learning (ICL) is a key mechanism for harnessing the capabilities of large vision–language models. |
| Approach: | They propose a transformer-based model with task-aware attention that dynamically configures ICL sequences. |
| Outcome: | Experiments on five LVLMs and nine datasets show that TACO surpasses baselines across diverse ICL tasks. |
Copied to clipboard
| Challenge: | a large number of e-commerce platforms require manual verification and specialized hardware. |
| Approach: | They propose a multimodal weight estimation framework that uses category-specific exemplars to infer discretized weight buckets. |
| Outcome: | The proposed approach outperforms strong multimodal KNN baselines in accuracy and near-bucket reliability. |
Copied to clipboard
| Challenge: | a new wave of large vision–language models (LVLMs) incorporate images as input in addition to text . a recent study examined potential gender and racial biases in such systems based on the perceived characteristics of the people in the input images. |
| Approach: | They examine potential gender and racial biases in large vision–language models . they query a dataset of AI-generated images of people to see whether they differ . |
| Outcome: | The proposed dataset shows that the images differ in gender and race according to the perceived characteristics of the person depicted. |
Copied to clipboard
| Challenge: | LVLMs are known for producing text that is factually inconsistent with visual input . factuality of generated captions for structured visuals has not been studied as much . |
| Approach: | They propose a typology of factual errors in captions generated by large vision-language models . they propose CHOCOLATE, a visual entailment model that outperforms current models based on this analysis . |
| Outcome: | The proposed model outperforms current models in evaluating caption factuality. |
Copied to clipboard
| Challenge: | Existing methods to mitigate object hallucination are impractical for proprietary LVLMs. |
| Approach: | They propose a framework to identify optimal visual prompts that enhance LVLM responses without access to model internals. |
| Outcome: | The proposed approach is model-agnostic and can be used on open-source and proprietary LVLMs. |
Copied to clipboard
| Challenge: | Existing document understanding benchmarks only handle a small number of pages . existing models are limited to handling only a limited number of documents . |
| Approach: | They propose a long document understanding benchmark that integrates three primary tasks and 20 sub-tasks based on different primary tasks. |
| Outcome: | The proposed model outperforms existing benchmarks on open-source and closed-source models . the model outpersforms other models on more than 33,000 pages of documents . |
Copied to clipboard
| Challenge: | Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating advanced capabilities in text generation and comprehension. |
| Approach: | They propose to use artwork explanation generation task to quantitatively assess the understanding and utilization of artworks knowledge. |
| Outcome: | The proposed task evaluates the understanding and utilization of knowledge about artworks from images and titles and generates explanations using only images. |
Copied to clipboard
| Challenge: | Existing methods for video temporal grounding suffer from limited temporal awareness and poor generalization. |
| Approach: | They propose a two-stage training framework that integrates supervised fine-tuning with reinforcement learning to improve both the accuracy and robustness of VTG models. |
| Outcome: | The proposed training framework outperforms existing models on multiple benchmarks on open-domain and challenging scenarios. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) perform outstandingly across multimodal tasks, but training them with preference data is computationally expensive. |
| Approach: | They propose to merge text-based reward models with LVLMs to create visionlanguage reward models (VLRMs) this approach offers an efficient method for incorporating textual preferences into LVRMs. |
| Outcome: | The proposed model improves over LVLMs’ scoring and text-based RMs, and offers an efficient method for incorporating textual preferences into LVRMs. |
Copied to clipboard
| Challenge: | Multimodal machine translation (MMT) aims to leverage additional modalities beyond text . current MMT systems rely heavily on monolingual English captioning data . |
| Approach: | They propose a reasoning-based framework to leverage large-scale vision-language models for MMT . they propose Detect, Disambiguate, and Translate framework to detect ambiguity in input sentence . |
| Outcome: | The proposed framework outperforms state-of-the-art models in disambiguation accuracy and translation quality. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are expensive and time-consuming to evaluate . however, they are limited in their use in industrial settings due to their limited availability and limited resources. |
| Approach: | They evaluate 13 open-source LVLMs as judges for diverse chart comprehension and reasoning tasks. |
| Outcome: | The proposed models can be used to assess chart comprehension and reasoning tasks, but they are expensive and time-consuming. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models (LLMs) have impressive capabilities in dealing with new tasks with the help of in-context learning (ICL). |
| Approach: | They propose to concate the image and text embeddings to enhance the retrieval performance of a visual-language task and to calculate a list-wise ranking loss for training the embeddable model. |
| Outcome: | The proposed framework fine-tunes the CLIP embedding model to better meet the needs of the large vision-language models. |
Copied to clipboard
| Challenge: | LVLMs have been shown to perform well on simple uni-modal benchmarks, but their detailed study on multi-modal models is still lacking. |
| Approach: | They propose a framework to analyze the impact of compression on LVLMs on multi-modal input driven tasks. |
| Outcome: | The proposed framework analyzes the impact of compression on generative performance of large vision language models on multi-modal input driven tasks. |
Copied to clipboard
| Challenge: | Proxy tuning is a decoding-time approach that fails to account for instance-specific variations in model certainty and domain shift. |
| Approach: | They propose a gray-box steering framework that dynamically modulates the logit contributions of a large base model, a fine-tuned expert, and an untune . |
| Outcome: | Adaptive Weighted Proxy Tuning achieves performance parity with fine-tuned models while remaining parameter-free. |
Copied to clipboard
| Challenge: | Recent advances in large vision-language models produce hallucinations that compromise output reliability. |
| Approach: | They propose a dual-stage framework for mitigating hallucinations without performance degradation . they propose semantic-aware component disentanglement and interpretable parameter updates . |
| Outcome: | The proposed model reduces hallucinations by 23.4% while maintaining 97.4% of general generative capability. |
Copied to clipboard
| Challenge: | Large Vision-Language Models have demonstrated remarkable capabilities in processing both visual and textual information. |
| Approach: | They examine the challenge of alignment and misalignment in LVLMs through an explainability lens. |
| Outcome: | The findings highlight the need for standardized evaluation protocols and in-depth explainability studies. |
Copied to clipboard
| Challenge: | Large Vision Language Models suffer from hallucinations, attributing incorrect or misleading features to images. |
| Approach: | They propose a test-time approach that recalibrates the influence of blind tokens . they identify blind token by analyzing layer-wise attention distributions over image tokens. |
| Outcome: | The proposed approach reduces hallucinations in large vision language models . it uses a contrastive decoding strategy to balance the influence of blind tokens . |
Copied to clipboard
| Challenge: | Fine-grained image classification is a challenge for vision-language models (VLMs) such as CLIP, which struggle to distinguish between semantically similar classes due to insufficient supervision for fine-grain tasks. |
| Approach: | They propose a framework that harnesses the complementary strengths of both CLIP-like and LVLMs to tackle these challenges. |
| Outcome: | The proposed framework outperforms existing models on multiple fine-grained datasets, particularly the Stanford Cars dataset. |
Copied to clipboard
| Challenge: | Existing work detects hallucination by directly judging whether an object exists in an image, overlooking the association between the object and semantics. |
| Approach: | They propose a framework that incorporates hallucination feedback at both object and sentence semantic levels to alleviate over 15% of hallucinism. |
| Outcome: | The proposed framework can alleviate over 15% of hallucination even with a marginal degree of training. |
Copied to clipboard
| Challenge: | Existing evaluation methods focus on object hallucinations, focusing on object outputs . current evaluation methods struggle to address subtle semantic distinctions between outputs and reference data . |
| Approach: | They propose a multi-dimensional benchmark covering objects, attributes, and relations . they propose metric that generalizes CHAIR metric and incorporates faithfulness and coverage . |
| Outcome: | The proposed evaluation framework is more comprehensive and better correlated with humans than existing evaluation methods. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) have unlocked many complex use cases that require Multi-Modal (MM) understanding and MM generation. |
| Approach: | They propose a plug-and-play technique that adds relevant retrieved information to prompts as few-shot examples during inference. |
| Outcome: | The proposed method significantly improves the output quality of large vision language models when input prompts are augmented with relevant information retrieved by Vision-Language retrievers like UniRAG. |
Copied to clipboard
| Challenge: | Large visionlanguage models (LVLMs) are a powerful visual-language reasoning tool. |
| Approach: | They propose to integrate attention analysis with LLaVA-CAM to determine interactions between visual representations. |
| Outcome: | The proposed approach can be used to determine interactions between visual representations. |
Copied to clipboard
| Challenge: | EmoGist is a training-free, in-context learning method for visual emotion classification . context-dependent definitions of emotion labels could allow more accurate predictions of emotions . |
| Approach: | They introduce EmoGist, a training-free, in-context learning method for performing visual emotion classification with LVLMs. |
| Outcome: | The proposed method improves micro F1 scores and macro F1 with LVLMs. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) with only 7B parameters perform poorly as judges in resource-constrained settings. |
| Approach: | They propose two approaches to ensure costefficient evaluation by combining multiple criteria into a single query and domainadaptive transfer learning to create a 2Bparameter VLM on a chart dataset. |
| Outcome: | The proposed model can effectively transfer knowledge from one dataset to another to make it a more specialized model. |
Copied to clipboard
| Challenge: | Existing benchmarks for large vision-language models (LVLMs) are limited to ophthalmology-specific applications. |
| Approach: | They introduce a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across five ocular imaging modalities and 13 state-of-the-art LVLM representatives from closed-source, open-source and medical domains. |
| Outcome: | The proposed model shows significant performance drop in ophthalmology compared to other domains. |
Copied to clipboard
| Challenge: | Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language. |
| Approach: | They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data . |
| Outcome: | The proposed model outperforms existing models in 14 tasks and 56 languages. |
Copied to clipboard
| Challenge: | Recent Large Vision Language Models demonstrate impressive abilities on image understanding and reasoning tasks. |
| Approach: | They propose a benchmark for fine-grained object classification that is difficult to evaluate . they benchmark 12 public LVLMs on and show CLIP models exhibit better performance . |
| Outcome: | The proposed model improves on 12 public LVLMs on image understanding and reasoning tasks. |
Copied to clipboard
| Challenge: | Useful answers require obvious landmarks as a reference point . a decomposed pipeline is the most effective strategy for generating a high-quality SKG . |
| Approach: | They propose to generate a spatial knowledge graph from a vehicle dashboard diagram . they use large vision-language models to generate the graph using a decomposed pipeline . |
| Outcome: | The proposed method identifies landmarks with 71.3% agreement with human annotators on a new vehicle dataset. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) often hallucinate and produce captions that mention concepts that cannot be found in the image. |
| Approach: | They propose to add grounding objectives to captions that explicitly align image regions or objects to text spans to reduce hallucination. |
| Outcome: | The proposed evaluation protocol reduces the amount of hallucination in LVLMs by adding grounding objectives. |
Copied to clipboard
| Challenge: | Existing methods focus on alignment training or decoding refinements but address symptoms at the generation stage without probing the underlying causes. |
| Approach: | They propose a training-free approach to mitigate hallucination by enhancing the role of vision-aware attention heads. |
| Outcome: | The proposed method achieves superior performance compared to state-of-the-art approaches in mitigating hallucinations while maintaining high efficiency with negligible additional time overhead. |
Copied to clipboard
| Challenge: | Current approaches focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overlooking a fundamental challenge: the same meme can express different intents depending on its conversational context. |
| Approach: | They propose a benchmark to evaluate how large vision language models understand memes in their original context. |
| Outcome: | The proposed benchmark evaluates how large vision language models understand meme intent in their original context. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated that large vision language models (LVLMs) are not multi-modal and lack multi-tasking capabilities. |
| Approach: | They evaluate the performance of large vision language models (LVLMs) for chart understanding and reasoning tasks and compare them to open-source models. |
| Outcome: | The proposed models demonstrate impressive abilities in generating fluent texts covering high-level data insights, but they also encounter common problems like hallucinations, factual errors, and data bias. |
Copied to clipboard
| Challenge: | Existing research has explored methods to enhance the performance of large vision-language models in spatial relations. |
| Approach: | They propose a constraint-aware prompting framework to reduce spatial relation hallucinations by incorporating two types of constraints into the prompt. |
| Outcome: | The proposed framework improves on three widely-used spatial relation datasets. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have shown exceptional performance in multimodal tasks, but their effectiveness in complex visual reasoning is constrained. |
| Approach: | They propose a training-free approach that enhances Reasoning in Large Vision-Language Models . they propose integrating Monte Carlo Tree Search and Self-Reward mechanisms into the reasoning tree . |
| Outcome: | The proposed approach surpasses current prompting methods and secures state-of-the-art performance across three multimodal reasoning benchmarks. |
Copied to clipboard
| Challenge: | LVLMs have shown impressive progress by integrating visual perception with linguistic understanding to produce contextually grounded outputs. |
| Approach: | They propose a visual evidence prompting method to mitigate hallucinations in large vision-language models by using small visual models to complement them. |
| Outcome: | The proposed method reduces hallucinations by reducing false activation and enhancing correct ones. |
Copied to clipboard
| Challenge: | Large-scale vision–language models have achieved remarkable progress on various reasoning tasks, but most studies focus on natural photographic images and pay limited attention to multi-panel visual narratives such as comics. |
| Approach: | They propose a benchmark dataset for chronological reasoning in multi-panel comics that covers six types of reasoning questions and spans both Western and Japanese comic styles. |
| Outcome: | The proposed dataset covers six types of reasoning questions and spans both Western and Japanese comic styles. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) are advanced models that process multiple modalities, such as images, audio, and video, alongside text. |
| Approach: | They propose to use a method to generate and verify draft tokens in parallel . they compare existing methods with small draft models and observe performance fluctuations . |
| Outcome: | The proposed method achieves an average walltime speedup of 1.74 over autoregressive decoding and a 5% improvement over single drafting methods. |
Copied to clipboard
| Challenge: | LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages. |
| Approach: | They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations. |
| Outcome: | The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages. |
Copied to clipboard
| Challenge: | Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures. |
| Approach: | They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step. |
| Outcome: | The proposed method outperforms slow-thinking methods while producing shorter responses. |
Copied to clipboard
| Challenge: | Large Vision-Language Models suffer from a problem known as language prior . such language priors can lead to undesirable biases and hallucinations when dealing with images that are out of distribution. |
| Approach: | They propose a benchmark to measure the language priors of Large Vision-Language Models. |
| Outcome: | The proposed benchmark is the first specifically designed to measure the language priors, or blindness, of LVLMs. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) generate detailed and coherent responses from visual inputs but are prone to generate hallucinations due to an over-reliance on language priors. |
| Approach: | They propose a method that reduces the text context and controls only the image-related POS tokens to maintain text quality by reducing the text contextualization. |
| Outcome: | The proposed method achieves state-of-the-art performance on object hallucination benchmarks and achieves Pareto optimality among the existing methods. |
Copied to clipboard
| Challenge: | Existing white-box jailbreak methods require full model accessibility and require computational costs. |
| Approach: | They propose a black-box jailbreak attack using Zeroth-Order optimization using ZO-SPSA. |
| Outcome: | The proposed method achieves highest jailbreak success rate on three LVLMs, including InstructBLIP, LLaVA and MiniGPT-4. |
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) primarily focus on real-world scenarios, leaving surreal, highly stylized, and semantically hybrid virtual-world situations significantly underexplored. |
| Approach: | They propose to use a manually annotated benchmark to evaluate LVLMs' ability to perceive and describe game character from the virtual-world. |
| Outcome: | The proposed task evaluates LVLMs’ ability to perceive and describe game character from the virtual-world. |
Copied to clipboard
| Challenge: | Large Vision–Language Models (LVLMs) integrate visual perception with language understanding, but how vision information contributes to the model’s decoding process remains under-explored. |
| Approach: | They propose a simple training-free decoding method that guides text generation in Large Vision–Language Models by Referencing Vision Tokens. |
| Outcome: | The proposed method leverages the semantic information embedded within vision tokens by projecting it into the text token distribution. |
Copied to clipboard
| Challenge: | Large vision-language models exhibit excellent ability in language understanding, question answering, and conversations of visual inputs, but they are prone to producing hallucinations. |
| Approach: | They propose to use supervised uncertainty quantification methods to detect hallucinations in large vision-language models. |
| Outcome: | The proposed methods outperform the others in detecting hallucinations on four representative LVLMs across two different tasks. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) are gaining traction in clinical tasks such as diagnostic support, report generation, and medical question answering. |
| Approach: | They present a systematic evaluation of nine DPO variants applied to two leading medical LVLMs. |
| Outcome: | The proposed model improves alignment and reduces severe hallucinations, but yields inconsistent gains over supervised fine-tuning. |
Copied to clipboard
| Challenge: | Existing fingerprinting methods for large vision-language models rely on backdoors to elicit abnormal outputs, but direct distortion of the model’s original outputs compromises modality alignment and degrades multimodal capabilities. |
| Approach: | They propose to embed a robust fingerprint while preserving the original normal outputs of the model. |
| Outcome: | The proposed fingerprint maintains multimodal performance and substantially enhances fingerprint robustness. |
Copied to clipboard
| Challenge: | Current multimodal benchmarks focus on facts within individual images, but neglect associative relations among multiple images. |
| Approach: | They propose a multi-image relational association task and a MMRA benchmark to evaluate LVLMs. |
| Outcome: | The proposed benchmarks show that entity-level multi-image perception tasks pose greater challenges than image-level tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) lack the capacity to handle multimodal inputs effectively. |
| Approach: | They introduce a reference-free and fine-grained evaluation metric that measures the faithfulness of the generated free-form answers from large vision-language models. |
| Outcome: | The proposed metric measures the faithfulness of free-form answers from large vision-language models. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) require instruction tuning on extensive data . training on large VL datasets can be prohibitively expensive . |
| Approach: | They propose a data selection technique that uses a small model as a reference model to select training data for efficient finetuning of a target LVLM. |
| Outcome: | The proposed method achieves superior performance and data selection efficiency against 8 strong baselines on two distinct datasets: LLaVA-1.5 and Vision-Flan. |
Copied to clipboard
| Challenge: | Existing methods for creating a vision question-answering with natural language explanations rely on human annotations that are time-consuming and costly. |
| Approach: | They propose a method that generates high-quality natural language explanations using LVLMs by using visual prompts. |
| Outcome: | The proposed method generates high-quality synthetic VQA-NLE datasets 20x faster than human annotations with minimal decrease in qualitative metrics. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have impressive capabilities in multi-modal context comprehension, but they still suffer from hallucination problems due to inconsistent outputs with the image content. |
| Approach: | They propose a training-free framework MVP to reduce hallucinations in Large Vision-Language Models . they propose multi-view information-seeking strategy to perceive the comprehensive information in the image . |
| Outcome: | The proposed framework reduces hallucinations in large vision-language models by combining multi-view multi-path reasoning with multi-vision multi-path reasoning. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. |
| Approach: | They propose large vision-Language Models to augment LLMs with visual inputs. |
| Outcome: | The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat. |
Copied to clipboard
| Challenge: | Existing data filtering methods rely on coarse-grained scores that lack granularity to identify nuanced semantic flaws. |
| Approach: | They propose a "Decomposition-then-Evaluation" paradigm that breaks model responses into constituent cognitive components. |
| Outcome: | The proposed model outperforms models trained on larger datasets in three key areas . the authors show that Logical Coherence is the most critical factor in data quality evaluation . |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are hardly comprehensively evaluated for their cognitive abilities. |
| Approach: | They propose to evaluate high-level cognitive abilities of Large Vision-Language Models (LVLMs) using images with rich semantics. |
| Outcome: | The proposed evaluation benchmark consists of 251 images along with comprehensive annotations. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) excel at visual understanding but face severe computational bottlenecks when processing high-resolution images and long videos due to massive visual token counts. |
| Approach: | They propose a taxonomy categorizing methods into vision-side, LLM-side and hybrid paradigms and analyze token selection mechanisms and pruning strategy. |
| Outcome: | The proposed method selectively removes less informative tokens while maintaining performance. |
Copied to clipboard
| Challenge: | Despite the promising performance of Large Vision Language Models, they sometimes generate incorrect outputs. |
| Approach: | They propose a multi-modal reward model that aligns LVLMs with human preferences. |
| Outcome: | The proposed model achieves excellent results on the latest multi-modal reward model benchmark and shows competitive performance on text-only reward model. |
Copied to clipboard
| Challenge: | LVLMs often mistakenly determine objects as present in images where they do not exist . authors propose a new benchmark to evaluate object hallucinations by removing objects from images and asking the model whether it can still see the removed objects. |
| Approach: | They propose a benchmark to evaluate object hallucinations by removing objects from images . they propose oDPO, a direct preference optimization objective based on visual objects . |
| Outcome: | The proposed benchmark reduces the likelihood of object hallucinations by removing objects from images and asking the model whether it can still see the removed objects. |
Copied to clipboard
| Challenge: | Hallucination is a critical challenge for large language models and large vision-language models (LVLMs) however, dedicated research on medical hallucinations remains unexplored. |
| Approach: | They provide a unified perspective on medical hallucination for both LLMs and LVLMs, and delve into its causes. |
| Outcome: | The proposed models have demonstrated impressive performance on a variety of medical benchmarks. |
Copied to clipboard
| Challenge: | Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. |
| Approach: | They evaluate LVLMs' ability to account for variation in color perception using the Ishihara Test. |
| Outcome: | The proposed models fail to reproduce the perceptual outcomes experienced by affected individuals and default to normative color perception. |
Copied to clipboard
| Challenge: | Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease. |
| Approach: | They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability. |
| Outcome: | The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) are evolving rapidly and require data with human supervision to achieve better alignment. |
| Approach: | They introduce VLFeedback, the first large-scale vision-language feedback dataset . they train Silkie, an LVLM fine-tuned via direct preference optimization . |
| Outcome: | The proposed model outperforms its base model in helpfulness, visual faithfulness, and safety metrics and exhibits enhanced resilience against red-teaming attacks. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) often produce object hallucinations due to their reliance on text cues and learned object co-occurrence biases. |
| Approach: | They propose a language-contrasting decoding algorithm that adjusts LVLM outputs based on LLM confidence levels to mitigate object hallucinations. |
| Outcome: | The proposed method shows up to %4 improvement in POPE F1 scores and %36 reduction in CHAIR scores on COCO validation set while improving captioning quality scores. |
Copied to clipboard
| Challenge: | Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes. |
| Approach: | They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models. |
| Outcome: | The proposed model outperforms existing benchmarks on multi-image models. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) are trained on large-scale datasets, which can pose privacy risks if training images contain sensitive information. |
| Approach: | They propose to detect whether a target image is used to train LVLMs by using image-text pairs and single-modality content to detect image-related data. |
| Outcome: | The proposed methods detect whether a target image is used to train the LVLM on large-scale datasets. |
Copied to clipboard
| Challenge: | Existing methods for rewriting text-to-image models require specialized vocabulary . a new approach uses large vision language models to optimize text-based models . |
| Approach: | They propose a prompt optimization framework that rephrases a user prompt into a text-to-image model by using large vision language models as solver and reward model. |
| Outcome: | The proposed model outperforms existing models on two popular datasets. |
Copied to clipboard
| Challenge: | Recent work has shown that pruning can reduce model performance, but it can also lead to degradation in safety performance. |
| Approach: | They propose a hierarchical safety realignment approach to prune large vision-Language Models . they quantify contribution of each attention head to safety and restore neurons . |
| Outcome: | The proposed approach achieves significant safety improvements in LVLMs pruned post pruning. |
Copied to clipboard
| Challenge: | generative AI agents cannot model common ground in a way that enables smooth communication . a recent study examined whether large language models and large vision language models engage in grounding as human discourse partners do . |
| Approach: | They propose to use referential communication to model common ground between a pair of directors and a picture matching system. |
| Outcome: | The proposed experiment shows that generative AI agents cannot model common ground . human conversation relies on common ground accrued and updated by interacting partners . |
Copied to clipboard
| Challenge: | Object hallucination has been an Achilles’ heel which hinders the broader applications of large vision-language models (LVLMs). |
| Approach: | They propose a logical closed loop-based framework for Object Hallucination Detection and Mitigation that uses logical consistency probing to raise questions with logical correlations to determine hallucinations. |
| Outcome: | The proposed method can be applied to all existing LVLMs and is effective and general. |
Copied to clipboard
| Challenge: | Existing methods for fingerprinting large vision-Language Models rely on explicit triggers, which have limitations in terms of stealthiness and robustness. |
| Approach: | They propose to use model fingerprints to verify the ownership of large vision-Language Models (LVLMs) they use implicit model fingerprinting techniques that leverage neighboring samples as implicit model . |
| Outcome: | The proposed fingerprinting technique is superior to existing methods, but has limitations in terms of stealthiness and robustness. |
Copied to clipboard
| Challenge: | Multimodal reasoning is a key capability for large vision-language models . however, the vanilla Chain-of-Thought method fails to address critical steps in multi-step reasoning tasks. |
| Approach: | They propose a bi-modal Behavioral Alignment method to augment multimodal reasoning . they use domain-specific language to integrate multimodal information into a precise alternative form . |
| Outcome: | The proposed method significantly improves GPT-4V(ision) on geometry problem solving, chess positional advantage prediction and molecular property prediction. |
Copied to clipboard
| Challenge: | Existing studies focus on the text modality or are limited to specific tasks. |
| Approach: | They propose a framework to teach Large Vision-Language Models to selectively utilize retrieved information and improve their robustness against irrelevant or misleading references. |
| Outcome: | The proposed framework improves LVLMs’ ability to utilize retrieved multimodal references and their robustness against irrelevant or misleading information. |
Copied to clipboard
| Challenge: | Existing methods for aligning LVLMs rely on external datasets, human annotations or complex post-processing. |
| Approach: | They propose a method that generates a debiased self-judgment score for LVLMs . this self-evaluation metric is created internally by the model without external resources . |
| Outcome: | The proposed approach outperforms existing methods in reducing hallucinations and safety concerns. |
Copied to clipboard
| Challenge: | Despite the success of Large Vision-Language Models, they suffer from hallucination. |
| Approach: | They propose a training-free strategy that "D**ive into" the attention of LVLMs to "R**educe" object hallucination by using classification tokens of ViT. |
| Outcome: | The proposed method reduces the impact of outlier tokens on LVLMs . the proposed method is based on LLaVA-1.5, LLvaVA-NeXT and InstructBLIP . |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have shown promising results on multimodal tasks, but remain prone to hallucinations due to their reliance on a single modality or memorizing training data without properly grounding their outputs. |
| Approach: | They propose a training-free, tri-layer contrastive decoding with watermarking that uses a watermark-related question to identify a pivot layer and apply tri-layered contrastive coding to generate the final output. |
| Outcome: | The proposed method reduces hallucinations and generates more visually grounded responses. |
Copied to clipboard
| Challenge: | Existing methods for acquiring large-scale intentions generate product-centric intentions without product images and incur high costs for scalability. |
| Approach: | They propose a multimodal framework that allows Large Vision-Language Models to infer purchase intentions from multimodal product metadata and prioritize human-centric ones. |
| Outcome: | The proposed framework shows that it is robust to different prompts and superior to previous methods. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have impressive capabilities across visual tasks, yet they remain hindered by the persistent challenge of hallucinations. |
| Approach: | They propose a novel approach that dynamically adapts decoding strategies by evaluating the correctness of the model’s attention on image tokens to distinguish the correct attention. |
| Outcome: | Extensive experiments show that the proposed approach outperforms existing decoding methods across multiple mainstream benchmarks, effectively mitigating hallucinations in LVLMs. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) generate detailed responses even when questions are ambiguous or unanswerable, leading to hallucinations and bias issues. |
| Approach: | They propose a three-tiered hierarchy for questions of invalid, ambiguous, and personalizable nature to measure the proactive engagement capabilities of LVLMs. |
| Outcome: | The proposed model generates contrastive response pairs for unlabeled questions, achieving 0.84 AAR, while maintaining comparable performance on general tasks. |
Copied to clipboard
| Challenge: | Large Vision–Language Models (LVLMs) suffer from object hallucination, generating descriptions for objects that are absent from the image, which undermines reliability and hinders real-world deployment. |
| Approach: | They propose a positional-alignment scheme that preserves pretrained weight order while globally—- visual–text distances, embeds an isotropic fused patch-distance metric, and applies a patch-delay causal mask to enforce spatial causality. |
| Outcome: | Extensive experiments on POPE, MMStar and SQA show that DAPE-BR reduces hallucinations and boosts performance. |
Copied to clipboard
| Challenge: | Large Visual Language Models (LVLMs) suffer from hallucinations due to limited training data, lack of * Equal contribution precise grounding, and over-reliance on language priors. |
| Approach: | They propose a framework to detect and mitigate hallucinations through claim verification using program-of-thought prompting and Python code to generate a graph. |
| Outcome: | The proposed framework improves over baseline LVLMs and existing methods across several benchmarks. |
Copied to clipboard
| Challenge: | Large vision-language models are prone to hallucinations, where contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. |
| Approach: | They propose to automate the generation of hallucination-related questions using images . they propose to use three image manipulation strategies to induce hallucinosity . |
| Outcome: | The proposed approach reduces human bias in crafting such examples and improves accuracy. |
Copied to clipboard
| Challenge: | Existing benchmarks that treat hallucinations as isolated errors neglect causal dependencies between visual perception and textual reasoning. |
| Approach: | They propose a Knowledge-Guided In-Context Probing framework that constructs a dual-perception ground truth to transform abstract priors into multi-granularity queries. |
| Outcome: | The proposed framework isolates deep reasoning failures from simple perceptual misses. |
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only. |
| Approach: | They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. |
| Outcome: | The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions. |
Copied to clipboard
| Challenge: | Existing datasets and models fail to consider critical aspects of medical diagnostics, authors argue . MMXU enables multi-image questions incorporating both current and historical patient data. |
| Approach: | They propose a dataset for MedVQA that focuses on identifying changes in specific regions between two patient visits. |
| Outcome: | The proposed dataset improves diagnostic accuracy by 20% by integrating historical data. |
Copied to clipboard
| Challenge: | Using retrieval augmentation, large vision language models can be used for diagnostic accuracy, but multimodal retrieval-augmented diagnosis is challenging. |
| Approach: | They propose a lightweight mechanism for enhancing diagnostic performance of retrieval-augmented LVLMs by fine-tuning a multimodal retriever and general-purpose backbone models. |
| Outcome: | The proposed mechanism achieves competitive results without medical training compared to pre-trained models with extensive training. |
Copied to clipboard
| Challenge: | Currently, large vision-language models are limited in their ability to provide correct answers for multimodal tasks . however, they can still provide correct responses for multiple images associated with a single image . a query-agnostic visual attack (QAVA) provides robust adversarial examples that generate incorrect responses to unspecified and unknown questions. |
| Approach: | They propose a query-agnostic visual attack to create adversarial examples that generate incorrect answers to unspecified and unknown questions. |
| Outcome: | The proposed model improves performance on images when the question is unknown compared to known target questions . |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. |
| Approach: | They introduce DivScene, a large-scale dataset with 4,614 houses across 81 scene types and 5,707 kinds of target objects. |
| Outcome: | The proposed dataset provides a much greater diversity of target objects and scene types than existing datasets, enabling a comprehensive task evaluation. |
Copied to clipboard
| Challenge: | Existing datasets focus on primary perception abilities and commonsense knowledge, or have low level of text comprehension difficulty, which are insufficient to reflect comprehensive capabilities of large vision-language models. |
| Approach: | They propose a multimodal benchmark based on the Chinese College Entrance Examination (GAOKAO) which sets human-level requirements for the model’s abilities, including perception, understanding, knowledge and reasoning. |
| Outcome: | The proposed model derives from native Chinese context and sets human-level requirements for its abilities, including perception, understanding, knowledge and reasoning. |
Copied to clipboard
| Challenge: | Visual instruction tuning is the predominant technology in eliciting multimodal task-solving capabilities of large vision-language models. |
| Approach: | They propose a visual instruction-free fine-tuning framework for large vision-language models . they require only text-only instructions and image caption data during training . |
| Outcome: | The proposed framework is based on visual instruction tuning, but requires images as input . it can achieve state-of-the-art performance on several downstream benchmarks with less training data. |
Copied to clipboard
| Challenge: | Existing approaches to address hallucinations in large vision-language models require substantial computational cost and time. |
| Approach: | They propose to leverage sparse autoencoders to identify semantic directions closely associated with faithfulness or hallucination, extracting more precise and disentangled hallucinian-related representations. |
| Outcome: | The proposed method outperforms existing decoding approaches while maintaining transferability across different model architectures with negligible additional time overhead. |
Copied to clipboard
| Challenge: | a perception bottleneck in large vision-language models is critical for chart understanding . instruction tuning improves the extraction capability of LVLMs, but the vision encoder remains a bottleneck . |
| Approach: | They propose to decompose the perception bottleneck into two components . the vision encoder bottleneck is where visual representation fails to encapsulate the correct information . |
| Outcome: | The proposed approach significantly mitigates the vision encoder bottleneck and improves the ability of LVLMs to comprehend charts. |
Copied to clipboard
| Challenge: | Advancements in Large Vision-Language Models (LVLMs) have demonstrated impressive performance in image-conditioned text generation, but hallucinated outputs pose a major barrier to their use in safety-critical applications. |
| Approach: | They propose a conformal-prediction-based framework that achieves finite-sample distribution-free statistical guarantees to the factuality of LVLM output. |
| Outcome: | The proposed framework reduces the error rate of LLaVa-1.5 claims from 87.8% to 10.0% while ensuring that the output is accurate. |
Copied to clipboard
| Challenge: | Existing security benchmarks only cover user-level prompts and environmental threats . however, these models are vulnerable to pop-up attacks and prompt injections . |
| Approach: | They propose a security benchmark that covers a set of six attack vectors that span both user-level and environment-level manipulations. |
| Outcome: | The proposed security benchmarks cover a set of six real-world web environments with 2,970 adversarial trajectories and a multi-layered evaluation protocol dissecting agent failures across internal reasoning, behavioral execution, and task outcomes. |
Copied to clipboard
| Challenge: | CEMTM is a context-enhanced multimodal topic model that can infer coherent topic structures from documents . traditional multimodal topics failed to capture deeper cross-modal interactions . large vision language models (LLMs) and LVLMs have shown remarkable capacity to encode rich semantic knowledge from vast corpora. |
| Approach: | They propose a context-enhanced multimodal topic model that uses tokens to weight contributions to topic inference. |
| Outcome: | The proposed model outperforms unimodal and multimodal benchmarks on six multimodal domains and captures semantics in scientific articles. |
Copied to clipboard
| Challenge: | Compared to domain-specific work in this task, this task proved particularly challenging due to the absence of domain- specific features. |
| Approach: | They propose an utterance generation model with a novel spatial graph that integrates spatial information to deal with the open-domain characteristics of the commentaries and significantly improves performance. |
| Outcome: | The proposed model significantly improves performance in the open-domain live commentary generation task. |
Copied to clipboard
| Challenge: | Chart question answering (CQA) is a key research challenge for large vision-language models . recent efforts focus on leveraging LVLMs directly on chart images . |
| Approach: | They propose a gaze-guided attention refinement that aligns image-text attention with human fixations to improve chart reasoning quality and interpretability. |
| Outcome: | The proposed approach improves answer accuracy and attention alignment yielding gains of up to 2.56 percentage points across multiple models. |
Copied to clipboard
| Challenge: | FinChart-Bench is the first benchmark specifically focused on real-world financial charts. |
| Approach: | They propose a benchmark specifically focused on real-world financial charts. |
| Outcome: | The proposed benchmark evaluates 26 state-of-the-art LVLMs on FinChart-Bench. |
Copied to clipboard
| Challenge: | HKVE selectively accepts gradient optimization results based on the distribution of attention scores across different layers, ensuring that every optimization step positively contributes to the attack. |
| Approach: | They propose a framework that selectively accepts gradient optimization results based on the distribution of attention scores across different layers and selectively takes them into account when calculating the attack success rate. |
| Outcome: | The proposed framework outperforms existing methods by achieving success rates of 75.08% on MiniGPT4, 85.84% on LLaVA and 81.00% on Qwen-VL. |
Copied to clipboard
| Challenge: | Current approaches to large vision-language models rely on costly annotations and are not comprehensive in terms of evaluating all aspects. |
| Approach: | They propose an automated method which can access LVLMs hallucination in an LLM-free and annotation-free way and model the dependency between different types of halluciNations. |
| Outcome: | The proposed model can model the dependency between different types of hallucinations and generate Q&A pairs on any image dataset at minimal cost. |
Copied to clipboard
| Challenge: | Large Vision-Language Models are susceptible to typographic attacks, which are misclassifications caused by an attack text that is added to an image. |
| Approach: | They propose a multi-image setting for studying typographic attacks by leveraging the difficulty of the target image, the strength of the attack text, and text-image similarity. |
| Outcome: | The proposed approach improves success rates by 21% over random, non-specific methods on the CLIP model while maintaining stealth in a multi-image scenario. |
Copied to clipboard
| Challenge: | Existing methods for reducing hallucinations incur a significant increase in latency. |
| Approach: | They propose a task-agnostic attention-guided head suppression strategy that can be seamlessly integrated during inference without incurring significant compute or latency overhead. |
| Outcome: | The proposed approach reduces hallucinations by 2.7x while maintaining F1 and improves throughput by 1.8% compared to existing methods. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) suffer from hallucination where generated textual descriptions fail to align accurately with visual semantics. |
| Approach: | They propose a training-free approach that mitigates hallucination through targeted intervention in the model’s intermediate activations by identifying directional patterns of hallucinism in the activation space using a small calibration set. |
| Outcome: | The proposed approach reduces hallucination across multiple benchmarks while maintaining performance on general visual understanding tasks. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have impressive multimodal abilities but remain prone to multilingual object hallucination. |
| Approach: | They propose a cross-lingual attention intervention method to mitigate multilingual object hallucination in LVLMs by aligning attention patterns. |
| Outcome: | The proposed method improves 13.56% (up to 30%) on the POPE and 21.75% on the hallucination subsets across languages. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) suffer from multimodal hallucinations . however, the generated hallucines could influence the models’ subsequent generation . |
| Approach: | They propose a framework to evaluate LVLMs' behaviors when encountering generated hallucinations and a method to revise the output distribution of LVLs with the one derived from the residual visual input. |
| Outcome: | The proposed framework reduces the performance of open-source LVLMs by 31%, indicating that they are prone to accept the generated hallucinations and make false claims that they would not have supported without distractions. |
Copied to clipboard
| Challenge: | AMIA is a lightweight, inference-only defense for Large Vision–Language Models . it automatically masks text-irrelevant image patches and conducts joint Intention Analysis . |
| Approach: | AMIA is a lightweight, inference-only defense for large vision–language models . it automatically masks a small set of text-irrelevant image patches to disrupt adversarial perturbations . |
| Outcome: | AMIA improves defense success rates across diverse LVLMs and jailbreak benchmarks . it preserves general utility with only 2% accuracy drop, incurs only modest inference overhead . |
Copied to clipboard
| Challenge: | Despite the significant success of large vision-language models, some studies have revealed that LVLMs suffer from the hallucination problem when given long-term misleading textual history. |
| Approach: | They propose a visual dialogue hallucination evaluation benchmark VisDiaHalBench to investigate the halluciation problem of large vision-language models when given long-term misleading textual history. |
| Outcome: | The proposed benchmark consists of samples with five-turn questions about an edited image and its original version. |
Copied to clipboard
| Challenge: | Large-scale vision language models excel at generating factual content, but their ability to rank images from multiple perspectives has not been explored. |
| Approach: | They propose a framework to evaluate large-scale vision-language models by measuring their ability to rank image texts from multiple perspectives. |
| Outcome: | The proposed evaluation framework measures how closely LVLMs' judgments align with human interpretations. |
Copied to clipboard
| Challenge: | Large vision-language models exhibit an imbalance in multilingual capabilities . |
| Approach: | They propose a training recipe that achieves efficient multilingual enhancement for LVLMs by Precise Language Specific layers fine-tuning. |
| Outcome: | The proposed training recipe achieves efficient multilingual enhancement for LVLMs by fine-tuning language specific layers. |
Copied to clipboard
| Challenge: | Large Vision/Language Models (LVLMs) are less capable of generating accompanying image sequences. |
| Approach: | They propose a method that integrates a Latent Diffusion Model (LDM) with an LLM to generate captions to maintain semantic coherence of the sequence. |
| Outcome: | The proposed method is preferred by humans in 46.6% of the cases against 26.6% for the second best method. |
Copied to clipboard
| Challenge: | Document Understanding is a foundational AI capability with broad applications . Large Vision-Language Models (LLMs) can't handle multi-page document comprehension . a logic-aware retrieval framework for multi-modal, multi- page document understanding is proposed . |
| Approach: | They propose a logic-aware retrieval framework for multi-modal, multi-page document understanding . MoLoRAG uses semantic and logical relevance to deliver more accurate retrieval . |
| Outcome: | The proposed framework improves on four DocQA datasets and demonstrates 9.68% accuracy improvement over existing methods. |
Copied to clipboard
| Challenge: | Recent advances in Large Vision Language Models (LVLMs) show promise for pathological diagnosis, yet their application in clinical settings faces critical challenges of multimodal hallucination and biased responses. |
| Approach: | They propose a framework that integrates medical expertise into preference alignment. |
| Outcome: | The proposed framework outperforms existing pathological LVLMs while maintaining pathological accuracy. |
Copied to clipboard
| Challenge: | Large vision-language models have been widely used but stereotypical biases are unexplored. |
| Approach: | They propose a framework to SCAN stereotypical bias within large vision-language models . they examine stereotype biases with respect to gender and race in three scenarios . |
| Outcome: | The proposed framework can reduce stereotypical biases in large vision-language models . the currently popular models show significant stereotype biase . |
Copied to clipboard
| Challenge: | Chain-of-thought reasoning improves performance of large language models, but is it faithfully reflecting internal processes? |
| Approach: | They propose a new evaluation pipeline for categorizing bias articulation patterns and a novel evaluation pipeline to examine CoT faithfulness in large vision-language models. |
| Outcome: | The proposed evaluation pipeline enables significantly more precise analysis of CoT reasoning than previous methods. |
Copied to clipboard
| Challenge: | Existing studies focus on posthoc alignment techniques, but the underlying safety mechanisms within LVLMs remain unexplored. |
| Approach: | They propose a tuning-free framework that leverages internal activations to enhance safety. |
| Outcome: | The proposed framework outperforms state-of-the-art methods in detecting jailbreak attacks against large vision-language models. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) generate responses that are plausible but incorrect or unsupported—commonly referred to as hallucinations. |
| Approach: | They propose a representation-level intervention framework that modulates hallucination-related features during inference by probing their encoded features. |
| Outcome: | The proposed framework reduces hallucinations while maintaining the performance and generalization capabilities of Large Vision-Language Models (LVLMs). |
Copied to clipboard
| Challenge: | Large vision–language models suffer from object-existence hallucinations when multi-step deliberation decouples from visual evidence. |
| Approach: | They propose a framework that allocates visual computation by uncertainty . they propose highlighting retains global context, while selective zoom-in performs local verification. |
| Outcome: | The proposed framework reduces the complexity of multimodal reasoning by minimizing the operator trade-off. |
Copied to clipboard
| Challenge: | Existing methods for hallucination mitigation are based on external dependency and require external annotations or auxiliary models for preference data collection. |
| Approach: | a new method is proposed to help model-generated hallucinations without external dependencies. |
| Outcome: | a new method that self-injects hallucinations into a generated response improves halluuutations mitigation. |
Copied to clipboard
| Challenge: | Existing training-free methods are vulnerable to the attention sink phenomenon . Existing methods include contrastive decoding and auxiliary expert models . |
| Approach: | They propose a training-free attention intervention that constructs a PAD map to identify semantically core visual regions and applies per-head Median Absolute Deviation Scaling to adaptively control the intervention strength. |
| Outcome: | The proposed intervention improves visual grounding and reduces hallucinations on multiple LVLMs and benchmarks. |
Copied to clipboard
| Challenge: | Existing scientific benchmarks lack human-annotated difficulty levels and structured taxonomies of scientific concepts. |
| Approach: | They propose a benchmark for evaluating mathematical and physical reasoning through text-only and text-image formats with human-annotated difficulty levels and detailed explanations. |
| Outcome: | The proposed model achieves only 63.77% accuracy and struggles with visual reasoning tasks. |
Copied to clipboard
| Challenge: | This work examines the alignment of large language models and large vision-language models with human perception. |
| Approach: | They use a dataset of *shitsukan* terms elicited from individuals in response to object images to evaluate their understanding of the Japanese concept of shitukan. |
| Outcome: | The proposed models demonstrated mixed accuracy across benchmark tasks, with limited overlap between model- and human-generated terms. |
Copied to clipboard
| Challenge: | Existing safeguards relying on pre-filtering or fine-tuning are costly and diminish overall utility. |
| Approach: | They propose a lightweight method that leverages LVLMs’ inherent multimodal alignment for zero-shot toxic image detection. |
| Outcome: | The proposed method achieves a 66.9% defense success rate with only 3.2% false positive rate and 7.2% overhead. |
Copied to clipboard
| Challenge: | Existing large vision-language models suffer from hallucination due to over-reliance on the Large Language Model (LLM) backbone. |
| Approach: | They propose a method to improve visual context learning by using a large-scale preference learning algorithm to improve hallucination. |
| Outcome: | The proposed method improves on human-annotated hallucination datasets. |
Copied to clipboard
| Challenge: | Existing LVLMs perform visual perception at multiple levels, but they are not able to perform multi-level tasks. |
| Approach: | They propose a visual–language benchmark to evaluate LVLMs' perceptions . they use manipulated images to examine how LVLs can perform multi-level tasks . |
| Outcome: | The proposed model performs poorly on high-level perception tasks, the authors show . they also show that current models do not generalize in understanding semantics of synthetic images . |
Copied to clipboard
| Challenge: | Existing safety research focuses on implicit social norms or text-only settings, overlooking the complexities of multimodal documents. |
| Approach: | They propose a benchmark to assess the safety of large vision-Language Models (LVLMs) they propose 'Document Policy Preservation Benchmark' to assess document policy compliance. |
| Outcome: | The proposed framework outperforms standard prompting defenses in the evaluation of multimodal documents. |
Copied to clipboard
| Challenge: | Existing attacks optimize image perturbations to maximize harmful output likelihood, but suffer from slow convergence due to gradient conflict between adversarial objectives and the model’s safety-retrieval mechanism. |
| Approach: | They propose a push-pull approach which suppresses attention to system-prompt tokens and anchors generation on adversarial image features to avoid collisions. |
| Outcome: | The proposed approach reduces gradient conflict by 45% and achieves 94.4% attack success rate on Qwen-VL (vs. 68.8% baseline) with 40% fewer iterations. |
Copied to clipboard
| Challenge: | Existing multilingual vision-language (VL) benchmarks typically only cover a handful of languages, underscoring the need for evaluation data for low-resource languages. |
| Approach: | They propose a multilingual vision-language benchmark that evaluates cross-modal and text-only topical matching across 205 languages. |
| Outcome: | The proposed model performs better in cross-modal and text-only topical matching in lower-resource languages than the most multilingual benchmarks. |
Copied to clipboard
| Challenge: | Existing models for large vision language models do not fully reflect their knowledge capacity and reliability, resulting in erroneous outputs that do not align with the image content or provide answers lacking knowledge evidence. |
| Approach: | They propose a Chinese-based benchmark for visual factuality across 8 major topics and 56 subtopics and a multi-hop question construction. |
| Outcome: | The proposed model decouples visual factuality into two parts: seeing the world and discovering knowledge. |
Copied to clipboard
| Challenge: | a crucial skill for embodied AI agents working with humans is grounding in referential communication. |
| Approach: | They use large vision language models to overhear spontaneous conversations between humans . they find that current LVLMs fail to show consistent performance improvement . |
| Outcome: | The proposed models fail to show consistent performance improvement over previous models . the authors release the results to facilitate future research . |
Copied to clipboard
| Challenge: | Existing methods to remove knowledge from large vision-Language Models often fail to provide quality and informative post-unlearning responses. |
| Approach: | They propose a task that requires models to provide privacy-preserving yet informative responses for LVLMs. |
| Outcome: | The proposed method reduces the risk of unlearning after naive suppression by providing informative and visually grounded responses. |
Copied to clipboard
| Challenge: | Large Vision-Language Models often produce hallucinations due to the limited ability to verify information in different regions of the image. |
| Approach: | a new decoding method improves factual grounding by modeling inter-region consistency . the method identifies salient regions using cross-attention and generates initial responses for each . |
| Outcome: | a training-free decoding method reduces hallucinations and improves response consistency . the proposed method generates initial responses for each region and weights reliability weights among responses . |
Copied to clipboard
| Challenge: | Existing Large Vision-Language Models (LVLMs) lack integrated commonsense knowledge . lack of integrated common knowledge limits their robustness and accuracy in VQA . |
| Approach: | They propose a framework to enhance multimodal inference by integrating commonsense reasoning. |
| Outcome: | MAGIC-VQA improves comprehensive benchmark datasets, surpassing existing models in tasks requiring advanced commonsense reasoning. |
Copied to clipboard
| Challenge: | a global shortage of healthcare workers has demanded the development of smart healthcare assistants. |
| Approach: | They analyze the healthcare knowledge of existing Large Vision Language Models (LVLMs) using an annotated open-ended task. |
| Outcome: | The study analyzes the knowledge of large vision language models using open-ended questions . the results highlight the need for specialized, domain-specific solutions . |
Copied to clipboard
| Challenge: | Existing safety guardrails fail to intercept latent intent, whereas LVLMs can implicitly synthesize holistic malicious semantics from fragmented visual cues. |
| Approach: | They propose an Emoji Chain Hinting Attack (ECHA) framework that decouples sensitive concepts into semantically related emoji chains and structural text masks. |
| Outcome: | The proposed framework outperforms existing baselines and bypasses safety guardrails in over 81% of instances with a single attempt. |
Copied to clipboard
| Challenge: | Current TIMT studies focus on providing translations for all text within an image, neglecting to provide bounding boxes and covering limited scenarios. |
| Approach: | They extend traditional TIMT into position-aware TIMt to support fine-grained translation . they introduce an Adaptive Image OCR Refinement Pipeline to refine results . |
| Outcome: | The proposed model supports fine-grained and layout-preserving translation . the experimental data highlight the scalability and generalizability of the model. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on coarse-grained hallucination detection and fail to capture hallucinics . vision encoders exhibit unique hallucinian characteristics, but suboptimal of simple feature fusion. |
| Approach: | They propose a visual encoder that employs different training paradigms to instill inductive biases in visual encoded models. |
| Outcome: | The proposed system reduces hallucinations and improves model performance. |
Copied to clipboard
| Challenge: | Recent research in large vision-language models has shown promising results, but the issue of hallucination remains. |
| Approach: | They propose an instruction-based method to reduce hallucinations in large vision-language models . they use disturbance instructions to exacerbate hallucinosity in multimodal fusion modules . |
| Outcome: | The proposed method reduces hallucinations in multimodal fusion modules by reducing alignment uncertainty and subtracting hallucines from the original distribution. |
Copied to clipboard
| Challenge: | Existing approaches to improve the performance of Large Visual Language Models (LVLMs) are limited by cross-modal interactions and representation disparities. |
| Approach: | They propose a Visual In-Context Learning method that retrieves images via a 'Retrieval & Rerank' paradigm and summarises images with task intent and task-specific visual parsing to compose language-based demonstrations that reduce token count. |
| Outcome: | The proposed method reduces token count and alleviates cross-modal interaction problem on visual reasoning datasets. |
Copied to clipboard
| Challenge: | Existing evaluation paradigms for geographic reasoning are outcome-centric and focus on label matching, leaving the underlying linguistic reasoning chains as unexamined black boxes. |
| Approach: | They propose a dynamic, human-preference-based evaluation framework for benchmarking open-world geographic reasoning. |
| Outcome: | The proposed framework reframes evaluation as a pairwise reasoning alignment task on in-the-wild images, where human judges compare model-generated explanations based on reasoning quality, evidence synthesis, and plausibility. |
Copied to clipboard
| Challenge: | Chart Question Answering systems are limited in their ability to interpret data visually and reason with visual representations. |
| Approach: | They propose a chart-based chart question-answering system that includes 1,341 charts from 99 diverse sources and 1,948 questions in various types. |
| Outcome: | The new benchmark includes 1,341 charts from 99 diverse sources and 1,948 questions in various types. |
Copied to clipboard
| Challenge: | Existing methods that adapt LVLMs to egocentric tasks overlook critical agent-environment interactions, limiting their ability to perform egoic reasoning. |
| Approach: | They propose a zero-shot paradigm to enhance egocentric reasoning by simulating human causal reasoning by formalizing ego-centric reasoning using a structural causal model. |
| Outcome: | The proposed method improves egocentric reasoning abilities on six tasks. |
Copied to clipboard
| Challenge: | Existing preference alignment methods focus on aligning model responses with human preferences while neglecting image-text modality alignment. |
| Approach: | They propose Entity-centric Multimodal Preference Optimization to improve modality alignment . they use open-source instruction datasets to automatically construct high-quality preference data . |
| Outcome: | The proposed approach reduces hallucination rates by 80.4% on Object HalBench and 52.6% on MM HalBech. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating a generalizable defense that is efficient for practical deployment. |
| Approach: | They propose a framework that uses a lightweight projection to separate benign and malicious inputs in safety-critical layers. |
| Outcome: | The proposed framework enables a simple yet powerful contrastive score that differentiates true malicious intent from mere distribution shift. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) have been criticized for their language bias. |
| Approach: | They propose to use a dual-attention mechanism to construct separate attention for visual and text inputs to enhance integration of visual inputs across models. |
| Outcome: | Experiments show that the proposed model debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without additional resources. |
Copied to clipboard
| Challenge: | Recent work shows that in-context learning for large language models exhibits compositional generalization capacity. |
| Approach: | They propose a method to exhibit in-context compositional generalization in large vision-language models by combining visual and linguistic modalities. |
| Outcome: | The proposed method reduces redundancy and complexity in in-context learning with LVLMs. |
Copied to clipboard
| Challenge: | Large vision-language models produce unfaithful visual hallucinations, also known as visual halluinations, which hinders their application in multimodal understanding and decision-making. |
| Approach: | They propose a plug-and-play train-free decoding algorithm for mitigating visual hallucinations . they leverage visual information to construct a coarse-to-fine visual view tree . |
| Outcome: | The proposed algorithm reduces visual hallucinations (VH) by leveraging visual information to construct a coarse-to-fine visual view tree (CFTree) |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have expanded capabilities beyond text understanding . a novel Chinese financial multimodal evaluation benchmark is used to evaluate LVLM capabilities . |
| Approach: | They propose a Chinese financial multimodal evaluation benchmark to evaluate LVLMs' capabilities . the model has an overall accuracy of 66.11% and an average score of 77.18 . |
| Outcome: | The proposed model achieves an overall accuracy of 66.11% on the question answering task and an average score of 77.18 on detection, recognition, and information extraction tasks. |
Copied to clipboard
| Challenge: | Large vision language models have impressive reasoning capabilities across complex multimodal tasks. |
| Approach: | They propose to use distribution-reshaping and trajectory-rebalancing to improve visual reasoning capabilities. |
| Outcome: | Experiments on Qwen2-VL-7B-Instruct and InternVL2.5-4B models show that their methods outperform baselines by 3.86 points. |
Copied to clipboard
| Challenge: | Existing studies have revealed that Large Vision-Language Models suffer from hallucinations in practice, including object hallucines, spatial hallucinos, attribute hallucinications, etc. |
| Approach: | They propose to use CLIP model to mitigate object hallucinations by using a data augmentation method to create negative samples with a variety of hallucinian issues. |
| Outcome: | The proposed method mitigates object hallucinations and can be used as a visual encoder, effectively alleviating the object halluination issue in LVLMs. |
Copied to clipboard
| Challenge: | LVLMs have shown remarkable performance in visual-language understanding for downstream multimodal tasks. |
| Approach: | They propose a method to alleviate hallucinations by masking the “image heads” in LVLMs . |
| Outcome: | The proposed method alleviates the phenomenon of hallucinations and retains the general capabilities of LVLMs. |
Copied to clipboard
| Challenge: | Despite advances in Large Vision Language Models, a gap remains in their interpretability and performance. |
| Approach: | They identify the Optical Character Recognition Head (OCR Head) heads that are more efficient at recognizing text from images. |
| Outcome: | The Optical Character Recognition Head (OCR Head) is identified as the most efficient head for recognizing text from images. |
Copied to clipboard
| Challenge: | Large Vision Language Models (LVLMs) have shown impressive performance on various vision-language tasks. |
| Approach: | They propose a benchmark framework for evaluating Visual Variation Robustness of Large Vision Language Models that incorporates automated evaluation dataset generation and principled metrics for thorough robustness assessment. |
| Outcome: | The proposed framework identifies a vulnerability to visual variations affecting even advanced models that excel at complex vision-language tasks but significantly underperform on simple tasks like object recognition. |
Copied to clipboard
| Challenge: | Existing approaches focus on action selection or use pre-trained models as world models to enhance planning capabilities. |
| Approach: | They propose a new learning framework that optimizes state prediction and action selection through preference learning. |
| Outcome: | The proposed method outperforms existing methods and GPT-4o on VoTa-Bench and Qwen2-VL (7B), LLaVA-1.6 (7B) and LLama-3.2 (11B). |
Copied to clipboard
| Challenge: | Large Vision-Language Models are hindered by a systemic efficiency barrier known as visual token dominance. |
| Approach: | They propose a systematic taxonomy of efficiency techniques structured around the inference lifecycle . they examine visual encoding, prefilling, and decoding to understand bottlenecks . |
| Outcome: | The proposed techniques reveal how upstream decisions dictate downstream bottlenecks . the proposed techniques include hybrid compression and modality-aware decoding . |
Copied to clipboard
| Challenge: | Existing pruning methods for large vision language models use visual tokens to prune . existing methods fail to balance efficiency and semantic alignment due to large number of visual token. |
| Approach: | They propose a cross-modal pruning framework that considers textual semantics and visual self-attention to combine them to achieve efficient inference acceleration. |
| Outcome: | The proposed pruning framework can retain only 25% of the visual tokens, with a minimal performance degradation of only 0.063% on LLaVA-1.5-13B. |
Copied to clipboard
| Challenge: | Existing methods for predicting hallucinations suffer from two drawbacks: Lack of scalable token-level rewards and Neglect of visual-anchored tokens. |
| Approach: | They propose a Token Preference Optimization model with self-calibrated rewards . they propose based on visual-anchored tokens and visual-aware training objective . |
| Outcome: | The proposed model improves hallucination performance by focusing on visual-anchored tokens without fine-grained annotations. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. |
| Approach: | They propose three confidence-based methods to enhance LVLMs' perception . they propose probabilistic and consistency-based signals are more reliable indicators . |
| Outcome: | Experiments on three LVLMs across three VQA datasets show that LVLs possess a reasonable perception level but there is room for improvement. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) generate contextually relevant responses by jointly interpreting visual and textual inputs. |
| Approach: | They propose a method to classify whether an input token is visually grounded by reinterpreting question prompts or replacing the detected absent tokens during generation. |
| Outcome: | The proposed method mitigates the models’ tendency to falsely presume the visual presence of text input and its generality across various LVLMs. |
Copied to clipboard
| Challenge: | Existing approaches to generating models rely on text and images, but video content is a rich source of multimodal knowledge. |
| Approach: | They propose a framework that dynamically retrieves videos based on their relevance with queries . they use large video language models to represent video content for retrieval . |
| Outcome: | The proposed framework retrieves videos based on relevance with queries and integrates both visual and textual information. |
Copied to clipboard
| Challenge: | Mainstream VLPs have significant security implications, but their security implications have not been thoroughly examined. |
| Approach: | a study evaluates the security of visual language projectors by comparing them to uncompressed projector. |
| Outcome: | The evaluation reveals significant differences in security profiles between compressed and uncompressed projectors. |
Copied to clipboard
| Challenge: | Existing methods for collecting medical data are expensive and time-consuming. |
| Approach: | They propose a method to train a large-scale LVLM capable of auto-generating medical visual instruction data to improve data efficiency. |
| Outcome: | The proposed method shows that it performs well across three major visual question answering (VQA) benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to prune redundant vision tokens struggle in shallow layers due to the lack of contextual information. |
| Approach: | They propose a layer-wise contextualized visual token pruning method that uses a plug-and-play Pruning Module to prune redundant vision tokens. |
| Outcome: | The proposed method outperforms training-free pruning methods under equal token budgets and surpasses training based methods with comparable supervision. |
Copied to clipboard
| Challenge: | According to psychological and neuroscientific research, a high-stress environment can restrict attentional resources and intensify negative affect, thereby impairing the ability to understand emotions. |
| Approach: | They constructed a large-vision language model that combines race, gender, and age group and used the Pretend prompt technique to induce LVLMs to interpret others’ emotions. |
| Outcome: | The results suggest that the effects of high-stress and demographic attributes identified in human research may also be reflected in LVLMs. |
Copied to clipboard
| Challenge: | Existing approaches to enhance multilingual reasoning capabilities rely on costly multilingual training or employ prompting with external translation tools. |
| Approach: | They propose a training-free inference-time method to enhance multilingual reasoning capabilities via Representation Engineering without additional training data or tools. |
| Outcome: | The proposed method outperforms existing methods on four reasoning benchmarks in English and Thai and Swahili. |
Copied to clipboard
| Challenge: | Existing research has focused on mitigating object hallucinations but often overlooks more complex relation hallucines, especially action relations involving interactions between objects. |
| Approach: | They propose a framework to locate action-relevant image regions and enhance the LVLM’s attention to those regions by using a Relation-aware Visual Enhancement method. |
| Outcome: | The proposed method achieves superior performance in mitigating action-relation hallucinations with negligible additional inference cost. |
Copied to clipboard
| Challenge: | Jailbreak attacks, where harmful prompts bypass generative models’ built-in safety, raise serious concerns about model vulnerability. |
| Approach: | They propose to reframe the standard generation task as a binary classification problem to assess model refusal tendencies for both harmful and benign queries. |
| Outcome: | The proposed defenses improve model safety or optimize the trade-off between safety and helpfulness. |
Copied to clipboard
| Challenge: | Existing defense strategies neglect visual threats and lack of fine-grained specificity regarding specific attack semantics. |
| Approach: | They propose a black-box defense framework that maps unsafe concepts to fine-grained, constructive Safe Concepts. |
| Outcome: | a new black-box defense framework enhances robustness against jailbreak attacks . it maps detected unsafe concepts to fine-grained, constructive Safe Concepts . the proposed framework is available for free at http://www.epa.org/recon/ . |
Copied to clipboard
| Challenge: | a recent study has found that large vision–language models are vulnerable to visual biases that inflate scores without altering semantic content. |
| Approach: | They propose a novel meta-evaluation benchmark that exhibits diverse score distributions. |
| Outcome: | The proposed model exhibits vulnerability across all domains, and combines multiple biases amplifies their effects, and pairwise evaluations are similarly susceptible. |
Copied to clipboard
| Challenge: | Existing approaches to detect toxicity in online multimodal environments require common-sense reasoning and contextual awareness. |
| Approach: | They propose a hybrid neurosymbolic framework that unifies distillation of implicit contextual knowledge from Large Vision-Language Models and infusion of explicit relational semantics through sub-graphs from Knowledge Graphs. |
| Outcome: | The proposed framework outperforms state-of-the-art models on two datasets with improvements of 0.5%, and 10.6% in HatefulMemes Benchmark. |
Copied to clipboard
| Challenge: | Existing methods for accelerating Large Vision-Language Models lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. |
| Approach: | They propose EffiVLM-BENCH framework for evaluating absolute performance and generalization and loyalty. |
| Outcome: | The proposed framework offers insights into optimal strategies for accelerating LVLMs. |
Copied to clipboard
| Challenge: | Recent advances in large vision-language models have improved causal reasoning abilities . however, current models struggle with tasks like causal reasoning . |
| Approach: | They propose a fine-grained and unified definition of causality involving interactions between humans and objects. |
| Outcome: | The proposed model surpasses traditional commonsense causality by including explicit causal graphs . it also shows that current LVLMs can benefit from a causally inspired prompting strategy . |
Copied to clipboard
| Challenge: | Large vision-language models have demonstrated strong capabilities in open-world visual understanding, but it is not clear how they address demographic biases in real life. |
| Approach: | They propose a method to assess visual fairness in LVLMs by question-answering/classification tasks. |
| Outcome: | The proposed approach improves transparency and offers a scalable solution for fairness mitigation. |
Copied to clipboard
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
Copied to clipboard
| Challenge: | ELENA is a framework for embodied emotion analysis using large vision language models . ELEna uses attention maps and a persistent bias towards the facial region . |
| Approach: | They propose a framework that utilizes large vision language models to generate ELENA . they propose to use attention maps to describe emotional reactions from body parts . |
| Outcome: | The proposed framework outperforms baseline models without fine-tuning . it uses large vision language models to generate embodied emotion narratives . |
Copied to clipboard
| Challenge: | Large Vision-Language Models have demonstrated impressive performance on vision-language reasoning tasks, but their potential for zero-shot fine-grained image classification remains underexplored. |
| Approach: | They propose a method that transforms zero-shot fine-grained image classification into a visual question-answering framework. |
| Outcome: | The proposed method outperforms the current state-of-the-art approach and outperformed existing methods. |
Copied to clipboard
| Challenge: | Existing visual token pruning methods leverage simple metrics derived from human experience, such as attention or similarity, to rank and select tokens within a highly entangled feature space. |
| Approach: | They propose a novel visual token pruning method that uses a concept-driven paradigm to quantify the Marginal Semantic Gain of each token's contribution to uncovered concepts. |
| Outcome: | The proposed method outperforms state-of-the-art methods in a concept-driven model while maintaining semantic completeness. |
Copied to clipboard
| Challenge: | Existing benchmarks overlook conflicts between visual and textual evidence and the importance of generating deflections when incomplete knowledge is retrieved. |
| Approach: | They propose a dynamic curation pipeline that preserves benchmark difficulty over time . they propose 'vlm-DeflectionBench' benchmark to probe model behaviour under conflicting evidence . |
| Outcome: | The proposed benchmarks overlook conflicts between visual and textual evidence and are prone to obsolescence . the proposed benchmark is based on 2,775 samples spanning diverse retrieval settings . |
Copied to clipboard
| Challenge: | Existing self-evaluation methods rely on a model’s ability to estimate the correctness of its own outputs, but they depend heavily on language priors and are therefore ill-suited for evaluating vision-conditioned predictions. |
| Approach: | They propose a vision-aware uncertainty quantification framework that measures how strongly a model’s output depends on visual evidence. |
| Outcome: | The proposed framework outperforms existing methods across multiple datasets. |
Copied to clipboard
| Challenge: | Existing models for large vision-Language Models lack reliable visual hallucination . despite advances in multimodal perception, they are still limited in real-world applications . |
| Approach: | They propose a framework that recasts attention calibration as a one-step score-based denoising process. |
| Outcome: | The proposed framework achieves superior performance on generative and discriminative tasks while maintaining efficiency comparable to standard decoding. |
Copied to clipboard
| Challenge: | Decode-Only models propagate information from left to right, but the model's attention still focuses on the visual representations, resulting in hallucinations. |
| Approach: | They propose to leverage the core information embedded in semantic representations to enhance the model's visual understanding by leveraging the attention distributions. |
| Outcome: | The proposed method reduces hallucinations by 80% by aligning the attention distribution with the actual information flow. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are being explored in medicine but their ability to conduct complex real-world telemedicine consultations remains underexplored. |
| Approach: | They propose to use large vision-language models to conduct telemedicine consultations using a framework that simulates patient variability and evaluates diagnostic accuracy and dialogue quality via Assessor Agent. |
| Outcome: | The proposed framework compares diagnostic strategies for open and closed-source LVLMs and shows that multimodal dialogue improves F1 score by 6.5% over non-dialogue settings. |
Copied to clipboard
| Challenge: | Existing pipelines for generating high-quality, ultra-detailed image captions are limited by the scarcity of image caption data. |
| Approach: | They propose a pipeline for generating high-quality, ultra-detailed image captions that integrates both pre-processing and post-processor stages. |
| Outcome: | The proposed pipeline improves LVLMs' perception and cognitive abilities across multiple vision-language benchmarks. |
Copied to clipboard
| Challenge: | Existing GUI benchmarks lack fine-grained diagnostics to identify which capabilities lead to task failures. |
| Approach: | They propose a multilingual P R GUI Benchmark to assess LVLMs' language capabilities . they propose XLI to align non-English hidden states with English ones during inference . |
| Outcome: | The proposed benchmark reveals consistent gaps between English and non-English settings . it reduces the cross-lingual gaps with an average gain of 6.5% in non- English settings compared to static benchmarks . |
Copied to clipboard
| Challenge: | Existing methods for medical visual question answering focus on image–caption pairs, limiting the model’s ability to learn relevant medical knowledge during training. |
| Approach: | They propose to synthesize instruction data from image–caption pairs and incorporate a multimodal medical knowledge graph to assist LVLMs in synthesizing knowledge-intensive instruction data. |
| Outcome: | The proposed model outperforms existing methods on the public datasets Slake and VQA-RAD by 4.16% and 4.50%. |
Copied to clipboard
| Challenge: | Existing video benchmarks often resemble image-based questions with scans of only a few key frames, without deep temporal reasoning. |
| Approach: | They propose a video benchmark to assess whether large vision-language models can genuinely think with videos rather than perform superficial frame-level analysis. |
| Outcome: | The proposed benchmark consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. |
Copied to clipboard
| Challenge: | Existing studies attribute object hallucinations to linguistic priors and data biases . MFCD method removes hallucinian distribution in the original output distribution . |
| Approach: | They propose a method that removes the hallucination distribution in the original output distribution . they propose MFCD to mitigate hallucinism in large visual-language models . |
| Outcome: | The proposed method reduces hallucination distributions without training or external tools . the proposed method can be applied to various LVLMs without modifying model architecture or training . |
Copied to clipboard
| Challenge: | Existing methods for enhancing understanding and reasoning abilities in graphbased tasks focus on specific graph types or tasks, posing challenges in designing versatile systems suitable for various tasks and graphs across diverse domains. |
| Approach: | They propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks. |
| Outcome: | Extensive evaluations on 14 LVLMs reveal that LVLs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information. |
Copied to clipboard
| Challenge: | Existing Large Vision-Language Models (LVLMs) learn visual capacity through visual instruction tuning. |
| Approach: | They propose a method for LVLMs to be trained by selective layers tuning . they propose removing non-critical layers outside the visual region . |
| Outcome: | The proposed approach preserves nearly 99% of visual performance and improves textual task results while reducing training time. |
Copied to clipboard
| Challenge: | Existing models assess spatial capabilities from a static, single-view and egocentric perspective, failing to capture the dynamic nature of real-world spatial cognition. |
| Approach: | They propose a benchmark to diagnose spatial reasoning capabilities using a 360 field of view. |
| Outcome: | The proposed benchmark evaluates allocentric and egocentric reasoning capabilities from multiple perspectives in high-quality 3D environments. |
Copied to clipboard
| Challenge: | Recent Large Vision-Language Models (LVLMs) have shown remarkable success in general semantic understanding, but struggle with 3D spatial reasoning tasks. |
| Approach: | They propose a framework to help vision encoders internalize 3D geometric information using only standard 2D images. |
| Outcome: | The proposed framework achieves State-of-the-Art (SOTA) performance across various model architectures. |
Copied to clipboard
| Challenge: | Large vision-language models have shown impressive ability in various language tasks, especially with their emergent in-context learning capability. |
| Approach: | They propose a causal reasoning benchmark for multi-modal in-context learning from large vision-language models that incorporates visual inputs. |
| Outcome: | The proposed model outperforms existing models on three visual causal reasoning tasks and demonstrates their strengths and weaknesses. |
Copied to clipboard
| Challenge: | Large vision-Language Models suffer from noisy supervision and semantic ambiguity in self-supervised settings. |
| Approach: | They propose a self-supervised framework that constructs reliable preference triplets . they propose 'trident' objective that enforces semantic separation between the triplet components . |
| Outcome: | The proposed framework outperforms state-of-the-art RLHF and RLAIF benchmarks on LLaVA-1.5-7B and achieves 95.70% precision on POPE using only 4k self-generated triplets and a single epoch. |
Copied to clipboard
| Challenge: | Large vision-language models often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation. |
| Approach: | They propose a visual reasoning framework that decouples vision-reasoning capabilities and multi-run proactive perception. |
| Outcome: | The proposed framework outperforms existing models on benchmarks for open-source and closed-source models with 13.2% performance gain. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are capable of learning from vast webscale datasets but pose privacy risks as they can unintentionally memorize sensitive information. |
| Approach: | They propose a Reliable Multi-hop and Multi-image Memorization Benchmark that ensures robust foundational learning through principled data scaling and reasoning-aware QA pairs. |
| Outcome: | Extensive experiments show that ReMem provides a reliable framework for diagnosing both learning and unlearning behaviors in Large Vision-Language Models. |
Copied to clipboard
| Challenge: | Existing work has explored unimodal biases in visual question answering, but the problem of selection bias in Multiple-Choice Question Answering (MCQA) remains underexplored. |
| Approach: | They propose a method that mitigates bias without retraining and is compatible with frozen LVLMs. |
| Outcome: | The proposed method mitigates bias without retraining and is compatible with frozen LVLMs. |
Copied to clipboard
| Challenge: | Existing decoding-based approaches do not explicitly decouple visual evidence from mixed vision–language representations. |
| Approach: | They propose to decouple visual evidence from mixed vision–language representations by dynamically identifying layers enriched with visual information and performing intra-layer decoupling to extract aggregated visual evidence. |
| Outcome: | Experiments show that DiVE achieves state-of-the-art performance on multiple benchmarks. |
Copied to clipboard
| Challenge: | MM-JudgeBench is the first large-scale benchmark for multilingual and multimodal judge model evaluation. |
| Approach: | They propose a multilingual benchmark for multilingual and multimodal judge model evaluation that includes over 60K pairwise preference instances spanning 25 typologically diverse languages. |
| Outcome: | The proposed benchmark includes over 60K pairwise preference instances spanning 25 languages. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are capable of processing visual inputs, but are susceptible to hallucinations. |
| Approach: | They propose a method to localize and localize specific visual tokens, which are defined as **Inert Tokens**, across layers, revealing a rigid semantic collapse. |
| Outcome: | The proposed approach reduces the likelihood of LVLMs being hijacked by visual inputs while maintaining general capabilities. |
Copied to clipboard
| Challenge: | Existing preference learning-based approaches rely on proprietary models to construct preference datasets, causing a distributional mismatch between the proprietary and target models. |
| Approach: | They propose a framework that aligns LVLMs using in-distribution data derived from the model's intrinsic knowledge. |
| Outcome: | The proposed framework surpasses baselines in hallucination mitigation while requiring only 5.2k samples. |
Copied to clipboard
| Challenge: | Existing work on large vision–language models focuses on point-and-click interaction, while remote-control interaction is underexplored. |
| Approach: | They propose a topology-aware training framework that injects topology awareness into LVLMs. |
| Outcome: | The proposed model achieves 68.3% success rate on TVWorld-N, surpassing closed-source benchmarks and state-of-the-art (SOTA) benchmarks show that existing agents lack topology awareness for focus-based, long-horizon TV navigation. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models and Large Vision Language Model (LVLMs) have demonstrated promising capabilities in complex reasoning tasks, but low-resource contexts like Bangla are underexplored. |
| Approach: | They propose a multilingual medical visual question answering dataset using Bangla. |
| Outcome: | The proposed model performs well on generalized visual tasks but struggles with fine-grained diagnostic reasoning, achieving low accuracy in specialized categories. |
Copied to clipboard
| Challenge: | Typical large vision-language models emphasize vision-to-language alignment while overlooking fine-grained visual information. |
| Approach: | They introduce autoregressive semantic visual reconstruction (ASVR) that enables joint learning of visual and textual modalities within a unified autoregression framework. |
| Outcome: | The proposed model improves baselines and multimodal understanding benchmarks by 2-3%. |
Copied to clipboard
| Challenge: | Current methods for radiology report generation rely on encoder-decoder based frameworks that fail to integrate multimodal clinical evidence with domain-specific knowledge. |
| Approach: | They propose a multimodal dual-path framework that synergistically integrates large vision-language models and large language models for radiology report generation. |
| Outcome: | The proposed framework improves on the public MIMIC-CXR benchmark and shows that it is superior to state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. |
| Approach: | They propose a multilingual chart question answering benchmark that enables efficient multilingual generation via data translation and code reuse. |
| Outcome: | The proposed benchmark systematically evaluates multilingual chart understanding on state-of-the-art LVLMs and shows a significant performance gap between English and other languages. |
Copied to clipboard
| Challenge: | a large vision-language model can generate hallucinations inconsistent with visual input . a lightweight method that embeds the last input token as a grounding signal reduces the likelihood of hallucinosity. |
| Approach: | They propose a training-free mitigation strategy that harnesses the hidden state of the last input token as a grounding signal to maintain visual fidelity throughout decoding and curb hallucinations. |
| Outcome: | The proposed method outperforms state-of-the-art methods on CHAIR, AMBER, and MMHal benchmarks. |
Copied to clipboard
| Challenge: | Existing verbalized confidence calibration methods for large vision language models optimize a single holistic confidence score using binary answer-level correctness. |
| Approach: | They propose a reinforcement learning framework that explicitly decouples confidence into visual and reasoning confidence. |
| Outcome: | Experiments show that the proposed framework decouples confidence into visual and reasoning confidence while suppressing ungrounded hallucinations while preserving valid perception. |
Copied to clipboard
| Challenge: | Hallucinations in Large Vision-Language Models (LVLMs) are a persistent challenge, stemming from inadequate integration of visual information during multimodal reasoning. |
| Approach: | They propose a visual feature incorporation method that encourages the model to learn visually-informed textual embeddings distinct from those of the base LLM and promotes a more balanced attention distribution. |
| Outcome: | The proposed method significantly reduces hallucinations and fosters more balanced multimodal reasoning. |
Copied to clipboard
| Challenge: | Existing multimodal reasoning benchmarks for large vision-language models emphasize single-image analysis and fail to exploit contextual information across multiple images. |
| Approach: | They propose a benchmark to evaluate Olympiad-level reasoning when evidence is distributed over multiple images. |
| Outcome: | The proposed model outperforms existing models on bi-image Olympiads and Gemini-3-Pro on multimodal Olympiad-level reasoning tasks. |
Copied to clipboard
| Challenge: | FineState-Bench evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. |
| Approach: | They propose a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. |
| Outcome: | The proposed benchmark evaluates whether an agent can ground an instruction to the intended UI control and reach the exact target state. |