Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing Vid-LLMs lack robust mechanisms for maintaining grounded spatiotemporal beliefs under conversational feedback. |
| Approach: | They propose a negation-based gaslighting evaluation framework and introduce a benchmark to investigate spatiotemporal sycophancy. |
| Outcome: | The proposed framework evaluates state-of-the-art Vid-LLMs across video understanding tasks. |
Similar Papers
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Current sycophancy research has largely overlooked its specific manifestations in the video-language domain. |
| Approach: | They propose a video-LLM sycophancy benchmarking and evaluation to evaluate scophancies in video-LLMs. |
| Outcome: | The proposed benchmark evaluates sycophantic behavior in state-of-the-art Video-LLMs across diverse question formats, prompt biases, and visual reasoning tasks. |
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models (2025.findings-emnlp)
Copied to clipboard
Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao, Yufan Zhou, Yixin Cao, Qifan Wang, Weifeng Ge, Lifu Huang
| Challenge: | Video Large Language Models (VLMs) have been praised for their performance in coarse-grained video understanding but still face ineffective temporal grounding and inadequate timestamp representations. |
| Approach: | They propose a novel Video-LLM that senses and reasoned over specific video moments with fine-grained temporal precision. |
| Outcome: | The proposed model surpasses existing models in fine-grained video understanding tasks and exhibits strong potential as a general video understanding assistant. |
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLM (2025.coling-main)
Copied to clipboard
| Challenge: | Existing video LLMs excel at capturing the overall description of a video but lack the ability to demonstrate an understanding of temporal dynamics and localized content within the video. |
| Approach: | They propose a Time-Perception Enhanced Video Grounding via Boundary Perception and Temporal Reasoning to improve LLMs' understanding of video temporality. |
| Outcome: | The proposed method improves on three datasets: ActivityNet, Charades, and DiDeMo (up to 11.2% improvement on R@0.3). |
Dynamic Attention-Guided Context Decoding for Mitigating Context Faithfulness Hallucinations in Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods, such as a n-terminal coding, do not provide accurate data for large language models. |
| Approach: | They propose a lightweight framework that leverages attention distributions and uncertainty signals in a single-pass decoding. |
| Outcome: | Experiments on open-book QA datasets show that DAGCD improves faithfulness and robustness while preserving computational efficiency. |
Chaos with Keywords: Exposing Large Language Models Sycophancy to Misleading Keywords and Evaluating Defense Strategies (2024.findings-acl)
Copied to clipboard
| Challenge: | sycophancy is a type of hallucination in Large Language Models, which can lead to false information being presented. |
| Approach: | They explore the sycophantic tendencies of Large Language Models where models provide accurate answers even if they are not entirely correct. |
| Outcome: | The proposed models generate factually correct statements even when they are not completely correct. |
Fairness Evaluation and Inference Level Mitigation in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drift, and the propagation of unwanted patterns during extended dialogues. |
| Approach: | They propose a pruning-based framework that detects context-aware neuron activations and applies adaptive masking to modulate their influence during generation. |
| Outcome: | The proposed framework detects context-aware neuron activations and applies adaptive masking to modulate their influence during generation. |
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects (2026.findings-acl)
Copied to clipboard
Jun Zhang, Yicheng Ji, Feiyang Ren, Yihang Li, Bowen Zeng, Zonghao Chen, Ke Chen, Lidan Shou, Gang Chen, Huan Li
| Challenge: | Large Vision-Language Models are hindered by a systemic efficiency barrier known as visual token dominance. |
| Approach: | They propose a systematic taxonomy of efficiency techniques structured around the inference lifecycle . they examine visual encoding, prefilling, and decoding to understand bottlenecks . |
| Outcome: | The proposed techniques reveal how upstream decisions dictate downstream bottlenecks . the proposed techniques include hybrid compression and modality-aware decoding . |
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs. |
| Approach: | They investigate why LLMs exhibit sycophancy when challenged in subsequent conversational turns, yet perform well when evaluating conflicting arguments presented simultaneously? |
| Outcome: | The proposed models are more likely to endorse a user’s counterargument when framed as a follow-up from a users, rather than when both responses are presented simultaneously for evaluation. |
Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multimodal large language models exhibit a pronounced form of visual sycophantic behavior when they process image inputs. |
| Approach: | They propose a technique that allows multimodal large language models to engage in reflective reasoning and determine whether a user’s instruction is misleading or corrective. |
| Outcome: | The proposed model resists misleading instructions but is stubborn even if it is wrong. |
Distorted or Fabricated? A Survey on Hallucination in Video LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Despite significant advances in video-language modeling, hallucinations remain a persistent challenge in video large language models. |
| Approach: | They present a systematic taxonomy that categorizes hallucinations into two core types: dynamic distortion and content fabrication. |
| Outcome: | The proposed taxonomy categorizes hallucinations into two core types: dynamic distortion and content fabrication. |