Papers by Jianfeng Dong

9 papers
Language Models as Inductive Reasoners (2024.eacl-long)

Copied to clipboard

Challenge: Inductive reasoning is a core component of human intelligence.
Approach: They propose a task to induce natural language rules from natural language facts using natural language as representation for knowledge instead of formal language.
Outcome: The proposed task surpasses baselines in both automatic and human evaluations.
Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective (2025.acl-long)

Copied to clipboard

Challenge: a number of attention-based large language models (LLMs) focus on individual head contributions, but the precise interaction mechanisms between attention heads remain poorly understood.
Approach: They propose a game-theoretic attention calibration method that uses the Harsanyi dividend . they selectively retain heads demonstrating significant cooperative gains and apply fine-grained adjustments to remaining heads .
Outcome: The proposed framework is based on the Harsanyi dividend, a concept from cooperative game theory.
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on enhancing multi-scale clip representations but lack robust data alignment . inherent data uncertainty renders PRVR vulnerable to distractor videos with spurious similarities .
Approach: proposed framework for partially relevant video retrieval aims to retrieve untrimmed videos partially relevant to a given query.
Outcome: The proposed framework can be seamlessly integrated into existing architectures.
Multimodal Language Models See Better When They Look Shallower (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that multimodal large language models extract visual features from the final layers of a pretrained Vision Transformer.
Approach: They propose a feature fusion method that strategically incorporates shallower layers . they propose MLLMs that extract visual features from the final layers of a pretrained Vision Transformer .
Outcome: The proposed method outperforms deep layers on fine-grained visual tasks . it is the first comprehensive study of visual layer selection for MLLMs .
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning (2025.coling-main)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have impressive capabilities in multi-modal context comprehension, but they still suffer from hallucination problems due to inconsistent outputs with the image content.
Approach: They propose a training-free framework MVP to reduce hallucinations in Large Vision-Language Models . they propose multi-view information-seeking strategy to perceive the comprehensive information in the image .
Outcome: The proposed framework reduces hallucinations in large vision-language models by combining multi-view multi-path reasoning with multi-vision multi-path reasoning.
Implicit Deep Latent Variable Models for Text Generation (D19-1)

Copied to clipboard

Challenge: Variational auto-encoders have been used for text generation but their representation power is limited due to two reasons.
Approach: They advocate sample-based representations of variational distributions for natural language . they further develop an LVM to directly match the aggregated posterior to the prior .
Outcome: The proposed model can be viewed as a natural extension of VAEs with a regularization of maximizing mutual information, mitigating the "posterior collapse" issue.
Towards Robust Temporal Activity Localization Learning with Noisy Labels (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for temporal activity localization are expensive and difficult to satisfy due to subjective labeling.
Approach: They propose a new TAL setting where a TAL model should be robust to mixed training data with noisy moment boundaries.
Outcome: The proposed method is significantly more robust to noisy training data than existing methods.
Adaptive Proposal Generation Network for Temporal Sentence Localization in Videos (2021.emnlp-main)

Copied to clipboard

Challenge: Temporal sentence localization in videos is an important yet challenging task in natural language processing.
Approach: They propose an Adaptive Proposal Generation Network to maintain the segment-level interaction while speeding up the efficiency.
Outcome: The proposed model outperforms state-of-the-art methods on three challenging benchmarks.
Reasoning Step-by-Step: Temporal Sentence Localization in Videos via Deep Rectification-Modulation Network (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for temporal sentence localization in videos focus on visual content, but they are insufficient to model complex video contents.
Approach: They propose a deep rectification-modulation network to correct attention misalignment . they use sentence information to capture frame-to-frame relation .
Outcome: The proposed method achieves state-of-the-art performance on three public datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations