Papers by Gang Yan

5 papers
CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Cross-modal retrieval tasks are used to retrieve data from one modality or another based on a query from another modality.
Approach: They propose a generative cross-modal retrieval framework based on coarse-to-fine semantic modeling . they propose combining K-Means and RQ-VAE to discretize multimodal data into token sequences that support autoregressive generation.
Outcome: The proposed framework achieves excellent performance and efficiency in multimodal retrieval tasks.
Self-Awareness before Action: Mitigating Logical Inertia via Proactive Cognitive Awareness (2026.findings-acl)

Copied to clipboard

Challenge: Existing work on abductive and long-context reasoning reports that current models still lack self-awareness of missing premises.
Approach: They propose a reasoning framework that introduces self-awareness of missing premises before making the final decision.
Outcome: SABA achieves best performance on all three difficulty splits of detective puzzle benchmark . a small early mistake can remain uncorrected and can guide later reasoning .
Visual Enhanced Entity-Level Interaction Network for Multimodal Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to generate concise summarizations rely on coarse-grained textual and visual information, but they are underutilized.
Approach: They propose a Visual Enhanced Entity-Level Interaction Network to address underutilization of multimodal inputs at a fine-grained level.
Outcome: The proposed model outperforms existing models on two MMS datasets and proposes new metrics to measure factual consistency of entities in the output.
Entity-level Interaction via Heterogeneous Graph for Multimodal Named Entity Recognition (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for name-based entity recognition neglect the integrity of entity semantics and conduct cross-modal interaction at token-level.
Approach: They propose a multimodal named entity recognition model that captures visual information and fuses it into tokens to rid non-entity tokens of visual noise.
Outcome: The proposed model captures entity-related visual information and fuses it into tokens . it eliminates visual noise and makes non-entity tokens easily misidentified as entities .
ASD-iLLM:An Intervention Large Language Model for Autistic Children based on Real Clinical Dialogue Intervention Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, leveraging large language models (LLMs) for autism intervention is a significant yet challenging task, especially when directly employing LLMs as an intervention doctor.
Approach: They propose a framework for training LLMs to conduct dialogue interventions in accordance with the principles of Applied Behavior Analysis (ABA) they also propose 'role-play' strategy in which LLM act as autistic children to comprehensively evaluate the doctor model's capabilities at the dialogue level.
Outcome: The proposed framework outperforms existing models in both automatic and human evaluation, with intervention strategies and dialogue style more closely resembling those of clinical intervention doctors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations