Papers by Yunxin Li
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment (2024.acl-long)
Copied to clipboard
| Challenge: | Recent Large Multimodal Models (LMMs) focus on visual knowledge-dimension alignment, but ignore visual knowledge. |
| Approach: | They propose a cognitive visual-language mapper that integrates visual-linguistic knowledge alignment with a fine-grained knowledge Adapter. |
| Outcome: | The proposed model significantly improves LMMs on knowledge-based visual question answering (VQA) it also improves the performance of other models, including GPT-4V and Gemini-Pro. |
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs (2026.acl-long)
Copied to clipboard
Zhenyu Liu, Xuanyu Zhang, Yunxin li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang
| Challenge: | despite significant progress, full-duplex SLMs are constrained by severe modality interference, authors say . modality interferes with acoustic and semantic modeling, making them unintelligent and unnatural . authors propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers . |
| Approach: | They propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel. |
| Outcome: | The proposed method significantly advances the state of the art on full-duplex benchmarks . it decouples conflicting modalities in deep layers while preserving cross-modality coherence . |
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for generating product descriptions from images are inaccurate and generic . e-commerce product descriptions are important for content marketing and increasing engagement . |
| Approach: | They propose a new setting for generating product descriptions from images, augmented by marketing keywords. |
| Outcome: | The proposed approach improves the accuracy and diversity of product descriptions by up to 3.3% on Rouge-L and 9.4% on D-5. |
Benchmarking LLM’s Capability in Reasoning over Conflicting Web References (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) integrated with retrieval-augmented generation (RAG) are a dominant framework for building intelligent assistants. |
| Approach: | They propose a benchmark to evaluate LLMs' reasoning capability over real-world conflicting documents retrieved from the web. |
| Outcome: | The proposed benchmark evaluates LLMs' reasoning capability over real-world conflicting documents retrieved from the web. |
U-CORE: A Unified Deep Cluster-wise Contrastive Framework for Open Relation Extraction (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing methods for Relation Extraction (RE) are limited due to the overlap between predefined and undefined relations. |
| Approach: | They propose a unified framework for both Zero-shot and Unsupervised Relation Extraction tasks by leveraging techniques from Contrastive Learning and Clustering. |
| Outcome: | The proposed framework improves on three well-known datasets showing an average improvement of 7.35% ARI on Zero-shot ORE tasks and 15.24% ARI for Unsupervised ORE. |
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget (2024.acl-long)
Copied to clipboard
| Challenge: | Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) but memory-constrained devices are a major concern in edge AI training and serving. |
| Approach: | They propose a framework for efficient serving of MoE-based large language models with tunable memory budgets. |
| Outcome: | Experiments show that SwapMoE can reduce memory consumption while maintaining reasonable accuracy. |
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts (2026.acl-long)
Copied to clipboard
Zhenyu Liu, Yunxin li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang
| Challenge: | Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. |
| Approach: | They propose a unified speech and music generation model built upon a novel framework . they propose specialized MoE architectures and curated training strategies to tackle data imbalances . |
| Outcome: | The proposed model achieves state-of-the-art performance on major speech and music generation benchmarks. |
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing GUI agents perform poorly on multi-application tasks, stalling at early sub-goals. |
| Approach: | They propose to assess GUI Agents on complex multi-step tasks that mirror real-world professions. |
| Outcome: | The proposed benchmark contains 181 tasks with an average of 5.0 sub-goals across 17 common desktop applications, of which 78% are inherently multi-application. |
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement (2025.acl-long)
Copied to clipboard
Maosongcao Maosongcao, Taolin Zhang, Mo Li, Chuyu Zhang, Yunxin Liu, Conghui He, Haodong Duan, Songyang Zhang, Kai Chen
| Challenge: | Existing high-quality human-annotated SFT data is a bottleneck for Large Language Models (LLMs). |
| Approach: | They propose a two-stage synthetic data generation framework that incorporates World Knowledge Trees and Self-Reflection Refinement to produce high-quality SFT data at scale. |
| Outcome: | The proposed model fine-tuned on 20K condor-generated samples achieves superior performance compared to instruct model trained with RLHF. |
A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for conditional inference on joint textual and visual clues lack multimodal context reasoning capability. |
| Approach: | They propose a multi-modal context reasoning approach that embeds textual semantics and objective image information into the pretrained language model to perform context reasoning. |
| Outcome: | The proposed approach improves on two data sets and shows 4.8% gain on the PMR. |
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint (2025.emnlp-main)
Copied to clipboard
Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu, Xiangyu Li, Hao Wen, Yizhen Yuan, Huiwen Zheng, Yan Liang, Yuanchun Li, Yunxin Liu
| Challenge: | Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints. |
| Approach: | They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking . |
| Outcome: | The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints. |
VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension (2025.acl-long)
Copied to clipboard
| Challenge: | Existing video evaluation benchmarks focus on a single language, typically English, and feature videos rooted in Western cultural contexts. |
| Approach: | They propose a video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension. |
| Outcome: | The proposed video evaluation benchmark bridges cultural, linguistic, and domain divides . existing benchmarks only feature videos from YouTube, Shutterstock, or established video datasets based on cultural diversity . |
A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text (2023.acl-long)
Copied to clipboard
| Challenge: | Pretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text, but their performance drops drastically when confronted with linguistically complex texts. |
| Approach: | They propose an end-to-end Neural Divide-and-Conquer Reasoning framework for linguistically complex texts that they struggle to comprehend. |
| Outcome: | The proposed framework significantly improves performance in complex image-text reasoning problem. |
Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision (2026.acl-long)
Copied to clipboard
Ge Chang, Jinbo Su, Jiacheng Liu, Pengfei Yang, Yuhao Shang, Huiwen Zheng, Hongli Ma, Yan Liang, Yuanchun Li, Yunxin Liu
| Challenge: | Existing methods for integrating textual graphs with LLMs are limited by symbolic inference and high annotation costs. |
| Approach: | They propose a textual graph reasoning framework that integrates textual diagrams with large language models. |
| Outcome: | The proposed approach achieves 15.6% accuracy and 17.2% in F1 score on three common datasets. |