Papers by Qiuchi Li
What Does Your Smile Mean? Jointly Detecting Multi-Modal Sarcasm and Sentiment Using Quantum Probability (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to model multi-modal sarcasm and sentiment are based on quantum probability . sarcasm and feelings embody intrinsic uncertainty of human cognition . |
| Approach: | They propose a quantum probability-driven multi-task learning framework for sarcasm and sentiment recognition using quantum superpositions and quantum interference. |
| Outcome: | The proposed model achieves state-of-the-art in multi-modal sarcasm and sentiment recognition. |
Towards the Law of Capacity Gap in Distilling Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Language model (LM) distillation aims at distilling knowledge in a large teacher LM to a small student one. |
| Approach: | They propose to use the law of capacity gap to distill knowledge from a large teacher to a small student model. |
| Outcome: | The proposed model outperforms other language models on a larger scale by using the law of capacity gap inducted from a preliminary study on small-scale (3B) LMs. |
Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) aims to mitigate the hallucination of Large Language Models (LLMs) however, external knowledge may contain noise and conflict with parametric knowledge of LLMs, leading to degraded performance. |
| Approach: | They propose a Dual-Stream Knowledge-Augmented Framework for Shared-Private Semantic Synergy that refines the traditional self-attention into a mixed-attention that distinguishes shared and private semantics for a controlled knowledge integration. |
| Outcome: | Extensive experiments show that the proposed framework achieves a superior performance over baselines. |
ConMA : Confidence-Guided Kernel Sampling with Multi-Stage Aggregation for LLM Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to test-time scaling rely on external verifiers and one-shot independent sampling. |
| Approach: | They propose a test-time scaling framework that reallocates a fixed inference budget into iterative sample–filter–diversify–select cycles. |
| Outcome: | ConMA outperforms baselines on multiple benchmarks while converging early with only 18 samples on average, substantially reducing inference cost. |
How does Attention Affect the Model? (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of attention in NLP do not consider changes in semantic capability of different components. |
| Approach: | They propose a framework that exploits a convex hull representation of sequence semantics in an n-dimensional Semantic Euclidean Space and defines indicators to capture the impact of attention on sequence semantic. |
| Outcome: | The proposed framework exploits a convex hull representation of sequence semantics in an n-dimensional Semantic Euclidean Space and defines indicators to capture the impact of attention on sequence semantic. |
Task-agnostic Distillation of Encoder-Decoder Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing distillation methods that focus on encoder-only LMs fail to handle the distillation of encoder decoder LM. |
| Approach: | They propose a method that finetunes pretrained language models (LMs) they propose 'MiniEnD' that allows for task-agnostic distillation of LMs. |
| Outcome: | The proposed distillation method is generally effective and competitive compared to other alternatives. |
CNM: An Interpretable Complex-valued Network for Matching (N19-1)
Copied to clipboard
| Challenge: | Existing work on quantum physics models language understanding using quantum probability . |
| Approach: | They propose a quantum-theoretic framework that unifies different linguistic units in a single complex-valued vector space and a complex-valuable network for semantic matching. |
| Outcome: | The proposed framework achieves comparable performances to strong CNN and RNN baselines on two benchmarking question answering (QA) datasets. |
A Multi-task Learning Framework for Opinion Triplet Extraction (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to Aspect-based sentiment analysis (ABSA) use aspect terms and their corresponding sentiment polarities as a reference, but they lack opinion terms as . |
| Approach: | They propose a multi-task learning framework to extract aspect terms and opinion terms and parse their sentiment dependencies with a biaffine scorer. |
| Outcome: | The proposed framework outperforms baseline and state-of-the-art approaches on four SemEval benchmarks. |
Aspect-based Sentiment Classification with Aspect-specific Graph Convolutional Networks (D19-1)
Copied to clipboard
| Challenge: | Existing aspects-based sentiment classification models lack a mechanism to account for relevant syntactical constraints and word dependencies. |
| Approach: | They propose to build a Graph Convolutional Network over the dependency tree of a sentence to exploit syntactical information and word dependencies. |
| Outcome: | The proposed model is comparable to state-of-the-art models on three benchmarking collections. |