Papers by Yingjie Li

29 papers
Multi-Task Stance Detection with Sentiment and Stance Lexicons (D19-1)

Copied to clipboard

Challenge: Recent studies show improvements in stance detection by using attention mechanism or sentiment information.
Approach: They propose a multi-task framework that incorporates attention mechanism and takes sentiment classification as an auxiliary task.
Outcome: The proposed model outperforms state-of-the-art deep learning methods on the SemEval-2016 dataset.
VerilogLAVD: LLM-Aided Pattern Generation for Verilog CWE Detection (2026.acl-long)

Copied to clipboard

Challenge: Existing static analysis tools focus on functional correctness and depend heavily on manual rules.
Approach: They propose a framework that generates executable Traversal Detection Patterns (TDPs) to help detect hardware vulnerabilities.
Outcome: The proposed framework improves the F1 score by 133% compared to LLM-based methods.
Stance Detection in COVID-19 Tweets (2021.acl-long)

Copied to clipboard

Challenge: a global pandemic of COVID-19 has forced major changes in our daily lives . a new stance detection dataset is being used to track the stances of Twitter users .
Approach: They use Twitter stance data to collect stances on topics related to the pandemic . they train models to take advantage of large amounts of unlabeled data .
Outcome: The proposed model improves on existing stance detection datasets and unlabeled data.
CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long Contexts (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable success in natural language processing (NLP), particularly in single-turn question answering (QA) on short-text.
Approach: They propose a framework that captures logical correlations across chunks of ELC and maintains coherence of multi-turn Questions.
Outcome: The proposed framework is able to capture logical correlations across chunks of ELC and maintain coherence of multi-turn Questions.
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack (2026.findings-acl)

Copied to clipboard

Challenge: Existing watermarking algorithms focus on defending against paraphrase and piggyback spoofing attacks, which can inject harmful content, compromise reliability, and undermine trust in attribution.
Approach: They propose an algorithm capable of defending against paraphrase and spoofing attacks.
Outcome: Experiments on large language models and language models show that DualGuard is the first watermarking algorithm capable of defending against both paraphrase and spoofing attacks.
Improving Stance Detection with Multi-Dataset Learning and Knowledge Distillation (2021.emnlp-main)

Copied to clipboard

Challenge: stance detection is a method to determine whether a text author is in favor of, against or neutral toward a specific target.
Approach: They propose a method that applies instance-specific temperature scaling to the teacher and student predictions.
Outcome: The proposed method outperforms the state-of-the-art on all datasets and on multiple datasets.
Distilling Calibrated Knowledge for Stance Detection (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for stance detection ignore meaningful signals among categories offered by hard labels.
Approach: They propose to use knowledge distillation to calibrate teacher predictions in each generation step.
Outcome: The proposed method can calibrate teacher predictions in each generation step and improves stance detection accuracy.
SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box Jailbreaking (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for jailbreak ignore the semantic differences between categories of harmful questions, leading to inconsistent success rates and reduced overall attack effectiveness.
Approach: They propose a category-aware jailbreak framework that incorporates the semantic category of harmful questions into prompt generation.
Outcome: The proposed framework improves attack success rates and category alignment and achieves better cross-category robustness compared to the state-of-the-art (SOTA) baselines.
ZeroStance: Leveraging ChatGPT for Open-Domain Stance Detection via Dataset Generation (2024.findings-acl)

Copied to clipboard

Challenge: Until recently, zero-shot stance detection was limited to in-domain tasks.
Approach: They propose a method for stance detection that trains a model that can generalize well to unseen targets across multiple domains.
Outcome: The proposed method generalizes well to unseen targets across multiple domains over baselines on most benchmarks.
Target-Aware Data Augmentation for Stance Detection (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for stance detection are not diversified or inconsistent with the given target and label information.
Approach: They propose to augment a text with a conditional masked word prediction task . they propose to replace a target mention with 'target-aware' sentences by replacing a reference word with .
Outcome: The proposed method outperforms existing methods on 11 targets.
Enhancing Argument Structure Extraction with Efficient Leverage of Contextual Information (2023.findings-emnlp)

Copied to clipboard

Challenge: Argument structure extraction (ASE) aims to identify the discourse structure of arguments within documents.
Approach: They propose an Efficient Context-aware ASE model that fully exploits contextual information by augmenting modeling capacity and augmenting training data.
Outcome: The proposed model can extract argumentative discourse structure from documents and reduce reliance on specific words or less informative sentences.
A Multi-Task Learning Framework for Multi-Target Stance Detection (2021.findings-acl)

Copied to clipboard

Challenge: Existing models fail to learn target-specific representations and are prone to overfitting.
Approach: They propose a multi-task learning network to train one model on all target pairs . their results show that their proposed model outperforms the best-performing baseline by 12.39% .
Outcome: The proposed model outperforms the best-performing baseline model by 12.39% in macro-averaged F1-score.
XAL: EXplainable Active Learning Makes Classifiers Better Low-resource Learners (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for active learning rely on model uncertainty or disagreement to pick unlabeled data, leading to over-confidence in superficial patterns and lack of exploration.
Approach: They propose to use a bi-directional encoder and a uni-directional decoder to generate and score an explanation for low-resource text classification.
Outcome: The proposed model improves on 9 strong baselines on six datasets and can generate explanations for its predictions.
CascadeFix: Multi-Location Program Repair via Cascading Planning and Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for automating program repair face insufficient bug dependency modeling and inadequate global repair planning when addressing semantically complex multi-location bugs.
Approach: They propose a multi-location automatic repair method via cascading planning and generation . they propose to model dependencies among bugs and cluster them to ensure rationality .
Outcome: The proposed method resolves 84 multi-location bugs, achieving a 31% improvement over current methods.
C-STANCE: A Large Dataset for Chinese Zero-Shot Stance Detection (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in zero-shot stance detection are limited to English and Chinese . stance can provide useful information for important events such as policymaking and presidential elections.
Approach: They present a Chinese dataset for zero-shot stance detection that is the first for ZSSD.
Outcome: The proposed dataset is the first Chinese dataset for zero-shot stance detection.
CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Multimodal large language models have demonstrated promising results in a variety of tasks that combine vision and language.
Approach: They propose a benchmark to assess the ability of models to use contextual information in free-form text to enhance visual comprehension.
Outcome: The proposed model fails to extract and utilize contextual information to improve understanding of images.
Declarative Techniques for NL Queries over Heterogeneous Data (2025.emnlp-industry)

Copied to clipboard

Challenge: In many industrial settings, users wish to ask questions in natural language . however, these applications do not cope with data source heterogeneity that typifies such environments.
Approach: They propose a declarative approach to handling data heterogeneity in industrial settings . they simulate the heterogenity of industrial environments by adding two extensions of the popular Spider benchmark dataset .
Outcome: The proposed approach copes with data source heterogeneity better than state-of-the-art systems.
DEIE: Benchmarking Document-level Event Information Extraction with a Large-scale Chinese News Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing event-based datasets mainly target sentence-level tasks . current models struggle with "document" annotation, a key feature of the current model .
Approach: They propose a large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments.
Outcome: The proposed dataset has over 56,000+ events and 242,000+ arguments.
Analyzing and Internalizing Complex Policy Documents for LLM Agents (2026.acl-long)

Copied to clipboard

Challenge: Large language model agents rely on in-context policy documents to act as effective user assistants.
Approach: They propose an agentic benchmark generator with Controllable Complexity in agent policy across four levels to evaluate agents under increasing complexity.
Outcome: The proposed method outperforms the baseline in data-sparse and high-complexity settings.
A Rationale-centric Counterfactual Data Augmentation Method for Cross-Document Event Coreference Resolution (2024.naacl-long)

Copied to clipboard

Challenge: Existing state-of-the-art event coreference resolution systems rely on spurious and spurious associations in the input mention pair text.
Approach: They propose a rationale-centric counterfactual data augmentation method that leverages the debiasing capability of counterfact data haussed by LLM-in-the-loop to mitigate spurious association while emphasizing causation.
Outcome: The proposed method achieves state-of-the-art on three popular cross-document benchmarks and demonstrates robustness in out-of domain scenarios.
Intra-Event and Inter-Event Dependency-Aware Graph Network for Event Argument Extraction (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models do not build dependency information among event argument roles . Existing methods do not learn the interactions between different roles based on event structure .
Approach: They propose an intra-event and inter-e event dependency-aware graph network to model dependencies between roles . they use event structure as the fundamental unit to construct role dependencies within events .
Outcome: The proposed model improves on the ACE05, RAMS, and WikiEvents datasets.
P-Stance: A Large Dataset for Stance Detection in Political Domain (2021.findings-acl)

Copied to clipboard

Challenge: stance detection is a method to determine whether a text author is in favor of, against or neutral toward a specific target.
Approach: They propose to use a large stance detection dataset in the political domain to detect stances on twitter.
Outcome: The proposed model achieves a macro-average F1-score of 80.53% and can be used to improve cross-domain stance detection.
Pro-Woman, Anti-Man? Identifying Gender Bias in Stance Detection (2024.findings-acl)

Copied to clipboard

Challenge: Gender bias has been widely observed in NLP models, which can perpetuate harmful stereotypes and discrimination.
Approach: They construct a dataset to measure gender bias in stance detection using 36k samples . they find that all models are gender-biased and prone to classify sentences that contain male nouns as Against and those with female noun as Favor .
Outcome: The proposed dataset shows that all models are gender-biased and prone to classify sentences that contain male nouns as Against and those with female noun as Favor.
SpARK: An Embarrassingly Simple Sparse Watermarking in LLMs with Enhanced Text Quality (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for detecting and monitoring generated text face a trade-off between the quality of the generated text and the effectiveness of the watermarking process.
Approach: They propose a new type of LLM watermark, Sparse WatermARK, which uses watermarks to a small subset of generated tokens distributed across the text.
Outcome: The proposed method outperforms existing methods in detectability and quality while maintaining generated text quality.
A New Direction in Stance Detection: Target-Stance Extraction in the Wild (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for stance detection assume that the target is known in advance . Existing tasks use implicit mentions in the source text and are infeasible to have manual annotations at a large scale.
Approach: They propose a task Target-Stance Extraction that aims to extract the (target, stance) pair from social media texts.
Outcome: The proposed task can facilitate future research in the field of stance detection.
Task Calibration: Calibrating Large Language Models on Inference Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive zero-shot performance on inference tasks, however, they may suffer from spurious correlations between input texts and output labels, which limits their ability to reason based purely on general language understanding.
Approach: They propose a zero-shot and inference-only calibration method inspired by mutual information which recovers LLM performance through task reformulation.
Outcome: The proposed calibration method improves on 13 benchmarks and prompt templates and can be integrated with other calibration methods.
PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that the main challenge lies in long context and perspective extraction.
Approach: They propose a benchmark to facilitate multi-faceted perspective retrieval and summarization . they propose measurable metrics to evaluate the comprehensiveness of the retrieval pipeline .
Outcome: The proposed system breaks free from information silos by combining two opposing claims . it can be used to extract multiple perspectives and improve performance on the platform .
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, but introduce a critical security vulnerability: Knowledge Base Leakage.
Approach: They propose a runtime defense mechanism inspired by stack canaries in software security . canaryRAG embeds carefully designed canary tokens into retrieved chunks and reformulates RAG extraction defense as a dual-path runtime integrity game.
Outcome: The proposed system can detect and prevent RAG Knowledge Base Leakage in real time . it can be integrated into arbitrary RAG pipelines without retraining or structural modifications .
CmEAA: Cross-modal Enhancement and Alignment Adapter for Radiology Report Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for automatic radiology report generation suffer from data bias.
Approach: They propose a method that connects a vision encoder with a frozen large language model by using a cross-modal enhancement and alignment adapter.
Outcome: The proposed model outperforms existing state-of-the-art methods on IU X-Ray and MIMIC-CXR datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations