Papers by Pan Deng

17 papers
SOUL: Towards Sentiment and Opinion Understanding of Language (2023.emnlp-main)

Copied to clipboard

Challenge: Sentiment analysis models often fail to capture the broader complexities of sentiment analysis.
Approach: They propose a task to evaluate sentiment understanding through two subtasks . they annotate a new dataset comprising 15,028 statements from 3,638 reviews .
Outcome: The proposed task evaluates sentiment understanding through two subtasks . it is a challenging task for both small and large language models, with performance gaps of up to 27% .
Automatic Prompt Engineering for Scalable Prompt Inversion in Text-to-Image Ad Generation (2026.acl-industry)

Copied to clipboard

Challenge: PRISM-DUEL is a black-box framework that formalizes prompt optimization as Automatic Prompt Engineering (APE) PRIMS-DUEl is motivated by advertising workflows requiring low-latency, diverse variants faithful to a human-designed ad.
Approach: They propose a black-box framework that formalizes prompt optimization as Automatic Prompt Engineering (APE) they obtain label-free pairwise preferences and rationales from an LLM judge over pairs of generated images and use a dueling-bandit optimizer to optimize a prompt for generating controlled variations while matching the reference ad's visual content.
Outcome: The proposed framework preserves visual similarity and semantic faithfulness while increasing diversity.
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on English-centric translation tasks have focused on multimodal large language models, but the exploration of many-to-many translation is limited by the scarcity of parallel data.
Approach: They propose a three-stage curriculum learning strategy that leverages the machine translation capabilities of large language models and adapts them to S2TT tasks.
Outcome: The proposed strategy achieves state-of-the-art average performance in 1514 language pairs, requiring fewer than 10 hours of speech data per language to achieve competitive results.
CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: OpenAI-o1 enables ‘slow thinking’ because it is closer to the human thought process .
Approach: They propose a new framework that integrates the Monte Carlo Tree Search algorithm and a dynamic mechanism for integrating new key information, termed ‘associative memory’.
Outcome: The proposed framework improves performance on open-source multi-hop reasoning datasets and more than 15% gain on proprietary CRB dataset.
TopWORDS-Poetry: Simultaneous Text Segmentation and Word Discovery for Classical Chinese Poetry via Bayesian Inference (2023.emnlp-main)

Copied to clipboard

Challenge: Experimental studies confirm that TopWORDS-Poetry can successfully segment poetry words without pre-given vocabulary or training corpus.
Approach: They propose an unsupervised method that can achieve reliable text segmentation and word discovery for classical Chinese poetry simultaneously without pre-given vocabulary or training corpus.
Outcome: Experimental results show that TopWORDS-Poetry can segment poetry lines into meaningful words with high quality without pre-given vocabulary or training corpus.
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have demonstrated that inference-time scaling increases performance of Large Language Models (LLMs) in various reasoning tasks such as mathematics and complex question answering by increasing the length of Chain-of-Thought (CoT).
Approach: They propose a model which synthesizes longer CoT data and iteratively improves performance through self-training by incorporating a few demonstration examples.
Outcome: The proposed model achieves an average improvement of more than +2.5 points across five reasoning tasks: MMLU, GSM8K, ARC-C, HellaSwag, and BBH on two backbone models.
Context Attribution with Multi-Armed Bandit Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to augmenting attribution with retrieval-augmented generation (RAG) focus on training models to explicitly cite context segments during generation, but their reliability remains unverifiable.
Approach: They propose a framework that formulates context attribution as a combinatorial multi-armed bandit problem by using Linear Thompson Sampling to efficiently identify the most influential context segments while minimizing the number of model queries.
Outcome: The proposed method reduces model queries by 30% while matching or exceeding the attribution quality of existing approaches.
Discourse Marker Augmented Network with Reinforcement Learning for Natural Language Inference (P18-1)

Copied to clipboard

Challenge: Existing approaches to natural language inference focus on interaction architectures of sentences . but, we propose to transfer knowledge from discourse markers to augment the model .
Approach: They propose to transfer knowledge from discourse markers to augment the quality of the NLI model.
Outcome: The proposed method achieves state-of-the-art performance on large-scale datasets.
Bidirectional Generative Framework for Cross-domain Aspect-based Sentiment Analysis (2023.acl-long)

Copied to clipboard

Challenge: Aspect-based sentiment analysis (ABSA) is a task of analyzing people's sentiments at the aspect level.
Approach: They propose a unified bidirectional generative framework to tackle cross-domain ABSA tasks . the framework trains a model in both text-to-label and label-totext directions .
Outcome: The proposed framework trains a model in both label-to-label and label- to-text directions to learn domain-agnostic features.
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations rely on point-wise confidence, which can mask brittle belief.
Approach: They propose a measure of belief robustness that evaluates coherence across a conceptual neighborhood.
Outcome: The proposed model is more resistant to interference than existing models.
An ensemble CNN method for biomedical entity normalization (D19-57)

Copied to clipboard

Challenge: Named entity recognition (NER) and entity normalization (entity linking) are two fundamental natural language processing tasks to achieve entity normalizing.
Approach: They propose a CNN method that normalizes microbiology-related entities to concepts in standard dictionaries.
Outcome: The proposed method performs well in the BioNLP-OST19 shared task Bacteria Biotope.
CoMave: Contrastive Pre-training with Multi-scale Masking for Attribute Value Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing .
Approach: They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce.
Outcome: The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks.
TopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference (2022.acl-long)

Copied to clipboard

Challenge: No existing methods can achieve effective text segmentation and word discovery in open domain Chinese texts.
Approach: They propose a Bayesian-based method that can achieve effective text segmentation and word discovery in open domain.
Outcome: The proposed method enjoys robust performance and transparent interpretation when no training corpus and domain vocabulary are available.
Sentiment Analysis in the Era of Large Language Models: A Reality Check (2024.findings-naacl)

Copied to clipboard

Challenge: Sentiment analysis (SA) has been a long-standing research area in natural language processing.
Approach: They propose a benchmark to evaluate LLMs' SA abilities and propose 'sentiEval' benchmark to be used for a more comprehensive evaluation.
Outcome: The proposed benchmark outperforms small language models on 26 datasets on 13 tasks and compared them with LLMs trained on domain-specific datasets.
ODL-TempLLM: Ontology-Guided and Description Logic-Reasoned Temporal Reasoning with LLMs (2026.acl-long)

Copied to clipboard

Challenge: Temporal reasoning is crucial for large language models to understand event concurrency and complex temporal interactions in natural language.
Approach: They propose an ontology-guided and description logic–constrained temporal reasoning paradigm that shifts focus from internal inference to the explicit modeling of temporal structure.
Outcome: The proposed method outperforms state-of-the-art methods by 2.07–31.83 F1 points and 1.00–30.73 EM points, exhibiting strong generalization and highlighting the potential of explicit temporal reasoning.
Reinforced Dynamic Reasoning for Conversational Question Generation (P19-1)

Copied to clipboard

Challenge: Empirical results on the recently released CoQA dataset demonstrate the effectiveness of our method . large-scale highquality conversational question answering datasets such as CoQA and QuAC can help train models to answer sequential questions.
Approach: They propose a task called Conversational Question Generation which generates a question based on a passage and a conversation history to generate the next question.
Outcome: The proposed method is based on a question-answering style conversation dataset . it can be used to generate meaningful questions on QA and SQuAD datasets .
ECC: An Emotion-Cause Conversation Dataset for Empathy Response (2025.emnlp-main)

Copied to clipboard

Challenge: Existing empathy dialogue datasets focus on emotion labels while cause annotations are added post hoc.
Approach: They propose an emotion-cause conversation dataset with 2.4K dialogues that can be scalable . they use a framework that utilizes knowledge and large language models to automatically generate dialogues .
Outcome: The proposed dataset can achieve comparable or even superior performance to existing empathy dialogue datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations