Papers by Qinliang Su
A Deep Neural Information Fusion Architecture for Textual Network Embeddings (D19-1)
Copied to clipboard
| Challenge: | Textual network embeddings aim to learn a low-dimensional representation for every node in the network while seeking to retain the original network information. |
| Approach: | They propose a deep neural architecture to fuse the two kinds of informations into one representation. |
| Outcome: | The proposed model outperforms the comparing methods on all three datasets. |
Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work mainly learns to map text into questions, lacking a mechanism to control results with desired complexity. |
| Approach: | They propose a novel controllable framework to generate QGs with desired complexity using contextual and commonsense clues from text. |
| Outcome: | The proposed framework can generate complex questions with desired complexity levels. |
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for sarcasm detection lack commonsense inferential ability when faced with complex situations. |
| Approach: | They propose a commonsense reasoning framework for sarcasm detection based on commonsensense augmentation to supplement commonsence knowledge and infer the incongruity. |
| Outcome: | The proposed framework is able to detect sarcasm in five datasets and is robust to complex scenarios. |
Domain Adaptation for Subjective Induction Questions Answering on Products by Adversarial Disentangled Learning (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods to answer subjective questions about products are often imbalanced across product domains. |
| Approach: | They propose a domain-adaptive model that integrates multiple viewpoints into a good answer by integrating these heterogeneous and inconsistent viewpoints. |
| Outcome: | The proposed model integrates multiple viewpoints into a single answer span and is able to integrate them into the answer. |
Low-Resource Generation of Multi-hop Reasoning Questions (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods to generate valid and fluent questions from text are limited and insufficient for training. |
| Approach: | They propose to generate multi-hop reasoning questions from the raw text in a low resource circumstance by deducing over multiple relations on several sentences in the text. |
| Outcome: | The proposed model can be applied to the task of machine reading comprehension and achieve significant performance improvements. |
Integrating Semantics and Neighborhood Information with Graph-Driven Generative Models for Document Retrieval (2021.acl-long)
Copied to clipboard
Zijing Ou, Qinliang Su, Jianxing Yu, Bang Liu, Jingwen Wang, Ruihui Zhao, Changyou Chen, Yefeng Zheng
| Challenge: | Existing methods for document hashing combine only one of semantics and neighborhood information, lacking a theoretical principle to guide the integration process. |
| Approach: | They propose to encode neighborhood information with a graph-induced Gaussian distribution and integrate it with generative models. |
| Outcome: | The proposed model can be trained as efficiently as state-of-the-art methods on benchmark datasets. |
HierPrompt: Zero-Shot Hierarchical Text Classification with LLM-Enhanced Prototypes (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for Hierarchical Text Classification are based on prototypes, but do not perform well due to ambiguity and impreciseness of category names. |
| Approach: | They propose a method that leverages hierarchy-aware prompts to instruct LLM to produce more representative and informative prototypes. |
| Outcome: | The proposed method outperforms existing methods on three benchmark datasets. |
Constituency Lattice Encoding for Aspect Term Extraction (2020.coling-main)
Copied to clipboard
| Challenge: | a challenge for aspect term extraction is to extract phrase-level aspect terms . a constituency lattice structure is constructed using the span annotations of constituents of a sentence . |
| Approach: | They propose to incorporate the span annotations of constituents of a sentence to leverage syntactic information in neural network models. |
| Outcome: | The proposed model outperforms existing models on two benchmark datasets. |
Learning to Answer Psychological Questionnaire for Personality Detection (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing text-based personality detection research relies on data-driven approaches to implicitly capture personality cues in online posts lacking the guidance of psychological knowledge. |
| Approach: | They propose a model to capture key information in texts and a questionnaire to help the user to make a personality assessment. |
| Outcome: | The proposed model captures key information in texts and a questionnaire and can be used to improve personality prediction. |
RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial Attacks (2023.acl-long)
Copied to clipboard
| Challenge: | Existing defenses focus on improving robustness of the victim model in training, but neglect to mitigate adversarial attacks during inference. |
| Approach: | They propose a framework that confuses attackers and corrects adversarial contexts . their framework helps improve the robustness of the victim model during inference . |
| Outcome: | The proposed framework improves the robustness of the victim model in training . it also corrects abnormal contexts in the representation level and filtering out examples . |
Document Hashing with Multi-Grained Prototype-Induced Hierarchical Generative Model (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing document hashing methods only consider flat semantics of documents, preserving hierarchical semantics. |
| Approach: | They propose a hierarchical generative model that can model and leverage hierarchic semantics . they introduce hierarchically-based prototypes into the model to construct a Hierarchical prior distribution . |
| Outcome: | The proposed model outperforms baseline methods on hierarchical and flat datasets. |
NASH: Toward End-to-End Neural Architecture for Generative Semantic Hashing (P18-1)
Copied to clipboard
Dinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang, Ricardo Henao, Lawrence Carin
| Challenge: | Existing approaches to fast similarity search require two-stage training and the binary constraints are handled ad-hoc. |
| Approach: | They propose an end-to-end neural architecture for semantic hashing where binary hash codes are treated as Bernoulli latent variables. |
| Outcome: | The proposed approach outperforms state-of-the-art models on unsupervised and supervised scenarios on three public datasets. |
Generating Deep Questions with Commonsense Reasoning Ability from the Text by Disentangled Adversarial Inference (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for commonsense question generation produce shallow questions that can be answered by simple word matching. |
| Approach: | They propose a task of commonsense question generation that aims to yield deep-level questions from the text. |
| Outcome: | The proposed model can yield deep-level and to-the-point questions from the text. |
Refining BERT Embeddings for Document Hashing via Mutual Information Maximization (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing unsupervised document hashing methods are mostly established on generative models . due to the difficulties of capturing long dependency structures, these methods rarely model the raw documents directly . |
| Approach: | They propose to learn hash codes from BERT embeddings by modifying existing models . they use mutual information maximization principle to maximize mutual information . |
| Outcome: | The proposed method outperforms existing methods learned from BERT embeddings on three benchmark datasets. |
Embedding Dynamic Attributed Networks by Modeling the Evolution Processes (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to embed nodes into low-dimensional vectors focus on static networks, but in practice, many networks are evolving over time and hence are dynamic, e.g., social networks. |
| Approach: | They propose to extract high-order neighborhood information at each given timestamp and then use an embedding prediction framework to capture the temporal correlations. |
| Outcome: | Extensive experiments on four real-world datasets show that the proposed method outperforms baseline methods for dynamic link prediction and node classification tasks. |
Targeting the Needle, Ignoring the Haystack: Anchoring Crucial Cues for Evolving Scam Call Detection via an LLM-Assisted Classifier (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for fraud detection on online service platforms often fail to generalize due to the scarcity of labeled data and the continuous evolution of conversational contexts. |
| Approach: | They propose a framework that anchors detection on Semantic Primitives . they prioritize stable evidence over conversational noise to ensure a verifiable fraud tactic . |
| Outcome: | The proposed framework achieves superior robustness and efficiency compared to baselines . it prioritizes stable evidence over diverse conversational noise . |
Syntax-Enhanced Pre-trained Model (2021.acl-long)
Copied to clipboard
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, Nan Duan
| Challenge: | Existing methods that use syntax of text in pre-training and fine-tuning suffer from discrepancy between the two stages. |
| Approach: | They propose a model that utilizes the syntactic structure of text in pre-training and fine-tuning stages. |
| Outcome: | The proposed model achieves state-of-the-art on six public benchmark datasets. |
Detecting Continuously Evolving Scam Calls under Limited Annotation: A LLM-Augmented Expert Rule Framework (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect scam calls rely on labeled data and assume static distribution of scam narratives. |
| Approach: | They propose a method leveraging large language models to detect continuously evolving scam calls . scammers continuously evolve their tactics, making these methods less effective . |
| Outcome: | The proposed approach is based on large language models to detect continuously evolving scam calls. |
Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text Clustering (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for text clustering use static pseudo-oracles, i.e., unidirectionally querying them for similarity assessment or data augmentation. |
| Approach: | They propose a training framework that enables bidirectional refinement between LLMs and embedding models by using task-aware prompts to guide the LLM in generating interpretations for the input texts. |
| Outcome: | Experiments on 14 benchmark datasets across 5 tasks demonstrate the effectiveness of the proposed training framework. |
Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms (P18-1)
Copied to clipboard
Dinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, Lawrence Carin
| Challenge: | Existing deep learning architectures to model compositionality in text sequences require a large number of parameters and expensive computations. |
| Approach: | They propose two additional pooling strategies over word embeddings for improved interpretability and hierarchical pooling for spatial (n-gram) information within text sequences. |
| Outcome: | The proposed pooling strategies improve interpretability and preserve spatial (n-gram) information within text sequences. |
Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing semantic hashing methods only learn a binary code for each document and use Hamming distance to evaluate document distances. |
| Approach: | They propose to leverage BERT embeddings to perform efficient retrieval based on product quantization technique . they transform original BERT embedded codewords and feed it into a probabilistic product quantizer module . |
| Outcome: | The proposed method outperforms current state-of-the-art methods on three benchmarks. |
Generative Semantic Hashing Enhanced via Boltzmann Machines (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for generative semantic hashing assume a factorized posterior distribution, enforcing independence among the bits of hash codes. |
| Approach: | They propose to use a Boltzmann machine distribution as the variational posterior to introduce correlations among the bits of hash codes. |
| Outcome: | The proposed method can achieve significant performance gains by combining two hash codes. |
Document Hashing with Mixture-Prior Generative Models (D19-1)
Copied to clipboard
| Challenge: | Existing generative hashing methods only consider the use of simple priors, which limits them to further improve their performance. |
| Approach: | They propose to use Gaussian and Bernoulli priors to generate hashing codes . they propose to cast a Gausssian latent representation into binary code . |
| Outcome: | The proposed models outperform existing methods on a benchmark dataset using Gaussian and Bernoulli priors. |
Leveraging BERT and TFIDF Features for Short Text Clustering via Alignment-Promoting Co-Training (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing clustering methods rely on keyword information, but they lack this information. |
| Approach: | They propose a CO**-**T**raining **C**lustering framework to make use of BERT and TFIDF features. |
| Outcome: | The proposed framework outperforms existing SOTA methods on eight datasets. |