Papers by Jian Su

25 papers
APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to generate reasoning programs that ignore the differences between facts treated all facts equally, leading to wrong punishment of programs that differed from the ground truth.
Approach: They propose an optimized training framework for long-form numerical reasoning that incorporates a number-aware negative sampling strategy and consistency-based reinforcement learning to increase execution accuracy.
Outcome: The proposed method improves the performance of long-form numerical reasoning on the FinQA and ConvFinQA leaderboards.
USB: A COMPREHENSIVE AND UNIFIED SAFETY EVALUATION BENCHMARK FOR MULTIMODAL LARGE LANGUAGE MODELS (2026.acl-long)

Copied to clipboard

Challenge: Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations.
Approach: They propose a framework that covers 61 risk categories across four modality interactions to address this gap.
Outcome: The proposed framework covers 61 risk categories across four distinct modality interactions.
Attentive Gated Lexicon Reader with Contrastive Contextual Co-Attention for Sentiment Classification (D18-1)

Copied to clipboard

Challenge: Existing sentiment lexicons do not handle word sense and the concept of semantic compositionality is non-existent in simple lexiconic approaches.
Approach: They propose a lexicon-driven contextual attention mechanism and a contrastive co-attention mechanism that models contrasting polarities between all positive and negative words in a sentence.
Outcome: The proposed model outperforms many other neural baselines on sentiment classification tasks on multiple benchmark datasets.
Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition (2025.coling-main)

Copied to clipboard

Challenge: Existing work mainly learns to map text into questions, lacking a mechanism to control results with desired complexity.
Approach: They propose a novel controllable framework to generate QGs with desired complexity using contextual and commonsense clues from text.
Outcome: The proposed framework can generate complex questions with desired complexity levels.
Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for sarcasm detection lack commonsense inferential ability when faced with complex situations.
Approach: They propose a commonsense reasoning framework for sarcasm detection based on commonsensense augmentation to supplement commonsence knowledge and infer the incongruity.
Outcome: The proposed framework is able to detect sarcasm in five datasets and is robust to complex scenarios.
MT3: A Synergistic Multi-Task RL Framework for Specializing MLLMs in Text Image Machine Translation (2026.acl-long)

Copied to clipboard

Challenge: Text Image Machine Translation (TIMT) is a critical subfield of machine translation . it requires accurate optical character recognition, robust visual-text reasoning, and high-quality translation a challenge .
Approach: They propose a multi-task optimization framework to specialize MLLMs into expert TIMT models.
Outcome: The proposed model outperforms baselines on the latest in-domain MIT-10M benchmark.
Soft Syntactic Reinforcement for Neural Event Extraction (2025.naacl-long)

Copied to clipboard

Challenge: Recent event extraction methods rely on pre-trained language models but still suffer from errors due to a lack of syntactic knowledge.
Approach: They propose a method to incorporate syntactic information into PLM-based models for event extraction (EE) this method uses a standard dependency corpus to select syntax-related dimensions of the model's representation.
Outcome: The proposed method outperforms baseline models and existing syntactic reinforcement methods on sentence-level and document-level EE benchmark datasets.
Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT - A Text-to-SQL Parsing Comparison (2023.findings-emnlp)

Copied to clipboard

Challenge: a number of open-source large language models claim to be performing better than commercial ones . however, these models fall short of the performance achieved by closed-source models like GPT-3.5 .
Approach: They evaluate six popular large language models against each other to evaluate their performance . authors say open-source models are not as effective as those built by commercial models .
Outcome: a new set of models claim to match or surpass the language understanding abilities of commercial models . the results show that the models performed far below the performance of closed-source models compared to open-source ones .
Domain Adaptation for Subjective Induction Questions Answering on Products by Adversarial Disentangled Learning (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to answer subjective questions about products are often imbalanced across product domains.
Approach: They propose a domain-adaptive model that integrates multiple viewpoints into a good answer by integrating these heterogeneous and inconsistent viewpoints.
Outcome: The proposed model integrates multiple viewpoints into a single answer span and is able to integrate them into the answer.
Low-Resource Generation of Multi-hop Reasoning Questions (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to generate valid and fluent questions from text are limited and insufficient for training.
Approach: They propose to generate multi-hop reasoning questions from the raw text in a low resource circumstance by deducing over multiple relations on several sentences in the text.
Outcome: The proposed model can be applied to the task of machine reading comprehension and achieve significant performance improvements.
Reasoning with Sarcasm by Reading In-Between (P18-1)

Copied to clipboard

Challenge: Sarcasm is a figurative speech act which manifests on social networks such as Twitter and Reddit.
Approach: They propose a model that looks in-between rather than across to explicitly model contrast and incongruity.
Outcome: The proposed model achieves state-of-the-art performance on all datasets and improves interpretability.
Humans Need Context, What about Machines? Investigating Conversational Context in Abusive Language Detection (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we examine the role of conversational context in abusive language detection . prior studies have ignored the contextual nature of abusive language, ignoring this aspect . toxicity, hate speech, harmful stereotypes are among the forms of harmful language .
Approach: They propose to use conversational context to analyze abusive language detection using two methods . they use "abusive language" as an umbrella term to refer to various forms of harmful language .
Outcome: The proposed approach is based on two datasets in English and a new dataset of French tweets annotated for hate speech and stereotypes.
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large-scale reinforcement learning (RL) methods have proven effective in enhancing the reasoning abilities of large language models.
Approach: They propose an open-source adaptation of the R1-Zero RL framework for machine translation (MT) their code is available at https://github.com/fzp0424/MT-R1-zero.
Outcome: The proposed framework surpasses towerinstruct-7B-v0.2 on the english-chinese benchmark by 1.26 points.
Generating Deep Questions with Commonsense Reasoning Ability from the Text by Disentangled Adversarial Inference (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for commonsense question generation produce shallow questions that can be answered by simple word matching.
Approach: They propose a task of commonsense question generation that aims to yield deep-level questions from the text.
Outcome: The proposed model can yield deep-level and to-the-point questions from the text.
Can LLM Safety Be Ensured by Constraining Parameter Regions? (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are often assumed to contain parameter subsets whose modification directly influences safety behaviors.
Approach: They evaluate four methods to identify parameter subsets with "safety regions" they find low overlap, but overlap drops when refinement is done using utility datasets .
Outcome: The proposed methods show low overlap and drop significantly when refined using utility datasets.
Exploring All-In-One Knowledge Distillation Framework for Neural Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge distillation methods only obtain one lightweight student each time . this could be resource-intensive and resulting in multiple students not being optimally utilized .
Approach: They propose a knowledge distillation framework which generates multiple satisfactory students at once.
Outcome: The proposed framework generates multiple satisfactory students at once.
An Exploratory Study on Model Compression for Text-to-SQL (2023.findings-acl)

Copied to clipboard

Challenge: Text-to-SQL translates user queries into SQL statements that can retrieve relevant answers from relational databases.
Approach: They propose to apply model compression techniques to sketch-based and sequence-to-sequence Text-toSQL models.
Outcome: The proposed models have higher inference efficiency and respond better to model compression than sequence-to-sequence models.
Exploring Better Text Image Translation with Multimodal Codebook (2023.acl-long)

Copied to clipboard

Challenge: Current studies on text image translation face bottlenecks due to lack of a publicly available dataset and poor optical character recognition.
Approach: They propose a text image translation model with a multimodal codebook and an OCR dataset for Chinese-English translation.
Outcome: The proposed model can associate the image with relevant texts, providing useful supplementary information for translation.
EmoTrans: Emotional Transition-based Model for Emotion Recognition in Conversation (2024.lrec-main)

Copied to clipboard

Challenge: Emotions are causally transmitted among communication participants, facilitating comprehension of intricate changes in emotional states during the conversation.
Approach: They propose an Emotional Transition-based Emotion Recognizer that captures ET features in an emotional conversation by concatenating the most recent utterances with their corresponding speakers.
Outcome: The proposed model is sensitive to emotions and captures ET features in the sample.
M2RC-EVAL: Massively Multilingual Repository-level Code Completion Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing repository-level code completion benchmarks focus on a limited number of languages . existing benchmarks report overall average scores of different languages ignoring fine-grained abilities .
Approach: They propose to use repository-level code completion benchmarks to evaluate general code intelligence abilities across languages for existing code Large Language Models.
Outcome: The proposed benchmarks improve the code completion abilities of existing LLMs by using two types of annotations on the parsed syntax tree.
Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing financial benchmarks suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation.
Approach: They propose a bilingual benchmark for financial LLMs that assesses models’ language understanding and generation capabilities.
Outcome: The proposed bilingual benchmark assesses models’ language understanding and generation capabilities.
Mitigating Linguistic Artifacts in Emotion Recognition for Conversations from TV Scripts to Daily Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Emotion Recognition in Conversations (ERC) focus on training and testing models on the same datasets and there is no prior work on adaptability.
Approach: They propose to use contrastive learning to prioritize emotional features over a linguistic style and refining emotion predictions with pseudo-emotion intensity score to improve model's robustness and accuracy in diverse conversational contexts.
Outcome: The proposed techniques reduce reliance on linguistic artifacts found in TV transcripts and improve model’s robustness and accuracy in diverse conversational contexts.
CDB: A Unified Framework for Hope Speech Detection Through Counterfactual, Desire and Belief (2025.findings-naacl)

Copied to clipboard

Challenge: Using algorithms to model user-generated desires on social media, we propose a new approach to understanding and detection of hope speech.
Approach: They propose a language-driven decomposition of the notional category hope and its automatic detection in a unified setting.
Outcome: The proposed model captures future-oriented hopes through desires and beliefs and the counterfactuality of past unfulfilled wishes and regrets.
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown their potential to deliver human-like judgments.
Approach: They propose a systematic LLM-based multi-agent framework for advanced LLM as-a-judge MT evaluation that integrates dimension-specific results into a final evaluation judgment.
Outcome: The proposed framework outperforms existing LLM-as-a-judge methods and competes with state-of-the-art automatic metrics even when powered by a suboptimal model like GPT-4o mini.
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Guide-Align is a guideline-oriented approach to augment the safety and quality of Large Language Models.
Approach: They propose a guideline-oriented method to augment the safety and quality of large language models.
Outcome: The proposed method outperforms existing methods on three benchmarks and shows significant improvements in security and quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations