Papers by Xiang Gao

64 papers
Evaluating and Enhancing the Robustness of Code Pre-trained Models through Structure-Aware Adversarial Samples Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained code models have made significant strides in the field of neural code intelligence, but they are susceptible to adversarial attacks that subtly modify the input sequence and can impair generalization.
Approach: They propose a set of novel robustness evaluation methods based on the intrinsic structure of the code to explore the impact of imperceptible perturbation.
Outcome: The proposed methods have demonstrated their effectiveness across a wide range of models and tasks, and are able to predict the performance of perturbed models.
Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigm (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on individual quality and do not assess the value of training data.
Approach: They propose a choice-based sample selection framework that evaluates sample quality . they use LLMs to evaluate the value of each option during the selection process .
Outcome: The proposed model outperforms the full dataset and recent studies on a larger medical dataset.
RockNER: A Simple Method to Create Adversarial Examples for Evaluating the Robustness of Named Entity Recognition Models (2021.emnlp-main)

Copied to clipboard

Challenge: Recent named entity recognition models have great performance on many conventional benchmarks, but it is not reliable in realistic applications.
Approach: They propose a method to create natural adversarial examples using Wikidata and pre-trained language models.
Outcome: The proposed method produces natural adversarial examples with a shifted distribution from training data.
Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown increasing power on NLP tasks. however, tuning these models for downstream tasks usually requires exorbitant costs.
Approach: They propose a black-box tuning technique that optimizes task-specific prompts without accessing gradients and hidden representations.
Outcome: The proposed method improves performance under few-shot learning scenarios.
Atoxia: Red-teaming Large Language Models with Target Toxic Answers (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) are still vulnerable to generation safety vulnerabilities.
Approach: They propose a method that A**tacks LLMs with target "toxi" given a particular harmful answer, the method generates a user query and a misleading answer opening to examine the internal defects of a given LLM.
Outcome: The proposed method detects safety risks in open-source models and state-of-the-art models such as GPT-4o.
Unlocking LLMs’ Self-Improvement Capacity with Autonomous Learning for Domain Adaptation (2025.findings-acl)

Copied to clipboard

Challenge: Existing models that use self-supervised and instruction fine-tuning can be trained using unlabeled corpora.
Approach: They propose to use unlabeled target corpora to adapt large language models to new domains . they propose to employ self-supervised pre-training and instruction fine-tuning methods .
Outcome: The proposed model can adapt to new domains using only a large amount of unlabeled target corpora.
Title2Event: Benchmarking Open Event Extraction with a Large-scale Chinese Title Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Existing EE datasets define fixed event types and design specific schemas for each of them, failing to cover diverse events emerging from the online text.
Approach: They propose to use a sentence-level dataset to benchmark Open Event Extraction without restricting event types.
Outcome: The proposed dataset contains more than 42,000 news titles in 34 topics collected from Chinese web pages.
Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading (P19-1)

Copied to clipboard

Challenge: a new approach to contentful neural conversation is proposed . end-to-end models are effective in learning fluent responses, but their responses are often vacuous and uninformative.
Approach: They propose a model that provides the conversation model with relevant text on the fly as a source of external knowledge.
Outcome: The proposed model improves the informativeness and diversity of generated output compared to previous methods.
Ciron: a New Benchmark Dataset for Chinese Irony Detection (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Chinese irony detection often lacks labeled benchmark datasets . despite its pervasive nature, irony is a trope whose actual meaning differs from what is literally enunciated.
Approach: They propose to use a Chinese benchmark dataset for automatic Chinese irony detection to provide a benchmark for machine learning models.
Outcome: The proposed dataset includes more than 8.7K posts, collected from Weibo, a micro blogging platform.
MixingBoard: a Knowledgeable Stylized Integrated Text Generation Platform (2020.acl-demos)

Copied to clipboard

Challenge: Neural text generation algorithms have seen great improvements over the past several years.
Approach: They propose a platform for quickly building demos with a focus on knowledge grounded stylized text generation.
Outcome: The proposed framework unifies existing text generation algorithms in a shared codebase and further adapts earlier algorithms for constrained generation.
Learning to Search Effective Example Sequences for In-Context Learning (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods address these factors in isolation, overlooking their interdependencies. Existing approaches focus on sequence selection, while focusing on the sequence of examples.
Approach: They propose a method that considers key factors involved in sequence selection and incrementally builds the sequence.
Outcome: Experiments across various datasets and language models show that the proposed method significantly reduces the search space and improves performance.
Farewell to Aimless Large-scale Pretraining: Influential Subset Selection for Language Model (2023.findings-acl)

Copied to clipboard

Challenge: Pretrained language models have achieved remarkable success in various natural language processing tasks.
Approach: They propose to use end-task knowledge to select a tiny subset of pretraining corpus to influence performance.
Outcome: The proposed model outperforms pretrained models on eight datasets covering four domains with 0.45% of the data and a three-orders-of-magnitude lower computational cost.
Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale (2024.emnlp-main)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) lack visual knowledge in medical applications due to data privacy concerns and high annotation costs.
Approach: They refined medical image-text pairs from PubMed and employed MLLMs (GPT-4V) to denoise and reformat the data.
Outcome: The proposed model significantly improves the MMMU Health & Medicine track and shows that it can be used in multimodal scenarios.
Microsoft Icecaps: An Open-Source Toolkit for Conversation Modeling (P19-3)

Copied to clipboard

Challenge: upcoming open-source natural language processing repository aims to train conversational agents for multi-turn situations.
Approach: They present the Intelligent Conversation Engine: Code and Pre-trained Systems (ICECAPS) the framework wraps TensorFlow functionality in a modular component-based architecture.
Outcome: The Intelligent Conversation Engine: Code and Pre-trained Systems (ICECAPS) is an open-source natural language processing repository.
Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) often experience “contextual hallucination” where they prioritize self-generated content over input context, leading to a disregard for pertinent details.
Approach: They propose a method that dynamically adjusts attention maps to enhance contextual relevance by using a trained classifier to identify attention maps likely to induce hallucinations.
Outcome: The proposed approach reduces hallucinations across open-source models on summarization and open-book QA tasks.
Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction (2021.acl-long)

Copied to clipboard

Challenge: Open-domain question answering is a task to answer questions using passages with diverse topics.
Approach: They propose a model that aggregates evidence from multiple passages to adaptively predict a single answer or a set of question-answer pairs for ambiguous questions.
Outcome: The proposed model achieves state-of-the-art performance on AmbigQA dataset and shows competitive performance on NQ-Open and TriviaQA.
Coarse-to-fine Few-shot Learning for Named Entity Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing few-shot NER solutions do not consider sub-class discrimination and various granularity of new classes during coarse training.
Approach: They propose a method that uses a cluster-based prototype loss to learn group-wise discriminative representations of coarse-grained classes and a mixture prototype loss for learning the representations.
Outcome: The proposed method shows superior performance over baseline methods on in-domain and cross-domain settings with various target granularity.
DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Chain-of-Thought prompting has improved the reasoning capabilities of Large Language Models (LLMs) but it is ineffective or detrimental to the performance on reasoning tasks in Smaller Language Model (SLMs) with less than 10 billion parameters.
Approach: They propose a Dialogue-guided Chain-of-Thought method to improve the reasoning capabilities of Large Language Models (LLMs) by generating intermediate reasoning steps in a dialogue format to guide the model to the final answer.
Outcome: The proposed method can achieve significant performance gains over state-of-the-art competitors on four arithmetic reasoning datasets.
A Neural Network Architecture for Program Understanding Inspired by Human Behaviors (2022.acl-long)

Copied to clipboard

Challenge: Existing studies for understanding programs do not take human behaviors as reference.
Approach: They propose a graph neural network model that takes human behaviors as reference in understanding programs.
Outcome: The proposed model performs better on code summarization and code clone detection tasks.
NICE: Neural Image Commenting with Empathy (2021.findings-emnlp)

Copied to clipboard

Challenge: Emotion and empathy are examples of human qualities lacking in many human-machine interactions.
Approach: They propose to generate images with human-generated comments with enhanced emotion and empathy while minimizing inappropriate or offensive outputs.
Outcome: The proposed model generates more human-like and engaging image comments on two images with human-generated comments and human annotations while minimizing inappropriate or offensive outputs.
HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches treat instruction-based text editing as a generic text generation problem. Existing methods either over-edit or fail to apply modifications consistently.
Approach: They propose a framework that processes each editing request to best align with it.
Outcome: The proposed framework achieves 9% improvement over the state-of-the-art model.
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to optimize RAG generators fail to align with RAG requirements thoroughly.
Approach: They propose a method for optimizing the RAG generator from multiple preference perspectives to align with RAG requirements comprehensively.
Outcome: The proposed method improves the performance of RAG generators by incorporating retrieved documents into the prompt.
DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation (2020.acl-demos)

Copied to clipboard

Challenge: DIALOGPT is a large, tunable neural conversational response generation model . trained on 147M conversation-like exchanges extracted from Reddit comment chains .
Approach: They present a large, tunable neural conversational response generation model, DIALOGPT . the model is trained on 147M conversation-like exchanges extracted from Reddit comment chains .
Outcome: The proposed model can generate more relevant, contentful and context-consistent responses than baseline systems.
LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Graph-based Retrieval-Augmented Generation (GraphRAG) is a new approach to document retrieval, but it is not suitable for legal reasoning.
Approach: They propose a framework for reliable legal reasoning that structures knowledge as relational graphs and uses a multi-agent system to verify validity.
Outcome: The proposed framework outperforms existing GraphRAG models in accurate and trustworthy legal analysis.
Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for language understanding and understanding can be trained to provide contextualized representations of words based on text data.
Approach: They propose a large-scale language VAE model Optimus that is pre-trained on large text corpus and fine-tuned for various language generation and understanding tasks.
Outcome: The proposed model achieves new state-of-the-art on VAE language modeling benchmarks.
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies show that LLM-based agents struggle to perform in zero-shot scenarios.
Approach: They propose a framework to quantify the behavior gap between AI agents and human experts . they propose to examine discrepancies in dialog acts, tool usage, and knowledge utilization .
Outcome: The proposed framework measures the behavior gap between AI agents and human experts on task-oriented dialogs.
Multilingual Generative Retrieval via Cross-lingual Semantic Compression (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for multilingual retrieval still face cross-lingual identifier misalignment and identifiere inflation.
Approach: They propose a framework that unifies semantically equivalent multilingual keywords into shared atoms to align semantics and compresses the identifier space.
Outcome: The proposed framework improves cross-lingual alignment and reduces redundancy.
UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Mobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users’ commands on mobile devices like Android.
Approach: They propose a UI Map-guided LLM-based approach to automate mobile tasks using static analysis and LLMs.
Outcome: The proposed approach achieves a 15.87% higher task execution success rate than SOTA approaches even when only APK is available.
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing prompt refinement methods suffer from semantic inconsistencies and fail to maintain users’ real intent.
Approach: They propose a self-instructed in-context learning framework that generates reliable derived prompts while keeping semantic consistency with original prompts.
Outcome: The proposed framework generates better derived prompts and significantly enhances LLMs’ ability to deliver more effective responses.
An Adaptive Prompt Generation Framework for Task-oriented Dialogue System (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing black-box large language models (LLMs) have excellent performance in task-oriented dialogue (TOD) tasks, but obtaining suitable prompts for specific tasks is challenging.
Approach: They propose a black-box large language model that generates domain and slot information in the belief state, which serves as prior knowledge for subsequent prompt generation.
Outcome: The proposed framework outperforms existing prompting methods on the MultiWOZ 2.0 dataset.
TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to fine-tune code intelligence models to individual tasks are costly and require large data sets.
Approach: They propose a Transferable fine-tuning strategy for Code representation learning that uses a tunable prefix encoder to capture cross-task and cross-language transferable knowledge and apply it to downstream adaptation.
Outcome: The proposed method can lead to superior performance on code-related tasks and encourage mutual reinforcement.
Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code Generation (2026.acl-long)

Copied to clipboard

Challenge: a benchmark is designed to evaluate the capability of Large Multimodal Models (LMMs) in converting complex, structured digital graphics into executable code.
Approach: They propose a benchmark to evaluate the capability of Large Multimodal Models to convert digital graphics into executable code.
Outcome: The proposed benchmark exposes the performance gap among leading LMMs . the benchmark features 1130 meticulously curated samples .
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks lack the ability to automatically evaluate from users’ perspective and lack the explainability of the results of LLM agents’ code generation capabilities.
Approach: They propose a new benchmark for LLM agents' automated evaluation by simulating user interaction.
Outcome: The proposed benchmark can evaluate the generated projects by user interaction simulation and by code similarity through existing objective indicators.
Huatuo-26M, a Large-scale Chinese Medical QA Dataset (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models are a powerful tool for medical research, but the data is a bottleneck.
Approach: They propose to use the largest ever medical Question Answering dataset with 26 Million QA pairs as a fine-tuning data for training large language models.
Outcome: The proposed dataset demonstrates that it can be used to train large language models and improves zero-shot performance on other datasets.
Open Set Relation Extraction via Unknown-Aware Training (2023.acl-long)

Copied to clipboard

Challenge: Existing supervised relation extraction methods can still misclassify unknown relations into known relations due to the lack of supervision signals.
Approach: They propose a method that regularizes the model by dynamically synthesizing negative instances that can provide the missing supervision signals.
Outcome: The proposed method achieves SOTA unknown relation detection without compromising the classification of known relations.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences.
Approach: They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge.
Outcome: The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria.
Boosting Language Models Reasoning with Chain-of-Knowledge Prompting (2024.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Chain-of-Thought (CoT) prompting can be effective on complex reasoning tasks but generates unfaithful and unfactual reasoning chains.
Approach: They propose a chain-of-knowledge prompting that elicits Large Language Models to generate explicit pieces of knowledge evidence in the form of structure triple.
Outcome: The proposed method improves commonsense, factual, symbolic, and arithmetic reasoning tasks by estimating the reliability of the reasoning chains in terms of factuality and faithfulness.
Uncertainty-aware Parameter-Efficient Self-training for Semi-supervised Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for pre-trained language models rely on noisy data, which can be expensive if all parameters are updated.
Approach: They propose a self-training framework that incorporates Monte Carlo dropouts into the model and judiciously selects reliable pseudo-labeled examples based on confidence and certainty.
Outcome: The proposed framework improves performance and efficiency over multiple tasks over multiple datasets.
Knowledge Prompting in Pre-trained Language Model for Natural Language Understanding (2022.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-enhanced pre-trained language models (PLMs) introduce redundant factual knowledge from knowledge bases and require complex modules.
Approach: They propose a knowledge prompting-based PLM framework that incorporates factual knowledge into PLMs.
Outcome: The proposed framework can be flexibly combined with existing mainstream PLMs.
Learning “O” Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NER (2023.acl-long)

Copied to clipboard

Challenge: Existing Named Entity Recognition systems are typically trained on a large-scale dataset with predefined entity classes, then deployed for entity recognition on the test data without further adaptation or refinement.
Approach: They propose a representation learning method that adaptively detects entity clusters in "O" and two effective distance-based relabeling strategies for better learning the old classes.
Outcome: The proposed method achieves 10.62% improvement over the baseline methods.
Dialogue Response Ranking Training with Large-Scale Human Feedback Data (2020.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain dialog models can minimize the perplexity of target human responses . however, some human responses are more engaging than others, spawning more followup interactions .
Approach: They train open-domain dialog models to minimize perplexity of target human responses . they use social media feedback data to train models to predict engaging dialog turns .
Outcome: The proposed model outperforms existing models on 133M human feedback pairs . it also outperformed the conventional dialog perplexity baseline model .
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have evaluated and shown limitations in specific capabilities such as visual understanding, but a systematic evaluation of VLMs’ fundamental WM abilities remains absent.
Approach: They propose a framework that assesses perception and prediction to provide an atomic evaluation of VLMs as WMs.
Outcome: The proposed framework assesses perception and prediction abilities on 15 latest VLMs and compares them to human-level models.
APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that RLHF improves performance of Large Language Models . BT-based RMs struggle to distinguish between similar preference responses .
Approach: They propose to enhance BT-based reward models by using an adaptive margin mechanism . they use semantic similarity and reward-predicted reward differences to adjust focus .
Outcome: Experimental results show that the proposed method outperforms existing methods in both in-distribution and OOD settings.
Graph Counselor: Adaptive Graph Exploration via Multi-Agent Synergy to Enhance LLM Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for enhancing LLM reliability suffer from inefficient information aggregation and rigid reasoning schemes.
Approach: They propose a method that explicitly models external knowledge integration capabilities by explicitly modeling knowledge relationships.
Outcome: The proposed method outperforms existing methods in multiple graph reasoning tasks.
Code Reffix: A Benchmark for Reflection-Guided Code Repair with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on the repair generation capability of LLMs, lacking fine-grained evaluation of reflection.
Approach: They propose a benchmark with oracle reflections and a dual-task protocol to decouple evaluation of reflection from repair.
Outcome: The proposed benchmarks show that underperforming reflection capabilities remain a bottleneck for code repair.
SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Large language models have a tendency to make confidently wrong predictions, highlighting the need for uncertainty quantification (UQ) . previous studies focused on aleatoric uncertainty, but the full spectrum of uncertainties, including epistemic, remains inadequately explored.
Approach: They propose a method to quantify uncertainty in large language models (LLMs) they use a set of perturbations and an aggregation module to generalize the method.
Outcome: The proposed method improves model uncertainty calibration and reduces expected calibration error by 50% on average.
Structure-aware Fine-tuning for Code Pre-trained Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing CodePTMs are mainly structure-free and structurebased, but how to fine-tune them remains a challenge.
Approach: They propose a plug-and-play fine-tuning method that incorporates structural knowledge into pre-trained code models.
Outcome: The proposed method can benefit CodePTMs more with limited training data.
Noise Robust Named Entity Understanding for Voice Assistants (2021.naacl-industry)

Copied to clipboard

Challenge: Named Entity Recognition and Entity Linking are challenging for voice assistants . utterances are relatively short, so there is not much context to help disambiguate .
Approach: They propose a Named Entity Understanding system that combines NER and EL in a joint reranking module.
Outcome: The proposed framework improves NER accuracy by up to 3.13% and EL accuracy by 3.6% in F1 score . it also leads to better accuracies in other natural language understanding tasks .
ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems (2020.acl-demos)

Copied to clipboard

Challenge: ConvLab-2 inherits Convlab's framework but integrates more powerful dialogue models and supports more datasets.
Approach: They present ConvLab-2, an open-source toolkit that enables researchers to build task-oriented dialogue systems with state-of-the-art models and perform an end-to-end evaluation.
Outcome: The new tool inherits ConvLab's framework and extends it by integrating many recently proposed state-of-the-art dialogue models.
Structuring Latent Spaces for Stylized Response Generation (D19-1)

Copied to clipboard

Challenge: Existing methods for generating responses in a targeted style are limited by the lack of parallel data.
Approach: They propose a method that bridges conversation modeling and non-parallel style transfer by sharing a structured latent space.
Outcome: The proposed system generates responses of the targeted style and outperforms baselines without sacrificing appropriateness.
Universal Prompt Optimizer for Safe Text-to-Image Generation (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies based on image checker, model fine-tuning and embedding blocking are impractical in real-world applications.
Approach: They propose a novel reward function measuring toxicity and text alignment of generated images and train the optimizer through Proximal Policy Optimization.
Outcome: The proposed model reduces the likelihood of various models in generating inappropriate images, with no significant impact on text alignment.
Pass-Tuning: Towards Structure-Aware Parameter-Efficient Tuning for Code Representation Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Code pre-trained models have been proposed and widely applied in the domain of code intelligence.
Approach: They propose a method that uses a plug-and-play graph neural network module as a tunable prefix to exploit structural information of source code.
Outcome: The proposed method exploits structural information of source code and could replace full fine-tuning.
Beyond Superficial Tests: Adversarial Refinement for Reliable Property-Based Testing (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation, yet their application to Property-Based Testing (PBT) remains fraught with a superficiality gap.
Approach: They propose an agentic framework that hardens software properties through Adversarial Refinement.
Outcome: a new framework hardens software properties through Adversarial Refinement that detects and fixes bugs in top-tier libraries.
CAT-probing: A Metric-based Approach to Interpret How Pre-trained Models for Programming Language Attend Code Structure (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing code pre-trained models fail to consider inherent characteristics of codes . Existing methods to interpret code pretrained model fail to take into account inherent characteristics .
Approach: They propose a probing method to quantitatively interpret how CodePTMs attend code structure.
Outcome: The proposed method denoises input code sequences and measures commonality between token-level attention scores and pair-wise distances between corresponding AST nodes.
DLM: A Decoupled Learning Model for Long-tailed Polyphone Disambiguation in Mandarin (2024.naacl-long)

Copied to clipboard

Challenge: Grapheme-to-phoneme conversion datasets suffer from the long-tail problem . context learning for polyphonic characters often stems from a single dimension .
Approach: They propose a model for long-tailed polyphone disambiguation in Mandarin that decouples representation and classification learnings.
Outcome: The proposed model can decouple representation and classification learnings . it achieves transition learning of context from local to global .
SCOTT: Self-Consistent Chain-of-Thought Distillation (2023.acl-long)

Copied to clipboard

Challenge: Large language models (LMs) generate free-text rationales for their predictions via chain-of-thought prompting, but there is little guarantee that the generated rationale is consistent with LM’s predictions or faithfully justify the decisions.
Approach: They propose a faithful knowledge distillation method to learn a small, self-consistent CoT model from a larger teacher model by contrastive decoding.
Outcome: The proposed method yields comparable performance but is less faithful than baselines.
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis (2025.findings-acl)

Copied to clipboard

Challenge: Existing large language models (LLMs) are proving to be effective in medical automatic diagnosis, but their interpretability remains unaddressed.
Approach: They propose to use a "Chain-of-Diagnosis" approach to enhance the interpretability of medical automatic diagnosis by outputting the disease confidence distribution.
Outcome: The proposed model outperforms other LLMs on automatic diagnostic tasks across three real-world benchmarks and provides interpretability while ensuring controllability in diagnostic rigor.
Conjoin after Decompose: Improving Few-Shot Performance of Named Entity Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing prompt-based NER models fail to detect entity boundaries, causing performance degradation.
Approach: They propose a model which consists of a BART encoder and a parabiotic decoder and propose ' boundary expansion strategy' to enhance the model's capability in entity type classification.
Outcome: The proposed model can achieve significant performance gains over state-of-the-art models.
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) struggle to use tools reliably in domain-specific settings.
Approach: They propose a neuro-symbolic approach to adapt large language models to task-specific tools . they propose reusable rules that are distilled from failure traces and injected into the prompt .
Outcome: Experiments show that the proposed approach outperforms prompting-based adaptation methods and complements finetuning.
When Gradient Descent Meets Derivative-Free Optimization: A Match Made in Black-Box Scenario (2023.findings-acl)

Copied to clipboard

Challenge: Large pre-trained language models (PLMs) are expensive and may not be open-sourced due to commercial considerations and potential risks of misuse.
Approach: They propose to introduce gradient descent into black-box tuning scenario . they propose a method which integrates gradient descent and derivative-free optimization .
Outcome: The proposed method achieves significant performance gains over previous state-of-the-art methods.
ConvLab: Multi-Domain End-to-End Dialog System Platform (P19-3)

Copied to clipboard

Challenge: ConvLab is an open-source multi-domain end-to-end dialog system platform . it allows researchers to quickly set up experiments with reusable components and compare a large set of different approaches in common environments.
Approach: They propose to use an open-source multi-domain end-to-end dialog system platform to train and evaluate dialog bots in common environments.
Outcome: The proposed system enables researchers to quickly set up experiments with reusable components and compare a large set of different approaches in common environments.
Jointly Optimizing Diversity and Relevance in Neural Response Generation (N19-1)

Copied to clipboard

Challenge: Recent neural conversation models often generate bland and generic responses . however, the improvement often comes at the cost of decreased relevance .
Approach: They propose a spacefusion model to jointly optimize diversity and relevance that fuses the latent space of a sequence-to-sequence model and that of an autoencoder model by leveraging novel regularization terms.
Outcome: The proposed model improves diversity and relevance compared to baselines in both diversity and diversity.
Mitigating Hallucination in Fictional Character Role-Play (2024.findings-emnlp)

Copied to clipboard

Challenge: Influence of parametric knowledge of large language models (LLMs) often causes role-playing characters to act out of character and hallucinate about things outside the scope of their knowledge.
Approach: They propose a method that modulates the influence of parametric knowledge using a pre-calibrated confidence threshold to mitigate hallucination in fictional character role-play.
Outcome: The proposed method reduces the factual accuracy of generated responses by 18% for adversarial questions and 44% in temporal hallucination for time-sensitive interviews.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations