Papers by Zhuang Li

75 papers
Towards Explainable Computerized Adaptive Testing with Large Language Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods focus on minimizing the number of questions required to assess ability, lacking clear and reliable explanations for the question selection process.
Approach: They propose to use large language models to enhance computer adaptive testing (CAT) by providing human-like interpretability and explanations.
Outcome: The proposed agent-based CAT performs comparably or superior to traditional CAT methods in accuracy and significantly improves student trust and satisfaction.
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control (2026.findings-acl)

Copied to clipboard

Challenge: Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, but academic research remains non-reproducible due to the lack of publicly available training data.
Approach: They propose a system for long-form song generation with fine-grained style conditioning that includes a licensed synthetic dataset and a song generation model, Muse.
Outcome: The proposed system achieves competitive performance on phoneme error rate, text–music style similarity, and audio aesthetic quality while enabling controllable segment-level generation across different musical structures.
SelF-Eval: Self-supervised Fine-grained Dialogue Evaluation (2022.coling-1)

Copied to clipboard

Challenge: Existing evaluation metrics are expensive and easy to conduct but ineffective to reflect dialogue quality.
Approach: They propose a self-supervised fine-grained dialogue evaluation framework which can automatically assign fine-granular scores for arbitrarily dialogue data.
Outcome: The proposed framework is highly consistent with human evaluations and better than the state-of-the-art models.
Are U a Joke Master? Pun Generation via Multi-Stage Curriculum Learning towards a Humor LLM (2024.findings-acl)

Copied to clipboard

Challenge: Existing research has demonstrated that the ability of large language models (LLMs) to generate humorous sentences is limited to producing 25 unique jokes.
Approach: They propose a multi-stage curriculum preference learning framework to optimize both pun structure preferences and humor preferences by a Chinese Pun dataset.
Outcome: The proposed method significantly outperforms baseline models on Chinese and English benchmark datasets.
Distributional Alignment for Large Language Models under Domain Shift (2026.findings-acl)

Copied to clipboard

Challenge: Existing distributional alignment models are unstable and degrade under cultural and domain shifts.
Approach: They propose a distributional alignment technique that improves distribution prediction under cultural and domain shift.
Outcome: The proposed method improves fidelity and robustness of LLM distribution estimation under domain and cultural shift.
Cyclical Contrastive Learning Based on Geodesic for Zero-shot Cross-lingual Spoken Language Understanding (2024.findings-acl)

Copied to clipboard

Challenge: zero-shot cross-lingual SLU is a challenging task in low-resource languages . a lack of labeled training data makes it difficult to align representations of similar sentences .
Approach: They propose a framework that uses cyclical contrastive learning to achieve consistency between languages . they propose to use geodesic to measure the similarity to construct positive and negative pairs .
Outcome: The proposed framework achieves state-of-the-art performance on multiATIS++ and MTOP datasets.
PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration (2025.acl-long)

Copied to clipboard

Challenge: Existing privacy protection methods for large language models suffer from performance degradation or large inference time overhead.
Approach: They propose a plug-and-play method to protect the privacy of user inputs during LLM inference . they use offline restoration vectors to train restoration vector for each privacy span type .
Outcome: The proposed method can prevent the linear growth of the privacy budget.
NAP2: A Benchmark for Naturalness and Privacy-Preserving Text Rewriting by Learning from Human (2025.findings-emnlp)

Copied to clipboard

Challenge: a large number of large language models are being used to protect user privacy . sanitizing sensitive text using two common strategies is the answer .
Approach: They propose sanitizing sensitive text using deleting expressions and abstracting them . they propose a tool for text rewriting that uses crowdsourcing and large language models .
Outcome: The proposed approach protects privacy before sending sensitive data to large language models . it combines crowdsourcing and large language modeling to create a text rewrite tool .
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach (2024.findings-eacl)

Copied to clipboard

Challenge: Large pre-trained language models such as GPT-3.5 and GPT-4 have gained significant attention in natural language research due to limited computational resources or inaccessible parameters.
Approach: They propose a neural programmer-interpreter approach that preserves the domain generalization ability of LLMs while editing their output.
Outcome: The proposed framework significantly improves GPT-3.5’s performance in logical form-to-text conversion and low-resource machine translation, surpassing other state-of-the-art (SOTA) LLM post-editing methods in cross-domain settings.
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching (2026.findings-acl)

Copied to clipboard

Challenge: Existing autoregressive models for dialogue generation suffer from high latency and stability issues.
Approach: They propose a non-autoregressive (NAR) zero-shot spoken dialogue generation model based on flow-matching.
Outcome: The proposed model outperforms existing models in speech generation due to poor speech intelligibility and turn-taking precision.
Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on neurons focus on emotion and rhetoric, neglecting their intrinsic connections.
Approach: They propose a framework for fine-grained steering of emotion and rhetoric in large language models . they propose 'neuro-based' masking method that integrates multi-dimensional screening .
Outcome: The proposed method achieves directed induction of non-target sentences and enhancement of emotion tasks via rhetoric neurons.
ClusterAttn: KV Cache Compression under Intrinsic Attention Clustering (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for sparse attention apply the same pattern across different attention heads and inputs, but fail to capture the intrinsic attention clustering in large language models.
Approach: They propose a training-free sparse attention method that provides an efficient prompt cache compression scheme under intrinsic attention clustering for efficient LLM inference.
Outcome: The proposed method reduces memory usage by 10%–65% and increases throughput by 2.6–4.8 times with no accuracy loss.
VOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder (2024.findings-naacl)

Copied to clipboard

Challenge: generative diversity is a critical yet underexplored issue in natural language generation . previous approaches to enhance diversity of Transformer models have been limited by their latent variables .
Approach: They propose a framework that bridges Transformer with VAE to enhance generative diversity.
Outcome: The proposed framework improves generative diversity while maintaining generative quality.
Align2LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Multi-modal Large Language Models (MLLMs) introduce significant variability in data quality.
Approach: They propose to use human and LLM preference alignment to compress large corpus of machine-generated multimodal instructions into a compact and high-quality form.
Outcome: The proposed algorithm outperforms LLaVA-series models in MLLM benchmarks by 90% . it uses human and LLM preference alignment to compress a large dataset .
TeamLoRA: Boosting Low-Rank Adaptation with Expert Collaboration and Competition (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning are resource-efficient, but performance often falls short . a new approach, TeamLoRA, integrates collaborative and competitive modules to improve performance.
Approach: They propose to introduce task-specific LoRA as domain experts to improve learning efficiency . teamLoRA integrates collaborative and competition modules to improve model learning .
Outcome: Experiments show that TeamLoRA improves performance in multi-task learning . teamLorea integrates collaborative and competitive modules to improve performance .
LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models struggle to detect lazy thinking in a zero-shot setting, but instruction-based fine-tuning significantly boosts performance by 10-20 performance points.
Approach: They propose to use LazyReview to train junior reviewers in the community to detect lazy thinking in peer-review sentences annotated with fine-grained lazy thinking categories.
Outcome: The proposed dataset shows that LLMs struggle to detect lazy thinking instances in a zero-shot setting, while instruction-based fine-tuning significantly boosts performance by 10-20 performance points.
Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training (2025.naacl-long)

Copied to clipboard

Challenge: Existing LLMs often rely on complex prompting or extensive fine-tuning to introduce new capabilities while preserving strong generalizability.
Approach: They propose a large-scale pre-training corpus to enhance LLM agents' capabilities . they use 103B agent-specific data encompassing 76,537 APIs .
Outcome: The proposed training corpus outperforms open-source LLMs and commercial LLM agents on three agent benchmarks.
Meta-Reflection: A Feedback-Free Reflection Learning Framework (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve large language models' ability to understand and reason are limited by external feedback.
Approach: They propose a feedback-free reflection mechanism that requires only a single inference pass without external feedback.
Outcome: The proposed method is based on an industrial e-commerce benchmark and public datasets.
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis (2025.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks fail to adequately evaluate the proficiency of Large Language Models (LLMs) Existing standards do not cover the skills needed to evaluate LLMs in scientific literature analysis.
Approach: They propose a benchmark to evaluate the proficiency of large language models in scientific literature analysis.
Outcome: SciAssess evaluates 11 LLMs on multiple tasks across scientific fields.
Incorporating Global Information in Local Attention for Knowledge Representation Learning (2021.findings-acl)

Copied to clipboard

Challenge: Graph Attention Networks (GATs) are a promising model that takes advantage of localized attention mechanism to perform knowledge representation learning (KRL) on graph-structure data.
Approach: They propose to incorporate global information into the GAT family of models by using an attention-based global random walk algorithm.
Outcome: Experimental results on KG entity prediction against the state-of-the-arts demonstrate the effectiveness of the proposed model.
Towards Explainable Chinese Native Learner Essay Fluency Assessment: Dataset, Tasks, and Method (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing GEC datasets in Chinese fail to consider specific grammatical error types and overlook cross-sentence grammamatical errors.
Approach: They propose to use Chinese essay fluency assessment to assess essay fluencies along with coarse and fine-grained errors and corrections to improve explainability.
Outcome: The proposed dataset encapsulates essay fluency scores along with both coarse and fine-grained errors and corrections.
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods such as Medusa lack adequate information interaction between different drafting heads.
Approach: They propose an enhanced speculative decoding framework that builds upon Medusa and integrates a drafting block capable of parallel inference.
Outcome: The proposed framework outperforms Medusa in terms of head accuracy and latency.
On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on Codex (2023.eacl-main)

Copied to clipboard

Challenge: Existing techniques for parsing natural-language utterances are vulnerable to adversarial attacks, requiring large amounts of labelled data and expensive human annotation.
Approach: They propose to enhance the adversarial robustness of a prompt-based semantic parser based on a language model trained on code by constructing a set of demonstration examples.
Outcome: The proposed method can be enhanced without significant amounts of labelled data or expensive human annotations on in-domain semantic parsing data.
On the Reliability of Large Language Models for Causal Discovery (2025.acl-long)

Copied to clipboard

Challenge: Existing statistical methods to identify causal relationships from observational data remain elusive.
Approach: They examine the impact of memorization for accurate causal relation prediction, the influence of incorrect causal relations in pre-training data and the contextual nuances that influence LLMs’ understanding of causal relations.
Outcome: The proposed models are effective in recognizing causal relations that occur frequently in pre-training data, but their ability to generalize to new or rare causal relations is limited.
Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for task-specific natural language generation do not contain any labeled examples.
Approach: They propose a variational autoencoder with disentanglement priors for task-specific natural language generation with none or a handful of task-related labeled examples.
Outcome: The proposed model outperforms baseline models in terms of data augmentation and text style transfer in the few-shot setting.
Active Learning for Multilingual Semantic Parser (2023.findings-eacl)

Copied to clipboard

Challenge: Existing multilingual semantic parsing datasets are limited in translation effort due to data imbalance.
Approach: They propose a first active learning procedure for multilingual semantic parsing (AL-MSP) it selects only a subset from existing datasets to be translated, they propose .
Outcome: The proposed method significantly reduces translation costs with ideal selection methods.
SOLAR: Serendipity Optimized Language Model Aligned for Recommendation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have shown strong potential in recommendation tasks . however, their application to serendipity-oriented recommendations remains challenging .
Approach: They propose a domain-adaptive instruction tuning method that aligns Large Language Models with recommendation tasks.
Outcome: The proposed framework bridges the domain gap between LLMs and recommendation tasks.
Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are designed as specific task solvers with sophisticated prompt engineering, but are inherently incapacitating to address complex dynamic scenarios.
Approach: They propose an LLM-based agent with policy-level reflection and optimization that can learn from interactive experiences and progressively elevate its behavioral policy.
Outcome: The proposed agent outperforms vanilla LLM and specialized models in blackjack and Texas hold’em.
How Robust Are Large Language Models for Clinical Numeracy? An Empirical Study on Numerical Reasoning Abilities in Clinical Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models for clinical numerical reasoning provide limited operation-level coverage and limited robustness of numerical understanding across clinical note formats.
Approach: They propose a benchmarking tool that evaluates four main types of clinical numeracy . they present longitudinal MIMIC-IV vital-sign records in three semantically equivalent representations .
Outcome: The proposed benchmark evaluates four main types of clinical numeracy: value retrieval, arithmetic computation, relational comparison, and aggregation.
MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning (PEFT) are limited by computational costs and performance degradation.
Approach: They propose a method that integrates Low-Rank Adaptation and Mixture-of-Experts (MoE) they propose combining expert load imbalance and representation collapse to improve LLM performance .
Outcome: The proposed method outperforms homogeneous MoE-LoRA architectures in performance and parameter efficiency.
Neural-based Mixture Probabilistic Query Embedding for Answering FOL queries on Knowledge Graphs (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods to embed entities and first-order logical queries in a vector space are often violated in real applications and limit their performance.
Approach: They propose a Neural-based Mixture Probabilistic Query Embedding Model that embeds entities and first-order logical queries in a vector space.
Outcome: The proposed model outperforms state-of-the-art methods on benchmark datasets.
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that manually ensuring a consistent response style and maintaining high data quality can significantly improve the performance of fine-tuned Large Language Models (LLMs).
Approach: They introduce a style-aware response ranking system that prioritizes instruction-response pairs based on their stylistic consistency.
Outcome: The proposed model matches or surpasses models trained on the entire dataset in coding and open-ended question-answering benchmarks.
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal Mixture-of-Experts models accurately perceive image content yet fail in subsequent reasoning . Seeing but not thinking phenomenon is a puzzling phenomenon .
Approach: They propose a routing-guided intervention method that enhances domain expert activation.
Outcome: The proposed method achieves consistent improvements on visual reasoning tasks.
Total Recall: a Customized Continual Learning Method for Neural Semantic Parsers (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for continual learning for semantic parsing fail to account for special properties of structured outputs . retraining from scratch is not feasible due to the fast growing number of tasks .
Approach: They propose a continual learning method that uses sequential learning to learn tasks without accessing full training data from previous tasks.
Outcome: The proposed method achieves a 3-6 times speedup compared to re-training from scratch.
AIGuard: A Benchmark and Lightweight Detection for E-commerce AIGC Risks (2025.findings-acl)

Copied to clipboard

Challenge: Existing detection methods lack real-world scenarios and corresponding risk datasets . current MLLMs lack knowledge and have limited capability to detect the risk of AIGC content.
Approach: They propose a benchmark for AIGC risk detection in real-world e-commerce . it includes 253,420 image-text pairs across four critical categories .
Outcome: The proposed method achieves 9.68% higher recall than leading multimodal models while using only 25% of training resources.
ReSel: N-ary Relation Extraction from Scientific Text and Tables by Learning to Retrieve and Select (2022.emnlp-main)

Copied to clipboard

Challenge: Our proposed method extracts N-ary relation tuples from scientific articles.
Approach: They propose a method that decomposes the task into two stages . they propose modal query and modal entity selection . their results show that ReSel outperforms state-of-the-art baselines significantly .
Outcome: The proposed method outperforms state-of-the-art baselines on three scientific information extraction datasets.
The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active Learning (2023.acl-long)

Copied to clipboard

Challenge: Prior studies have focused on translating utterances from high-resource languages to low-resourced languages.
Approach: They propose an active learning approach that exploits the strengths of both human and machine translations by iteratively adding small batches of human translations into the machine-translated training set.
Outcome: The proposed approach reduces errors and bias in the translated data, resulting in higher parser accuracies than the current model trained on machine translations.
SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social Media (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for opinion survey research exhibit severe biases and lack traceability.
Approach: They propose a finite-state orchestrated agentic framework that automates the collection and analysis of human opinions from social media platforms.
Outcome: The proposed framework achieves close alignment with authentic survey results across multiple domains, with average relative improvements of 68,98% and 51,37% when compared to opinion synthesis and agent-based approaches.
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Current TIMT studies focus on providing translations for all text within an image, neglecting to provide bounding boxes and covering limited scenarios.
Approach: They extend traditional TIMT into position-aware TIMt to support fine-grained translation . they introduce an Adaptive Image OCR Refinement Pipeline to refine results .
Outcome: The proposed model supports fine-grained and layout-preserving translation . the experimental data highlight the scalability and generalizability of the model.
ScEdit: Script-based Assessment of Knowledge Editing (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge Editing (KE) has gained increasing attention, yet current evaluation frameworks do not integrate KE into real-world application scenarios.
Approach: They propose a script-based benchmark which encompasses both counterfactual and temporal edits and integrates token-level and text-level evaluation methods.
Outcome: The proposed method combines token-level and text-level evaluation methods with a new fact-based evaluation framework.
TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in natural language processing tasks but are vulnerable to harmful content and being exploited for malicious purposes.
Approach: They propose a framework to measure the risk coverage of alignment datasets across three dimensions: Lexical Diversity, Malicious Intent, and Jailbreak Tactics.
Outcome: The proposed framework measures risk coverage across Lexical Diversity, Malicious Intent, and Jailbreak Tactics.
RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis (2025.emnlp-main)

Copied to clipboard

Challenge: Existing solutions to fine-tune large language models for domain-specific tasks are ineffective in addressing privacy concerns.
Approach: They propose a privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.
Outcome: The proposed framework fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.
DORM: Preference Data Weights Optimization for Reward Modeling in LLM Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to align large language models with human preferences are noisy and varying in importance of preference samples.
Approach: a new method enhances reward modeling by learning to dynamically weigh preference data.
Outcome: a new method improves the performance of large language models with human preferences . it initializes data importance and iteratively refines them to maximize validation performance.
DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement (2025.emnlp-main)

Copied to clipboard

Challenge: Current approaches typically merge sentence-level parsing outputs for discourse input, resulting in fragmented graphs and degraded downstream performance.
Approach: They propose a task for discourse-level text scene graph parsing that merges sentence-level outputs for discourse input and propose 'DiscoSG' a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence caption-graph pairs is used to test the new task.
Outcome: The proposed task improves SPICE by 30% over the baseline while achieving 86 faster inference than existing models.
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training (2024.emnlp-main)

Copied to clipboard

Challenge: Existing text-to-image models struggle to generate images with legible visual texts . current models lack support for Chinese texts, misspelling, and lack of diversity .
Approach: They propose to empower backbone models to generate visual texts in Chinese and English . they propose to augment conventional training objective with glyph-aware training losses .
Outcome: The proposed methods can generate visual texts in English and Chinese while maintaining image generation quality.
T2S-GPT: Dynamic Vector Quantization for Autoregressive Sign Language Production from Text (2024.acl-long)

Copied to clipboard

Challenge: Existing vector quantization methods are fixed-length encodings, overlooking the uneven information density in sign language.
Approach: They propose a two-stage sign language production paradigm that encodes sign language sequences into discrete codes and autoregressively generates sign languages from text.
Outcome: The proposed model can dynamically adjust the encoding length based on the information density in sign language to achieve accurate and compact encoded enccoding.
RENOVI: A Benchmark Towards Remediating Norm Violations in Socio-Cultural Conversations (2024.findings-naacl)

Copied to clipboard

Challenge: Norm violations occur when individuals fail to conform to culturally accepted behaviors, which may lead to potential conflicts.
Approach: They propose to use a large corpus of 9,258 multi-turn dialogues annotated with social norms to equip AI systems with a remediation ability.
Outcome: The proposed system can understand and remediate norm violations step by step.
IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models (2024.acl-long)

Copied to clipboard

Challenge: IMO is a machine learning model that learns invariant features from unseen domains.
Approach: They propose IMO: Invariant features Masks for Out-of-Distribution text classification to achieve OOD generalization by learning invariant feature masks.
Outcome: The proposed model outperforms baseline models in various evaluation metrics and settings.
Audio-centric Video Understanding Benchmark without Text Shortcut (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in multimodal large language models (MLLMs) focus on visual abilities, but audio is essential for video understanding.
Approach: They propose an audio-centric video understanding benchmark to evaluate video comprehension capabilities of multimodal LLMs with a particular focus on auditory information.
Outcome: The proposed video understanding benchmarks evaluate video comprehension capabilities of multimodal models with a particular focus on auditory information.
Beyond Benchmarks: A Capability-Based Maturity Model for Systematic AI Integration in Hospitals (2026.findings-acl)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) excel in standardized tests focused on medical knowledge recall, but not in real-world healthcare scenarios.
Approach: They propose a "capability-based hospital AI Maturity Model" framework that categorizes capabilities into distinct maturity levels . medical artificial intelligence is currently at a critical transition stage from technical verification to deep clinical integration .
Outcome: The proposed model provides a clear, stepwise evolutionary path for hospitals from foundational infrastructure construction to ubiquitous intelligence.
DiffusionNER: Boundary Diffusion for Named Entity Recognition (2023.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) tasks are fundamental to many structured information extraction tasks.
Approach: They propose a named entity recognition task that uses a boundary-denoising diffusion process to denoise noisy spans.
Outcome: The proposed method achieves comparable or even better performance than previous state-of-the-art models on flat and nested datasets.
QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Existing review-based product question answering systems generate only a single answer, ignoring the diversity of viewpoints.
Approach: They propose a task which aims to summarize diverse customer opinions into representative Key Points and quantify their prevalence to effectively answer user queries.
Outcome: The proposed task summarizes diverse customer opinions into representative Key Points and quantifies their prevalence to answer user queries.
Multi-level Relevance Document Identifier Learning for Generative Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Existing methods generate DocIDs based on textual content, which may result in weak semantic connections for similar documents due to variations in expression.
Approach: They propose a new retrieval paradigm that generates unique document identifiers . they propose to use queries as a bridge to connect documents with varying relevance levels .
Outcome: The proposed approach outperforms existing methods on multilingual e-commerce search datasets.
FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing (2023.findings-acl)

Copied to clipboard

Challenge: Existing parsers that convert image captions into scene graphs often suffer from errors and inconsistency.
Approach: They propose a dataset that re-annotates image captions using a new intermediate representation called FACTUAL-MR and a metric to measure scene graph similarity.
Outcome: The proposed parser outperforms existing parsers in terms of faithfulness and consistency on multiple benchmark datasets.
FaithfulPersona: Balancing Faithfulness and Personalization in Code Explanations through Self-Critique (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for generating faithful code explanations face challenges balancing faithfulness to the original code and personalization for diverse user needs.
Approach: They propose a benchmark and method for generating faithful personalized code explanations using code samples and user profiles.
Outcome: The proposed method achieves 3.7% improvement in Pass@5 compared to the strong baseline method, Self-Consistency, while maintaining high personalization with a 61.08% win rate in the LLM-as-a-Judge evaluation.
Boosting LLM’s Molecular Structure Elucidation with Knowledge Enhanced Tree Search Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Molecular structure elucidation involves deducing a molecule’s structure from various types of spectral data, which is crucial in chemical experimental analysis.
Approach: They propose a Knowledge-enhanced reasoning framework for Molecular Structure Elucidation that leverages Monte Carlo Tree Search for test-time scaling as a plugin to extend the LLMs’ coverage of the chemical structure space.
Outcome: The proposed framework significantly improves on both GPT-4o-mini and GPT4o, and a specialized molecule-spectrum scorer improves performance.
BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for deep search agents rely on blackbox web search APIs . dynamic and opaque web APIs hinder reproducibility and fair comparisons - authors .
Approach: They propose a benchmark that employs a fixed corpus for controlled retrieval for deep search agents.
Outcome: The new benchmark shows that agents that combine large language models with retrieval tools excel at complex, reasoning-intensive queries.
InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing grounding approaches work well for simple queries, but many real-world information needs require synthesizing multiple pieces of evidence.
Approach: They introduce "integrative grounding" to evaluate the ability to ground large language models in external knowledge sources.
Outcome: The proposed approach is robust to redundant evidence, but rationalizes using internal knowledge when information is incomplete.
TMID: A Comprehensive Real-world Dataset for Trademark Infringement Detection in E-Commerce (2023.emnlp-industry)

Copied to clipboard

Challenge: Annually, e-commerce platforms incur substantial financial losses due to trademark infringements.
Approach: They propose a dataset to detect trademark infringement in merchant registrations . they use legal rules and contextual information from Alipay to gather contextual information with annotations from legal experts.
Outcome: The proposed dataset is sourced from Alipay, one of the world’s largest e-commerce and digital payment platforms.
AHEAD: Attention Head Energy-Aware Dynamics for Hallucination Mitigation in MLLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to hallucination mitigation ignore heterogeneous behaviors of attention heads . hallucinosity is a critical barrier to multimodal large language models' reliability, authors say .
Approach: They propose a framework that quantifies the energetic properties of each attention head during object generation through two potential networks and dynamically adjusts their contributions at inference time.
Outcome: The proposed framework reduces hallucination rates without fine-tuning the base model while maintaining generation quality.
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a promising new approach to understanding biological sequences such as proteins.
Approach: They propose an LLM that can generate protein sequences in human and protein languages by pre-training an Lm on protein and natural language corpora and supervised instruction tuning to facilitate alignment.
Outcome: The proposed model outperforms state-of-the-art LLMs on protein-text generation tasks by a large margin.
Few-shot In-context Learning on Knowledge Base Question Answering (2023.acl-long)

Copied to clipboard

Challenge: KB-BINDER enables few-shot in-context learning over knowledge base questions . KBQA is a difficult problem due to the heterogeneity of knowledge bases .
Approach: They propose a framework that enables few-shot in-context learning over KBQA tasks.
Outcome: The proposed framework can outperform state-of-the-art models on GraphQA and MetaQA datasets.
PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments (2026.findings-acl)

Copied to clipboard

Challenge: evaluating therapeutic competence of large language models remains challenging due to unstructured and longitudinal nature of counseling.
Approach: They propose a framework that calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.
Outcome: The proposed framework calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments.
Let’s Negotiate! A Survey of Negotiation Dialogue Systems (2024.findings-eacl)

Copied to clipboard

Challenge: Recent research has focused on negotiation dialogue systems, but no systematic review of this task has been conducted.
Approach: They propose to provide a systematic review of negotiation dialogue systems and to provide an overview of current research.
Outcome: The proposed systems are based on the literature and are compared against existing systems.
From Scores to Preferences: Redefining Evaluation Paradigm for Speech Quality Reward Modeling (2026.findings-acl)

Copied to clipboard

Challenge: Experimental results show that the MOS-aware GRM significantly improves fine-grained speech quality discrimination.
Approach: They propose a MOS-aware reward model that incorporates MOS gap into reward function during reinforcement learning.
Outcome: The proposed model significantly improves fine-grained speech quality discrimination.
Context Dependent Semantic Parsing: A Survey (2020.coling-main)

Copied to clipboard

Challenge: Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations.
Approach: They propose to use contextual information to translate natural language utterances into machine-readable meaning representations.
Outcome: The proposed methods do not utilize contextual information, which could boost the semantic parsing systems.
Few-Shot Semantic Parsing for New Predicates (2021.eacl-main)

Copied to clipboard

Challenge: a recent study shows that state-of-the-art neural semantic parsers are less accurate when there is only a handful of utterance-logical form pairs per predicate.
Approach: They propose to use a meta-learning method to train a few-shot learning problem . they also propose to regularize attention scores with alignment statistics and apply a smoothing technique .
Outcome: The proposed method outperforms baselines in one and two-shot settings.
POLYIE: A Dataset of Information Extraction from Polymer Material Scientific Literature (2024.naacl-long)

Copied to clipboard

Challenge: SciIE datasets for polymer materials are lacking for this class of materials . POLYIE is curated from 146 full-length polymer scholarly articles .
Approach: They propose a SciIE dataset for polymer materials that uses entity annotations from 146 full-length articles.
Outcome: The proposed dataset is curated from 146 full-length polymer scholarly articles . it presents challenges due to diverse lexical formats of entities and ambiguity between entities .
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in reasoning models have demonstrated remarkable capabilities on mathematical and coding tasks, but their effectiveness in embodied domains remains largely unexplored.
Approach: They propose a reasoning model for interactive embodied tasks that synthesizes 9.3k coherent Observation-Thought-Action trajectories containing 64k ego-centric images and 90k diverse reasoning processes.
Outcome: The proposed model outperforms existing visual reasoning models by +9%, 24%, and +13% on long-horizon tasks.
Neural-DINF: A Neural Network based Framework for Measuring Document Influence (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to measure scholarly impact of documents without citations only consider word frequency change.
Approach: They propose a neural network framework that measures document influence without citations by using word frequency changes and word semantic shifts.
Outcome: The proposed model outperforms existing models on document influence evaluation without citations.
On Robustness of Neural Semantic Parsers (2021.eacl-main)

Copied to clipboard

Challenge: Semantic parsing maps natural language (NL) utterances into logical forms (LFs) adversarial examples are created by adding tiny perturbations to inputs but can severely deteriorate model performance.
Approach: They propose to construct robustness test sets based on existing benchmark corpora and to evaluate the effect of data augmentation.
Outcome: The proposed method measures the performance of the proposed parsers on robustness test sets and evaluates the effect of data augmentation.
iTAG: Inverse Design for Natural Text Generation with Accurate Causal Graph Annotations (2026.acl-long)

Copied to clipboard

Challenge: Lack of causally annotated text data for use as ground truth hinders causal discovery . early template-based generation methods sacrifice text naturalness in exchange for high annotation costs .
Approach: They propose a method which performs real-world concept assignment to nodes before converting causal graphs into text.
Outcome: The proposed method shows high annotation accuracy and naturalness across extensive tests.
CultureInstruct: Curating Multi-Cultural Instructions at Scale (2025.naacl-long)

Copied to clipboard

Challenge: Large language models exhibit severe cultural bias, despite their success in recent years . a critical challenge of LLMs is integration of cultural knowledge into these models .
Approach: They propose a large-scale instruction-tuning dataset to reduce cultural bias in large language models.
Outcome: The proposed model outperforms GPT-4o Mini and GPT-42 with 18.47% and 13.07% relative improvements on cultural benchmarks.
Example Quality Matters: Multi-Aspects Example Augmentation for Private Library Programming (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to code generation fail to consider the quality of retrieved examples.
Approach: They propose a retrieval-augmented generation method that combines existing API examples to improve complexity and readability.
Outcome: The proposed method achieves up to 22% accuracy improvement over baseline methods.
ModularMoE: Fast LLM Customization with Parameter-Sharing Mixture-of-Experts for Low-Resource Settings (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models impose significant computational and storage burdens on personal devices . existing customization approaches incur excessive computational costs or lead to suboptimal performance .
Approach: They propose a training framework that converts pre-trained LLMs into parameter-sharing MoE models for lightweight deployment.
Outcome: The proposed training framework outperforms state-of-the-art training frameworks at the same sparsity level while delivering up to 2.71 inference speedup.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations