Papers by Rui Lu

41 papers
TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method (2022.emnlp-main)

Copied to clipboard

Challenge: a new lyric-to-melody generation system bridges the gap between lyrics and melodies . previous generation systems lack paired data and lack of control on generated melodie.
Approach: They develop a lyric-to-melody generation system with music template to bridge the gap between lyrics and melodies.
Outcome: The proposed system bridges the gap between lyrics and melodies by using music template.
Optimizing Code Retrieval: High-Quality and Scalable Dataset Annotation through Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for code retrieval struggle to balance scalability and annotation quality.
Approach: They propose a method that integrates functions called within the repository and information on third-party APIs to enhance the annotation context.
Outcome: The proposed method improves the annotation context by incorporating functions called within the repository and information on third-party API functionalities.
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing (2024.emnlp-main)

Copied to clipboard

Challenge: a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities .
Approach: They present a comparative analysis to identify and distinguish LLM activities from human activities.
Outcome: The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities.
Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event Forecasting (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods focus on entities and structural dependencies but overlook implicitly relevant information.
Approach: They propose a method that leverages event semantics for relevance modeling and incorporates a self-supervised semantic filter based on factual event associations to capture implicitly relevant historical information.
Outcome: The proposed method outperforms existing methods on three public benchmark datasets and is highly effective on two structured temporal knowledge graph forecasting datasets.
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods rely on ranking losses to teach reward model to assess preferences, but they are susceptible to noise and ambiguous data, often failing to deeply understand human intentions.
Approach: They propose a method that incorporates contrastive learning into the reward modeling process to enhance generalization and stabilize the reinforcement learning training process.
Outcome: The proposed method enhances generalization of the reward model, stabilizes the reinforcement learning training process, and improves the final alignment with human preferences.
SDA: Simple Discrete Augmentation for Contrastive Sentence Representation Learning (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for data augmentation have not been well explored.
Approach: They propose to use punctuation insertion, modal verbs, and double negation to produce diverse forms of sentences.
Outcome: The proposed methods perform better on diverse datasets with semantic similarity and standard negation.
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing tools for detecting safety issues in LLMs are expensive and inefficient.
Approach: They propose an LLM-based safety detector which annotates the safety of queries and provides explanations for its decisions.
Outcome: The proposed detector outperforms baselines on four sets of query-response pairs and is effective as a safety evaluator for advanced LLMs.
Deeply Coupled Cross-Modal Prompt Learning (2023.findings-acl)

Copied to clipboard

Challenge: Existing prompt-tuning methods focus on language branch or learn vision-language interaction in a shallow mechanism.
Approach: They propose a Deeply coupled Cross-modal Prompt learning method based on CLIP to facilitate the interplay between vision and language with a Cross-Modal Prompting Attention mechanism.
Outcome: The proposed method enables the interplay between vision and language with a Cross-Modal Prompt Attention mechanism.
AgentGym: Evaluating and Training Large Language Model-based Agents across Diverse Environments (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are promising foundations to build generally-capable agents . however, the community lacks a unified interactive framework that covers diverse environments for comprehensive evaluation of agents.
Approach: They propose a framework that features 7 real-world scenarios, 14 environments, and 89 tasks for unified, real-time, and concurrent agent interaction.
Outcome: The proposed framework features 7 real-world scenarios, 14 environments, and 89 tasks for unified, real-time, and concurrent agent interaction.
CLEAN–EVAL: Clean Evaluation on Contaminated Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to evaluate large language models are prone to data contamination.
Approach: They propose a method which parses contaminated data and back-translates it into a candidate set.
Outcome: The proposed method reduces data contamination and evaluates the LLMs more cleanly.
AgentTuning: Enabling Generalized Agent Abilities for LLMs (2024.findings-acl)

Copied to clipboard

Challenge: Open large language models (LLMs) with great performance in various tasks are far inferior to commercial models such as ChatGPT and GPT-4 when acting as agents to tackle complex tasks in the real world.
Approach: They propose a method to enhance the agent capabilities of LLMs while maintaining their general abilities.
Outcome: The AgentLM-70B is comparable to GPT-3.5-turbo on unseen agent tasks, demonstrating generalized agent capabilities.
Model Surgery: Modulating LLM’s Behavior Via Simple Parameter Editing (2025.naacl-long)

Copied to clipboard

Challenge: Current approaches for detoxification or preventing jailbreaking involve fine-tuning billions of parameters through gradient descent with substantial computational cost.
Approach: They propose to use supervised fine-tuning and Reinforcement Learning from human feedback to modify LLMs' behavior by directly editing a small subset of parameters.
Outcome: Experiments show that editing a small subset of parameters can modulate specific behaviors of LLMs, such as detoxification and resistance to jailbreak, with only inference-level computational resources.
Word Sense Disambiguation with Knowledge-Enhanced and Local Self-Attention-based Extractive Sense Comprehension (2022.coling-1)

Copied to clipboard

Challenge: Word sense disambiguation (WSD) is one of the most challenging tasks in natural language processing.
Approach: They propose a method to extract the right sense from a sentence context . they propose to incorporate additional examples and definitions of related senses in WordNet .
Outcome: The proposed method achieves better performance than baseline models on public benchmark datasets.
ModRWKV: Transformer Multimodality in Linear Time (2025.emnlp-main)

Copied to clipboard

Challenge: Currently, multimodal studies are based on large language models with quadratic-complexity Transformer architectures.
Approach: They propose a decoupled multimodal framework built upon the RWKV7 architecture as its LLM backbone and a lightweight architecture to achieve multi-source information fusion.
Outcome: The proposed framework achieves multi-source information fusion through dynamically adaptable heterogeneous modality encoders.
Eliminating Out-of-Domain Recommendations in LLM-based Recommender Systems: A Unified View (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to reduce OOD recommendations fall into three grounding paradigms: retrieval, constrained generation and discrete item tokenizer generation.
Approach: They propose a framework that instantiates three grounding paradigms under a single architecture . embedding-based retrieval, constrained generation and discrete item-tokenizer methods are implemented .
Outcome: The proposed framework eradicates OOD recommendations across all variants and achieves state-of-the-art accuracy compared to strong ID-based and LLM-based baselines.
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal Mixture-of-Experts models accurately perceive image content yet fail in subsequent reasoning . Seeing but not thinking phenomenon is a puzzling phenomenon .
Approach: They propose a routing-guided intervention method that enhances domain expert activation.
Outcome: The proposed method achieves consistent improvements on visual reasoning tasks.
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to visual chain-of-thought are limited by external tools or fail to generate high-fidelity diagrams.
Approach: They propose a framework to enable large multimodal models with VCoT capabilities . they pre-train a model on a 15.2M-pair corpus and teach it how to leverage visual aids .
Outcome: The proposed framework unlocks complex, human-like visual reasoning in large language models . it pre-trains the model on a 15.2M-pair corpus and fine-tunes it on MathCanvas-Instruct .
CLEAR: Can Language Models Really Understand Causal Graphs? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing language models lack a conceptual framework for understanding causal graphs, but there is still potential for improvement.
Approach: They develop a framework to define causal graph understanding by assessing language models’ behaviors through four practical criteria derived from diverse disciplines.
Outcome: The proposed framework defines three complexity levels and encompasses 20 causal graph-based tasks across 20 different levels.
Uncovering Main Causalities for Long-tailed Information Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Information Extraction (IE) aims to extract structural information from unstructured texts.
Approach: They propose a framework that aims to uncover the main causalities behind data in the view of causal inference.
Outcome: The proposed framework can detect the main causalities behind data in the view of causal inference.
Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for long-form speech are limited to limited domains, creating a significant gap with the diverse downstream applications.
Approach: They propose a benchmark that decomposes "long-form speech quality" into specific, disentangled dimensions.
Outcome: The proposed benchmark decomposes “long-form speech quality” into specific, disentangled dimensions.
CA-GAR: Context-Aware Alignment of LLM Generation for Document Retrieval (2025.findings-acl)

Copied to clipboard

Challenge: Recent techniques such as Generation-Augmented Retrieval (GAR) and Generative Document Retrieleval (GDR) leverage LLMs to enhance retrieval performance but face key challenges: GAR’s generated content may not always align with the target document corpus, while GDR limits the generative capacity of LLM.
Approach: They propose a Context-Aware Generation-Augmented Retrieval approach which integrates corpus information into their generation process.
Outcome: Experimental results show that CA-GAR outperforms existing methods on seven tasks and four non-English languages.
Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the Uyghur (2025.coling-main)

Copied to clipboard

Challenge: Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data.
Approach: They propose a direct preference optimization based on translation self-evolution to expand low-resource languages into large language models by using Uyghur as an example.
Outcome: The proposed strategy expands low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data.
Fair Abstractive Summarization of Diverse Perspectives (2024.naacl-long)

Copied to clipboard

Challenge: Existing work on summarization metrics and large language models has not explored fair abstractive summarizing.
Approach: They propose four reference-free automatic metrics to measure the differences between target and source perspectives.
Outcome: The proposed methods alleviate fair abstractive summarization on user-generated data.
The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story Characters (2025.acl-long)

Copied to clipboard

Challenge: Theory-of-Mind (ToM) is a psychological capability that allows humans to understand and interpret the mental states of others.
Approach: They propose a CharToM-QA benchmark to assess the importance of comprehensive contextual understanding about personal backgrounds in ToM.
Outcome: The proposed model outperforms existing models on 1,035 ToM questions based on classic novels and shows that educated participants perform better when they have read the novels than non-educated participants.
Reward Modeling Requires Automatic Adjustment Based on Data Quality (2024.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement Learning from Human Feedback (RLHF) is a method for aligning language models with human values.
Approach: They propose a method that automatically adjusts reward modeling based on data quality . they use preference data to train a reward model that is more aligned with human values .
Outcome: The proposed method stabilizes reward model training and significantly improves alignment performance on human preference datasets.
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing reinforcement learning methods rely on sparse outcome rewards, which fail to credit correct intermediate steps in partially successful solutions.
Approach: They propose a process reward model that rewards correct steps only when they detect errors . they propose VPPO, which rewards the correct prefix and an erroneous suffix .
Outcome: a new approach outperforms sparse-reward RL and prior PRM-guided baselines on Pass@1 and Pass@K . a process reward model (PRM) outperformed sparser-rebound RL on multiple reasoning benchmarks .
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on textual plan generation only focus on LLMs, enabling applications in robotics, virtual assistants, and instruc.
Approach: They propose a framework that generates and refines text-image plans step-by-step . they collect a new benchmark consisting of 1,100 tasks and their text- image pair solutions covering 11 daily topics.
Outcome: The proposed framework generates and refines text-image plans step-by-step and improves on existing models.
Robust Unsupervised Neural Machine Translation with Adversarial Denoising Training (2020.coling-main)

Copied to clipboard

Challenge: Unsupervised neural machine translation (UNMT) has attracted great interest in the machine translation community.
Approach: They propose to explicitly take noisy data into consideration to improve the robustness of UNMT based systems.
Outcome: The proposed methods significantly improved the robustness of the conventional UNMT systems in noisy scenarios.
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that fine-tuning with benign data can compromise safety of aligned LLMs.
Approach: They propose a Layer-Aware Representation Filtering method that detects safety-degrading layers within the LLM and leverages their representations to detect them.
Outcome: The proposed method can detect safety-degrading features in benign data and remove them from the model.
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for On-Policy LLM RL typically train a separate process reward model, which suffers from distribution mismatch and reward hacking.
Approach: They propose a reinforcement learning framework that directly incorporates on-policy tree search for RL training.
Outcome: Experiments on math and code reasoning benchmarks show that tree search achieves superior performance compared to traditional ChainRL.
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities.
Approach: They propose a comprehensive benchmark covering 29 languages, built on an English benchmark.
Outcome: The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark.
Hierarchical Reward Modeling for Fault Localization in Large Code Repositories (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limited fault localization capabilities due to limited context length.
Approach: They propose a hierarchical localization reward model to evaluate and select the most accurate fault localization candidates from the outputs of LLMs.
Outcome: The proposed model improves the final line-level localization recall by 12% on the SWE-Bench-Lite dataset.
Beyond Query Bias: Candidate-Aware Iterative Refinement for Zero-Shot Composed Image Retrieval (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to retrieve target images suffer from inherent cognitive bias due to unknown candidate distribution.
Approach: They propose a training-free framework that reframes ZS-CIR as a self-correcting process . they propose to use retrieved results as feedback to perceive the candidate distribution .
Outcome: Experiments on public benchmarks show that CoRR outperforms other SOTA methods.
Evaluating Large Language Models on Wikipedia-Style Survey Generation (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that large language models can perform well in general tasks, but their effectiveness and limitations in domainspecific tasks remain unclear.
Approach: They examine the proficiency of Large Language Models (LLMs) in generating succinct survey articles specific to the niche field of NLP in computer science.
Outcome: The LLMs perform better in generating succinct survey articles specific to the niche field of NLP in computer science, compared to human-authored surveys, but they exhibit bias in evaluation.
R-CHAR: A Metacognition-Driven Framework for Role-Playing in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing role-playing structures lack cognitive consistency in complex scenarios . Existing models excel in math and coding tasks but lack coherent reasoning .
Approach: They propose a metacognition-driven framework that enhances role-playing performance . experimental results show performance improvements across varying scenario complexities .
Outcome: The proposed framework outperforms existing models in social intelligence tasks and shows strength in long-context comprehension and group-level social interactions.
Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs)-based Role-Playing Language Agents (RPLAs) have attracted broad attention in various applications.
Approach: They propose a benchmark for evaluating character thought generation using literature . they propose 'MIRROR' which generates character thoughts by retrieving memories, predicting character reactions, and synthesizing motivations.
Outcome: The proposed benchmark outperforms existing methods in evaluating character thought generation.
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference (2024.findings-emnlp)

Copied to clipboard

Challenge: Text summarization tasks employ Pre-trained Language Models (PLMs) to fit diverse datasets.
Approach: They propose a human summarization preference alignment framework to align PLMs with human preferences.
Outcome: The proposed framework narrows the gap between automatic and human evaluations by integrating three components.
What Makes Pre-trained Language Models Better Zero-shot Learners? (2023.acl-long)

Copied to clipboard

Challenge: Current methods for prompt learning in zero-shot scenarios rely on a development set with sufficient human-annotated data to select the best-performing prompt template.
Approach: They propose a method for screening reasonable prompt templates in zero-shot text classification using language discrepancy.
Outcome: The proposed method improves prediction performance in a realistic zero-shot setting, eliminating the need for labelled examples.
CEHA: A Dataset of Conflict Events in the Horn of Africa (2025.coling-main)

Copied to clipboard

Challenge: Existing datasets categorizing conflict events do not cover all of the fine-grained types of conflict relevant to areas like the Horn of Africa.
Approach: They propose to use online news articles to categorize violent conflict events . they propose to extract event-relevance and event-types from 500 English event descriptions .
Outcome: The proposed dataset categorizes conflict risk according to specific areas required by stakeholders in the Humanitarian-Peace-Development Nexus.
MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities on general text, but their proficiency in specialized scientific domains remains uncharacterized.
Approach: They evaluate the capabilities of large language models in metabolomics research using MetaBench . they found that models perform well on text generation tasks, but cross-database identifier grounding remains challenging .
Outcome: The evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations