Papers by Weiwei Sun

42 papers
Context-DPO: Aligning Language Models for Context-Faithfulness (2025.findings-acl)

Copied to clipboard

Challenge: Context-DPO is the first alignment method specifically designed to enhance contextfaithfulness for large language models.
Approach: They propose a benchmark that simulates Retrieval-Augmented Generation scenarios with knowledge conflicts to evaluate context-faithfulness.
Outcome: The proposed method improves LLMs' context-faithfulness by 35% to 280% over open-source models.
Universal Semantic Tagging for English and Mandarin Chinese (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to generating semantic annotations for different languages are attracting more and more interest.
Approach: They propose to extend Universal Semantic Tagging to Mandarin Chinese and evaluate its performance.
Outcome: The proposed scheme is only tested in four Indo–European languages . accuracies are 92.7% and 94.6% for Chinese and English respectively .
Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to solve multi-hop question are constrained by the retriever and the noise in the retrieved documents.
Approach: They propose a framework that integrates parametric knowledge of large language models with external documents to solve a multi-hop question.
Outcome: The proposed framework is based on the parametric knowledge of LLMs and external documents to solve a multi-hop question.
MAIN: Mutual Alignment Is Necessary for instruction tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Instruction tuning has enabled large language models to achieve remarkable performance, yet its success heavily depends on the availability of high-quality instruction-response pairs.
Approach: They propose a mutual alignment framework which enforces coherence between instructions and responses through mutual constraints.
Outcome: The proposed framework generalizes well across model architectures and sizes, achieving state-of-the-art performance on LLaMA, Mistral, and Qwen models across diverse benchmarks.
HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are a promising alternative to expensive human evaluations.
Approach: They propose a framework that iteratively aligns LLM-based evaluators with human preference . they decompose a given evaluation task into finer-grained criteria .
Outcome: The proposed framework iteratively aligns LLM-based evaluators with human preference . it decomposes a given evaluation task into finer-grained criteria . the framework is efficient to train and more explainable than relying solely on prompts .
Se2: Sequential Example Selection for In-Context Learning (2024.findings-acl)

Copied to clipboard

Challenge: Prior work has explored the selection of examples for in-context learning, neglecting the internal relationships between examples and exist an inconsistency between training and inference.
Approach: They propose a sequential-aware method that leverages the LLM’s feedback on varying context, aiding in capturing inter-relationships and sequential information among examples.
Outcome: Experiments on 23 NLP tasks show that Se2 surpasses baselines and achieves 42% relative improvement over random selection.
To Copy Rather Than Memorize: A Vertical Learning Paradigm for Knowledge Graph Completion (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding knowledge graphs implicitly memorize relation rules to infer missing links, but they are difficult to memorize due to the inherent deficiencies of such implicit memorization strategy.
Approach: They propose a vertical learning paradigm that allows to explicitly copy target information from related factual triples for more accurate prediction.
Outcome: The proposed model improves generalization ability and makes distant link prediction significantly easier.
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability (2026.acl-long)

Copied to clipboard

Challenge: Existing rerankers perform poorly in complex ranking scenarios due to the scarcity of reasoning-intensive training data.
Approach: They propose an automated reasoning-intensive training framework which generates high-quality training labels from training queries and passages.
Outcome: The proposed model outperforms baselines significantly and achieves much lower latency than the pointwise reranker.
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities (2025.naacl-long)

Copied to clipboard

Challenge: Recent research shows that Large Language Models (LLMs) are vulnerable to automated jailbreak attacks.
Approach: They propose a framework that crafts adversarial LLMs with enhanced jailbreak ability.
Outcome: ADV-LLM significantly reduces the computational cost of generating adversarial suffixes while achieving nearly 100% ASR on various open-source LLMs.
RADE: Reference-Assisted Dialogue Evaluation for Open-Domain Dialogue (2023.acl-long)

Copied to clipboard

Challenge: Evaluating open-domain dialogue systems is challenging because of the one-to-many problem.
Approach: They propose a reference-based dialogue evaluation approach that leverages the pre-created utterance as reference other than the gold response to relieve the one-to-many problem.
Outcome: The proposed method outperforms state-of-the-art evaluation methods on three datasets and two existing benchmarks.
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work utilizes generative LLMs for Information Retrieval (IR) rather than direct passage ranking.
Approach: They investigate generative LLMs such as ChatGPT and GPT-4 for relevance ranking in IR and use a test set to verify the model’s ability to rank unknown knowledge.
Outcome: The proposed model outperforms a 3B supervised model on the BEIR benchmark.
GeAR: Generation Augmented Retrieval (2025.findings-acl)

Copied to clipboard

Challenge: Document retrieval techniques are used to compute semantic similarity between a query and documents, but the scalar similarity fails to reflect enough information, hindering the interpretation of retrieval results.
Approach: They propose a method which improves the global document-query similarity through contrastive learning and integrates well-designed fusion and decoding modules.
Outcome: The proposed method improves the global document-query similarity through contrastive learning and integrates well-designed fusion and decoding modules.
Democratizing Reasoning Ability: Tailored Learning from Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit impressive emergent abilities in natural language processing, but their democratization is hindered due to huge computation requirements and closed-source nature.
Approach: They propose a tailored learning approach to distill the exclusive reasoning ability to smaller LMs to facilitate democratization.
Outcome: The proposed approach enables the democratization of the exclusive reasoning ability by leveraging the black-box model as a reasoning teacher.
Graph-Based Meaning Representations: Design and Processing (P19-4)

Copied to clipboard

Challenge: This tutorial focuses on representing and processing sentence meaning in the form of labeled directed graphs.
Approach: This tutorial will briefly review relevant background in formal and linguistic semantics . it will also briefly define a unified abstract view on different flavors of semantic graphs - and associated terminology .
Outcome: The tutorial will briefly review relevant background in formal and linguistic semantics .
Auto Search Indexer for End-to-End Document Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Generative retrieval heavily relies on the “preprocessed” document identifiers, thus limiting its retrieval performance and ability to retrieve new documents.
Approach: They propose a fully end-to-end retrieval paradigm that can learn the best docids for existing and new documents automatically via a semantic indexing module.
Outcome: The proposed model outperforms baselines on public and industrial datasets and can handle new documents.
Semantic Parsing for English as a Second Language (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on domain adaptation in NLP focus on learning challenges at the syntax-semantics interface during second language acquisition.
Approach: They propose to use English Resource Grammar and TLE to parse ESL data using a reranking model to evaluate the quality of the annotations.
Outcome: The proposed model can obtain a very promising quality in comparison to human annotations.
Improving the Robustness of Large Language Models via Consistency Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Large language models have shown tremendous success in following user instructions and generating helpful responses, but their robustness is still far from optimal.
Approach: They propose a two-stage training framework that helps a model generalize on following instructions via similar instruction augmentations.
Outcome: The proposed training framework improves diversity and aligns the model with human expectations by differentiating subtle differences in similar responses.
NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement Learning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing formal proof assistants rely on instruction tuning and lack fine-grained structural and semantic alignment.
Approach: They propose a reinforcement learning framework that enables LLMs to translate natural language into formal language such as Lean 4 . they use a model with basic translation ability to refine the model's reinforcement learning .
Outcome: The proposed method outperforms baseline models on NL-to-Lean 4 tasks.
DUET: Joint Exploration of User–Item Profiles in Recommendation System (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLMs are opaque and difficult to interpret, resulting in limited interpretability.
Approach: They propose an interaction-aware profile generator that jointly produces user and item profiles conditioned on both user history and item evidence.
Outcome: The proposed model outperforms baselines on three real-world datasets.
Compositional Syntactico-SemBanking for English as a Second or Foreign Language (2025.findings-acl)

Copied to clipboard

Challenge: Despite the widespread use of English as a Second or Foreign Language (ESFL), developing syntactico-semantic representations for it is limited.
Approach: They propose a Synchronous Hyperedge Replacement Grammar-based constructivist approach to address the challenges in ESFL.
Outcome: The proposed approach bridges the gap between literal cues and intended meaning by using constructions as fundamental units.
Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for retrieving historical messages are based on similarity-based mechanisms.
Approach: They propose a system that integrates System-1 similarity search with a complementary System-2 mechanism, termed Global Selection.
Outcome: The proposed framework achieves state-of-the-art on long-term memory benchmarks and 93.9 on LoCoMo and 91.6 on LongMemEval-S.
Semantic Role Labeling for Learner Chinese: the Importance of Syntactic Parsing and L2-L1 Parallel Data (D18-1)

Copied to clipboard

Challenge: a learner language (interlanguage) is an idiolect developed by a learning of a second or foreign language.
Approach: They propose to use semantic role labeling as a case task to parse interlanguages . they then evaluate three off-the-shelf SRL systems to gauge how successful they are .
Outcome: The proposed model achieves an F-score of 72.06, a 2.02 point improvement over the baseline.
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method (2024.naacl-long)

Copied to clipboard

Challenge: Recent literature reveals that Large Language Models (LLMs) hallucinate intermittently, which impedes their reliability for further utilization.
Approach: They propose a self-detection method to detect which questions an LLM does not know by combining the two components to identify whether the model generates a non-factual response to the question.
Outcome: The proposed method can detect which questions an LLM does not know across factoid question-answering, arithmetic reasoning, and commonsense reasoning tasks.
Parsing into Variable-in-situ Logico-Semantic Graphs (2020.acl-main)

Copied to clipboard

Challenge: a new type of graph-based meaning representation allows analysis for scope-related phenomena.
Approach: They propose variable-in-situ logico-semantic graphs to bridge gap between semantic graph and logical form parsing.
Outcome: The proposed graph-based meaning representation achieves 92.39% accuracy in terms of elementary dependency match . the output of the proposed parser is highly coherent .
ResLoRA: Identity Residual Mapping in Low-Rank Adaption (2024.findings-acl)

Copied to clipboard

Challenge: Low-rank adaptation (LoRA) is one of the most popular parameter-efficient fine-tuning methods.
Approach: They propose a low-rank adaptation method that adds residual paths during training and merges them together during inference to achieve better results.
Outcome: The proposed method achieves 2.5x faster convergence speed and improves performance by 14.3% on NLG, NLU, and text-to-image tasks.
Language Generation via DAG Transduction (P18-1)

Copied to clipboard

Challenge: Existing formal frameworks for graph manipulation are underexploited.
Approach: They propose a DAG transducer to perform graph-to-program transformation using a declarative programming language.
Outcome: The proposed transducer achieves a BLEU-4 score of 68.07 for natural language generation from type-logical semantic graphs.
Generative Knowledge Selection for Knowledge-Grounded Dialogues (2023.findings-eacl)

Copied to clipboard

Challenge: Knowledge selection is the key in knowledge-grounded dialogues (KGD), which aims to select an appropriate knowledge snippet to be used in the utterance based on dialogue history.
Approach: They propose a generative approach for knowledge selection called GenKS that learns to select snippets by generating their identifiers with a sequence-to-sequence model.
Outcome: The proposed approach captures intra-knowledge interaction inherently through attention mechanisms while generating their identifiers with a sequence-to-sequence model.
Token-level Proximal Policy Optimization for Query Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved search engines and recommendation systems through their text understanding capabilities.
Approach: They propose a token-level proximal policy optimization approach to empower LLMs to perform better in query generation through fine-tuning.
Outcome: The proposed approach outperforms existing LLMs on an open-source and industrial dataset.
Pre- and In-Parsing Models for Neural Empty Category Detection (P18-1)

Copied to clipboard

Challenge: Existing studies on empty category detection have shown positive effects on syntactic parsing . empty categories are used to indicate long-distance dependencies, discontinuous constituents, and certain dropped elements.
Approach: They propose to use ECD to detect empty categories without syntactic analysis.
Outcome: The proposed models outperform the prior state-of-the-art by significant margins.
How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on enhancing the factualness of large language models using context knowledge.
Approach: They propose to use ChatGPT to construct probing datasets that provide diverse and coherent evidence corresponding to various facts.
Outcome: The proposed model can encode knowledge across different layers, and it is compared with existing models.
Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context Compression (2026.acl-long)

Copied to clipboard

Challenge: Existing reinforcement learning pipelines suffer from degraded instruction following, excessive rollout costs, and strict context limits.
Approach: They propose a reinforcement learning (RL) fine-tuning of large language model (LLM) agents for long-horizon multi-turn tool use where context length quickly becomes a bottleneck.
Outcome: The proposed framework improves the success rate while maintaining the same or even lower working context length compared to baselines.
Coding Textual Inputs Boosts the Accuracy of Neural Networks (2020.emnlp-main)

Copied to clipboard

Challenge: a new approach to natural language processing uses arbitrary symbols to represent meaning . Soundex, MetaPhone, NYSIIS, logogram are used as inputs for NLP .
Approach: They propose to use arbitrary symbols to represent linguistic meaning of a word . they propose to integrate codewords with text to provide more reliable inputs .
Outcome: The proposed approach outperforms state-of-the-art models on machine translation, language modeling, and part-of speech tagging.
DiQAD: A Benchmark Dataset for Open-domain Dialogue Quality Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on dialogue quality assessment are uncapable of providing an end-to-end and human-epistemic assessment dataset . open-domain dialogue assessment is complicated and costly, but it can be done by recruiting human evaluators.
Approach: They propose a large-scale dialogue quality assessment dataset for automatically assessing open-domain dialogue quality.
Outcome: The proposed dataset is openly accessible at https://github.com/yukunZhao/Dialogue_quality_evaluation.
Accurate SHRG-Based Semantic Parsing (P18-1)

Copied to clipboard

Challenge: Graph-structured semantic representations can encode rich semantic information of natural language sentences.
Approach: They propose a SHRG-based parser that relates synchronous production rules to syntacto-semantic composition processes.
Outcome: The proposed model improves on the best existing model by 4.87 points . it relates synchronous production rules to syntacto-semantic composition process .
A Computational Simulation of Language Production in First Language Acquisition (2025.emnlp-main)

Copied to clipboard

Challenge: Existing computational studies of child language acquisition focus on isolated mechanisms, such as spreading activation in retrieval, sentence planning, or production efficiency.
Approach: They propose a computational framework for modeling child language production using graphs to formalize meaning and Synchronous Hyperedge Replacement Grammar to formalized the syntax–semantics interface.
Outcome: The proposed framework is based on graphs to formalize meaning and Synchronous Hyperedge Replacement Grammar (SHRG) resulting interpretable grammars are evaluated by their ability to generate utterances .
MEFT: Memory-Efficient Fine-Tuning through Sparse Adapter (2024.acl-long)

Copied to clipboard

Challenge: Parameter-Efficient Fine-tuning (PEFT) methods are limited on knowledge-intensive tasks due to the limited number of trainable parameters.
Approach: They propose a mechanism that fine-tunes Large Language Models with larger adapters . they store and update the parameters of larger adapter adapters on the CPU .
Outcome: The proposed method achieves comparable results to those obtained with larger memory capacities over the limited bandwidth of PCI Express (PCIe).
Answering Ambiguous Questions via Iterative Prompting (2023.acl-long)

Copied to clipboard

Challenge: Empirical studies show that AmbigPrompt achieves state-of-the-art or competitive results while using less memory and having a lower inference latency than competing approaches.
Approach: They propose an answering model with a prompting model to address imperfections in open-domain question answering . Empirical studies show AmbigPrompt achieves state-of-the-art or competitive results .
Outcome: The proposed framework improves on two commonly-used open benchmarks and achieves state-of-the-art or competitive results while using less memory and having a lower inference latency.
Exact yet Efficient Graph Parsing, Bi-directional Locality and the Constructivist Hypothesis (2020.acl-main)

Copied to clipboard

Challenge: Existing algorithms for graph parsing are exponential or high-degree polynomial w.r.t. grammars, and there are few systems that can parse large but frequent MRs with a realistic, wide-coverage grammar in a reasonable time.
Approach: They propose an exact graph parsing algorithm that exploits locality as terminal edge-adjacency in HRG rules and categorizes a subclass of HRG.
Outcome: The proposed method can parse graphs with a (competence) grammar in a time-efficient manner.
Calibrating LLM-Based Evaluator (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for large language models lack the ability to calibrate their outputs towards human preference.
Approach: They propose a multi-stage, gradient-free approach to calibrate an LLM-based evaluator toward human preference.
Outcome: The proposed approach improves correlation with expert evaluation on multiple text quality evaluation datasets.
UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have impressive capabilities but need for task-specific prompt engineering can hinder their generalization.
Approach: They propose a lightweight and versatile retriever that automatically retrieves prompts for a given zero-shot task input.
Outcome: The proposed model is universally applicable across tasks and models . it mitigates hallucination problem in chatGPT, and it improves even the strongest LLMs.
Alleviating Performance Degradation Caused by Out-of-Distribution Issues in Embedding-Based Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies reveal query out-of-distribution issues degrading ANN performance . a distribution regularizer is introduced into the encoder training objective to encourage alignment between query and base embeddings.
Approach: They introduce a distribution regularizer into the encoder training objective to encourage alignment between query and base embeddings.
Outcome: The proposed method consistently improves retrieval performance across multiple datasets.
MAIR: A Massive Benchmark for Evaluating Instructed Retrieval (2024.emnlp-main)

Copied to clipboard

Challenge: Existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models.
Approach: They propose a multi-task instruction-tuned IR benchmark that includes 126 distinct IR tasks across 6 domains.
Outcome: The proposed model performs better on instruction-tuned models than non-instruction-tunned models on MAIR.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations