Papers by Wei Jing

57 papers
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models have created significant safety concerns . factuality ability is crucial in determining whether they can be deployed and applied safely and compliantly within specific regions.
Approach: They propose a benchmark to evaluate the factuality of large language models in China . they evaluate the models' ability to provide accurate and reliable information .
Outcome: The proposed benchmark evaluates the factuality abilities of existing LLMs and compares them to LLM abilities.
WavLLM: Towards Robust and Adaptive Speech Large Language Model (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have expanded their scope to encompass multimodal functions.
Approach: They propose a robust and adaptive speech large language model with dual encoders . they validate the model on universal speech benchmarks and apply it to specialized speech-question-answer datasets based on a CoT approach .
Outcome: The proposed model achieves state-of-the-art performance across a range of speech tasks on the same model size.
Reinforcement Tuning for Detecting Stances and Debunking Rumors Jointly with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Social media has become a fertile ground for nurturing rumors and misinformation due to its lack of systematic moderation.
Approach: They propose a framework to enhance the joint predictive capabilities of LLMs for stance detection and rumor verification tasks.
Outcome: The proposed framework outperforms state-of-the-art methods and generalizes to non-LLMs accommodated as task models.
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations (2026.acl-long)

Copied to clipboard

Challenge: Previous work shows that large language models generate hallucinations, yet the origins and mechanisms of these signals remain unclear.
Approach: They propose to validate and disentangle two different pathways for truthfulness cues . they also propose to use the same mechanism to derive self-contained evidence from the generated answer .
Outcome: The proposed applications improve hallucination detection performance by integrating two different inputs.
USB: A COMPREHENSIVE AND UNIFIED SAFETY EVALUATION BENCHMARK FOR MULTIMODAL LARGE LANGUAGE MODELS (2026.acl-long)

Copied to clipboard

Challenge: Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations.
Approach: They propose a framework that covers 61 risk categories across four modality interactions to address this gap.
Outcome: The proposed framework covers 61 risk categories across four distinct modality interactions.
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing (2024.emnlp-main)

Copied to clipboard

Challenge: a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities .
Approach: They present a comparative analysis to identify and distinguish LLM activities from human activities.
Outcome: The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities.
Adaptive Policy with Wait-k Model for Simultaneous Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to simultaneous machine translation require a robust read/write policy . a standalone multi-path wait-k model performs competitively with adaptive policies .
Approach: They propose a more flexible approach by decoupling the adaptive policy model from the translation model.
Outcome: The proposed approach outperforms baseline approaches in translation tasks.
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending (2024.acl-long)

Copied to clipboard

Challenge: Existing models that use self-attention and position embedding have anomalous behavior that hinder long context window extrapolation.
Approach: They propose a collinear constraint between Q and K to integrate RoPE and self-attention.
Outcome: The proposed model integrates self-attention and position embedding into LLMs without fine-tuning.
Enhancing Self-Attention with Knowledge-Assisted Attention Maps (2022.naacl-main)

Copied to clipboard

Challenge: Existing works of knowledge infusion depend on multi-task learning frameworks, which are inefficient and require large-scale retraining when new knowledge is considered.
Approach: They propose a method which integrates knowledge-generated attention maps into the self-attention mechanism and integrates it into the model.
Outcome: The proposed model outperforms existing methods on academic datasets and industry-scale ad relevance applications.
The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for difficulty estimation rely on repeated response sampling, auxiliary models, or fine-tuning the target model itself.
Approach: They propose a method that leverages only the hidden representations produced by large language models.
Outcome: The proposed method outperforms baselines in difficulty estimation on textual and multimodal tasks and improves adaptive reasoning strategies with fewer generated tokens.
Span-based Localizing Network for Natural Language Video Localization (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to NLVL are either ranking tasks or regressing the target video span.
Approach: They propose a video span localizing network to solve a natural language video localization task using a span-based QA approach.
Outcome: The proposed network outperforms the state-of-the-art methods on three benchmark datasets.
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are evolving towards autonomous agents . retrieval capabilities are well-benchmarked, but post-retrieval synthesis is under-evaluated due to open-ended writing.
Approach: They propose a benchmark to evaluate information consolidation capabilities using survey papers as gold standards.
Outcome: The proposed benchmark analyzes the post-retrieval synthesis stage of large language models . it leverages high-quality survey papers as gold standards and reverse-engineers research requests . the proposed benchmark outperforms single-turn generation and reduces hallucinations .
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing models generate explanations that appear coherent while containing unfaithful intermediate steps.
Approach: They propose a causality-inspired framework for evaluating CoT quality using controlled perturbations as an instrumental signal to separate genuine step-to-step dependence from bias-driven artifacts.
Outcome: Experiments on GSM8K, MATH, and CommonsenseQA show that FACT-E improves reasoning-trajectory selection and yields stronger in-context learning exemplars.
AgentV-RL: Scaling Reward Modeling with Agentic Verifier (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to improve LLM reasoning are limited in complex domains and lack external grounding makes verifiers unreliable on computation-intensive tasks.
Approach: They propose a framework that transforms reward modeling into a multi-turn, tool-augmented deliberative process.
Outcome: The proposed framework surpasses state-of-the-art ORMs by 25.2% under parallel and sequential TTS.
Federated LoRA Fine-Tuning with Pipelined Error-Mitigated Aggregation and Matrix-Wise Freezing (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models often suffer from biased model aggregation and are hindered by significant communication and computation burden.
Approach: They propose a Federated low-rank adaptation system for large language models that leverages pipelined error-mitigated model aggregation and adaptive matrix-wise parameter freezing to mitigate aggregations.
Outcome: The proposed system improves time-to-target by 2.17-8.48 on real-world datasets.
FAITH: Factuality Alignment through Integrating Trustworthiness and Honestness (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to correct factually inaccurate outputs are lacking the semantic richness needed to properly understand its internal states of trustworthiness and honesty.
Approach: They propose a framework for factuality alignment that integrates natural-language uncertainty signals with external knowledge and computes confidence scores and semantic entropy from LLM outputs.
Outcome: Extensive experiments on four knowledge-intensive benchmarks show that FAITH improves the factual accuracy and truthfulness of Large Language Models (LLMs).
Better Simultaneous Translation with Monotonic Knowledge Distillation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to train offline MT models require generating target tokens before source sentence is fully consumed.
Approach: They propose a method that leverages traditional translation models as teachers to generate monotonic yet accurate reference translations for sequence-level knowledge distillation.
Outcome: The proposed approach improves on strong baselines and on a monotonic version of the WMT15 De-En test set.
Leveraging Web-Crawled Data for High-Quality Fine-Tuning (2024.findings-emnlp)

Copied to clipboard

Challenge: Currently, large language models are fine-tuned using expensive human-annotated data or GPT-4 generated data.
Approach: They propose to use web-crawled data to train a language model on a smaller set of data . their results show that the model can convert web data with irregular formats into high-quality ones .
Outcome: The proposed model outperforms open-source models larger than 32B and outperformed open-sourced models such as GPT-3.5.
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Unlike chatbots, autonomous agents act directly on external environments, making tool invocation safety critical for reliable deployment.
Approach: They develop a benchmark for step-level tool invocation safety detection in LLM agents and a guardrail model that proactively detects unsafe tool invoking actions before execution using multi-task reinforcement learning.
Outcome: The proposed model reduces harmful tool invocations of ReAct-style agents by 65% on average and improves benign task completion by 10% under prompt injection attacks.
Modeling Evolution of Message Interaction for Rumor Resolution (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for rumor resolution ignore local interactions during the message diffusion which is important for the identification of rumors.
Approach: They propose to model confrontation and reciprocity between message pairs via discrete variational autoencoders which effectively reflects the diversified opinion interactivity.
Outcome: Experiments on a PHEME dataset show that the proposed model achieves higher accuracy than existing methods.
Enhancing the Reasoning Capabilities of Small Language Models via Solution Guidance Fine-Tuning (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks.
Approach: They propose a new reasoning strategy Solution Guidance (SG) and a plug-and-play training paradigm Solution-Guidance Fine-Tuning (SGFT) which focuses on problem understanding and decomposition at the semantic and logical levels, rather than specific computations.
Outcome: The proposed reasoning strategy Solution Guidance (SG) and plug-and-play training paradigm Solution-Guidance Fine-Tuning (SGFT) improves the reasoning capabilities of small language models on various reasoning tasks.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
WSDMS: Debunk Fake News via Weakly Supervised Detection of Misinforming Sentences with Contextualized Social Wisdom (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for debunking fake news rely on blending of authentic and fabricated content by creators.
Approach: They propose a model that detects misinformation at sentence-level using social media conversations . they use a bag-level annotation system to train the model .
Outcome: The proposed model outperforms existing state-of-the-art models on three real-world benchmarks and outperformed existing state of the art models in debunking fake news at sentence and article levels.
Bi-directional CognitiveThinking Network for Machine Reading Comprehension (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for reading comprehension are still in their infancy at the level of cognitive intelligence.
Approach: They propose a bi-directional cognitive knowledge framework to simulate reverse thinking and inertial thinking in the brain to answer questions.
Outcome: The proposed framework shows that bi-directional knowledge helps the QA task.
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on pre-trained LLMs to better understand and improve their trustworthiness.
Approach: They apply linear probing to LLMs to explore five key dimensions of trustworthiness: reliability, privacy, toxicity, fairness, and robustness.
Outcome: The proposed model can distinguish concepts in each trustworthiness dimension, suggesting that it can be trained in early pre-training.
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents (2025.acl-long)

Copied to clipboard

Challenge: Existing models that use Large Language Models (LLMs) show superior performance in various tasks, but lack of controllability leads to unfocused conversations or task failure.
Approach: They propose a standard operating procedure (SOP) framework to regulate dialogue flow by integrating Chain of Thought reasoning and supervised fine-tuning for SOP prediction.
Outcome: The proposed method achieves a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also shows notable gains for open-source models.
MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline (2024.findings-acl)

Copied to clipboard

Challenge: Large language models lack mathematical reasoning, a hurdle on the path to true artificial general intelligence.
Approach: They propose a protocol for fine-tuning large language models with a Python code interpreter to enhance the text analysis of the LLMs.
Outcome: The proposed protocol improves the performance of a 7B-parameter LLM on the GSM8K and MATH datasets while allowing for an outlier-free value model-based inference method.
Sentence-Level Evidence Embedding for Claim Verification with Hierarchical Attention Networks (P19-1)

Copied to clipboard

Challenge: Claim verification is cumbersome and inefficient for human fact-checkers to find consistent pieces of evidence.
Approach: They propose an end-to-end hierarchical attention network that learns to represent coherent evidence and their semantic relatedness with the claim.
Outcome: The proposed model outperforms state-of-the-art models on three datasets . it is based on a coherence-based attention layer and entailment-based one .
Dialectical Structured Reasoning for Explainable Multimodal Fake News Detection (2026.findings-acl)

Copied to clipboard

Challenge: Existing fake news detection models are opaque and lack deductive transparency . a framework for dialectical structured reasoning is proposed to address this limitation .
Approach: They propose a framework that model fake news detection as an explicit dialectical process over multimodal social context.
Outcome: The proposed framework achieves state-of-the-art while producing transparent explanations that mirror human reasoning process.
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for automated fact-checking often overlook deceptive misinformation styles in generated explanations.
Approach: They propose a framework that explicitly controls reasoning style by anchoring explanations to the predicted verdict.
Outcome: The proposed framework achieves state-of-the-art under LLaMA-series models with 465 samples.
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF).
Approach: a new study proposes a domain-informed self-consistency policy optimization extension to GRPO that addresses inter-group imbalance.
Outcome: a new extension of GRPO addresses inter-group imbalance with two key innovations . the proposed method outperforms existing GR PO variants by 5% on Qwen3 models .
Improve Speech Translation Through Text Rewrite (2025.coling-industry)

Copied to clipboard

Challenge: Recent advances in speech translation (ST) research have focused on the unique characteristics of spontaneous speech, including accents and presentation quality.
Approach: They propose to transform transcribed speech into a cleaner style more in line with the expectations of translation models built from written text.
Outcome: Experiments on public and in-house translation models show that the proposed model can be effectively distilled into a standalone translation model.
Towards Universal Debiasing for Language Models-based Tabular Data Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing large language models have exacerbated fairness issues in tabular data generation . inherent historical biases in tabulated data cause LLMs to exacerbate fairness problems .
Approach: They propose a universal debiasing framework that minimizes group-level dependencies . it leverages the autoregressive structure and analytic sampling distributions of LLM-based tabular data generators .
Outcome: The proposed framework minimizes group-level dependencies while reducing mutual information between advantaged and protected attributes.
Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework (2026.acl-industry)

Copied to clipboard

Challenge: a production-grade pricing system for tourism is challenging due to unstructured nature of travel orders and ever-evolving pricing policies.
Approach: They propose a production-grade pricing system with a strict decision boundary . they propose to combine structured extraction and bounded policy/path selection with interpretable condition trees .
Outcome: The proposed system processed 3,960 orders in six months and reduced the order management team from 15-20 to 3 . the system reduced the per-order handling time from 10 minutes to 2 minutes.
Rumor Detection on Twitter with Tree-structured Recursive Neural Networks (P18-1)

Copied to clipboard

Challenge: Existing methods for detecting rumors are difficult to implement and require a lot of effort.
Approach: They propose two recursive neural models that follow tweets' propagation layouts to learn discriminative features from tweets and generate more powerful representations for rumors detection.
Outcome: The proposed models perform better than state-of-the-art approaches on two public Twitter datasets and show superior performance on detecting rumors at very early stage.
Answer-focused and Position-aware Neural Question Generation (D18-1)

Copied to clipboard

Challenge: Recent neural network-based approaches generate interrogative words that do not match the answer type.
Approach: They propose an answer-focused and position-aware neural question generation model to address these issues.
Outcome: The proposed model outperforms the baseline and outperformed the state-of-the-art system.
Controllable Natural Language Generation with Contrastive Prefixes (2022.findings-acl)

Copied to clipboard

Challenge: Existing work on controllable natural language generation has focused on fine-tuning existing models or using attribute discriminators.
Approach: They propose a lightweight framework for controllable GPT2 generation that utilizes attribute-specific vectors to steer natural language generation.
Outcome: The proposed framework can guide generation towards desired attributes while keeping high linguistic quality.
Discrete Argument Representation Learning for Interactive Argument Pair Identification (2021.naacl-main)

Copied to clipboard

Challenge: Existing research on monological argumentation covers claims generation, argument structure prediction, and essay scoring.
Approach: They propose to identify argument pairs from two posts with opposite stances to a certain topic.
Outcome: The proposed framework outperforms competing models on a large-scale dataset . it also proves that it is useful for analyzing argument pairs from two posts .
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity (2025.findings-emnlp)

Copied to clipboard

Challenge: Using real-world datasets, we conduct the most comprehensive study to date, auditing various state-of-the-art reward models across nine sensitive attributes, including age, gender, ethnicity, etc.
Approach: They propose a method to mitigate group disparities in reward modeling by using real-world data.
Outcome: The proposed method is based on a population-based dataset with nine demographic attributes, including gender, ethnicity, age, gender, and ethnicity.
Parallel Attention Network with Sequence Matching for Video Grounding (2021.findings-acl)

Copied to clipboard

Challenge: Existing approaches to video grounding are sensitive to quality of proposals and inefficient because all proposal-query pairs are compared.
Approach: They propose a Parallel Attention Network with Sequence matching to capture selfmodal contexts and cross-modal attentive information between video and text.
Outcome: The proposed approach is superior to state-of-the-art methods on three datasets.
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that fine-tuning with benign data can compromise safety of aligned LLMs.
Approach: They propose a Layer-Aware Representation Filtering method that detects safety-degrading layers within the LLM and leverages their representations to detect them.
Outcome: The proposed method can detect safety-degrading features in benign data and remove them from the model.
LLM Inductive Reasoning Through Multi-Agent Enhanced Monte Carlo Tree Search (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for enhancing inductive reasoning of large language models often lack explicit optimization guidance and effective error correction.
Approach: They propose a plug-and-play test-time framework that integrates multi-agent coordination with Monte Carlo Tree Search to improve inductive reasoning.
Outcome: The proposed framework outperforms existing methods on four benchmarks and shows consistent improvements on QWQ-32B and Deepseek-V3 .
FrontCoder: Scaling Visual Fidelity in Front-End Code Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing work on front-end code generation fails to provide visual fidelity and rendering quality for front- end developers.
Approach: They propose a three-stage pipeline to enhance front-end code generation capabilities in LLMs . they use synthetic data, quality-controlled supervised fine-tuning, and reinforcement learning .
Outcome: The proposed model achieves competitive performance with frontier models while maintaining generation efficiency.
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models (2026.acl-industry)

Copied to clipboard

Challenge: Small language models (SLMs) are promising for real-world deployment but struggle with high-stakes legal reasoning tasks.
Approach: They propose a diagnostic-driven synthesis framework that extracts and refines reasoning trajectories from a capable teacher via fine-grained prompting and a self-reflective verification is employed to adaptively select the most effective data for the SLM student.
Outcome: The proposed framework extracts and refines reasoning trajectories from a capable teacher via fine-grained prompting, then a self-reflective verification is employed to adaptively select the most effective data for the student.
Multi-Passage Machine Reading Comprehension with Cross-Passage Answer Verification (P18-1)

Copied to clipboard

Challenge: Recent years have seen rapid growth in the MRC community . MRC is believed to be a crucial step in building a general intelligent agent .
Approach: They propose an end-to-end neural model that enables multiple passages to verify each other based on their content representations.
Outcome: The proposed model outperforms the baseline on the English MS-MARCO dataset and the Chinese DuReader dataset, and achieves state-of-the-art performance on both datasets.
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for storytelling lack coherence and consistency, compromising the overall storytelling experience.
Approach: They propose a novel approach that improves the coherence and consistency of automatically generated stories by managing plot nodes and enabling dynamic interactions between different parts of the story.
Outcome: The proposed approach outperforms existing methods in 84.33% of the trials.
Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation (2025.coling-main)

Copied to clipboard

Challenge: Current named entity recognition methods struggle with text-image mismatch problem due to a lack of visual context.
Approach: They propose an adaptive mixup image augmentation method that generates augmented images based on matching score between text and image .
Outcome: The proposed method can be integrated into existing models and demonstrate consistent performance improvements.
Multiplex Graph Neural Network for Extractive Text Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for extractive text summarization do not consider multiple types of inter-sentential relationships, nor model intra-sententential relationships.
Approach: They propose a novel method to combine different types of relationships among sentences and words to model sentence embedding.
Outcome: The proposed model is compared with existing methods on CNN/DailyMail benchmark dataset to demonstrate its effectiveness.
Debunking Rumors on Twitter with Tree Transformer (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for rumor detection follow tree edges or treat all posts fully-connected during feature learning.
Approach: They propose a new rumor detection model based on tree transformer to better utilize user interactions in the dialogue . they propose to use post-level self-attention to aggregate the intra-/inter-subtree stances .
Outcome: The proposed model improves rumor detection performance on social media conversations . it is based on a conversation tree that encodes important information indicative of credibility .
Supervised Gradual Machine Learning for Aspect-Term Sentiment Analysis (2023.tacl-1)

Copied to clipboard

Challenge: Recent work shows that Aspect-Term Sentiment Analysis (ATSA) can be performed by Gradual Machine Learning (GML) but the current unsupervised solution is limited by inaccurate knowledge conveyance.
Approach: They propose a supervised approach which leverages binary polarity relations between instances to enable supervised knowledge conveyance.
Outcome: The proposed approach outperforms pure DNN solutions on real benchmark data.
KnowLA: Enhancing Parameter-efficient Finetuning with Knowledgeable Adaptation (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for parameter-efficient finetuning (PEFT) are limited and only finetune a small number of parameters using limited instruction data.
Approach: They propose a method that inserts an adaptation layer into an LLM to integrate embeddings of entities appearing in the input text.
Outcome: The proposed method can activate parameterized knowledge in an LLM without changing its parameters or input prompts.
Multi-Stage Balanced Distillation: Addressing Long-Tail Challenges in Sequence-Level Knowledge Distillation (2024.findings-emnlp)

Copied to clipboard

Challenge: Knowledge distillation (KD) is a promising solution for large language models, but their deployment remains computationally expensive.
Approach: They propose a framework which iteratively balances training data within a fixed computational budget and enables the transfer of knowledge from expensive teacher LLMs to smaller student models.
Outcome: The proposed framework achieves state-of-the-art performance across diverse long-tailed datasets, enhancing both the efficiency and efficacy of the distilled models.
Multi-Grained Knowledge Distillation for Named Entity Recognition (2021.naacl-main)

Copied to clipboard

Challenge: Pre-trained big models have delivered top performance in Seq2seq modeling, but their deployments in real-world applications are often hindered by excessive computations and memory demands.
Approach: They propose a distillation scheme to efficiently transfer knowledge from big models to their cheaper counterparts.
Outcome: The proposed scheme maximizes the assimilation of knowledge from the teacher model to the student model.
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to reduce hallucinations in large language models are inaccurate and inaccuracies in the generated feedback.
Approach: They propose a method that helps LLMs determine whether to utilize multiple generated feedback responses and how to identify the most useful ones.
Outcome: Extensive experiments show that the proposed method outperforms baselines on encyclopedic and commonsense knowledge QA tasks.
Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show superior cross-model transferability . Prompt injection attacks are among the most critical threats .
Approach: They propose an activations-guided prompt injection attack framework to address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box approaches.
Outcome: The proposed framework achieves 49.6% success rate and 34.6% improvement over human-crafted prompts on five mainstream LLMs.
MASTER: A Multi-Agent System with LLM Specialized MCTS (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being explored for problem-solving tasks . their strategic planning capability is often viewed with skepticism due to their limited planning capabilities.
Approach: They propose a framework that coordinates agent recruitment and communication through LLM specialized MCTS.
Outcome: The proposed framework achieves 76% accuracy on HotpotQA and 80% on WebShop . it relies on extensive sampling simulations to approximate the true reward distribution .
The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to assess instruction quality require trained raters to observe classrooms based on established criteria.
Approach: They propose to use Natural Language Processing techniques to assess multiple high-inference instructional practices in in-person K-12 classrooms and simulated performance tasks for pre-service teachers.
Outcome: The proposed method is able to assess multiple high-inference instructional practices in two educational settings: in-person K-12 classrooms and simulated performance tasks for pre-service teachers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations