Papers with LS

13 papers
Chinese Lexical Substitution: Dataset and Method (2023.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for lexical substitution (LS) are limited and limited in coverage . despite extensive research on Lexical Substitution in various languages, there is limited evidence for LS in Chinese.
Approach: They propose to use human and machine collaboration to construct a Chinese LS dataset . they combine four unsupervised LS methods to generate candidate substitutes .
Outcome: The proposed method outperforms existing benchmarks on the Chinese lexical substitution task.
Personalizing Lexical Simplification (C18-1)

Copied to clipboard

Challenge: Experimental results show that even a simple personalized CWI model can help the system avoid some unnecessary simplifications and produce more readable output.
Approach: They evaluate the performance of a state-of-the-art LS system on individual learners of English at different proficiency levels and measure the benefits of using complex word identification models to personalize the system.
Outcome: The proposed system produces a more readable output for learners with special needs and those with language disabilities.
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to boost performance in multilingual models but scalability is difficult to manage.
Approach: They propose a method that incorporates language-specific (LS) modules to boost model performance.
Outcome: The proposed method outperforms state-of-the-art methods while outperforming existing methods.
An LLM-Enhanced Adversarial Editing System for Lexical Simplification (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to simplify text rely heavily on annotated data, making it challenging to apply in low-resource scenarios.
Approach: They propose a Lexical Simplification method without parallel corpora that uses an Adversarial Editing System and an LLM-enhanced loss to distill knowledge into a small-size LS system.
Outcome: The proposed method uses an LLM-enhanced loss to distill knowledge from Large Language Models (LLMs) into a small-size LS system.
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on LLM confidence estimations in languages other than English have been limited to English.
Approach: They propose to use question-related language to prompt LLMs to assess their confidence in large language models.
Outcome: The proposed model improves on question-related language prompts for LS tasks, while English exhibits notable linguistic dominance in confidence estimations.
Label Smoothing for Text Mining (2022.coling-1)

Copied to clipboard

Challenge: Existing text mining models are trained with 0-1 hard label that indicates whether an instance belongs to a class, ignoring rich information of the relevance degree.
Approach: They propose a keyword-based method to automatically generate soft labels from hard labels . they exploit relevance between labels and instances to incorporate them into models .
Outcome: The proposed method improves models under balanced and unbalanced conditions.
ParaLS: Lexical Substitution via Pretrained Paraphraser (2023.acl-long)

Copied to clipboard

Challenge: Lexical substitution (LS) is an extremely powerful technology that can be used as a backbone of various NLP applications such as writing assistance.
Approach: They propose two simple decoding strategies that focus on the variations of the target word during decoding to generate substitutes from a paraphraser.
Outcome: The proposed methods outperform state-of-the-art LS methods based on pre-trained language models on three benchmarks.
Knowledge Distillation ≈ Label Smoothing: Fact or Fallacy? (2023.emnlp-main)

Copied to clipboard

Challenge: Knowledge distillation (KD) is a method for knowledge transfer from one model to another . recent studies suggest it is based on label smoothing, but it is not .
Approach: They propose to compare the predictive confidences of models trained with knowledge distillation . they propose to use a method that is similar to label smoothing to train models .
Outcome: Experiments on four text classification tasks show that knowledge distillation and label smoothing drive model confidence in opposite directions.
ReLearn: Unlearning via Learning for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for unlearning large language models often rely on reverse optimization to reduce target token probabilities.
Approach: They propose a data augmentation and fine-tuning pipeline for effective unlearning . they propose augmentation, evaluation frameworks to measure contextual forgetting .
Outcome: The proposed framework achieves targeted forgetting while preserving high-quality outputs.
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for regularizing deep neural networks rely on weight decay, dropout, batch/layer normalization to converge faster and generalize.
Approach: They propose a framework for training with label regularization which includes conventional LS but can also model instance-specific variants.
Outcome: The proposed approach consistently yields better results than conventional regularization on seven machine translation and three image classification tasks while maintaining training efficiency.
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification (2022.coling-1)

Copied to clipboard

Challenge: Lexical simplification (LS) is the task of replacing complex words with simpler alternatives to make texts more accessible to various target populations.
Approach: They propose to use a Brazilian Portuguese multi-candidate dataset to test LS systems.
Outcome: The proposed model outperforms existing models on Brazilian Portuguese and Brazilian newspaper articles.
DYNTEXT: Semantic-Aware Dynamic Text Sanitization for Privacy-Preserving LLM Inference (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to protect privacy of sensitive data are differential privacy (DP) and DP is used to protect users from privacy leakage.
Approach: They propose an LDP-based Dynamic Text sanitization for privacy-preserving LLM inference that dynamically constructs semantic-aware adjacency lists of sensitive tokens to sample non-sensitive tokens for perturbation.
Outcome: The proposed model excels on three datasets.
RALS: Resources and Baselines for Romanian Automatic Lexical Simplification (2025.emnlp-main)

Copied to clipboard

Challenge: Text simplification is the process of transforming texts into variants that are simpler to understand by larger audiences or easier to process by existing NLP systems.
Approach: They propose a method for ordering simplification suggestions using a pairwise ranking approximation method, arranging candidates from simple to complex based on a separate set of human judgments.
Outcome: The proposed system is the first to combine lexical simplification and complexity prediction in Romanian with human lexicals.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations