Papers with LS
Chinese Lexical Substitution: Dataset and Method (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks for lexical substitution (LS) are limited and limited in coverage . despite extensive research on Lexical Substitution in various languages, there is limited evidence for LS in Chinese. |
| Approach: | They propose to use human and machine collaboration to construct a Chinese LS dataset . they combine four unsupervised LS methods to generate candidate substitutes . |
| Outcome: | The proposed method outperforms existing benchmarks on the Chinese lexical substitution task. |
Personalizing Lexical Simplification (C18-1)
Copied to clipboard
| Challenge: | Experimental results show that even a simple personalized CWI model can help the system avoid some unnecessary simplifications and produce more readable output. |
| Approach: | They evaluate the performance of a state-of-the-art LS system on individual learners of English at different proficiency levels and measure the benefits of using complex word identification models to personalize the system. |
| Outcome: | The proposed system produces a more readable output for learners with special needs and those with language disabilities. |
Condensing Multilingual Knowledge with Lightweight Language-Specific Modules (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to boost performance in multilingual models but scalability is difficult to manage. |
| Approach: | They propose a method that incorporates language-specific (LS) modules to boost model performance. |
| Outcome: | The proposed method outperforms state-of-the-art methods while outperforming existing methods. |
An LLM-Enhanced Adversarial Editing System for Lexical Simplification (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to simplify text rely heavily on annotated data, making it challenging to apply in low-resource scenarios. |
| Approach: | They propose a Lexical Simplification method without parallel corpora that uses an Adversarial Editing System and an LLM-enhanced loss to distill knowledge into a small-size LS system. |
| Outcome: | The proposed method uses an LLM-enhanced loss to distill knowledge from Large Language Models (LLMs) into a small-size LS system. |
MlingConf: A Comprehensive Study of Multilingual Confidence Estimation on Large Language Models (2025.findings-acl)
Copied to clipboard
Boyang Xue, Hongru Wang, Rui Wang, Sheng Wang, Zezhong Wang, Yiming Du, Bin Liang, Wenxuan Zhang, Kam-Fai Wong
| Challenge: | Existing studies on LLM confidence estimations in languages other than English have been limited to English. |
| Approach: | They propose to use question-related language to prompt LLMs to assess their confidence in large language models. |
| Outcome: | The proposed model improves on question-related language prompts for LS tasks, while English exhibits notable linguistic dominance in confidence estimations. |
Label Smoothing for Text Mining (2022.coling-1)
Copied to clipboard
| Challenge: | Existing text mining models are trained with 0-1 hard label that indicates whether an instance belongs to a class, ignoring rich information of the relevance degree. |
| Approach: | They propose a keyword-based method to automatically generate soft labels from hard labels . they exploit relevance between labels and instances to incorporate them into models . |
| Outcome: | The proposed method improves models under balanced and unbalanced conditions. |
ParaLS: Lexical Substitution via Pretrained Paraphraser (2023.acl-long)
Copied to clipboard
| Challenge: | Lexical substitution (LS) is an extremely powerful technology that can be used as a backbone of various NLP applications such as writing assistance. |
| Approach: | They propose two simple decoding strategies that focus on the variations of the target word during decoding to generate substitutes from a paraphraser. |
| Outcome: | The proposed methods outperform state-of-the-art LS methods based on pre-trained language models on three benchmarks. |
Knowledge Distillation ≈ Label Smoothing: Fact or Fallacy? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a method for knowledge transfer from one model to another . recent studies suggest it is based on label smoothing, but it is not . |
| Approach: | They propose to compare the predictive confidences of models trained with knowledge distillation . they propose to use a method that is similar to label smoothing to train models . |
| Outcome: | Experiments on four text classification tasks show that knowledge distillation and label smoothing drive model confidence in opposite directions. |
ReLearn: Unlearning via Learning for Large Language Models (2025.acl-long)
Copied to clipboard
Haoming Xu, Ningyuan Zhao, Liming Yang, Sendong Zhao, Shumin Deng, Mengru Wang, Bryan Hooi, Nay Oo, Huajun Chen, Ningyu Zhang
| Challenge: | Existing methods for unlearning large language models often rely on reverse optimization to reduce target token probabilities. |
| Approach: | They propose a data augmentation and fine-tuning pipeline for effective unlearning . they propose augmentation, evaluation frameworks to measure contextual forgetting . |
| Outcome: | The proposed framework achieves targeted forgetting while preserving high-quality outputs. |
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for regularizing deep neural networks rely on weight decay, dropout, batch/layer normalization to converge faster and generalize. |
| Approach: | They propose a framework for training with label regularization which includes conventional LS but can also model instance-specific variants. |
| Outcome: | The proposed approach consistently yields better results than conventional regularization on seven machine translation and three image classification tasks while maintaining training efficiency. |
ALEXSIS-PT: A New Resource for Portuguese Lexical Simplification (2022.coling-1)
Copied to clipboard
| Challenge: | Lexical simplification (LS) is the task of replacing complex words with simpler alternatives to make texts more accessible to various target populations. |
| Approach: | They propose to use a Brazilian Portuguese multi-candidate dataset to test LS systems. |
| Outcome: | The proposed model outperforms existing models on Brazilian Portuguese and Brazilian newspaper articles. |
DYNTEXT: Semantic-Aware Dynamic Text Sanitization for Privacy-Preserving LLM Inference (2025.findings-acl)
Copied to clipboard
Juhua Zhang, Zhiliang Tian, Minghang Zhu, Yiping Song, Taishu Sheng, Siyi Yang, Qiunan Du, Xinwang Liu, Minlie Huang, Dongsheng Li
| Challenge: | Existing methods to protect privacy of sensitive data are differential privacy (DP) and DP is used to protect users from privacy leakage. |
| Approach: | They propose an LDP-based Dynamic Text sanitization for privacy-preserving LLM inference that dynamically constructs semantic-aware adjacency lists of sensitive tokens to sample non-sensitive tokens for perturbation. |
| Outcome: | The proposed model excels on three datasets. |
RALS: Resources and Baselines for Romanian Automatic Lexical Simplification (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text simplification is the process of transforming texts into variants that are simpler to understand by larger audiences or easier to process by existing NLP systems. |
| Approach: | They propose a method for ordering simplification suggestions using a pairwise ranking approximation method, arranging candidates from simple to complex based on a separate set of human judgments. |
| Outcome: | The proposed system is the first to combine lexical simplification and complexity prediction in Romanian with human lexicals. |