Papers by Peichao Lai
Enhancing Unsupervised Sentence Embeddings via Knowledge-Driven Data Augmentation and Gaussian-Decayed Contrastive Learning (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for data augmentation neglect fine-grained knowledge, such as entities and quantities, leading to insufficient diversity and high data noise. |
| Approach: | They propose a pipeline-based data augmentation method via LLMs and introduce the Gaussian-decayed gradient-assisted Contrastive Sentence Embedding (GCSE) model to enhance unsupervised sentence embeddings. |
| Outcome: | The proposed method achieves state-of-the-art performance in semantic textual similarity tasks using fewer data samples and smaller LLMs. |
Improving Low-Resource Sequence Labeling with Knowledge Fusion and Contextual Label Explanations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to sequence labeling are limited due to the scarcity of domain-specific data and semantic distribution biases in domain-based contexts. |
| Approach: | They propose a framework that integrates an LLM-based knowledge enhancement workflow with a span-based Knowledge Fusion for Rich and Efficient Extraction model. |
| Outcome: | The proposed model achieves state-of-the-art performance on multiple domain-specific sequence labeling datasets and is highly efficient. |
Quantum-inspired Language Model with Lindblad Master Equation and Interference Measurement for Sentiment Analysis (2024.naacl-long)
Copied to clipboard
| Challenge: | Quantum-inspired models have demonstrated superior performance in many downstream language tasks, such as question answering and sentiment analysis. |
| Approach: | They propose a quantum-inspired neural network that integrates the Lindblad Master Equation to model the evolution process and the interferometry to the measurement process, providing more physical meaning to strengthen the interpretability. |
| Outcome: | The proposed model outperforms existing models on sentiment analysis datasets and shows that it is more accurate and performs better than existing models. |
PCBERT: Parent and Child BERT for Chinese Few-shot NER (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches to improve model performance on few-shot or zero-shot datasets are not effective for Chinese few- shot NER. |
| Approach: | They propose a prompt-based Parent and Child BERT for Chinese few-shot NER to train an annotating model on high-resource datasets and then discover more implicit labels on low-resourced datasets. |
| Outcome: | The proposed model can be used on Weibo and other Chinese NER datasets and it is shown to be effective in few-shot learning. |