TopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference (2022.acl-long)
Copied to clipboard
| Challenge: | No existing methods can achieve effective text segmentation and word discovery in open domain Chinese texts. |
| Approach: | They propose a Bayesian-based method that can achieve effective text segmentation and word discovery in open domain. |
| Outcome: | The proposed method enjoys robust performance and transparent interpretation when no training corpus and domain vocabulary are available. |
Similar Papers
TopWORDS-Poetry: Simultaneous Text Segmentation and Word Discovery for Classical Chinese Poetry via Bayesian Inference (2023.emnlp-main)
Copied to clipboard
| Challenge: | Experimental studies confirm that TopWORDS-Poetry can successfully segment poetry words without pre-given vocabulary or training corpus. |
| Approach: | They propose an unsupervised method that can achieve reliable text segmentation and word discovery for classical Chinese poetry simultaneously without pre-given vocabulary or training corpus. |
| Outcome: | Experimental results show that TopWORDS-Poetry can segment poetry lines into meaningful words with high quality without pre-given vocabulary or training corpus. |
CWSeg: An Efficient and General Approach to Chinese Word Segmentation (2023.acl-industry)
Copied to clipboard
| Challenge: | Existing methods for Chinese word segmentation have achieved state-of-the-art performance, but they pose challenges in the deployment. |
| Approach: | They propose to augment PLM-based Chinese word segmentation schemes by developing cohort training and versatile decoding strategies. |
| Outcome: | The proposed model can be used to augment existing PLM-based models and improve their performance on Chinese LLaMA and Alpaca datasets. |
Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and Denoising (2023.acl-long)
Copied to clipboard
| Challenge: | Recent research on multi-criteria Chinese word segmentation focuses on building complex private structures, adding more handcrafted features, or introducing complex optimization processes. |
| Approach: | They propose a model that fits multiple Chinese word segments using input-hint inputs. |
| Outcome: | The proposed model achieves state-of-the-art (SoTA) performance on multiple datasets simultaneously. |
Multi-Tiered Cantonese Word Segmentation (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing work on word segmentation for Chinese does not have conventional word boundaries as English does. |
| Approach: | They propose a linguistically motivated, multi-tiered word segmentation system for Cantonese . they propose linguisticly motivated, linguistic-motivated system that can cater to different needs . |
| Outcome: | The proposed system can be adapted to Cantonese corpus data. |
State-of-the-art Chinese Word Segmentation with Bi-LSTMs (D18-1)
Copied to clipboard
| Challenge: | A wide variety of neural-network architectures have been proposed for the task of Chinese word segmentation. |
| Approach: | They propose a bidirectional LSTM model with standard deep learning techniques and best practices for the task of Chinese word segmentation. |
| Outcome: | The proposed model outperforms models based on standard deep learning techniques and best practices on Chinese word segmentation datasets. |
Adaptive Multi-Task Transfer Learning for Chinese Word Segmentation in Medical Text (C18-1)
Copied to clipboard
| Challenge: | Chinese word segmentation (CWS) tools face a performance drop when dealing with domain text . domain-specific CWS requires extremely high annotation cost due to ambiguity caused by domain terms and writing style . |
| Approach: | They propose to exploit domain-invariant knowledge from high resource to low resource domains to build Chinese word segmentation models. |
| Outcome: | The proposed model achieves higher accuracy than single-task CWS and other transfer learning baselines . the model is based on domain-invariant knowledge from high resource to low resource domains based in the biomedical domain . |
Deep Bayesian Natural Language Processing (P19-4)
Copied to clipboard
| Challenge: | Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks. |
| Approach: | This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues . |
| Outcome: | This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding. |
Is Word Segmentation Necessary for Deep Learning of Chinese Representations? (P19-1)
Copied to clipboard
| Challenge: | Using word-based models, we compare word-oriented models with char-based ones . word-driven models are more vulnerable to data sparsity and the presence of out-of-vocabulary words . |
| Approach: | They benchmark word-based models with char-based model which does not involve word segmentation in four NLP benchmark tasks. |
| Outcome: | The proposed model outperforms char-based models in four NLP benchmark tasks. |
Deep Bayesian Learning and Understanding (C18-3)
Copied to clipboard
| Challenge: | COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning. |
| Approach: | a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding . |
| Outcome: | This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 . |
Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence Representation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent Chinese word segmentation models tend to learn the segmentation knowledge through in-vocabulary words rather than understanding the meaning of the entire context. |
| Approach: | They propose a context-aware approach that incorporates unsupervised sentence representation learning over different dropout masks into the multi-criteria training framework. |
| Outcome: | The proposed approach achieves state-of-the-art (SoTA) performance on six of the nine CWS benchmark datasets and out-of vocabulary (OOV) recalls for eight of nine. |