Challenge: Mandarin Alphabetical Words (MAWs) are a key component of Modern Chinese . they are characterized by unique code-mixing idiosyncrasies influenced by language exchanges .
Approach: They propose to construct a large collection of Mandarin Alphabetic Words from Sina Weibo . they propose to use a web-based technique to identify and validate MAWs .
Outcome: The proposed method identifies 16,207 Mandarin Alphabetic Words (MAWs) using a web-based technique . the results show that the proposed method is useful for linguistic research and inquiries .

Similar Papers

MHE: Code-Mixed Corpora for Similar Language Identification (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of code-mixed data-sets is presented for similar language identification . the data-settings are based on a more-resourced minority language, Magahi .
Approach: They propose a Magahi-Hindi-English code-mixed corpus for similar language identification . they discuss the complexity of the data-set and provide a few baselines .
Outcome: The proposed corpus provides a language id at two levels: word and sentence.
TopWORDS-Poetry: Simultaneous Text Segmentation and Word Discovery for Classical Chinese Poetry via Bayesian Inference (2023.emnlp-main)

Copied to clipboard

Challenge: Experimental studies confirm that TopWORDS-Poetry can successfully segment poetry words without pre-given vocabulary or training corpus.
Approach: They propose an unsupervised method that can achieve reliable text segmentation and word discovery for classical Chinese poetry simultaneously without pre-given vocabulary or training corpus.
Outcome: Experimental results show that TopWORDS-Poetry can segment poetry lines into meaningful words with high quality without pre-given vocabulary or training corpus.
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)

Copied to clipboard

Challenge: Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count.
Approach: They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP.
Outcome: The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task .
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien (2022.findings-emnlp)

Copied to clipboard

Challenge: CM is a challenging task when mixed languages include dialects.
Approach: They propose to construct a Hokkien-Mandarin CM dataset to overcome the limitation . they propose to use a linguistics-based toolkit to train the model for translation tasks .
Outcome: The proposed model achieves good results on CM data translation while maintaining monolingual translation quality.
TopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference (2022.acl-long)

Copied to clipboard

Challenge: No existing methods can achieve effective text segmentation and word discovery in open domain Chinese texts.
Approach: They propose a Bayesian-based method that can achieve effective text segmentation and word discovery in open domain.
Outcome: The proposed method enjoys robust performance and transparent interpretation when no training corpus and domain vocabulary are available.
“Is Whole Word Masking Always Better for Chinese BERT?”: Probing on Chinese Grammatical Error Correction (2022.findings-acl)

Copied to clipboard

Challenge: a Chinese model with whole word masking has no subword because each token is an atomic character.
Approach: They propose to use whole word masking to mask all subwords corresponding to a word at once . they ask models to revise or insert tokens in a masked language modeling manner .
Outcome: The proposed model performs better when one character is inserted or replaced . the model trained with standard character-level masking performs best when one token is masked .
A Pragmatic Approach for Classical Chinese Word Segmentation (L18-1)

Copied to clipboard

Challenge: Classical Chinese word segmentation is largely neglected due to its obsoleteness . a new approach to segmentation using a marked-up corpus is needed .
Approach: They propose a pragmatic approach to deal with Classical Chinese word segmentation without any marked-up corpus.
Outcome: The proposed method makes the CCWS without any marked-up corpus more accurate compared with collocation-based methods.
Revisiting Classical Chinese Event Extraction with Ancient Literature Information (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on classical Chinese event extraction focus on grafting the complex modeling from English or modern Chinese works, neglecting the unique characteristic of this language.
Approach: They propose a Literary Vision-Language Model (VLM) for classical Chinese event extraction . they integrate annotations, historical background and character glyphs to capture the inner- and outer-context information from the sequence.
Outcome: The proposed model can capture the inner- and outer-context information at nearly zero cost.
Enriching Linguistic Representation in the Cantonese Wordnet and Building the New Cantonese Wordnet Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Currently, our wordnet includes a little over 5,200 concepts and 16,300 senses .
Approach: They propose to improve the Cantonese Wordnet by increasing the general coverage, adding functional categories, enriching verbal representations and creating the Cannese WordNet Corpus .
Outcome: The new version includes a little over 5,200 concepts and 16,300 senses .
Advancing Multi-Criteria Chinese Word Segmentation Through Criterion Classification and Denoising (2023.acl-long)

Copied to clipboard

Challenge: Recent research on multi-criteria Chinese word segmentation focuses on building complex private structures, adding more handcrafted features, or introducing complex optimization processes.
Approach: They propose a model that fits multiple Chinese word segments using input-hint inputs.
Outcome: The proposed model achieves state-of-the-art (SoTA) performance on multiple datasets simultaneously.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations