| Challenge: | a limited number of Cantonese-specific datasets are available for PyCantones. |
| Approach: | They introduce PyCantonese, an open-source Python library for Cantonesi linguistics and natural language processing. |
| Outcome: | The proposed library is open-source and available for free for all purposes, including commercial ones. |
Similar Papers
CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing (2020.lrec-1)
Copied to clipboard
Ossama Obeid, Nasser Zalmout, Salam Khalifa, Dima Taji, Mai Oudah, Bashar Alhafni, Go Inoue, Fadhl Eryani, Alexander Erdmann, Nizar Habash
| Challenge: | CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Approach: | They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Outcome: | The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing. |
Stanza: A Python Natural Language Processing Toolkit for Many Human Languages (2020.acl-demos)
Copied to clipboard
| Challenge: | Existing tools that support only a few major languages are under-optimized for accuracy due to a focus on efficiency or use of less powerful models. |
| Approach: | They introduce a Python natural language processing toolkit that supports 66 languages . they train Stanza on 112 datasets and show it generalizes well on all languages compared to other tools . |
| Outcome: | The proposed toolkit performs well on 112 datasets and is compatible with the popular Java CoreNLP software. |
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)
Copied to clipboard
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos, Vivek Srikumar, Nicholas Rizzolo, Lev Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew, Zhili Feng, John Wieting, Xiaodong Yu, Yangqiu Song, Shashank Gupta, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth
| Challenge: | a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks. |
| Approach: | They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community . |
| Outcome: | The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges. |
CantoMap: a Hong Kong Cantonese MapTask Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of connected spoken Hong Kong Cantonese is constructed to study the phonology and semantics of the language. |
| Approach: | They propose to build a corpus of connected spoken Hong Kong Cantonese with phonemic transcription and controlled elicitation tasks. |
| Outcome: | The proposed corpus contains 768 minutes of recordings and transcripts of forty speakers. |
elfen: A Python Package for Efficient Linguistic Feature Extraction for Natural Language Datasets (2026.eacl-demo)
Copied to clipboard
| Challenge: | elfen is a Python library for efficient linguistic feature extraction for text datasets. |
| Approach: | They propose a Python library for efficient linguistic feature extraction for text datasets. |
| Outcome: | The proposed library enables linguistic feature extraction on thousands of items even on limited computing resources. |
N-LTP: An Open-source Neural Language Technology Platform for Chinese (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Existing tools that teach an independent model for each task are not supported in Chinese. |
| Approach: | They propose an open-source neural language platform supporting six Chinese NLP tasks . source code, documentation, and pre-trained models are available at https://github.com/hit-SCIR/ltp . |
| Outcome: | The proposed platform supports six Chinese NLP tasks. |
GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek (2025.coling-demos)
Copied to clipboard
Lefteris Loukas, Nikolaos Smyrnioudis, Chrysa Dikonomaki, Spiros Barbakos, Anastasios Toumazatos, John Koutsikakis, Manolis Kyriakakis, Mary Georgiou, Stavros Vassos, John Pavlopoulos, Ion Androutsopoulos
| Challenge: | GR-NLP-TOOLKIT is an open-source natural language processing toolkit for modern Greek. |
| Approach: | They present GR-NLP-TOOLKIT, an open-source natural language processing toolkit for Greek. |
| Outcome: | The toolkit provides state-of-the-art performance in five core NLP tasks . it can be easily installed in Python and is accessible through a demonstration platform on HuggingFace . |
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)
Copied to clipboard
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
| Challenge: | Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources. |
| Approach: | They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools. |
| Outcome: | The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large. |
Cifu: a Frequency Lexicon of Hong Kong Cantonese (2020.lrec-1)
Copied to clipboard
| Challenge: | lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC. |
| Approach: | They introduce a lexical database for Hong Kong Cantonese that offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC. |
| Outcome: | The proposed lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information. |
When Cantonese NLP Meets Pre-training: Progress and Challenges (2022.aacl-tutorials)
Copied to clipboard
| Challenge: | Cantonese is an influential Chinese variant with a large population of speakers worldwide. |
| Approach: | This tutorial will review Cantonese's progress in linguistics and NLP . it will introduce transformer-based pre-training methods for a wide range of downstream tasks . |
| Outcome: | This tutorial will present the main challenges for Cantonese NLP in relation to Cantonesian language idiosyncrasies of colloquialism and multilingualism. |