PyCantonese: Cantonese Linguistics and NLP in Python (2022.lrec-1)

Copied to clipboard

Challenge: a limited number of Cantonese-specific datasets are available for PyCantones.
Approach: They introduce PyCantonese, an open-source Python library for Cantonesi linguistics and natural language processing.
Outcome: The proposed library is open-source and available for free for all purposes, including commercial ones.

Similar Papers

CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing (2020.lrec-1)

Copied to clipboard

Challenge: CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Approach: They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Outcome: The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing.
Stanza: A Python Natural Language Processing Toolkit for Many Human Languages (2020.acl-demos)

Copied to clipboard

Challenge: Existing tools that support only a few major languages are under-optimized for accuracy due to a focus on efficiency or use of less powerful models.
Approach: They introduce a Python natural language processing toolkit that supports 66 languages . they train Stanza on 112 datasets and show it generalizes well on all languages compared to other tools .
Outcome: The proposed toolkit performs well on 112 datasets and is compatible with the popular Java CoreNLP software.
CogCompNLP: Your Swiss Army Knife for NLP (L18-1)

Copied to clipboard

Challenge: a corpus-reader module supports popular corpora, feature extraction and annotation modules for semantic and syntactic tasks.
Approach: They propose a library that provides modules to address different challenges . they provide a corpus-reader module that supports popular corpora in the NLP community .
Outcome: The proposed library simplifies the process of design and development of NLP applications by providing modules to address different challenges.
CantoMap: a Hong Kong Cantonese MapTask Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of connected spoken Hong Kong Cantonese is constructed to study the phonology and semantics of the language.
Approach: They propose to build a corpus of connected spoken Hong Kong Cantonese with phonemic transcription and controlled elicitation tasks.
Outcome: The proposed corpus contains 768 minutes of recordings and transcripts of forty speakers.
elfen: A Python Package for Efficient Linguistic Feature Extraction for Natural Language Datasets (2026.eacl-demo)

Copied to clipboard

Challenge: elfen is a Python library for efficient linguistic feature extraction for text datasets.
Approach: They propose a Python library for efficient linguistic feature extraction for text datasets.
Outcome: The proposed library enables linguistic feature extraction on thousands of items even on limited computing resources.
N-LTP: An Open-source Neural Language Technology Platform for Chinese (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing tools that teach an independent model for each task are not supported in Chinese.
Approach: They propose an open-source neural language platform supporting six Chinese NLP tasks . source code, documentation, and pre-trained models are available at https://github.com/hit-SCIR/ltp .
Outcome: The proposed platform supports six Chinese NLP tasks.
GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek (2025.coling-demos)

Copied to clipboard

Challenge: GR-NLP-TOOLKIT is an open-source natural language processing toolkit for modern Greek.
Approach: They present GR-NLP-TOOLKIT, an open-source natural language processing toolkit for Greek.
Outcome: The toolkit provides state-of-the-art performance in five core NLP tasks . it can be easily installed in Python and is accessible through a demonstration platform on HuggingFace .
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)

Copied to clipboard

Challenge: Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources.
Approach: They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools.
Outcome: The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large.
Cifu: a Frequency Lexicon of Hong Kong Cantonese (2020.lrec-1)

Copied to clipboard

Challenge: lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC.
Approach: They introduce a lexical database for Hong Kong Cantonese that offers phonological and orthographic information, frequency measures, and lexically neighborhood information for lexicals in HKC.
Outcome: The proposed lexical database for Hong Kong Cantonese offers phonological and orthographic information, frequency measures, and lexically neighborhood information.
When Cantonese NLP Meets Pre-training: Progress and Challenges (2022.aacl-tutorials)

Copied to clipboard

Challenge: Cantonese is an influential Chinese variant with a large population of speakers worldwide.
Approach: This tutorial will review Cantonese's progress in linguistics and NLP . it will introduce transformer-based pre-training methods for a wide range of downstream tasks .
Outcome: This tutorial will present the main challenges for Cantonese NLP in relation to Cantonesian language idiosyncrasies of colloquialism and multilingualism.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations