Building Parallel Monolingual Gan Chinese Dialects Corpus (L18-1)

Copied to clipboard

Challenge: In particular, we manually annotate a Gan Chinese Dialects Corpus (GCDC) including 131.5 hours and 310 documents with 6 different genres, containing news, official document, story, prose, poet, letter and speech, from 19 different Gan regions.
Approach: They propose a scheme to represent Gan Chinese dialects using Chinese character, Chinese Pinyin and Chinese audio forms.
Outcome: The proposed scheme is based on a Gan Chinese Dialects Corpus (GCDC) with 131.5 hours and 310 documents with 6 different genres, containing news, official document, story, prose, poet, letter and speech, from 19 different Gan regions.

Similar Papers

Would LLMs be Good Historical Linguists and Chinese Dialect Learners? (2026.acl-long)

Copied to clipboard

Challenge: Large language models struggle with low-resource Chinese dialects due to substantial phonological divergence.
Approach: They propose to incorporate Middle Chinese, the common historical ancestor of modern Chinese dialects, into LLMs to improve dialectal pronunciation modeling.
Outcome: The proposed approach improves on standard Chinese but struggles with low-resource Chinese dialects . the proposed model improves over baselines while revealing variation across dialects.
OCNLI: Original Chinese Natural Language Inference (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent efforts to extend natural language understanding to other languages have focused on (automatically) translating existing English datasets.
Approach: They propose to use a Chinese dataset to generate annotated sentences from native speakers specializing in linguistics to elicit annotations.
Outcome: The proposed dataset does not rely on automatic translation or non-expert annotation. instead, it elicits annotations from native speakers specializing in linguistics.
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem (2026.findings-acl)

Copied to clipboard

Challenge: despite its linguistic significance, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models.
Approach: They propose to use WenetSpeech-Wu as a large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect of Chinese.
Outcome: The proposed dataset includes 8,000 hours of speech data and strong open-source models . the proposed dataset is competitive and empirically validated .
Building an English-Chinese Parallel Corpus Annotated with Sub-sentential Translation Techniques (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that human translators often resort to different non-literal translation techniques besides literal translation . however, they receive less attention in developing natural language processing (NLP) applications.
Approach: They propose to have a better semantic control of extracting paraphrases from bilingual parallel corpora.
Outcome: The proposed method can automatically recognize different non-literal translation techniques . the results confirm the hypothesis of the proposed method .
GCDT: A Chinese RST Treebank for Multigenre and Multilingual Discourse Parsing (2022.aacl-short)

Copied to clipboard

Challenge: GCDT is the largest hierarchical discourse treebank for Mandarin Chinese in the framework of Rhetorical Structure Theory (RST).
Approach: They propose to use a Chinese hierarchical discourse treebank to parse Mandarin Chinese using relation inventory and a multilingual training program.
Outcome: The proposed dataset includes state-of-the-art scores for Chinese RST parsing and RST Parsing on the English GUM dataset, using cross-lingual training in Chinese and English with multilingual embeddings.
MC2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China (2024.acl-long)

Copied to clipboard

Challenge: MC2 is the largest open-source corpus of minority languages in china . MC2, however, includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
Approach: They propose a multilingual corpus of minority languages in China that includes four underrepresented languages . they prioritize accuracy while enhancing diversity by using a quality-centric approach .
Outcome: The proposed model prioritizes accuracy while enhancing diversity, the authors say . MC2 includes four underrepresented languages: Tibetan, Uyghur, Kazakh, and Mongolian .
The ManDi Corpus: A Spoken Corpus of Mandarin Regional Dialects (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods of remote speech data collection were limited by the telephone bandwidth and were therefore of low quality for phonetic research.
Approach: They introduce a spoken corpus of regional Mandarin dialects and Standard Mandarin.
Outcome: The proposed corpus contains 357 recordings (about 9.6 hours) of monosyllabic words, disyllable words, short sentences, a short passage and a poem, produced in standard Mandarin and in one of six regional Mandarin dialects.
End-to-End Chinese Speaker Identification (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for speaker identification in texts are incomplete and introduce errors that propagate and seriously affect the final output.
Approach: They propose to use speaker identification (SI) in texts to identify the speaker(s) for each utterance in texts.
Outcome: The proposed model can achieve comparable or better than previous state-of-the-art methods on all public SI datasets for Chinese.
FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction (2022.findings-emnlp)

Copied to clipboard

Challenge: grammatical error correction (GEC) is a complex task that requires high-quality data from native speakers.
Approach: They propose a human-annotated corpus to detect, identify and correct grammatical errors in Chinese examinations.
Outcome: The proposed model outperforms other models in low-resource settings, but there is a significant gap between the models and humans that encourages future models to bridge it.
From Shijing to English and German: Resources and Evaluation for LLM Translation of Early Chinese Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) show promise in literary translation, but their performance in poetry remains unexplored.
Approach: They propose a framework that integrates knowledge-driven, rule-based, and LLM-as-judge metrics into a Shijing corpus . their code, lexical KB, and corpus reconstruction protocols are available at https://github.com/ML-KULeuven/ShijingLLMTrans.
Outcome: The proposed framework achieves higher human correlation than traditional metrics and high statistical stability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations