Challenge: fnqiè spellings in the Gungyùn provide pronunciations for more than 20000 characters . historical glossing practices are a crucial part of the pronunciation of ancient Chinese characters - but no integrated resources have been developed on the topic .
Approach: They propose to standardize digital versions of fnqiè spellings in the Gungyùn, one of the early rhyme books in the history of Chinese, providing pronunciations for more than 20000 characters.
Outcome: The proposed resource can predict historical spellings with high precision and shed light on ancient glossing practices.

Similar Papers

From Shijing to English and German: Resources and Evaluation for LLM Translation of Early Chinese Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) show promise in literary translation, but their performance in poetry remains unexplored.
Approach: They propose a framework that integrates knowledge-driven, rule-based, and LLM-as-judge metrics into a Shijing corpus . their code, lexical KB, and corpus reconstruction protocols are available at https://github.com/ML-KULeuven/ShijingLLMTrans.
Outcome: The proposed framework achieves higher human correlation than traditional metrics and high statistical stability.
ZiNet: Linking Chinese Characters Spanning Three Thousand Years (2022.findings-acl)

Copied to clipboard

Challenge: tens of thousands of ancient characters must be deciphered by experts to interpret unearthed documents.
Approach: They propose a diachronic Chinese knowledge base to help researchers discover glyph similar characters by measuring glyph similarities between ancient Chinese characters.
Outcome: The proposed method shows strong correlations between the scores obtained from the method and from human experts.
A Pragmatic Approach for Classical Chinese Word Segmentation (L18-1)

Copied to clipboard

Challenge: Classical Chinese word segmentation is largely neglected due to its obsoleteness . a new approach to segmentation using a marked-up corpus is needed .
Approach: They propose a pragmatic approach to deal with Classical Chinese word segmentation without any marked-up corpus.
Outcome: The proposed method makes the CCWS without any marked-up corpus more accurate compared with collocation-based methods.
Rethinking Dictionaries and Glyphs for Chinese Language Pre-training (2023.findings-acl)

Copied to clipboard

Challenge: Large-scale pre-trained language models (PLMs) such as BERT and GPT have revolutionized various research fields in natural language processing (NLP)
Approach: They propose a new learning paradigm that enhances the semantics understanding ability of Chinese PLMs with dictionary knowledge and structure of Chinese characters.
Outcome: The proposed model improves on both modern Chinese understanding benchmark CLUE and ancient Chinese understanding.
CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs (2025.acl-long)

Copied to clipboard

Challenge: CKnowEdit is the first-ever knowledge editing dataset designed to correct linguistic, factual, and logical errors in Large Language Models.
Approach: They propose a Chinese knowledge editing dataset to correct linguistic, factual, and logical errors in Large Language Models.
Outcome: The proposed dataset highlights the challenges that LLMs face in mastering Chinese . CKnowEdit can correct linguistic, factual, and logical errors in Chinese, the authors show .
Driving Chinese Spelling Correction from a Fine-Grained Perspective (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluations for Chinese spelling correction lack nuanced typology for spelling errors, creating an "invisible" bottleneck .
Approach: They propose a fine-grained evaluation principle for Chinese spelling correction (CSC) they categorize spelling errors into six different types and use it to evaluate models .
Outcome: The proposed evaluation principle can be leveraged to enhance CSC training models.
VisCGEC: Benchmarking the Visual Chinese Grammatical Error Correction (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on Chinese grammatical error correction ignore multi-modality and faked errors, which pushes techniques far away from real-world scenarios.
Approach: They propose to benchmark Chinese grammatical error correction for Chinese as a foreign language learner (CFL) using a dataset, they propose to use two CGEC frameworks to conduct experiments .
Outcome: The proposed approach achieves an F 0.5 score of only 28.9%.
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)

Copied to clipboard

Challenge: Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count.
Approach: They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP.
Outcome: The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task .
Automatic Reconstruction of Ancient Chinese Pronunciations (2024.findings-emnlp)

Copied to clipboard

Challenge: A human language is comprised of a pronunciation system and a writing system, both evolving and changing over time.
Approach: They reformulate existing phonetic rules into a dataset of 70,943 entries for 17,001 Chinese characters and use it to perform a temporal prediction task.
Outcome: The transformer-based model significantly advances the digitization and computational reconstruction of ancient Chinese phonology, providing a more complete and temporally contextualized resource for computational linguistics and historical research.
Kanbun-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Classical Chinese was introduced to Japan approximately 2,000 years ago . it was gradually adapted to a Japanese form called Kanbun-Kundoku (Kanbun) in Japanese reading and translating methods .
Approach: They construct a dataset that compares Classical Chinese and Kanbun in Japan using character reordering and machine translation tasks.
Outcome: The proposed dataset compares the current language models with human scores and compared them with human-level models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations