First Steps Towards the Integration of Resources on Historical Glossing Traditions in the History of Chinese: A Collection of Standardized Fǎnqiè Spellings from the Guǎngyùn (2024.lrec-main)
Copied to clipboard
| Challenge: | fnqiè spellings in the Gungyùn provide pronunciations for more than 20000 characters . historical glossing practices are a crucial part of the pronunciation of ancient Chinese characters - but no integrated resources have been developed on the topic . |
| Approach: | They propose to standardize digital versions of fnqiè spellings in the Gungyùn, one of the early rhyme books in the history of Chinese, providing pronunciations for more than 20000 characters. |
| Outcome: | The proposed resource can predict historical spellings with high precision and shed light on ancient glossing practices. |
Similar Papers
From Shijing to English and German: Resources and Evaluation for LLM Translation of Early Chinese Poetry (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) show promise in literary translation, but their performance in poetry remains unexplored. |
| Approach: | They propose a framework that integrates knowledge-driven, rule-based, and LLM-as-judge metrics into a Shijing corpus . their code, lexical KB, and corpus reconstruction protocols are available at https://github.com/ML-KULeuven/ShijingLLMTrans. |
| Outcome: | The proposed framework achieves higher human correlation than traditional metrics and high statistical stability. |
ZiNet: Linking Chinese Characters Spanning Three Thousand Years (2022.findings-acl)
Copied to clipboard
| Challenge: | tens of thousands of ancient characters must be deciphered by experts to interpret unearthed documents. |
| Approach: | They propose a diachronic Chinese knowledge base to help researchers discover glyph similar characters by measuring glyph similarities between ancient Chinese characters. |
| Outcome: | The proposed method shows strong correlations between the scores obtained from the method and from human experts. |
A Pragmatic Approach for Classical Chinese Word Segmentation (L18-1)
Copied to clipboard
| Challenge: | Classical Chinese word segmentation is largely neglected due to its obsoleteness . a new approach to segmentation using a marked-up corpus is needed . |
| Approach: | They propose a pragmatic approach to deal with Classical Chinese word segmentation without any marked-up corpus. |
| Outcome: | The proposed method makes the CCWS without any marked-up corpus more accurate compared with collocation-based methods. |
Rethinking Dictionaries and Glyphs for Chinese Language Pre-training (2023.findings-acl)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models (PLMs) such as BERT and GPT have revolutionized various research fields in natural language processing (NLP) |
| Approach: | They propose a new learning paradigm that enhances the semantics understanding ability of Chinese PLMs with dictionary knowledge and structure of Chinese characters. |
| Outcome: | The proposed model improves on both modern Chinese understanding benchmark CLUE and ancient Chinese understanding. |
CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | CKnowEdit is the first-ever knowledge editing dataset designed to correct linguistic, factual, and logical errors in Large Language Models. |
| Approach: | They propose a Chinese knowledge editing dataset to correct linguistic, factual, and logical errors in Large Language Models. |
| Outcome: | The proposed dataset highlights the challenges that LLMs face in mastering Chinese . CKnowEdit can correct linguistic, factual, and logical errors in Chinese, the authors show . |
Driving Chinese Spelling Correction from a Fine-Grained Perspective (2025.coling-main)
Copied to clipboard
| Challenge: | Existing evaluations for Chinese spelling correction lack nuanced typology for spelling errors, creating an "invisible" bottleneck . |
| Approach: | They propose a fine-grained evaluation principle for Chinese spelling correction (CSC) they categorize spelling errors into six different types and use it to evaluate models . |
| Outcome: | The proposed evaluation principle can be leveraged to enhance CSC training models. |
VisCGEC: Benchmarking the Visual Chinese Grammatical Error Correction (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies on Chinese grammatical error correction ignore multi-modality and faked errors, which pushes techniques far away from real-world scenarios. |
| Approach: | They propose to benchmark Chinese grammatical error correction for Chinese as a foreign language learner (CFL) using a dataset, they propose to use two CGEC frameworks to conduct experiments . |
| Outcome: | The proposed approach achieves an F 0.5 score of only 28.9%. |
WikiHan: A New Comparative Dataset for Chinese Languages (2022.coling-1)
Copied to clipboard
| Challenge: | Currently, there are 1.3 billion speakers of Sinitic varieties, making the family one of the largest in terms of speaker count. |
| Approach: | They have collected a single constituent and structured form of Chinese varieties for comparative linguistics and Chinese NLP. |
| Outcome: | The proposed dataset contains 67,943 entries across 8 varieties and Middle Chinese . it achieves 54.11% accuracy and 17.69% error rate on a protoform reconstruction task . |
Automatic Reconstruction of Ancient Chinese Pronunciations (2024.findings-emnlp)
Copied to clipboard
| Challenge: | A human language is comprised of a pronunciation system and a writing system, both evolving and changing over time. |
| Approach: | They reformulate existing phonetic rules into a dataset of 70,943 entries for 17,001 Chinese characters and use it to perform a temporal prediction task. |
| Outcome: | The transformer-based model significantly advances the digitization and computational reconstruction of ancient Chinese phonology, providing a more complete and temporally contextualized resource for computational linguistics and historical research. |
Kanbun-LM: Reading and Translating Classical Chinese in Japanese Methods by Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Classical Chinese was introduced to Japan approximately 2,000 years ago . it was gradually adapted to a Japanese form called Kanbun-Kundoku (Kanbun) in Japanese reading and translating methods . |
| Approach: | They construct a dataset that compares Classical Chinese and Kanbun in Japan using character reordering and machine translation tasks. |
| Outcome: | The proposed dataset compares the current language models with human scores and compared them with human-level models. |