Papers by Xingyi Duan
CCTC: A Cross-Sentence Chinese Text Correction Dataset for Native Speakers (2022.coling-1)
Copied to clipboard
| Challenge: | Chinese text correction datasets focus on detecting and correcting Chinese spelling errors and grammatical errors. |
| Approach: | They propose a Chinese text correction dataset for native speakers . they manually annotated 1,500 Chinese texts written by native speakers. |
| Outcome: | The proposed dataset can detect and correct Chinese spelling errors and grammatical errors. |
IFlyLegal: A Chinese Legal System for Consultation, Law Searching, and Document Analysis (D19-3)
Copied to clipboard
| Challenge: | Legal Tech is a system that performs legal consulting, multi-way law searching, and legal document analysis using deep contextual representations and various attention mechanisms. |
| Approach: | They propose a Chinese legal system that performs legal consulting, multi-way law searching, and legal document analysis using deep contextual representations and various attention mechanisms. |
| Outcome: | The proposed system performs legal consulting, multi-way law searching, and legal document analysis by exploiting techniques such as deep contextual representations and various attention mechanisms. |
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem (2026.findings-acl)
Copied to clipboard
Chengyou Wang, Mingchen Shao, Jingbin Hu, Zeyu Zhu, Hongfei Xue, Bingshen Mu, Xin Xu, Xingyi Duan, Binbin Zhang, Zhu Pengcheng, Chuang Ding, Xiaojun Zhang, Hui Bu, Lei Xie
| Challenge: | despite its linguistic significance, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. |
| Approach: | They propose to use WenetSpeech-Wu as a large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect of Chinese. |
| Outcome: | The proposed dataset includes 8,000 hours of speech data and strong open-source models . the proposed dataset is competitive and empirically validated . |