Papers with phonetics
Multi-layer Annotation of the Rigveda (L18-1)
Copied to clipboard
| Challenge: | Using a multi-level annotation, we present a corpus of the R. GVEDA . |
| Approach: | They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links. |
| Outcome: | The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments. |
Learning from the Dictionary: Heterogeneous Knowledge Guided Fine-tuning for Chinese Spell Checking (2022.findings-emnlp)
Copied to clipboard
Yinghui Li, Shirong Ma, Qingyu Zhou, Zhongli Li, Li Yangning, Shulin Huang, Ruiyang Liu, Chao Li, Yunbo Cao, Haitao Zheng
| Challenge: | Chinese Spell Checking (CSC) aims to detect and correct Chinese spelling errors. |
| Approach: | They propose a framework which renders Chinese Spell Checking model to learn heterogeneous knowledge from the dictionary in terms of phonetics, vision, and meaning. |
| Outcome: | The proposed framework renders the CSC model to learn heterogeneous knowledge from the dictionary in terms of phonetics, vision, and meaning. |
The French-Algerian Code-Switching Triggered audio corpus (FACST) (L18-1)
Copied to clipboard
| Challenge: | The French Algerian Code-Switching Triggered corpus is a corpus of spontaneous CS utterances . it is used to support linguistic and phonetic studies in phonetics and prosody . |
| Approach: | They propose to use a triggering protocol to elicit CS in natural conversations . they propose to do data segmentation and annotation in each language . |
| Outcome: | The proposed corpus is based on a code-switching protocol and is well-suited for linguistic and acoustic-phonetic studies. |
Enhancing Chinese Offensive Language Detection with Homophonic Perturbation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Detecting offensive language in Chinese is challenging due to homophonic substitutions used to evade detection. |
| Approach: | They propose to use HED-COLD to build a large-scale homophonic dataset for Chinese offensive language detection and a homophone-aware pretraining strategy to learn phonetics and orthography. |
| Outcome: | The proposed framework achieves state-of-the-art performance on the COLD test set and the toxicity benchmark ToxiCloakCN. |