Papers with CAPT
Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss (2025.naacl-long)
Copied to clipboard
| Challenge: | APA and MDD are two of the main tasks of computer-assisted pronunciation training (CAPT) systems. |
| Approach: | They propose a computer-assisted pronunciation training approach that integrates APA and MDD tasks in parallel. |
| Outcome: | The proposed approach improves on APA and MDD tasks, and achieves an F1 score of 63.85%. |
Unlocking Large Audio-Language Models for Interactive Language Learning (2026.findings-eacl)
Copied to clipboard
| Challenge: | Computer-Assisted Pronunciation Training (CAPT) systems provide unintuitive feedback that lacks actionable guidance. |
| Approach: | They propose to use audio-language models to provide more user-friendly feedback for pronunciation training. |
| Outcome: | The proposed model outperforms baselines on mispronunciation detection and suggestion generation. |
Rethinking Denoised Auto-Encoding in Language Pre-Training (2021.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained models such as BERT have achieved success in learning sequence representations, but they tend to learn representations that are covariant with the noise of pre-training. |
| Approach: | They propose to train self-trained models to learn noise invariant sequence representations . they encourage consistency between original sequence and corrupted version via unsupervised instance-wise training signals. |
| Outcome: | The proposed model improves on 11 natural language understanding and cross-modal tasks and achieves 0.6% gain on GLUE benchmarks and 0.8% increment on NLVR2 . |
MIAPARLE: Online training for the discrimination of stress contrasts (L18-1)
Copied to clipboard
| Challenge: | Second language learners tend to imprint the prosody of their mother language onto the second language (L2) . this can hamper communication between learners and natives, and can also affect the credibility of learners and how they are evaluated by others. |
| Approach: | They propose a tool that focuses on stress perception for speakers whose L1 is a fixed-stress language, such as French. |
| Outcome: | The tool is particularly useful for speakers whose L1 is a fixed-stress language, such as French. |
Automatic Pronunciation Assessment - A Review (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Pronunciation assessment and its application in computer-aided pronunciation training (CAPT) have seen impressive progress in recent years. |
| Approach: | They review methods employed in computer-aided pronunciation training for both phonemic and prosodic pronunciations. |
| Outcome: | The proposed system should be able to automatically score non-native speech segments and give meaningful feedback. |
FAB: The French Absolute Beginner Corpus for Pronunciation Training (2020.lrec-1)
Copied to clipboard
| Challenge: | French Absolute Beginner corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners. |
| Approach: | They introduce the French Absolute Beginner (FAB) speech corpus which is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners. |
| Outcome: | The proposed corpus is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners. |
Data-Efficient Adaptation to Contextual Shifts in LLM-based Conversational Recommendation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing data selection methods struggle to distinguish learnable samples under contextual shifts. |
| Approach: | They propose a framework agnostic to underlying large language model-based conversational recommender systems (CRSs) that captures user preferences through free-form conversations and generates contextually relevant recommendations. |
| Outcome: | The proposed framework outperforms baselines on three CRS benchmarks with real-world temporal splits. |
Synergizing Semantic Anchors and Ordinal Smoothed Cross-Entropy for Speech Fluency Classification (2026.findings-acl)
Copied to clipboard
Mulati Kahaer, Sirajahmat Ruzmamat, XuDong Pang, Subinuer Maimaitituerxun, Zaokere Kadeer, Abudurexiti Reheman, Wenwen Lu, Panpan Zheng, Aishan Wumaier
| Challenge: | Existing methods fail to bridge the semantic gap between static expert priors and dynamic temporal representations while overlooking the inherent ordinal nature of fluency scores. |
| Approach: | They propose a set of expert features targeting fluency disruptions and rhythmic regularity to provide explicit linguistic priors. |
| Outcome: | The proposed model outperforms baseline models in both macroscopic and microscopic speech flow trends and local anomalies. |