A Corpus of Spontaneous L2 English Speech for Real-situation Speaking Assessment (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing automated scoring systems rely on highly controlled elicitation protocols, such as reading aloud isolated words or short sentences, limiting their ability to evaluate spontaneous speech. |
| Approach: | They propose to collect a corpus of spontaneous L2 English speech from university students as part of a French national certificate in English. |
| Outcome: | The results show that only 35.4% of the 6,350 targeted words had stress detected on the expected syllable, revealing a common stress shift to the final s. |
Similar Papers
Constructing Korean Learners’ L2 Speech Corpus of Seven Languages for Automatic Pronunciation Assessment (2024.lrec-main)
Copied to clipboard
| Challenge: | Multilingual L2 speech corpora for automatic speech assessment are currently available, but lack comprehensive annotations of L2 from non-native speakers of various languages. |
| Approach: | They propose to use Korean learners’ L2 speech corpus of seven languages to develop automatic speech assessment. |
| Outcome: | The proposed corpus contains 1,200 hours of L2 speech data from Korean learners (400 hours for English, 200 hours each for Japanese and Chinese, 100 hours each in French, German, Spanish, and Russian). |
Unraveling Spontaneous Speech Dimensions for Cross-Corpus ASR System Evaluation for French (2024.lrec-main)
Copied to clipboard
| Challenge: | 'spontaneous speech' is a catch-all term used for situations like speaking with a friend, being interviewed on radio/TV or giving a lecture. |
| Approach: | They propose to use four dimensions to describe spontaneous speech variation in automatic speech recognition systems. |
| Outcome: | The proposed system can be used to predict the WER of speech recognition systems on face-to-face interactions. |
CBFC: a parallel L2 speech corpus for Korean and French learners (L18-1)
Copied to clipboard
| Challenge: | Using corpora for second language acquisition has become more and more common . corporata are used to study morpho-syntactic phenomena in English as a foreign language . |
| Approach: | They propose to use a bilingual corpus of French learners of Korean and Korean learners of French to provide a translated and annotated corpus to the scientific community. |
| Outcome: | The proposed corpus can be used for a wide array of purposes in the field of theoretical but also applied linguistics. |
Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases (2025.acl-long)
Copied to clipboard
Rena Wei Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau
| Challenge: | Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge. |
| Approach: | They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data. |
| Outcome: | The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages. |
Urdu Pitch Accents and Intonation Patterns in Spontaneous Conversational Speech (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies of Urdu intonation describe scripted and laboratory speech . |
| Approach: | They summarise Urdu pitch accents and their intonation patterns using a simplified version of the Rhythm and Pitch labelling system and a simple RAP system. |
| Outcome: | The analysis of a hand-labelled telephone conversation shows that low pitch accents play an important role in Urdu spontaneous speech. |
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting. |
| Approach: | They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language. |
| Outcome: | The proposed model learns to generate speech signal based on pronunciation of source language. |
Developing a Benchmark for Pronunciation Feedback: Creation of a Phonemically Annotated Speech Corpus of isiZulu Language Learner Speech (2024.lrec-main)
Copied to clipboard
Alexandra O’Neil, Nils Hjortnaes, Francis Tyers, Zinhle Nkosi, Thulile Ndlovu, Zanele Mlondo, Ngami Phumzile Pewa
| Challenge: | Existing corpora for computer-assisted pronunciation training (CAPT) do not apply well to research in pronunciation feedback. |
| Approach: | They propose to create a corpus of isiZulu language learner speech that has been annotated for phoneme errors and suprasegmental errors in tone. |
| Outcome: | The proposed corpus is comprised of gold standard recordings from isiZulu teachers and recordings from students that have been annotated for pronunciation errors. |
Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking Styles (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Arabic speech recordings has been built to allow comparisons between Arabic and other languages. |
| Approach: | They propose to build a corpus of Arabic speech recordings that can be compared with other languages. |
| Outcome: | The proposed corpus can be used for forensic phonetic research and casework applications. |
Investigating the effect of auxiliary objectives for the automated grading of learner English speech transcriptions (2020.acl-main)
Copied to clipboard
| Challenge: | a growing demand for the ability to communicate in English means automated tutoring and assessment systems are becoming more popular. |
| Approach: | They propose to use automatic speech recognition transcripts to grade spontaneous speech based on textual features. |
| Outcome: | The proposed system improves on a transformer encoder with native language identification as an auxiliary task. |
TLT-school: a Corpus of Non Native Children Speech (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of speech utterances collected in schools of northern italy is being used to assess the performance of students learning both English and German. |
| Approach: | a corpus of speech utterances collected in schools of northern italy is described . the corpus is going to be freely distributed to scientific community . |
| Outcome: | The corpus of speech utterances collected in schools of northern italy is a "Trentino Language Testing" in schools" the data are used to assess the performance of students learning English and German . |