Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods to analyze speech representations using pretraining data are difficult to achieve for endangered languages. |
| Approach: | They propose an unsupervised method to examine the level of abstraction in vector representations of speech from a pretrained model to determine their level of abstractness. |
| Outcome: | The proposed method is fully unsupervised and could be used in comparative studies on under-documented languages. |
Similar Papers
Quantifying Language Variation Acoustically with Few Resources (2022.naacl-main)
Copied to clipboard
| Challenge: | acoustic models represent linguistic information based on massive amounts of data. |
| Approach: | They examine the model's ability to distinguish low-resource (Dutch) regional varieties by extracting embeddings from hidden layers and dynamic time warping. |
| Outcome: | The proposed model outperforms transcription-based models without phonetic transcriptions on the basis of only six seconds of speech. |
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations (2025.emnlp-main)
Copied to clipboard
| Challenge: | a recent study evaluated the extent to which SLMs encode nuanced syntactic and conceptual features . acoustic and phonetic features are shallow, but the extent of nuance is unclear . |
| Approach: | a new study evaluates contextual syntactic and semantic features in transformer-based speech language models . authors compare SLMs to linguistic competence assessments for large language models. |
| Outcome: | a new study compares SLMs with linguistic competence assessments to assess speech recognition and understanding . the results show that SLM models encode grammatical features more robustly than conceptual ones . |
The Geometry of Multilingual Language Model Representations (2022.emnlp-main)
Copied to clipboard
| Challenge: | XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning. |
| Approach: | They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language. |
| Outcome: | The proposed model can extract features for downstream tasks and cross-lingual transfer learning. |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
What Do Language Models Hear? Probing for Auditory Representations in Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | a linear probe is used to retrieve the correct text representation of an object given a snippet of audio related to that object. |
| Approach: | They develop a linear probe that retrieves the correct text representation of an object . they then test the probe's generalization to objects that were not seen during training . |
| Outcome: | The proposed model generalizes to objects that were not seen during training, the study finds . the model can learn representations of perceptual concepts that plausibly mirror the grounded representations . |
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)
Copied to clipboard
| Challenge: | Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. |
| Approach: | They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. |
| Outcome: | The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks. |
Towards a Deep Understanding of Multilingual End-to-End Speech Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent years have witnessed the rapid development of end-to-end speech-totext translation (ST) which has demonstrated remarkable performance and outperformed conventional cascaded systems. |
| Approach: | They employ Singular Value Canonical Correlation Analysis to analyze representations learnt in a multilingual end-to-end speech translation model trained over 22 languages. |
| Outcome: | The proposed approach outperforms existing cascaded systems in predicting phonetic features and improves translation quality. |
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages (2025.naacl-long)
Copied to clipboard
| Challenge: | In the brains of human bilinguals, syntax processing may occur in similar regions for their first and second language, depending on factors like when the second language was learned and language proficiency. |
| Approach: | They propose to use sparse autoencoders to train Llama-3-8B and Aya-23-8B models to train multilingual models that share morphsyntactic representations of grammatical concepts. |
| Outcome: | The proposed model can predict plural verbs in different languages by activating the same plural feature. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)
Copied to clipboard
| Challenge: | a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information. |
| Approach: | They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection . |
| Outcome: | The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations. |