BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTS (2025.emnlp-main)
Copied to clipboard
Alexandre Costa Ferro Filho, Rafaello Virgilli, Lucas Alcantara Souza, F S de Oliveira, Marcelo Henrique Lopes Ferreira, Daniel Tunnermann, Gustavo Dos Reis Oliveira, Anderson Da Silva Soares, Arlindo Rodrigues Galvão Filho
| Challenge: | ADD detection is a key area of research for low-resource languages like Portuguese, which lacks high-quality datasets. |
| Approach: | They propose to provide the first publicly available ADD dataset for Portuguese, encompassing both Brazilian and European variants. |
| Outcome: | The proposed dataset contains over 458,000 utterances, including a smaller portion of real speech from 62 speakers and a large collection of synthetic samples generated using multiple zero-shot text-to-speech (TTS) models, each conditioned on the original speaker’s voice. |
Similar Papers
Cross-Domain Audio Deepfake Detection: Dataset and Analysis (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing audio deepfake detection datasets are outdated and lack generalization capabilities. |
| Approach: | They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. |
| Outcome: | The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models. |
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods (2025.acl-long)
Copied to clipboard
| Challenge: | Existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakkes. |
| Approach: | They propose a large-scale speech deepfake dataset that includes over 3 million deepfak samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools. |
| Outcome: | The proposed dataset includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools. |
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset (2024.findings-acl)
Copied to clipboard
Sheza Munir, Wassay Sajjad, Mukeet Raza, Emaan Abbas, Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza
| Challenge: | Automatic speaker verification systems are facing escalating challenges due to deepfake attacks. |
| Approach: | They propose a Urdu deepfake audio dataset for deepfak detection focusing on two spoofing attacks – Tacotron and VITS TTS. |
| Outcome: | The proposed dataset evaluates two spoofing attacks in Urdu with a human evaluation to gauge whether people are able to distinguish deepfake audios from real (bonafide) audios. |
A Data-Centric Approach to Generalizable Speech Deepfake Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Speech deepfake detection (SDD) is a critical research area as speech synthesis technologies become more sophisticated. |
| Approach: | They propose a data-centric approach to generalize SDD data from two perspectives . they propose naive aggregation strategies for mixing heterogeneous data and diversity-optimized sampling strategy for a single dataset and multiple datasets. |
| Outcome: | The proposed approach outperforms the naive aggregation baseline on a 12k-hour data pool while using only 3% of the total available data. |
IndicSynth: A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in synthetic speech generation technology have enabled the generation of high-quality synthetic (fake) speech that emulates human voices. |
| Approach: | They propose a dataset that contains 4,000 hours of synthetic speech from 989 target speakers for 12 low-resourced Indian languages. |
| Outcome: | The proposed dataset contains 4,000 hours of synthetic speech from 989 target speakers, including 456 females and 533 males for 12 low-resourced Indian languages. |
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)
Copied to clipboard
| Challenge: | a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information. |
| Approach: | They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection . |
| Outcome: | The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations. |
RTCFake: Speech Deepfake Detection in Real-Time Communication (2026.findings-acl)
Copied to clipboard
Jun Xue, Zhuolin Yi, Yihuan Huang, Yanzhen Ren, Yujie Chen, Cunhang Fan, Zicheng Su, Yongcheng Zhang, Bo Cai
| Challenge: | Existing detection studies focus on offline simulations and struggle to cope with complex distortions introduced during RTC transmission. |
| Approach: | They propose a large-scale speech deepfake dataset tailored for RTC scenarios . the dataset is constructed by transmitting speech through multiple social media and conferencing platforms . |
| Outcome: | The proposed dataset is constructed by transmitting speech through multiple mainstream social media and conferencing platforms, enabling precise pairing between offline and online speech. |
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Speech deepfakes are highly realistic and can generate a few seconds of recorded speech. |
| Approach: | They propose an ALM that integrates semantic and prosodic representations from Whisper and TRILLsson to generate a speech deepfake dataset. |
| Outcome: | The proposed framework outperforms existing ALMs on the ICF benchmark in Indic languages. |
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)
Copied to clipboard
| Challenge: | a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized . |
| Approach: | They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content . |
| Outcome: | The proposed corpus is based on a pipeline methodology and is available for querying and downloading. |
Generation-Based Data Augmentation for Offensive Language Detection: Is It Worth It? (2023.eacl-main)
Copied to clipboard
| Challenge: | generative data augmentation has been shown to be effective in offensive language detection but the potential for bias injection has not been investigated. |
| Approach: | They propose to investigate the robustness of models trained on generated data in a variety of data augmentation setups and analyze models using the HateCheck suite. |
| Outcome: | The proposed model training setups on four English offensive language datasets are robust and robust, while the generative DA setups do not present bias injection issues. |