Challenge: ADD detection is a key area of research for low-resource languages like Portuguese, which lacks high-quality datasets.
Approach: They propose to provide the first publicly available ADD dataset for Portuguese, encompassing both Brazilian and European variants.
Outcome: The proposed dataset contains over 458,000 utterances, including a smaller portion of real speech from 62 speakers and a large collection of synthetic samples generated using multiple zero-shot text-to-speech (TTS) models, each conditioned on the original speaker’s voice.

Similar Papers

Cross-Domain Audio Deepfake Detection: Dataset and Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Existing audio deepfake detection datasets are outdated and lack generalization capabilities.
Approach: They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models.
Outcome: The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models.
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods (2025.acl-long)

Copied to clipboard

Challenge: Existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakkes.
Approach: They propose a large-scale speech deepfake dataset that includes over 3 million deepfak samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools.
Outcome: The proposed dataset includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools.
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Automatic speaker verification systems are facing escalating challenges due to deepfake attacks.
Approach: They propose a Urdu deepfake audio dataset for deepfak detection focusing on two spoofing attacks – Tacotron and VITS TTS.
Outcome: The proposed dataset evaluates two spoofing attacks in Urdu with a human evaluation to gauge whether people are able to distinguish deepfake audios from real (bonafide) audios.
A Data-Centric Approach to Generalizable Speech Deepfake Detection (2026.acl-long)

Copied to clipboard

Challenge: Speech deepfake detection (SDD) is a critical research area as speech synthesis technologies become more sophisticated.
Approach: They propose a data-centric approach to generalize SDD data from two perspectives . they propose naive aggregation strategies for mixing heterogeneous data and diversity-optimized sampling strategy for a single dataset and multiple datasets.
Outcome: The proposed approach outperforms the naive aggregation baseline on a 12k-hour data pool while using only 3% of the total available data.
IndicSynth: A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in synthetic speech generation technology have enabled the generation of high-quality synthetic (fake) speech that emulates human voices.
Approach: They propose a dataset that contains 4,000 hours of synthetic speech from 989 target speakers for 12 low-resourced Indian languages.
Outcome: The proposed dataset contains 4,000 hours of synthetic speech from 989 target speakers, including 456 females and 533 males for 12 low-resourced Indian languages.
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)

Copied to clipboard

Challenge: a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information.
Approach: They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection .
Outcome: The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations.
RTCFake: Speech Deepfake Detection in Real-Time Communication (2026.findings-acl)

Copied to clipboard

Challenge: Existing detection studies focus on offline simulations and struggle to cope with complex distortions introduced during RTC transmission.
Approach: They propose a large-scale speech deepfake dataset tailored for RTC scenarios . the dataset is constructed by transmitting speech through multiple social media and conferencing platforms .
Outcome: The proposed dataset is constructed by transmitting speech through multiple mainstream social media and conferencing platforms, enabling precise pairing between offline and online speech.
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages (2026.findings-acl)

Copied to clipboard

Challenge: Speech deepfakes are highly realistic and can generate a few seconds of recorded speech.
Approach: They propose an ALM that integrates semantic and prosodic representations from Whisper and TRILLsson to generate a speech deepfake dataset.
Outcome: The proposed framework outperforms existing ALMs on the ICF benchmark in Indic languages.
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized .
Approach: They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content .
Outcome: The proposed corpus is based on a pipeline methodology and is available for querying and downloading.
Generation-Based Data Augmentation for Offensive Language Detection: Is It Worth It? (2023.eacl-main)

Copied to clipboard

Challenge: generative data augmentation has been shown to be effective in offensive language detection but the potential for bias injection has not been investigated.
Approach: They propose to investigate the robustness of models trained on generated data in a variety of data augmentation setups and analyze models using the HateCheck suite.
Outcome: The proposed model training setups on four English offensive language datasets are robust and robust, while the generative DA setups do not present bias injection issues.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations