Challenge: Audio deepfake detection systems do not generalize well to realistic in-the-wild deepfakkes.
Approach: They propose a novel In-Context Learning paradigm with comparison-guidance for Audio Deepfake detection framework that uses audio language models for training-free generalization to unseen deepfakes.
Outcome: The proposed framework improves macro F1 over specialized detectors on in-the-wild datasets with up to 2 relative improvement over existing models.

Similar Papers

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Existing algorithms for audio deepfake detection are based on layer-wise analysis of self-supervised learning (SSL) models.
Approach: They conduct a layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts.
Outcome: The proposed models achieve competitive equal error rate (EER) scores even when employing a reduced number of layers.
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)

Copied to clipboard

Challenge: a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information.
Approach: They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection .
Outcome: The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations.
Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for deepfake detection suffer from two limitations: modality fragmentation and shallow inter-modal reasoning.
Approach: They propose a framework for multimodal deepfake detection that uses contrastive learning and large language models to mitigate modality fragmentation and refine embeddings to address shallow inter-modal reasoning.
Outcome: ConLLM reduces audio deepfake EER by 50%, improves video accuracy by 8%, and achieves approximately 9% accuracy gains in audio-visual tasks.
RTCFake: Speech Deepfake Detection in Real-Time Communication (2026.findings-acl)

Copied to clipboard

Challenge: Existing detection studies focus on offline simulations and struggle to cope with complex distortions introduced during RTC transmission.
Approach: They propose a large-scale speech deepfake dataset tailored for RTC scenarios . the dataset is constructed by transmitting speech through multiple social media and conferencing platforms .
Outcome: The proposed dataset is constructed by transmitting speech through multiple mainstream social media and conferencing platforms, enabling precise pairing between offline and online speech.
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning (2024.emnlp-main)

Copied to clipboard

Challenge: In-context learning has shown high efficacy in several NLP tasks, especially in few-shot settings.
Approach: They propose a backdoor attack method that poisons demonstration examples and poisons the demonstration context, preserving the model's generality.
Outcome: The proposed method can make models behave in alignment with predefined intentions without fine-tuning the model.
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in audio generation led to an increasing number of deepfakes . however, these methods are typically tested in an in-domain setup .
Approach: They propose a large-scale cross-domain audio deepfake benchmark comprising 668.8 hours of real and deepfak speech.
Outcome: The proposed benchmark compares audio deepfake detectors with existing methods in the wild . the results show that the proposed methods perform better in different languages than existing methods .
A Survey to Recent Progress Towards Understanding In-Context Learning (2025.findings-naacl)

Copied to clipboard

Challenge: Existing research on In-Context Learning (ICL) is unclear, despite empirical success . a data generation perspective is used to interpret ICL .
Approach: They propose to use data generation to reinterpret recent efforts from a systematic angle to demonstrate the potential broader usage of ICL.
Outcome: The proposed model can learn from examples provided in the prompt, enabling downstream generalization without the need for gradient updates.
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods (2025.acl-long)

Copied to clipboard

Challenge: Existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakkes.
Approach: They propose a large-scale speech deepfake dataset that includes over 3 million deepfak samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools.
Outcome: The proposed dataset includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools.
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Automatic speaker verification systems are facing escalating challenges due to deepfake attacks.
Approach: They propose a Urdu deepfake audio dataset for deepfak detection focusing on two spoofing attacks – Tacotron and VITS TTS.
Outcome: The proposed dataset evaluates two spoofing attacks in Urdu with a human evaluation to gauge whether people are able to distinguish deepfake audios from real (bonafide) audios.
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages (2026.findings-acl)

Copied to clipboard

Challenge: Speech deepfakes are highly realistic and can generate a few seconds of recorded speech.
Approach: They propose an ALM that integrates semantic and prosodic representations from Whisper and TRILLsson to generate a speech deepfake dataset.
Outcome: The proposed framework outperforms existing ALMs on the ICF benchmark in Indic languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations