Papers by Muskaan Singh

7 papers
HCFD: A Benchmark for Audio Deepfake Detection in Healthcare (2026.findings-acl)

Copied to clipboard

Challenge: a new task for detecting codec-fakes under pathological speech conditions is presented . we focus on codec based synthetic speech since neural codec decoding is a core building block in speech generation pipelines.
Approach: They propose a new task for detecting codec-fakes under pathological speech conditions . they focus on codec based synthetic speech since neural codec decoding is a core building block in speech pipelines .
Outcome: The proposed framework outperforms speech-based models on Healthcare CodecFake . it achieves the strongest performance on the task across clinical conditions and codecs .
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
DIVINE : Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment (2026.eacl-long)

Copied to clipboard

Challenge: Existing frameworks for diagnosing oro-facial neurological disorders are based on shared and modality-specific representations, but they are not fully disentangled.
Approach: They propose a fully disentangled multimodal framework that captures vocal and facial cues.
Outcome: The proposed framework achieves 98.26% accuracy and 97.51% F1-score under modality-constrained scenarios.
ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Summarization is a challenging problem, and it is difficult to create, correct, and evaluate the summaries manually.
Approach: They propose an open-source tool for meeting annotation, alignment, and evaluation . the tool aims to provide an efficient and clear interface for fast annotation .
Outcome: The proposed tool is open-source and installable from PyPI.
Prosody as Supervision: Bridging the Non-Verbal–Verbal for Multilingual Speech Emotion Recognition (2026.acl-long)

Copied to clipboard

Challenge: Existing paradigms for low-resource multilingual speech emotion recognition rely on labeled verbal speech and lack cross-lingual transfer.
Approach: They propose a paralinguistic supervision paradigm for low-resource multilingual speech emotion recognition that leverages non-verbal vocalizations to exploit prosody-centric emotion cues.
Outcome: The proposed framework outperforms Euclidean counter parts and strong SSL baselines in the language-based evaluation of low-resource multilingual speech emotion recognition (LRM-SER)
ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and Czech (2022.lrec-1)

Copied to clipboard

Challenge: Automated minuting is a rather unstructured writing activity and can be difficult due to a variety of factors including the quality of automatic speech recorders, availability of public meeting data, subjective knowledge of the minuter, etc.
Approach: They propose a dataset on automatic minuting which includes transcripts from ASRs and minuted by annotators.
Outcome: The proposed dataset covers more than 160 hours of meeting content.
Bridging Attribution and Open-Set Detection using Graph-Augmented Instance Learning in Synthetic Speech (2026.eacl-long)

Copied to clipboard

Challenge: Synthetic speech detection is a critical part of safeguarding digital communication, enabling systems to identify and mitigate the risks posed by highly realistic, machine-generated voices.
Approach: They propose a framework that combines SFMs with graph-based modeling and open-set generalization to capture meaningful relationships between utterances and recognize speech that doesn’t belong to any known generator.
Outcome: The proposed framework improves performance across both tasks, with Mamba-based embeddings delivering particularly strong results.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations