Papers by Siddhant Arora
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback (2026.findings-acl)
Copied to clipboard
Siddhant Arora, Jinchuan Tian, Jiatong Shi, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe
| Challenge: | Existing studies on reinforcement learning from human or AI feedback have focused on semantic rewards at the utterance level. |
| Approach: | They propose a multi-reward RLAIF framework for speech-in/speech-out dialogue systems . they combine semantic, audio-quality, and emotion-consistency rewards . |
| Outcome: | The proposed framework improves speech-in/speech-out dialogue system quality . it combines semantic, audio-quality, and emotion-consistency rewards . the proposed framework is available to download from the cdc. |
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions (2024.naacl-long)
Copied to clipboard
Siddhant Arora, Hayato Futami, Jee-weon Jung, Yifan Peng, Roshan Sharma, Yosuke Kashiwagi, Emiru Tsunoo, Karen Livescu, Shinji Watanabe
| Challenge: | Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model’s behavior and surpassing performance of task-specific models. |
| Approach: | They adapt a pre-trained automatic speech recognition model to additional tasks using single-token task specifiers. |
| Outcome: | The proposed model can generalize to new datasets and languages for seen task types. |
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner (2026.acl-short)
Copied to clipboard
Guan-Ting Lin, Shih-Yun Shan Kuan, Jiatong Shi, Kai-Wei Chang, Siddhant Arora, Shinji Watanabe, Hung-yi Lee
| Challenge: | Full-duplex speech agents are often half-duplice, alternating turns between user and system. |
| Approach: | They propose a streaming framework that integrates with an examiner that enforces staged goals under two pacing setups. |
| Outcome: | The framework reports fluency, multi-turn instruction following, and task-specific competence. |
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)
Copied to clipboard
Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O’Toole, Siddhant Arora, Thomas Norton, Norman Sadeh, Shomir Wilson
| Challenge: | a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection . |
| Approach: | They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
| Outcome: | The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding (2026.findings-acl)
Copied to clipboard
Masao Someki, Chien-yu Huang, Siddhant Arora, Samuele Cornell, Markus Müller, Nathan Susanj, Rupak Vignesh Swaminathan, Grant Strimel, Jing Liu, Shinji Watanabe
| Challenge: | Long-form audio understanding poses significant challenges due to the extreme length of audio sequences and the need to reason over heterogeneous acoustic cues distributed over time. |
| Approach: | They propose a retrieval-augmented generation framework for scalable long-form audio understanding . planRAG-Audio explicitly plans which modalities and temporal spans are required for a given query . |
| Outcome: | Experiments show that planRAG-Audio reduces the length of inputs for long-form audio models . the proposed framework can efficiently reason over long-term speech data . |
Memory-QA: Answering Recall Questions Based on Multimodal Memories (2025.emnlp-main)
Copied to clipboard
Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiun-Zu Kuo, Jiayang Xu, Aaron Colak, Xin Luna Dong
| Challenge: | Memory-QA is a real-world task that involves answering recall questions about visual content from previously stored multimodal memories. |
| Approach: | They propose a memory-QA task that involves answering recall questions about visual content from previously stored multimodal memories. |
| Outcome: | The proposed solution improves memory recording, compression, storage, and search accuracy over state-of-the-art solutions. |
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding Tasks (2023.acl-long)
Copied to clipboard
Suwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad, Felix Wu, Roshan S Sharma, Wei-Lun Wu, Hung-yi Lee, Karen Livescu, Shinji Watanabe
| Challenge: | Spoken language understanding (SLU) tasks have received little attention and resources compared to lower-level tasks like speech and speaker recognition. |
| Approach: | They propose annotated SLU benchmark tasks based on freely available speech data to complement existing benchmarks and address gaps in the evaluation landscape. |
| Outcome: | The proposed benchmarks complement existing benchmarks and address gaps in the evaluation landscape. |
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to connectionist temporal classification (CTC) are based on pre-trained language models (LMs) |
| Approach: | They propose a formulation of connectionist temporal classification that relaxes the conditional independence assumptions used in conventional CTC and incorporates linguistic knowledge through explicit output dependency. |
| Outcome: | The proposed model improves over conventional approaches across variations in speaking styles and languages while maintaining CTC’s training efficiency. |
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)
Copied to clipboard
Siddhant Arora, Ankita Pasad, Chung-Ming Chien, Jionghao Han, Roshan Sharma, Jee-weon Jung, Hira Dhamyal, William Chen, Suwon Shon, Hung-yi Lee, Karen Livescu, Shinji Watanabe
| Challenge: | Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech. |
| Approach: | They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs . |
| Outcome: | The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks . |
Token-level Sequence Labeling for Spoken Language Understanding using Compositional End-to-End Models (2022.findings-emnlp)
Copied to clipboard
| Challenge: | End-to-end spoken language understanding systems model sequence labeling as a sequence prediction task causing a divergence from its well-established token-level tagging formulation. |
| Approach: | They propose to model sequence labeling as a sequence prediction task . their systems explicitly separate the added complexity of recognizing spoken mentions from the NLU task of sequence labelling . |
| Outcome: | The proposed systems outperform both cascaded and direct models on a labeling task of named entity recognition across SLU benchmarks. |