Papers by Anil Nelakanti
Interactive Post-Editing for Verbosity Controlled Translation (2022.coling-1)
Copied to clipboard
| Challenge: | Recent machine translation models have shown to excel with aspects of translation quality like adequacy and fluency but these models still suffer notable shortcomings like out-of-domain data, low-resource languages, rare words and longer sentences. |
| Approach: | They propose to use human-in-loop interactive post-editing models to improve translation quality and rephrase the text with a desired style variation. |
| Outcome: | The proposed model achieves BERTScore over state-of-the-art machine translation models while maintaining the desired token-level and verbosity preference. |
Empathic Machines: Using Intermediate Features as Levers to Emulate Emotions in Text-To-Speech Systems (2022.naacl-main)
Copied to clipboard
| Challenge: | a method to control affective prosody of text-to-speech systems is proposed to use phoneme-level intermediate features as levers . DS is used to disentangle features relating to affective proody from those due to acoustics conditions and speaker identity . |
| Approach: | They propose a method to control the emotional prosody of Text to Speech systems by using phoneme-level intermediate features as levers. |
| Outcome: | The proposed method improves over the prior art in emulating emotion in speech . it adds the much-coveted "human touch" in machine dialogue, the authors say . |
ParrotTTS: Text-to-speech synthesis exploiting disentangled self-supervised representations (2024.findings-eacl)
Copied to clipboard
| Challenge: | ParrotTTS can train a multi-speaker variant using transcripts from a single speaker in low resource setup and generalizes to languages not seen while training the self-supervised backbone. |
| Approach: | They propose a modular text-to-speech synthesis model that can train a multi-speaker variant using transcripts from a single speaker. |
| Outcome: | The proposed model outperforms state-of-the-art multi-lingual text-to-speech models using only a fraction of paired data as latter. |