Papers by Shiva Sundaram

1 papers
Multimodal and Multiresolution Speech Recognition with Transformers (2020.acl-main)

Copied to clipboard

Challenge: Existing audio visual automatic speech recognition systems rely on audio input to produce transcriptions.
Approach: They propose an audio visual automatic speech recognition system using a transformer-based architecture and incorporate a multitask training criterion for multiresolution ASR.
Outcome: The proposed system can generate character and subword transcriptions with visual information.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations