Papers by Tyler Miller
Text-Free Image-to-Speech Synthesis Using Learned Segmental Units (2021.acl-long)
Copied to clipboard
| Challenge: | Existing models for synthesising fluent, natural-sounding spoken audio captions do not require natural language text as an intermediate representation or source of supervision. |
| Approach: | They propose a model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supervision. |
| Outcome: | The proposed model captures diverse visual semantics of images and can replace text with a set of discrete, sub-word speech units. |