Papers by Jeonghun Yeo
Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Visual speech processing requires context modeling due to the ambiguous nature of lip movements. |
| Approach: | They propose a framework to maximize the context modeling capability by bringing the power of LLMs. |
| Outcome: | The proposed framework maximizes the power of visual speech processing by bringing it to the forefront of the field. |