Papers by Egor Lakomkin

2 papers
KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos (D18-2)

Copied to clipboard

Challenge: KT-Speech-Crawler is an automated dataset building tool for speech recognition.
Approach: They propose an approach for automatic dataset construction for speech recognition by crawling YouTube videos.
Outcome: The proposed algorithm can obtain 150 hours of transcribed speech in a day with an estimated 3.5% word error rate.
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs (2024.naacl-long)

Copied to clipboard

Challenge: a new model for speech processing and reasoning uses curated data instead of text.
Approach: They extend the instruction-tuned Llama-2 model with end-to-end speech processing and reasoning abilities without using any carefully curated paired data.
Outcome: The proposed model outperforms or outperfects existing models on synthesized and recorded speech QA tests.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations