Papers by Tanvir Mahmud
OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for audio separation are limited due to over-separation, under-serparation and dependence on predefined training sources. |
| Approach: | They propose a framework that leverages large language models (LLMs) for automated audio separation, eliminating the need for manual intervention and overcoming source limitations. |
| Outcome: | The proposed framework outperforms existing methods in separating new, unseen, and variable sources in real-world mixtures, and is available on github. |