Papers by Fabian Retkowski
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)
Copied to clipboard
| Challenge: | despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation . |
| Approach: | They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture . |
| Outcome: | The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs. |
Summarizing Speech: A Comprehensive Survey (2025.emnlp-main)
Copied to clipboard
Fabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau, Shinji Watanabe, Jan Niehues, Alexander Waibel
| Challenge: | Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice. |
| Approach: | They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets. |
| Outcome: | The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets. |
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for text segmentation are small in scale, synthesized, or only contain well-structured documents. |
| Approach: | They propose a benchmark YTSeg focusing on spoken content that is unstructured and unstructures . they also introduce an efficient hierarchical segmentation model MiniSeg that outperforms state-of-the-art benchmarks. |
| Outcome: | The proposed model outperforms state-of-the-art models on unstructured spoken content . the proposed model could be used for "smart chaptering" tasks . |
Zero-Shot Strategies for Length-Controllable Summarization (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large language models struggle with precise length control, particularly in zero-shot settings. |
| Approach: | They propose to use length approximation, target adjustment, sample filtering and automated revisions to improve LLMs' length control capabilities. |
| Outcome: | The proposed methods improve length control in large language models while maintaining or enhancing summary quality without the need for model fine-tuning or architectural changes. |
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation (2023.emnlp-demo)
Copied to clipboard
Christian Huber, Tu Anh Dinh, Carlos Mullov, Ngoc-Quan Pham, Thai Binh Nguyen, Fabian Retkowski, Stefan Constantin, Enes Ugan, Danni Liu, Zhaolin Li, Sai Koneru, Jan Niehues, Alexander Waibel
| Challenge: | a framework to evaluate low-latency speech translations is currently only limited to specific aspects and is not able to compare different approaches. |
| Approach: | They propose a framework to perform and evaluate low-latency speech translation in realistic conditions. |
| Outcome: | The proposed framework evaluates various aspects of low-latency speech translation under realistic conditions. |
BOOM: Beyond Only One Modality KIT’s Multimodal Multilingual Lecture Companion (2026.eacl-demo)
Copied to clipboard
Sai Koneru, Fabian Retkowski, Christian Huber, Lukas Hilgert, Seymanur Akti, Enes Yavuz Ugan, Alexander Waibel, Jan Niehues
| Challenge: | a multimodal multilingual lecture companion is needed to preserve lecture content in its entirety . globalization of education and rapid growth of online learning have made localizing educational content a challenge . |
| Approach: | They propose a multimodal multilingual lecture companion that translates lecture audio and slides to produce synchronized outputs across three modalities. |
| Outcome: | The proposed solution preserves the original content in its entirety while preserving translations across three modalities. |