Papers by Alexander Waibel

14 papers
Few-Shot Learning Translation from New Languages (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows strong transfer learning capability to unseen languages in sequence-to-sequence neural networks . current transfer learning methods require much less downstream task data than would otherwise be required.
Approach: They first train word embeddings models on varying amounts of data and plug them into a machine translation model.
Outcome: The proposed model can learn Flores with only 500 parallel sentences and 31,250 sentences of monolingual data, and it can exceed 15 BLEU on unseen languages.
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)

Copied to clipboard

Challenge: despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation .
Approach: They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture .
Outcome: The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs.
Summarizing Speech: A Comprehensive Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Podcasts and other audiovisual content are becoming more and more a part of everyday communication and the digital age is changing from text to voice.
Approach: They synthesize the current state of the field and highlight the need for realistic evaluation benchmarks and multilingual datasets.
Outcome: The proposed frameworks are based on evaluation protocols and datasets and highlight the need for realistic benchmarks and multilingual datasets.
Fluent Translations from Disfluent Speech in End-to-End Speech Translation (N19-1)

Copied to clipboard

Challenge: Disfluency removal is an intermediate step between speech recognition and machine translation (MT) with the rise of end-to-end speech translation systems, disfluency recognition and removal needs to be incorporated into the model architectures or handled as a post-processing step.
Approach: They propose to use a sequence-to-sequence model to translate from noisy, disfluent speech to fluent text with disfluencies removed using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
Outcome: The proposed model generates fluent translations from disfluent speech using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions (2024.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for text segmentation are small in scale, synthesized, or only contain well-structured documents.
Approach: They propose a benchmark YTSeg focusing on spoken content that is unstructured and unstructures . they also introduce an efficient hierarchical segmentation model MiniSeg that outperforms state-of-the-art benchmarks.
Outcome: The proposed model outperforms state-of-the-art models on unstructured spoken content . the proposed model could be used for "smart chaptering" tasks .
Zero-Shot Strategies for Length-Controllable Summarization (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models struggle with precise length control, particularly in zero-shot settings.
Approach: They propose to use length approximation, target adjustment, sample filtering and automated revisions to improve LLMs' length control capabilities.
Outcome: The proposed methods improve length control in large language models while maintaining or enhancing summary quality without the need for model fine-tuning or architectural changes.
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation (2023.emnlp-demo)

Copied to clipboard

Challenge: a framework to evaluate low-latency speech translations is currently only limited to specific aspects and is not able to compare different approaches.
Approach: They propose a framework to perform and evaluate low-latency speech translation in realistic conditions.
Outcome: The proposed framework evaluates various aspects of low-latency speech translation under realistic conditions.
DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Code-switched (CSW) speech is a linguistic phenomenon that occurs when spoken utterances switch languages between sentences.
Approach: They propose to use a dataset to evaluate German-English CSW speech . they show that the dataset includes splits with varying degrees of CSW .
Outcome: The proposed dataset includes spontaneous speech from diverse domains, enabling realistic CSW evaluation in German-English.
BOOM: Beyond Only One Modality KIT’s Multimodal Multilingual Lecture Companion (2026.eacl-demo)

Copied to clipboard

Challenge: a multimodal multilingual lecture companion is needed to preserve lecture content in its entirety . globalization of education and rapid growth of online learning have made localizing educational content a challenge .
Approach: They propose a multimodal multilingual lecture companion that translates lecture audio and slides to produce synchronized outputs across three modalities.
Outcome: The proposed solution preserves the original content in its entirety while preserving translations across three modalities.
KIT Lecture Translator: Multilingual Speech Translation with One-Shot Learning (C18-2)

Copied to clipboard

Challenge: In today's globalized world, communication is difficult and often the language barrier still prevents communication.
Approach: They have developed a low-latency translation system that is adapted to lectures and covers several language pairs.
Outcome: The proposed system improves performance but also covers several European languages.
Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages (2024.acl-long)

Copied to clipboard

Challenge: Multilingual neural machine translation systems learn to map sentences of different languages into a common representation space.
Approach: They propose a setup where we decouple learning of vocabulary and syntax and train to translate while keeping those word representations frozen.
Outcome: The proposed setup achieves near parity with a supervised setting on the TED domain with varying number of languages seen by the encoder.
Paraphrases as Foreign Languages in Multilingual Neural Machine Translation (P19-2)

Copied to clipboard

Challenge: Unlike previous studies that use paraphrases at the word/phrase level, we train on parallel paraphrase training on closely related languages.
Approach: They train on parallel paraphrases in the style of multilingual Neural Machine Translation (NMT) they train on translations of the whole corpus that are consistent in structure as paraphrase versions at the corpus level.
Outcome: The proposed training on paraphrases outperforms the baselines on two languages and improves lexical choice and entropy.
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space.
Approach: They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data.
Outcome: The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures.
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are rapidly developing and are becoming more and more useful in scientific tasks.
Approach: They propose to use LLM-as-a-judge to grade LLMs on SciEx to assess their ability on scientific tasks.
Outcome: The proposed benchmarks show that the LLMs perform decently on free-form exams, achieving 0.948 Pearson correlation with expert grading.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations