Papers by Gallil Maimon
Slamming: Training a Speech Language Model on One GPU in a Day (2025.findings-acl)
Copied to clipboard
| Challenge: | *Slam* is a recipe for training high-quality Speech Language Models (SLMs) on a single academic GPU in 24 hours. |
| Approach: | They propose a recipe for training high-quality Speech Language Models on a single academic GPU in 24 hours. |
| Outcome: | The proposed training recipe outperforms predicted compute optimal performance, giving an optimistic view to SLM feasibility. |
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)
Copied to clipboard
| Challenge: | DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Approach: | They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Outcome: | The proposed method outperforms baseline methods on quantitative and qualitative evaluations. |
StressTest: Can YOUR Speech LM Handle the Stress? (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent speech-aware language models (SLMs) have enabled direct audio processing, allowing models to access the full expressive range of spoken language. |
| Approach: | They propose a data generation pipeline that simulates change of meaning implied by stress variation and propose 'stresstest' to evaluate models' ability to distinguish between meanings of speech based on stress pattern. |
| Outcome: | The proposed model outperforms existing models on sentence stress reasoning and detection. |