Papers by Jean Lahoud

3 papers
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches do not emphasize step-wise problem-solving.
Approach: They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step.
Outcome: The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling.
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM (2025.findings-acl)

Copied to clipboard

Challenge: Existing speech-enabled LLMs degrade conversational quality by modifying the LLM, compromising its linguistic capabilities.
Approach: They propose a lightweight 30M-parameter, LLM-agnostic, autoregressive streaming TTS system that generates high-quality speech with low latency.
Outcome: The proposed system achieves a significantly lower word error rate compared to speech-enabled LLMs while operating at comparable latency.
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)

Copied to clipboard

Challenge: Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content.
Approach: They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set .
Outcome: The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations