Papers by Mohamed Elhoseiny
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows (2025.emnlp-main)
Copied to clipboard
Kirolos Ataallah, Eslam Mohamed Bakr, Mahmoud Ahmed, Chenhui Gou, Khushbu Pahwa, Jian Ding, Mohamed Elhoseiny
| Challenge: | Existing benchmarks fail to test the full range of cognitive skills needed to process long-form videos . |
| Approach: | They propose a benchmark to evaluate models' ability to process long-form videos rigorously. |
| Outcome: | The benchmark measures the cognitive skills of models in understanding long-form videos . it offers the largest set of question-answer pairs for long video comprehension . |
Towards AI-Assisted Psychotherapy: Emotion-Guided Generative Interventions (2025.emnlp-main)
Copied to clipboard
Kilichbek Haydarov, Youssef Mohamed, Emilio Goldenhersch, Paul OCallaghan, Li-jia Li, Mohamed Elhoseiny
| Challenge: | Large language models (LLMs) lack rich non-verbal emotional cues essential to real-world therapy. |
| Approach: | They propose a multimodal dataset of 1,441 publicly sourced therapy session videos containing both dialogue and non-verbal signals such as facial expressions and vocal tone. |
| Outcome: | The proposed model improves the quality of generated interventions and evaluators misalign with expert assessments in this domain, highlighting the need for human-centered evaluation. |
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages (2024.emnlp-main)
Copied to clipboard
Youssef Mohamed, Runjia Li, Ibrahim Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny
| Challenge: | Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo emphasizes diversity of opinions over languages and cultures. |
| Approach: | They propose a vision-language benchmark that spans 28 languages and encompasses approximately 200,000 annotations. |
| Outcome: | The proposed benchmark spans 28 languages and encompasses approximately 200,000 annotations . the challenge is to build machine learning systems that assign emotional captions to images . |
ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture (2022.emnlp-main)
Copied to clipboard
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny
| Challenge: | ArtELingo is a benchmark and dataset designed to encourage work on diversity across languages and cultures. |
| Approach: | They introduce a benchmark and dataset designed to encourage work on diversity across languages and cultures. |
| Outcome: | The new benchmark and dataset compared artELingo annotations across languages and cultures and found that diversity improves the performance of baseline models. |