Papers with Baseline
All-in-One: A Deep Attentive Multi-task Learning Framework for Humour, Sarcasm, Offensive, Motivation, and Sentiment on Memes (2020.aacl-main)
Copied to clipboard
| Challenge: | Empirical results show the efficacy of our proposed multi-task framework over existing state-of-the-art systems. |
| Approach: | They propose a multi-task, multi-modal deep learning framework to solve multiple tasks simultaneously. |
| Outcome: | The proposed framework performs better than existing state-of-the-art systems on a complicated form of information, i.e., memes. |
RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks treat multi-turn conversation and reasoning-intensive retrieval separately, yet real-world information seeking requires both. |
| Approach: | They propose a framework that transforms complex queries into fact-grounded multi-turn dialogues through multi-level validation. |
| Outcome: | The proposed framework outperforms existing systems in a number of domains and can be used to improve multi-turn conversation retrieval. |
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation (2025.findings-naacl)
Copied to clipboard
João Matos, Shan Chen, Siena Kathleen V. Placino, Yingya Li, Juan Carlos Climent Pardo, Daphna Idan, Takeshi Tohyama, David Restrepo, Luis Filipe Nakayama, José María Millet Pascual-Leone, Guergana K Savova, Hugo Aerts, Leo Anthony Celi, An-Kwok Ian Wong, Danielle Bitterman, Jack Gallifant
| Challenge: | Existing multiple-choice question and answer (QA) datasets are text-only and available in a limited subset of languages and countries. |
| Approach: | They propose a multilingual, multimodal benchmarking dataset to evaluate multimodal/vision language models in healthcare. |
| Outcome: | The WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries. |