Papers by Shogo Okada
Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users’ Interest Level (L18-1)
Copied to clipboard
Masahiro Araki, Sayaka Tomimasu, Mikio Nakano, Kazunori Komatani, Shogo Okada, Shinya Fujie, Hiroaki Sugiyama
| Challenge: | a group of researchers is building a corpus for evaluating elements of multimodal dialogue systems. |
| Approach: | They propose to build a corpus for evaluating elements of the multimodal dialogue system . they use the Wizard of Oz method to record chat dialogue data between a human and a virtual agent . |
| Outcome: | The proposed method annotates chat dialogue data between a human and a virtual agent and measures their interest level in the data. |
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines (2025.naacl-long)
Copied to clipboard
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Cheng Ching Lam, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Christabelle Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo
| Challenge: | Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts. |
| Approach: | They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset. |
| Outcome: | The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages. |