Papers by Shin’ichi Satoh
Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation (D18-1)
Copied to clipboard
| Challenge: | Existing methods for object detection only handle pre-specified classes, requiring large amounts of visual samples for training. |
| Approach: | They propose a method to retrieve and localize objects specified by a textual query from one million images in 0.5 seconds with high precision. |
| Outcome: | The proposed method can retrieve and localize objects specified by a textual query from one million images in 0.5 seconds with high precision. |
Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for referring image segmentation may encounter limitations in maintaining focus on relevant information during specific stages and rectifying errors propagated from early stages. |
| Approach: | They propose a network that integrates a Learnable Contextual Embedding module and a Progressive Alignment Network to enhance the cascade framework. |
| Outcome: | The proposed network achieves state-of-the-art results on three commonly used benchmarks. |
J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)
Copied to clipboard
| Challenge: | Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora. |
| Approach: | They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions. |
| Outcome: | The proposed model is effective for training models and can be used for future research across a wide range of tasks. |
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)
Copied to clipboard
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin’ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan
| Challenge: | Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content. |
| Approach: | They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set . |
| Outcome: | The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu. |