Papers by Shin’ichi Satoh

4 papers
Discriminative Learning of Open-Vocabulary Object Retrieval and Localization by Negative Phrase Augmentation (D18-1)

Copied to clipboard

Challenge: Existing methods for object detection only handle pre-specified classes, requiring large amounts of visual samples for training.
Approach: They propose a method to retrieve and localize objects specified by a textual query from one million images in 0.5 seconds with high precision.
Outcome: The proposed method can retrieve and localize objects specified by a textual query from one million images in 0.5 seconds with high precision.
Referring Image Segmentation via Joint Mask Contextual Embedding Learning and Progressive Alignment Network (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for referring image segmentation may encounter limitations in maintaining focus on relevant information during specific stages and rectifying errors propagated from early stages.
Approach: They propose a network that integrates a Learnable Contextual Embedding module and a Progressive Alignment Network to enhance the cascade framework.
Outcome: The proposed network achieves state-of-the-art results on three commonly used benchmarks.
J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)

Copied to clipboard

Challenge: Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora.
Approach: They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions.
Outcome: The proposed model is effective for training models and can be used for future research across a wide range of tasks.
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)

Copied to clipboard

Challenge: Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content.
Approach: They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set .
Outcome: The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations