Papers by Kazuhiro Nakadai
Multilingual Gloss-free Sign Language Translation: Towards Building a Sign Language Foundation Model (2025.acl-short)
Copied to clipboard
| Challenge: | Existing studies focus on translating a single SL into a spoken language (one-to-one SLT) however, multilingual SLT remains unexplored due to language conflicts and alignment difficulties across SLs and spoken languages. |
| Approach: | They propose a multilingual gloss-free model that can be used to translate a single SL into a spoken language and generate a token-level SL identification and spoken text. |
| Outcome: | The proposed model supports 10 SLs and handles one-to-one, many-to-1, and many- to-many SLT tasks. |
Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions (L18-1)
Copied to clipboard
| Challenge: | Existing technologies for CG-supported data display are not able to depict all relevant features of a natural signing sequence such as facial expression, spatial references or inter-sign movement. |
| Approach: | They collected a corpus of Japanese Sign Language sentences for deep neural network learning. |
| Outcome: | The proposed model could be used to train language features in Japanese Sign Language (JSL) |
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing multi-turn methods for large language models exploit conversational context to bypass safety constraints gradually. |
| Approach: | They propose a framework of five conversation patterns to construct multi-turn jailbreaks through natural dialogue. |
| Outcome: | The proposed framework exploits conversational contexts to construct multi-turn jailbreaks . it reveals that models exhibit distinct weakness profiles and model families share similar failure modes . |
Improvement in Sign Language Translation Using Text CTC Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Current sign language translation (SLT) approaches rely on gloss-based supervision with Connectionist Temporal Classification (CTC) limiting their ability to handle non-monotonic alignments between sign language video and spoken text. |
| Approach: | They propose a method that integrates CTC/Attention with the attention mechanism during decoding and integrates it with the sign language video and spoken text. |
| Outcome: | The proposed method outperforms the pure-attention baseline and achieves comparable results to state-of-the-art methods. |