Papers by Kentaro Yamada
Action Inference for Destination Prediction in Vision-and-Language Navigation (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing work on vision-and-language navigation focuses on spatial reasoning and semantic grounding of visual information, but there is still scope for improvement. |
| Approach: | They propose a VLN task of destination prediction for picking up a pedestrian that requires action inference from a crowd-sourced dataset. |
| Outcome: | The proposed model can reason about the effect of the next action and the next on the destination to a certain extent. |
Transformer-based Lexically Constrained Headline Generation (2021.emnlp-main)
Copied to clipboard
Kosuke Yamada, Yuta Hitomi, Hideaki Tamori, Ryohei Sasano, Naoaki Okazaki, Kentaro Inui, Koichi Takeda
| Challenge: | Existing automatic headline generation methods cannot include a given phrase in the generated headline. |
| Approach: | They propose a Transformer-based method that guarantees to include a given phrase in a generated headline. |
| Outcome: | The proposed method achieves ROUGE scores comparable to previous methods with Japanese news corpus. |
GesNavi: Gesture-guided Outdoor Vision-and-Language Navigation (2024.eacl-srw)
Copied to clipboard
| Challenge: | Existing datasets for outdoor Vision-and-Language Navigation (VLN) tasks do not include verbal instructions for communicating with mobility. |
| Approach: | They propose a dataset for gesture-guided outdoor VLN instructions with demonstrative expressions that incorporates gestures and linguistic commands. |
| Outcome: | The proposed datasets are compared against existing datasets and analysed in detail. |