Papers by Kentaro Yamada

3 papers
Action Inference for Destination Prediction in Vision-and-Language Navigation (2024.acl-srw)

Copied to clipboard

Challenge: Existing work on vision-and-language navigation focuses on spatial reasoning and semantic grounding of visual information, but there is still scope for improvement.
Approach: They propose a VLN task of destination prediction for picking up a pedestrian that requires action inference from a crowd-sourced dataset.
Outcome: The proposed model can reason about the effect of the next action and the next on the destination to a certain extent.
Transformer-based Lexically Constrained Headline Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing automatic headline generation methods cannot include a given phrase in the generated headline.
Approach: They propose a Transformer-based method that guarantees to include a given phrase in a generated headline.
Outcome: The proposed method achieves ROUGE scores comparable to previous methods with Japanese news corpus.
GesNavi: Gesture-guided Outdoor Vision-and-Language Navigation (2024.eacl-srw)

Copied to clipboard

Challenge: Existing datasets for outdoor Vision-and-Language Navigation (VLN) tasks do not include verbal instructions for communicating with mobility.
Approach: They propose a dataset for gesture-guided outdoor VLN instructions with demonstrative expressions that incorporates gestures and linguistic commands.
Outcome: The proposed datasets are compared against existing datasets and analysed in detail.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations