Papers by Taichi Nishimura

5 papers
Visual Grounding Annotation of Recipe Flow Graph (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies have ground visual observations with procedural texts with graphs to understand which objects are aligned with textual descriptions.
Approach: They propose to provide visual grounding annotations to recipe flow graphs by adding bounding boxes to image sequences of recipes and annotating two types of event attributes with each bounding box.
Outcome: The proposed dataset gives visual grounding with workflow’s contextual information between procedural text and visual observation in an indirect manner.
Visual Recipe Flow: A Dataset for Learning Visual State Changes of Objects with Recipe Flows (2022.coling-1)

Copied to clipboard

Challenge: a new dataset enables us to learn a cooking action result for each object in a recipe text.
Approach: They propose a multimodal dataset that enables us to learn a cooking action result for each object in a recipe text.
Outcome: The proposed dataset reduces human annotation costs by allowing multimodal information retrieval.
Image Description Dataset for Language Learners (2022.lrec-1)

Copied to clipboard

Challenge: Language learners are limited by the number of texts or speech they are asked to answer . automatic assessment of image descriptions requires a system that depends on both the learner's native language and the target language.
Approach: They propose a dataset that consists of images, their descriptions, and assessment annotations . they propose 'automatic error correction' task that encodes multimodal information from a learner sentence with an image and accurately decodes a corrected sentence.
Outcome: The proposed model can revise errors that cannot be revised without an image.
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing studies have focused on reproducible video moment retrieval and highlight detection . lack of reproducible experiments means that researchers set up individual environments .
Approach: They propose a user-friendly library for reproducible video moment retrieval and highlight detection . they propose MR and highlight retrieval methods that can be used to find specific moments .
Outcome: The proposed library reproduces the reported results in the reference papers.
Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to evaluate geoparsing systems are small-scale and lack coverage of location expressions on general domains.
Approach: They propose a method to construct a large-scale corpus for geoparsing from Wikipedia articles.
Outcome: The proposed method can annotate multiple location expressions with coordinates even with ambiguous expressions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations