Papers by Shohei Higashiyama

8 papers
Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation (2024.findings-eacl)

Copied to clipboard

Challenge: et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing.
Approach: They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions.
Outcome: The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities)
A Text Embedding Model with Contrastive Example Mining for Point-of-Interest Geocoding (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on coarse-grained locations, but we focus on fine-grain POIs, which have many candidates with similar names.
Approach: They develop a text embedding-based geocoding model and investigate (1) entry encoding representations and (2) hard negative mining approaches suitable for enhancing the model’s disambiguation ability.
Outcome: The proposed model significantly improves its disambiguation ability and entry encoding representations.
Incorporating Word Attention into Character-Based Word Segmentation (N19-1)

Copied to clipboard

Challenge: Word segmentation models are used to minimize the effort in feature engineering.
Approach: They propose a character-based model that learns the importance of multiple candidate words for a corresponding character on the basis of an attention mechanism and makes use of it for segmentation decisions.
Outcome: The proposed model outperforms the state-of-the-art models on Japanese and Chinese benchmark datasets.
Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languages (2025.findings-emnlp)

Copied to clipboard

Challenge: Lexical normalization research has sought to tackle the challenge of processing informal expressions in user-generated text.
Approach: They focus on Japanese normalization and developing methods based on state-of-the-art pre-trained models .
Outcome: The proposed methods achieve high accuracy and efficiency across multiple evaluation perspectives.
Overview of the 6th Workshop on Asian Translation (D19-52)

Copied to clipboard

Challenge: The 6th workshop on Asian translation (WAT2019) was held in hong kong, hongkong, and hong kong.
Approach: They present the results of the shared tasks from the 6th workshop on Asian translation (WAT2019) 25 teams participated in the shared task and 10 research paper submissions were accepted .
Outcome: The results of the 6th workshop on Asian translation (WAT2019) include JaEn, JaZh scientific paper translation subtasks, Ja'En, ja'Ko, Ja’En patent translation sub tasks, Hi'En and My'En patent subtask and Ru'Ja news commentary translation task.
Automatic Generation of a Compositional QA Benchmark for Geospatial Reasoning under Spatial and Entity Constraints (2026.eacl-srw)

Copied to clipboard

Challenge: Recent advances in large language models have enhanced their ability to perform reasoning tasks that integrate linguistic, visual, and factual information.
Approach: They propose a method for constructing compositional geographic question answering datasets that jointly consider spatial and entity constraints.
Outcome: The proposed method performs well on questions involving rich entity grounding, but its accuracy drops on quantitative spatial reasoning questions.
Graph-Structured Trajectory Extraction from Travelogues (2025.acl-long)

Copied to clipboard

Challenge: Existing studies treat travelogues as sequences of visited locations, but they lack a benchmark dataset.
Approach: They propose to represent the trajectory as a graph that can capture the hierarchy as well as the visiting order and construct a benchmark dataset for the extraction.
Outcome: The proposed dataset shows that even naive baseline systems can predict visited locations and the visiting order between them, while it is more challenging to predict the hierarchical relations.
User-Generated Text Corpus for Evaluating Japanese Morphological Analysis and Lexical Normalization (2021.naacl-main)

Copied to clipboard

Challenge: Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text.
Approach: They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information.
Outcome: The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations