Papers by Shoko Wakamiya

10 papers
Single-Agent Generation Surpasses Multi-Agent Systems in Semantic Diversity (2026.findings-acl)

Copied to clipboard

Challenge: Multi-Agent Systems (MAS) are used to improve reasoning diversity and robustness by simulating interactions among agents with distinct roles.
Approach: They find that a Multi-Output strategy produces the highest diversity without degrading logical validity.
Outcome: The proposed approach outperforms multi-agent systems in semantic diversity . the results point to a more efficient and effective way to expand diversity - the authors say .
Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation (2024.findings-eacl)

Copied to clipboard

Challenge: et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing.
Approach: They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions.
Outcome: The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities)
A Herd of Language Models Makes a Better Zero-shot Annotator for Clinical Named Entity Recognition (2026.findings-acl)

Copied to clipboard

Challenge: Clinical named entity recognition (NER) is a core task in clinical NLP.
Approach: They propose a label-modeling method for M**ulti-LLM **A**nnotation using **R**epresentation learning to capture contextual similarity.
Outcome: The proposed method improves the average F1 score by 8.6% over zero-shot baselines while reducing annotation costs.
Enhancing Hate Speech Classifiers through a Gradient-assisted Counterfactual Text Generation Strategy (2025.findings-emnlp)

Copied to clipboard

Challenge: Strong attribute control can distort meaning, while prioritizing semantic preservation may weaken attribute alignment.
Approach: They propose a method that restricts accepted samples to text meeting a minimum BERTScore threshold and applies gradient-assisted proposal generation to improve attribute alignment.
Outcome: a new method for counterfactual text generation improves attribute alignment and semantic preservation . the proposed method achieved the best macro F1-score in two of three test sets .
Annotation-Scheme Reconstruction for “Fake News” and Japanese Fake News Dataset (2022.lrec-1)

Copied to clipboard

Challenge: Contemporary research focuses on the factuality aspect of the news, but this aspect alone is insufficient to explain “fake news.”
Approach: They propose to use Japanese fake news datasets to classify whether news content is false . they propose to do this by using existing fake news data to investigate fake news .
Outcome: The proposed scheme will provide an in-depth understanding of fake news in Japan and other languages.
Exploring LLM Annotation for Adaptation of Clinical Information Extraction Models under Data-sharing Restrictions (2025.findings-acl)

Copied to clipboard

Challenge: In-hospital text data often contains valuable clinical information, yet fine-tuned small language models (SLMs) for information extraction remain challenging due to differences in formatting and vocabulary across institutions.
Approach: They leverage large language models to annotate the target domain data for adaptation . they use in-hospital text data to extract clinical information .
Outcome: The proposed model outperforms manual annotation on four clinical information extraction tasks with a larger number of annotated data.
Offensive Language Detection on Video Live Streaming Chat (2020.coling-main)

Copied to clipboard

Challenge: a prototype of a live chat room that detects offensive expressions in live streaming chats is presented . offensive expression detection on social media platforms can provide more protection for users .
Approach: They propose a live chat room that detects offensive expressions in live streaming chats in real time . they used a dataset from Twitch to analyze offensive expression patterns .
Outcome: The proposed chat room detects offensive expressions in live streaming chats in real time.
MultiMSD: A Corpus for Multilingual Medical Text Simplification from Online Medical References (2025.findings-acl)

Copied to clipboard

Challenge: Medical texts contain technical terms, and non-experts often cannot use information effectively.
Approach: They propose a method for training medical text simplification models to actively paraphrase medical terms.
Outcome: The proposed method improves the performance of medical text simplification in nine languages.
RecordTwin: Towards Creating Safe Synthetic Clinical Corpora (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to generate high-quality synthetic corpus from clinical documents require learning from the original clinical documents.
Approach: They propose a method to generate synthetic corpus from clinical documents using a large language model.
Outcome: The proposed method generates synthetic documents from in-hospital clinical documents.
J-MeDic: A Japanese Disease Name Dictionary based on Real Clinical Usage (L18-1)

Copied to clipboard

Challenge: a study finds that medical texts are written mostly in natural language, requiring NLP for medical texts.
Approach: They develop a Japanese disease name dictionary to fill the gap between medical names and clinical words . they allocated the standard disease code to the names using manual, semi-automatic or automatic methods .
Outcome: The Japanese disease name dictionary fills the gap between standard medical names and real clinical words . the study found that 55.3% of the names covered by the dictionary were SDNs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations