Papers by Yuta Nakashima

11 papers
Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers (2021.acl-srw)

Copied to clipboard

Challenge: a growing body of research has been focused on what attention heads learn during the pre-training of visual grounded language models.
Approach: They propose to use visual grounding to supervise attention directly to learn visual ground.
Outcome: The proposed method improves the performance of a state-of-the-art visual grounded language model on vision-and-language tasks.
Efficient Vocabulary Reduction for Small Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult.
Approach: They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance.
Outcome: The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance.
A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently.
Approach: They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity.
Outcome: The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives.
WRIME: A New Dataset for Emotional Intensity Estimation with Subjective and Objective Annotations (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies on emotion analysis use subjective emotional intensity labels by the writers and objective ones by the readers.
Approach: They annotate 17,000 SNS posts with both the writer's subjective emotional intensity and the reader's objective emotional intensity to construct a Japanese emotion analysis dataset.
Outcome: The results show that the reader cannot fully detect the emotions of the writer, especially anger and trust.
iParaphrasing: Extracting Visually Grounded Paraphrases via an Image (C18-1)

Copied to clipboard

Challenge: iParaphrasing extracts visually grounded paraphrases, which are different phrasal expressions describing the same visual concept in an image.
Approach: They propose a task to extract visually grounded paraphrases from images . they propose to model the similarity between the extracted VGPs using existing methods .
Outcome: The proposed task extracts visually grounded paraphrases from images . the proposed method has the potential to improve multimodal language and image tasks .
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training (2025.coling-main)

Copied to clipboard

Challenge: Recent approaches for visually-rich document understanding use manually annotated semantic groups.
Approach: They propose a new variant of the VrDU task that does not use manually annotated semantic groups.
Outcome: The proposed method improves on the existing methods while sacrificing performance.
Constructing a Public Meeting Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora are created from text that has already been digitized.
Approach: They propose a full pipeline of analysis of a large corpus about a century of public meeting in historical Australian news papers, from construction to visual exploration.
Outcome: The proposed method achieves a high recall rate and an F-score of 87.8% on a historical Australian newspaper database.
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes (2024.emnlp-main)

Copied to clipboard

Challenge: Traditional approaches only target labeled attributes, ignoring biases from unlabeled ones.
Approach: They propose a method that ensures protected group independence from all attributes and mitigates inpainting biases through data filtering.
Outcome: The proposed approach achieves an average reduction of 46.1% in leakage-based bias metrics for multi-label classification and 74.8% for image captioning.
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences (2025.acl-industry)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have transformed image captioning . existing evaluations lack standardized criteria and a standardized evaluation framework .
Approach: They propose a leaderboard for evaluating detailed captions that addresses three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations.
Outcome: The proposed model evaluates caption quality, descriptiveness, risks, and societal biases while tailoring criteria to user preferences.
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text.
Approach: They compare standard-format captions and recent GCE processes from the perspectives of gender bias and hallucination.
Outcome: The proposed methods amplify gender bias by 30.9% and increase hallucination by 59.5%.
Emotional Intensity Estimation based on Writer’s Personality (2022.aacl-srw)

Copied to clipboard

Challenge: Existing emotion analysis models are difficult to accurately estimate the writer’s subjective emotions behind the text.
Approach: They propose a method for personalized emotional intensity estimation based on a writer's personality test for Japanese SNS posts.
Outcome: The proposed method improves on the existing method and the proposed hybrid model achieved state-of-the-art performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations