Papers by Yuta Nakashima
Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers (2021.acl-srw)
Copied to clipboard
| Challenge: | a growing body of research has been focused on what attention heads learn during the pre-training of visual grounded language models. |
| Approach: | They propose to use visual grounding to supervise attention directly to learn visual ground. |
| Outcome: | The proposed method improves the performance of a state-of-the-art visual grounded language model on vision-and-language tasks. |
Efficient Vocabulary Reduction for Small Language Models (2025.coling-industry)
Copied to clipboard
| Challenge: | Large language models (LLMs) have high computational costs and energy consumption, making their deployment in industrial settings difficult. |
| Approach: | They propose a small language model that compresses the embedding layer and reduces model size without significant loss of performance. |
| Outcome: | The proposed model reduces the embedding layer while maintaining performance while improving accuracy and performance. |
A Japanese Dataset for Subjective and Objective Sentiment Polarity Classification in Micro Blog Domain (2022.lrec-1)
Copied to clipboard
Haruya Suzuki, Yuto Miyauchi, Kazuki Akiyama, Tomoyuki Kajiwara, Takashi Ninomiya, Noriko Takemura, Yuta Nakashima, Hajime Nagahara
| Challenge: | Existing studies on emotion analysis have studied the analysis of basic emotions and sentiment polarity independently. |
| Approach: | They extend the WRIME dataset with basic emotion intensity from both the writer's subjective and reader's perspective to include the Japanese sentiment polarity. |
| Outcome: | The proposed dataset is the first large-scale corpus to annotate both basic emotions and sentiment polarity labels from both the writer’s and reader’s perspectives. |
WRIME: A New Dataset for Emotional Intensity Estimation with Subjective and Objective Annotations (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing studies on emotion analysis use subjective emotional intensity labels by the writers and objective ones by the readers. |
| Approach: | They annotate 17,000 SNS posts with both the writer's subjective emotional intensity and the reader's objective emotional intensity to construct a Japanese emotion analysis dataset. |
| Outcome: | The results show that the reader cannot fully detect the emotions of the writer, especially anger and trust. |
iParaphrasing: Extracting Visually Grounded Paraphrases via an Image (C18-1)
Copied to clipboard
| Challenge: | iParaphrasing extracts visually grounded paraphrases, which are different phrasal expressions describing the same visual concept in an image. |
| Approach: | They propose a task to extract visually grounded paraphrases from images . they propose to model the similarity between the extracted VGPs using existing methods . |
| Outcome: | The proposed task extracts visually grounded paraphrases from images . the proposed method has the potential to improve multimodal language and image tasks . |
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training (2025.coling-main)
Copied to clipboard
| Challenge: | Recent approaches for visually-rich document understanding use manually annotated semantic groups. |
| Approach: | They propose a new variant of the VrDU task that does not use manually annotated semantic groups. |
| Outcome: | The proposed method improves on the existing methods while sacrificing performance. |
Constructing a Public Meeting Corpus (2020.lrec-1)
Copied to clipboard
Koji Tanaka, Chenhui Chu, Haolin Ren, Benjamin Renoust, Yuta Nakashima, Noriko Takemura, Hajime Nagahara, Takao Fujikawa
| Challenge: | Existing corpora are created from text that has already been digitized. |
| Approach: | They propose a full pipeline of analysis of a large corpus about a century of public meeting in historical Australian news papers, from construction to visual exploration. |
| Outcome: | The proposed method achieves a high recall rate and an F-score of 87.8% on a historical Australian newspaper database. |
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes (2024.emnlp-main)
Copied to clipboard
Yusuke Hirota, Jerone Andrews, Dora Zhao, Orestis Papakyriakopoulos, Apostolos Modas, Yuta Nakashima, Alice Xiang
| Challenge: | Traditional approaches only target labeled attributes, ignoring biases from unlabeled ones. |
| Approach: | They propose a method that ensures protected group independence from all attributes and mitigates inpainting biases through data filtering. |
| Outcome: | The proposed approach achieves an average reduction of 46.1% in leakage-based bias metrics for multi-label classification and 74.8% for image captioning. |
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences (2025.acl-industry)
Copied to clipboard
Yusuke Hirota, Boyi Li, Ryo Hachiuma, Yueh-Hua Wu, Boris Ivanovic, Marco Pavone, Yejin Choi, Yu-Chiang Frank Wang, Yuta Nakashima, Chao-Han Huck Yang
| Challenge: | Large Vision-Language Models (LVLMs) have transformed image captioning . existing evaluations lack standardized criteria and a standardized evaluation framework . |
| Approach: | They propose a leaderboard for evaluating detailed captions that addresses three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations. |
| Outcome: | The proposed model evaluates caption quality, descriptiveness, risks, and societal biases while tailoring criteria to user preferences. |
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. |
| Approach: | They compare standard-format captions and recent GCE processes from the perspectives of gender bias and hallucination. |
| Outcome: | The proposed methods amplify gender bias by 30.9% and increase hallucination by 59.5%. |
Emotional Intensity Estimation based on Writer’s Personality (2022.aacl-srw)
Copied to clipboard
| Challenge: | Existing emotion analysis models are difficult to accurately estimate the writer’s subjective emotions behind the text. |
| Approach: | They propose a method for personalized emotional intensity estimation based on a writer's personality test for Japanese SNS posts. |
| Outcome: | The proposed method improves on the existing method and the proposed hybrid model achieved state-of-the-art performance. |