Papers by Yoshinari Fujinuma

8 papers
Diable: Efficient Dialogue State Tracking as Operations on Tables (2023.findings-acl)

Copied to clipboard

Challenge: Existing systems for dialogue state tracking use the full dialogue history as input and generate the entire state from scratch at each dialogue turn.
Approach: They propose a task formalisation that represents the dialogue state as a table and formalises it as 'table manipulation task' they represent the dialogue as if it were a list with all the slots and generate the entire state from scratch at each dialogue turn.
Outcome: The proposed system outperforms existing systems while maintaining competitive accuracy.
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes.
Approach: They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender.
Outcome: The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning.
A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity (P19-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings encode the meaning of words from different languages into a shared low-dimensional space.
Approach: They show that modularity can be used to improve unsupervised word embeddings . they show that it is useful for low-resource languages where one often has few bilingual pairs .
Outcome: The proposed model improves unsupervised cross-lingual word embeddings on distant language pairs in low-resource settings.
A Multi-Modal Multilingual Benchmark for Document Image Classification (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing document image classification datasets have several limitations and we present two new datasets that overcome these limitations.
Approach: They propose to use two newly curated multilingual datasets that overcome these limitations and propose to develop multilingual Document AI models.
Outcome: The proposed datasets overcome limitations in document image classification and open the door for future research into improving Document AI models.
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used as evaluators but the reliability of the outcomes remains a challenge.
Approach: They show that LLM judge outputs are highly sensitive to pre-defined score ranges and that similar biases exist among models from the same family.
Outcome: The contrastive decoding of LLM judge outputs achieves 11.7% relative improvement in Spearman correlation with human judgments, averaged across score ranges.
Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability (2022.acl-long)

Copied to clipboard

Challenge: Pretrained multilingual models enable zero-shot learning even for unseen languages . current multilingual model covers only a small subset of the world's languages - due to data sparsity, they are not likely to obtain good results for many lowresource languages.
Approach: They ask: how does the number of pretraining languages influence zero-shot learning for unseen languages? do the findings change if the languages used for pretraining are all related?
Outcome: The results show that pretrained models can zero-shot learn for unseen languages even for limited amounts even for low-resource languages.
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation (2024.naacl-short)

Copied to clipboard

Challenge: Document translation is a challenge for machine translation systems that focus on textual content at the sentence level, ignoring global context and visual layout structure.
Approach: They propose a benchmark dataset to evaluate document-level NMT systems . they use visual cues to preserve reading order and contiguous blocks of text .
Outcome: The proposed benchmarks assess document-level NMT systems on the comprehensive task of translating semi-structured documents.
Why Overfitting Isn’t Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries (2020.acl-main)

Copied to clipboard

Challenge: Recent studies only evaluate cross-lingual word embeddings on bilingual lexicon induction (BLI) however, underfitting can hinder generalization to other downstream tasks.
Approach: They retrofit cross-lingual word embeddings to the training dictionary and a synthetic dictionary to improve their results.
Outcome: The proposed method improves accuracy on two downstream tasks, despite underfitting the training dictionary.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations