Papers by Yoshinari Fujinuma
Diable: Efficient Dialogue State Tracking as Operations on Tables (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing systems for dialogue state tracking use the full dialogue history as input and generate the entire state from scratch at each dialogue turn. |
| Approach: | They propose a task formalisation that represents the dialogue state as a table and formalises it as 'table manipulation task' they represent the dialogue as if it were a list with all the slots and generate the entire state from scratch at each dialogue turn. |
| Outcome: | The proposed system outperforms existing systems while maintaining competitive accuracy. |
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)
Copied to clipboard
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth
| Challenge: | Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes. |
| Approach: | They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender. |
| Outcome: | The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning. |
A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity (P19-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings encode the meaning of words from different languages into a shared low-dimensional space. |
| Approach: | They show that modularity can be used to improve unsupervised word embeddings . they show that it is useful for low-resource languages where one often has few bilingual pairs . |
| Outcome: | The proposed model improves unsupervised cross-lingual word embeddings on distant language pairs in low-resource settings. |
A Multi-Modal Multilingual Benchmark for Document Image Classification (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing document image classification datasets have several limitations and we present two new datasets that overcome these limitations. |
| Approach: | They propose to use two newly curated multilingual datasets that overcome these limitations and propose to develop multilingual Document AI models. |
| Outcome: | The proposed datasets overcome limitations in document image classification and open the door for future research into improving Document AI models. |
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used as evaluators but the reliability of the outcomes remains a challenge. |
| Approach: | They show that LLM judge outputs are highly sensitive to pre-defined score ranges and that similar biases exist among models from the same family. |
| Outcome: | The contrastive decoding of LLM judge outputs achieves 11.7% relative improvement in Spearman correlation with human judgments, averaged across score ranges. |
Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability (2022.acl-long)
Copied to clipboard
| Challenge: | Pretrained multilingual models enable zero-shot learning even for unseen languages . current multilingual model covers only a small subset of the world's languages - due to data sparsity, they are not likely to obtain good results for many lowresource languages. |
| Approach: | They ask: how does the number of pretraining languages influence zero-shot learning for unseen languages? do the findings change if the languages used for pretraining are all related? |
| Outcome: | The results show that pretrained models can zero-shot learn for unseen languages even for limited amounts even for low-resource languages. |
M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation (2024.naacl-short)
Copied to clipboard
Benjamin Hsu, Xiaoyu Liu, Huayang Li, Yoshinari Fujinuma, Maria Nadejde, Xing Niu, Ron Litman, Yair Kittenplon, Raghavendra Pappagari
| Challenge: | Document translation is a challenge for machine translation systems that focus on textual content at the sentence level, ignoring global context and visual layout structure. |
| Approach: | They propose a benchmark dataset to evaluate document-level NMT systems . they use visual cues to preserve reading order and contiguous blocks of text . |
| Outcome: | The proposed benchmarks assess document-level NMT systems on the comprehensive task of translating semi-structured documents. |
Why Overfitting Isn’t Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries (2020.acl-main)
Copied to clipboard
| Challenge: | Recent studies only evaluate cross-lingual word embeddings on bilingual lexicon induction (BLI) however, underfitting can hinder generalization to other downstream tasks. |
| Approach: | They retrofit cross-lingual word embeddings to the training dictionary and a synthetic dictionary to improve their results. |
| Outcome: | The proposed method improves accuracy on two downstream tasks, despite underfitting the training dictionary. |