Papers by Yusuke Oda

12 papers
Completely Modular Fine-tuning for Dynamic Language Adaptation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing studies on multilingual fine-tuning with a fixed set of languages lack dynamic adaptability to new languages.
Approach: They propose a modular fine-tuning pipeline that enables dynamic language adaptation for LLMs by first training English-centric adapters for each language separately and then merging them for arbitrary-direction translation.
Outcome: The proposed pipeline achieves 86% performance over traditional fine-tuning on four languages, while training only 0.1% parameters and relying on English as a bridge language without catastrophic forgetting.
Detecting Sensitive Personal Information in Japanese Pre-Training Corpora for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large-scale pre-training corpora are essential for large language models, but if such content remains unfiltered, there is a risk that LLMs may memorize it and leak it through their outputs.
Approach: They construct a Japanese text corpora dataset and train machine learning models to detect SCPI in text.
Outcome: The proposed classifier can detect information related to SCPI in Japanese text.
TDDC: Timely Disclosure Documents Corpus (2020.lrec-1)

Copied to clipboard

Challenge: TDDC was prepared by manually aligning the sentences from past Japanese and English timely disclosure documents . tens of thousands of original Japanese documents are disclosed every year, but the availability of English disclosure documents is limited.
Approach: They describe the details of the Timely Disclosure Documents Corpus (TDDC) TDDC was prepared by manually aligning the sentences from past Japanese and English timely disclosure documents .
Outcome: The timely disclosure documents corpus (TDDC) was created by aligning sentences from past documents in Japanese and English.
Are Prompt-based Models Clueless? (2022.acl-long)

Copied to clipboard

Challenge: Prompting has reduced the data requirement by reusing the language model head and formatting the task input to match the pre-training objective.
Approach: They propose to examine whether few-shot prompt-based models exploit superficial cues by reusing the model head and formatting the input to match the pre-training objective.
Outcome: The proposed models perform well on instances with superficial cues, but often outperform random accuracy on instances without superficial cuing.
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have greatly improved natural language understanding and generation.
Approach: They train a wide range of base models on a variety of datasets including code generation, mathematical reasoning, and general-domain tasks.
Outcome: The results show that training–task synergies persist across all models while others vary substantially, emphasizing the importance of model-specific strategies.
Scaling Data-Constrained Language Models with Synthetic Data (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) improve with more training data, but practical limitations on data collection constrain further scaling.
Approach: They compare three strategies to generate Japanese text, repeat the limited Japanese Web text, and use English Web text to fill the data shortfall.
Outcome: The proposed model outperforms baselines and achieves the performance achieved when the entire token budget is filled with additional organic Japanese Web text.
Developing Japanese CLIP Models Leveraging an Open-weight LLM for Large-scale Dataset Translation (2025.naacl-srw)

Copied to clipboard

Challenge: lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of vision-language models.
Approach: They construct large-scale Japanese image-text pairs using machine translation and pre-trained CLIP models on a Japanese dataset.
Outcome: The results show that pre-trained models achieve competitive average scores on Japanese culture tasks compared to models of similar size.
Comprehensive Study of Bilingual and Multi-category Instruction Pre-training (2026.findings-eacl)

Copied to clipboard

Challenge: Instruction pre-training (IPT) has recently gained attention as an intermediate stage between pre- and post-training for large language models.
Approach: They study the optimal balance between raw and instruction-response data, languages, and task categories in an LLM instruction-respondence dataset.
Outcome: The proposed model improves on English-centric and bilingual models using bilingual instruction-response datasets.
Overview of the 6th Workshop on Asian Translation (D19-52)

Copied to clipboard

Challenge: The 6th workshop on Asian translation (WAT2019) was held in hong kong, hongkong, and hong kong.
Approach: They present the results of the shared tasks from the 6th workshop on Asian translation (WAT2019) 25 teams participated in the shared task and 10 research paper submissions were accepted .
Outcome: The results of the 6th workshop on Asian translation (WAT2019) include JaEn, JaZh scientific paper translation subtasks, Ja'En, ja'Ko, Ja’En patent translation sub tasks, Hi'En and My'En patent subtask and Ru'Ja news commentary translation task.
Findings of the Third Workshop on Neural Generation and Translation (D19-56)

Copied to clipboard

Challenge: The 3rd Workshop on Neural Machine Translation and Generation (WNGT) was held in concert with the annual conference of the Empirical Methods in Natural Language Processing (EMNLP 2019).
Approach: They describe the results of the third workshop on Neural Generation and Translation held in concert with the annual conference of the Empirical Methods in Natural Language Processing (EMNLP 2019).
Outcome: The results of the 3rd Workshop on Neural Machine Translation and Generation (WNGT) were summarized in Sections 3 and 4.
Leveraging High-Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-training (2025.findings-emnlp)

Copied to clipboard

Challenge: low-resource language corpora in professional domains like medicine hinder cross-lingual domain adaptation of pre-trained large language models.
Approach: They examine how linguistic features affect performance on a Japanese–English medical knowledge benchmark.
Outcome: The proposed model can leverage English-language resources in medical domains while ensuring sufficient coverage of language-specific expressions in a target language.
Instability in Downstream Task Performance During LLM Pretraining (2025.findings-emnlp)

Copied to clipboard

Challenge: a study of large language models shows that task scores fluctuate throughout training .
Approach: They empirically analyze the stability of downstream task performance in an LLM .
Outcome: The proposed methods improve performance stability without changes to the training procedure.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations