Papers by Masato Mita

21 papers
Language Acquisition Device in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are less data-efficient than humans, and pre-pretraining on synthetic languages has been proposed to close this gap.
Approach: They propose to pre-pretrain on MP-STRUCT, a formal language whose strings encode hierarchical composition, feature-based dependencies, and long-distance displacement via MERGE, AGREE, and MOVE.
Outcome: The proposed model outperforms k-Shuffle Dyck despite not being definable in C-RASP despite being hierarchically expressive and circuit-theoretically learnable .
Developmentally-plausible Working Memory Shapes a Critical Period for Language Acquisition (2025.acl-long)

Copied to clipboard

Challenge: Large language models possess general linguistic abilities comparable to humans but their efficiency in language acquisition remains far inferior.
Approach: They propose a method that initially constrains working memory during the early stages of training and gradually relaxes this constraint as learning progresses.
Outcome: The proposed method outperforms conventional methods without memory constraints or with static memory constraints.
An Empirical Study of Incorporating Pseudo Data into Grammatical Error Correction (D19-1)

Copied to clipboard

Challenge: incorporating pseudo data in the training of grammatical error correction models has been a key factor in improving performance of such models.
Approach: They investigate the choice of how pseudo data should be generated or used in a grammatical error correction model and show that the results are state-of-the-art.
Outcome: The proposed method achieves state-of-the-art on the CoNLL-2014 test set and the official test set of the BEA-2019 shared task without making any modifications to the model architecture.
Construction of a Quality Estimation Dataset for Automatic Evaluation of Japanese Grammatical Error Correction (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on automatic evaluation of grammatical error correction (GEC) have shown that quality estimation models built from manual evaluation can achieve high performance in automatic evaluation in English.
Approach: They used a dataset with manual evaluation to build an automatic evaluation model for Japanese GEC.
Outcome: The proposed model is based on a Japanese dataset with manual evaluation and meta-evaluation.
ClozEx: A Task toward Generation of English Cloze Explanation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing tasks and datasets specifically designed for generating language learner explanations for cloze questions are lacking . clozing questions are used to assess language proficiency and enhance language learning .
Approach: They propose a task ClozEx to generate explanations for cloze questions in LA . they use a curated dataset of clozing questions paired with explanations .
Outcome: The proposed task generates fluent explanations for cloze questions in English as a second language learners.
CAMERA³: An Evaluation Dataset for Controllable Ad Text Generation in Japanese (2024.lrec-main)

Copied to clipboard

Challenge: Despite numerous efforts in ad text generation, the aspect of diversifying a text has received limited attention, particularly in non-English languages like Japanese.
Approach: They present a dataset for ad text generation in Japanese using annotators to examine the capabilities of recent NLG models.
Outcome: The proposed dataset includes 3,980 ad texts written by experts taking into account various aspects of ade appeals.
Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation (2020.coling-main)

Copied to clipboard

Challenge: a paper aims to improve performance measures for grammatical error correction . conventional measures treat all errors equally, but some are easier to correct .
Approach: They propose a way to determine the difficulty of error correction and to motivate researchers . paper examines performance measures for grammatical error correction using a scorer and weighting algorithm .
Outcome: The proposed measures agree with our intuition of correction difficulty . the results show that the measures are more complex than conventional measures .
AdTEC: A Unified Benchmark for Evaluating Text Quality in Search Engine Advertising (2025.naacl-long)

Copied to clipboard

Challenge: Existing pre-trained language models outperform them in certain domains, indicating that there is significant potential for further improvement in this area.
Approach: They propose to use pre-trained language models to evaluate ad texts from multiple perspectives within real-world advertising operations to define five tasks and construct a Japanese dataset.
Outcome: The proposed benchmark outperforms existing pre-trained language models in several tasks, but humans outperformed them in certain domains.
Token-length Bias in Minimal-pair Paradigm Datasets (2024.lrec-main)

Copied to clipboard

Challenge: Minimal-pair paradigm datasets are used to evaluate the linguistic knowledge of language models and provide an unsupervised method of acceptability judgment.
Approach: They propose a debiased minimal pair generation method that allows MPP datasets to evaluate the linguistic knowledge of a language model correctly.
Outcome: The proposed method is based on the percentage of minimal pairs in the MPP dataset where the model assigns a higher sentence log-likelihood than an unacceptable sentence.
Cloze Quality Estimation for Language Assessment (2023.findings-eacl)

Copied to clipboard

Challenge: Cloze tests are widely used in language proficiency tests, but they suffer from low quality and low reliability.
Approach: They propose a task to evaluate whether a cloze test is of sufficient "high-quality" they use a dataset that includes English clozing tests and corresponding evaluations by native English speakers.
Outcome: The proposed method could contribute to the CQE task, but the task is still challenging.
Cross-Corpora Evaluation and Analysis of Grammatical Error Correction Models — Is Single-Corpus Evaluation Enough? (N19-1)

Copied to clipboard

Challenge: Existing studies have evaluated grammatical error correction models on a single corpus, but the evaluation is incomplete because the task difficulty varies depending on the corpus and conditions such as proficiency levels of the writers and essay topics.
Approach: They evaluate the performance of several GEC models against various learner corpora and compare their rankings against the corpus.
Outcome: The evaluation of several models against learner corpora shows that the models’ rankings vary depending on the corpus, indicating that single-corpus evaluation is insufficient for GEC models.
Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error Correction (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for incorporating a masked language model into an EncDec model have potential drawbacks when applied to GEC.
Approach: They propose to incorporate a pre-trained masked language model (MLM) into an encoder-decoder model for grammatical error correction.
Outcome: The proposed method achieves state-of-the-art on BEA-2019 and CoNLL-2014 benchmarks.
Preventing Critical Scoring Errors in Short Answer Scoring with Confidence Estimation (2020.acl-srw)

Copied to clipboard

Challenge: Recent Short Answer Scoring systems use Quadratic Weighted Kappa (QWK) but it is unsatisfactory when measuring their effectiveness in actual usage.
Approach: They propose a task formulation of Short Answer Scoring (SAS) that matches actual usage and extracts as many scoring predictions that are not critical scoring errors (CSEs).
Outcome: The proposed system predicts scores with zero critical scoring errors (CSEs) for 50% of test data at maximum by filtering out low-reliability predictions on the basis of a certain confidence estimation.
PheMT: A Phenomenon-wise Dataset for Machine Translation Robustness on User-Generated Contents (2020.coling-main)

Copied to clipboard

Challenge: Existing studies suggest that Neural Machine Translation still struggles with certain kinds of input with considerable noise, such as User-Generated Contents (UGC) on the Internet.
Approach: They propose to evaluate the robustness of Neural Machine Translation models against specific linguistic phenomena in Japanese-English translation.
Outcome: The proposed model can handle user-generated content (UGC) on the Internet, but it is difficult to translate clean inputs.
A Self-Refinement Strategy for Noise Reduction in Grammatical Error Correction (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for grammatical error correction (GEC) rely on supervised learning with manually created datasets.
Approach: They propose to denoise GEC datasets by leveraging prediction consistency of existing models.
Outcome: The proposed method outperforms baseline methods on CoNLL-2014, JFLEG, and BEA-2019 benchmarks.
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and problem sets for automatic ad text generation are lacking . however, the growing volume of search queries has fueled research on the automatic generation of ads.
Approach: They propose to standardize the task of automatic ad text generation (ATG) using a benchmark dataset, CAMERA, to enable the utilization of multi-modal information and facilitate industry-wise evaluations.
Outcome: The proposed dataset standardizes the task of automatic ad text generation (ATG) it shows that existing metrics align with human evaluations and that the proposed methods can be used to improve the quality of the results.
GitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical Errors (2020.lrec-1)

Copied to clipboard

Challenge: Lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction.
Approach: They propose to make GitHub Typo Corpus a multilingual dataset of misspellings and grammatical errors available for use in NLP.
Outcome: The proposed dataset contains more than 350k edits and 65M characters in more than 15 languages.
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks .
Approach: a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models.
Outcome: The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models.
Do Grammatical Error Correction Models Realize Grammatical Generalization? (2021.findings-acl)

Copied to clipboard

Challenge: Existing models for grammatical error correction use pseudo data, but they are inconvenient for realworld deployment due to large amounts of training data.
Approach: They propose a method to evaluate whether GEC models can generalize to unseen errors by using synthetic and real GEC datasets with controlled vocabularies.
Outcome: The proposed model fails to realize grammatical generalization even in simple settings with limited vocabulary and syntax, suggesting it lacks the generalization ability required to correct errors from provided training examples.
ProQE: Proficiency-wise Quality Estimation dataset for Grammatical Error Correction (2022.lrec-1)

Copied to clipboard

Challenge: Prior work has shown that QE models of grammatical error correction are biased toward data by learners with relatively high proficiency levels.
Approach: They investigated whether learners' proficiency affects supervised quality estimation models of grammatical error correction (GEC) . they created a QE dataset that includes multiple proficiency levels and explored the necessity of performing proficiency-wise evaluation for QE of GEC.
Outcome: The proposed model is based on multiple proficiency levels and can be performed in real-world scenarios.
Targeted Syntactic Evaluation for Grammatical Error Correction (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation datasets based on learner-produced texts are insufficient for evaluating models . Currently, sequence-to-sequence models and sequence tagging models perform well on beginner-level grammar items .
Approach: They propose a new evaluation paradigm that assesses GEC models using minimal pairs of ungrammatical and grammatically paired sentences for each grammar item.
Outcome: The proposed evaluation paradigm assesses models using minimal pairs of ungrammatical and grammatically-spaced sentences for each grammar item.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations