Papers by Ryo Nagata
Creating Corpora for Research in Feedback Comment Generation (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpus of learner corpora with feedback comments is limited due to the lack of public access to this task. |
| Approach: | They describe two corpora that have been manually annotated with feedback comments . they describe how the principle and guidelines for feedback comment annotation work . |
| Outcome: | The proposed corpus is available on the web and will facilitate research in feedback comment generation. |
Exploring the Capacity of a Large-scale Masked Language Model to Recognize Grammatical Errors (2022.findings-acl)
Copied to clipboard
| Challenge: | a language model-based error detection method can learn errors with a small training sample. |
| Approach: | They propose a language model-based method for grammatical error detection with feedback comments. |
| Outcome: | The proposed method can learn errors with a little training data and improve recall faster than non-language models. |
Toward a Task of Feedback Comment Generation for Writing Learning (D19-1)
Copied to clipboard
| Challenge: | Existing work on feedback comment generation has been limited . despite its usefulness, there is no publicly available dataset for research on feedback comments . |
| Approach: | They introduce a task of automatically generating feedback comments such as a hint or an explanatory note for writing learning for non-native learners of English. |
| Outcome: | The proposed task is based on a corpus of 1,900 essays with all preposition errors annotated with feedback comments. |
Exploring Methods for Generating Feedback Comments for Writing Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating explanatory notes for language learners are inadequate . nagata et al. demonstrates that neural-retrieval-based methods can generate feedback comments for preposition use . |
| Approach: | They investigate three different methods for generating feedback comments for preposition use . grammatical and writing items can also be used to generate feedback comments . |
| Outcome: | The proposed methods outperform neural-retrieval-based methods in generating feedback comments for preposition use. |
Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation (2020.coling-main)
Copied to clipboard
| Challenge: | a paper aims to improve performance measures for grammatical error correction . conventional measures treat all errors equally, but some are easier to correct . |
| Approach: | They propose a way to determine the difficulty of error correction and to motivate researchers . paper examines performance measures for grammatical error correction using a scorer and weighting algorithm . |
| Outcome: | The proposed measures agree with our intuition of correction difficulty . the results show that the measures are more complex than conventional measures . |
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new method for finding semantic differences in words appears in two corpora, but it requires a variance of word vectors . a word covers more meanings in a corpus, and its mean word vector becomes shorter . |
| Approach: | They propose a method to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector. |
| Outcome: | The proposed methods rival the best-performing system in the SemEval-2020 Task 1 . they are robust for the skew in corpus sizes and capable of detecting infrequent words . |
Cross-Corpora Evaluation and Analysis of Grammatical Error Correction Models — Is Single-Corpus Evaluation Enough? (N19-1)
Copied to clipboard
| Challenge: | Existing studies have evaluated grammatical error correction models on a single corpus, but the evaluation is incomplete because the task difficulty varies depending on the corpus and conditions such as proficiency levels of the writers and essay topics. |
| Approach: | They evaluate the performance of several GEC models against various learner corpora and compare their rankings against the corpus. |
| Outcome: | The evaluation of several models against learner corpora shows that the models’ rankings vary depending on the corpus, indicating that single-corpus evaluation is insufficient for GEC models. |
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting semantic change only measure the level of individual usage instances. |
| Approach: | They propose to use unbalanced optimal transport to capture semantic change through excess and deficit in the alignment between usage instances. |
| Outcome: | The proposed method captures semantic change through excess and deficit in the alignment between usage instances. |
Exploring the Influence of Spelling Errors on Lexical Variation Measures (C18-1)
Copied to clipboard
| Challenge: | Lexical richness measures such as Type-Token Ratio and Yule's K are often used for learner English analysis and assessment but are unstable because of spelling errors. |
| Approach: | They propose to use a dictionary to calculate the difference between TTR and Yule’s K caused by spelling errors and to deepen the understanding of the influence of spelling errors on them. |
| Outcome: | The proposed measures are based on English learner corpora of three groups and estimate their values before and after spelling errors are manually corrected. |
Revisiting Statistical Laws of Semantic Shift in Romance Cognates (2022.coling-1)
Copied to clipboard
| Challenge: | Despite their shared etymology, some cognate pairs have experienced semantic shift. |
| Approach: | They examine the relationship between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy, and examine the effect of morphologically complex etyma on semantic shift. |
| Outcome: | The results show that frequency and polysemy have positive effects on semantic shift and that morphologically complex etyma are more resistant to it. |
Cross-lingual and Word-Independent Methods for Quantifying Degree of Grammaticalization (2026.eacl-long)
Copied to clipboard
Ryo Nagata, Daichi Mochihashi, Misato Ido, Yusuke Kubota, Naoki Otani, Yoshifumi Kawasaki, Hiroya Takamura
| Challenge: | Existing methods for quantifying the degree of grammaticalization are language- and word-dependent . existing methods are language dependent and lack training data . |
| Approach: | They propose to use Positive-Unlabeled learning or Cross-Validation-like learning to quantify degree of grammaticalization. |
| Outcome: | The proposed method achieves high correlations to human judgments in English deverbal prepositions and Japanese nouns being grammaticalized. |
A New Formulation of Zipf’s Meaning-Frequency Law through Contextual Diversity (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have examined Zipf's meaning-frequency law as a relationship between word frequency and the number of meanings based on contextualized word vectors . |
| Approach: | They propose to use word frequency as a relationship between word frequency and contextual diversity to examine Zipf's meaning-frequency law for a wider variety of words and corpora than previous studies have shown. |
| Outcome: | The proposed formulation gives a new interpretation of Zipf's meaning-frequency law and enables us to examine it for a wider variety of words and corpora than previous studies have shown. |
A Computational Approach to Quantifying Grammaticization of English Deverbal Prepositions (2024.lrec-main)
Copied to clipboard
| Challenge: | Linguistic studies have revealed important aspects of grammaticization of deverbal prepositions. |
| Approach: | They propose a computational approach to measure the degree of grammaticization of deverbal prepositions based on corpus data. |
| Outcome: | The proposed method correlates well with human judgements and supports previous findings in linguistics. |