Papers by Ryo Nagata

13 papers
Creating Corpora for Research in Feedback Comment Generation (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpus of learner corpora with feedback comments is limited due to the lack of public access to this task.
Approach: They describe two corpora that have been manually annotated with feedback comments . they describe how the principle and guidelines for feedback comment annotation work .
Outcome: The proposed corpus is available on the web and will facilitate research in feedback comment generation.
Exploring the Capacity of a Large-scale Masked Language Model to Recognize Grammatical Errors (2022.findings-acl)

Copied to clipboard

Challenge: a language model-based error detection method can learn errors with a small training sample.
Approach: They propose a language model-based method for grammatical error detection with feedback comments.
Outcome: The proposed method can learn errors with a little training data and improve recall faster than non-language models.
Toward a Task of Feedback Comment Generation for Writing Learning (D19-1)

Copied to clipboard

Challenge: Existing work on feedback comment generation has been limited . despite its usefulness, there is no publicly available dataset for research on feedback comments .
Approach: They introduce a task of automatically generating feedback comments such as a hint or an explanatory note for writing learning for non-native learners of English.
Outcome: The proposed task is based on a corpus of 1,900 essays with all preposition errors annotated with feedback comments.
Exploring Methods for Generating Feedback Comments for Writing Learning (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for generating explanatory notes for language learners are inadequate . nagata et al. demonstrates that neural-retrieval-based methods can generate feedback comments for preposition use .
Approach: They investigate three different methods for generating feedback comments for preposition use . grammatical and writing items can also be used to generate feedback comments .
Outcome: The proposed methods outperform neural-retrieval-based methods in generating feedback comments for preposition use.
Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation (2020.coling-main)

Copied to clipboard

Challenge: a paper aims to improve performance measures for grammatical error correction . conventional measures treat all errors equally, but some are easier to correct .
Approach: They propose a way to determine the difficulty of error correction and to motivate researchers . paper examines performance measures for grammatical error correction using a scorer and weighting algorithm .
Outcome: The proposed measures agree with our intuition of correction difficulty . the results show that the measures are more complex than conventional measures .
Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment (2023.emnlp-main)

Copied to clipboard

Challenge: a new method for finding semantic differences in words appears in two corpora, but it requires a variance of word vectors . a word covers more meanings in a corpus, and its mean word vector becomes shorter .
Approach: They propose a method to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector.
Outcome: The proposed methods rival the best-performing system in the SemEval-2020 Task 1 . they are robust for the skew in corpus sizes and capable of detecting infrequent words .
Cross-Corpora Evaluation and Analysis of Grammatical Error Correction Models — Is Single-Corpus Evaluation Enough? (N19-1)

Copied to clipboard

Challenge: Existing studies have evaluated grammatical error correction models on a single corpus, but the evaluation is incomplete because the task difficulty varies depending on the corpus and conditions such as proficiency levels of the writers and essay topics.
Approach: They evaluate the performance of several GEC models against various learner corpora and compare their rankings against the corpus.
Outcome: The evaluation of several models against learner corpora shows that the models’ rankings vary depending on the corpus, indicating that single-corpus evaluation is insufficient for GEC models.
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting semantic change only measure the level of individual usage instances.
Approach: They propose to use unbalanced optimal transport to capture semantic change through excess and deficit in the alignment between usage instances.
Outcome: The proposed method captures semantic change through excess and deficit in the alignment between usage instances.
Exploring the Influence of Spelling Errors on Lexical Variation Measures (C18-1)

Copied to clipboard

Challenge: Lexical richness measures such as Type-Token Ratio and Yule's K are often used for learner English analysis and assessment but are unstable because of spelling errors.
Approach: They propose to use a dictionary to calculate the difference between TTR and Yule’s K caused by spelling errors and to deepen the understanding of the influence of spelling errors on them.
Outcome: The proposed measures are based on English learner corpora of three groups and estimate their values before and after spelling errors are manually corrected.
Revisiting Statistical Laws of Semantic Shift in Romance Cognates (2022.coling-1)

Copied to clipboard

Challenge: Despite their shared etymology, some cognate pairs have experienced semantic shift.
Approach: They examine the relationship between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy, and examine the effect of morphologically complex etyma on semantic shift.
Outcome: The results show that frequency and polysemy have positive effects on semantic shift and that morphologically complex etyma are more resistant to it.
Cross-lingual and Word-Independent Methods for Quantifying Degree of Grammaticalization (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for quantifying the degree of grammaticalization are language- and word-dependent . existing methods are language dependent and lack training data .
Approach: They propose to use Positive-Unlabeled learning or Cross-Validation-like learning to quantify degree of grammaticalization.
Outcome: The proposed method achieves high correlations to human judgments in English deverbal prepositions and Japanese nouns being grammaticalized.
A New Formulation of Zipf’s Meaning-Frequency Law through Contextual Diversity (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have examined Zipf's meaning-frequency law as a relationship between word frequency and the number of meanings based on contextualized word vectors .
Approach: They propose to use word frequency as a relationship between word frequency and contextual diversity to examine Zipf's meaning-frequency law for a wider variety of words and corpora than previous studies have shown.
Outcome: The proposed formulation gives a new interpretation of Zipf's meaning-frequency law and enables us to examine it for a wider variety of words and corpora than previous studies have shown.
A Computational Approach to Quantifying Grammaticization of English Deverbal Prepositions (2024.lrec-main)

Copied to clipboard

Challenge: Linguistic studies have revealed important aspects of grammaticization of deverbal prepositions.
Approach: They propose a computational approach to measure the degree of grammaticization of deverbal prepositions based on corpus data.
Outcome: The proposed method correlates well with human judgements and supports previous findings in linguistics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations