Papers by Masato Hagiwara
Machine Learning–Driven Language Assessment (2020.tacl-1)
Copied to clipboard
| Challenge: | Language proficiency tests are cumbersome to create and maintain, and items may be copied and leaked or simply used too often. |
| Approach: | They propose a method that uses machine learning and natural language processing to induce proficiency scales and linguistic models to estimate item difficulty directly for computer-adaptive testing. |
| Outcome: | The proposed method produces scores that are reliable and reliable while generating item banks large enough to satisfy security requirements. |
Project MOSLA: Recording Every Moment of Second Language Acquisition (2024.lrec-main)
Copied to clipboard
| Challenge: | Second language acquisition (SLA) is a complex and dynamic process. |
| Approach: | They created a longitudinal, multimodal, multilingual, and controlled dataset by inviting participants to learn one of three target languages from scratch over a span of two years, exclusively through online instruction. |
| Outcome: | The proposed dataset sheds light on the complex and dynamic nature of the acquisition of a second language and its implications for proficiency assessment, language and speech processing, and multimodal learning analytics. |
TEASPN: Framework and Protocol for Integrated Writing Assistance Environments (D19-3)
Copied to clipboard
| Challenge: | TEASPN is an open-source protocol for integrated writing assistance environments . authors propose that developers and researchers can integrate the latest developments in natural language processing with low cost. |
| Approach: | They propose a protocol and framework for integrating writing aids with writing software. |
| Outcome: | The proposed protocol standardizes the way writing software communicates with servers that implement such technologies, allowing developers and researchers to integrate the latest developments in natural language processing (NLP) with low cost. |
GitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical Errors (2020.lrec-1)
Copied to clipboard
| Challenge: | Lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction. |
| Approach: | They propose to make GitHub Typo Corpus a multilingual dataset of misspellings and grammatical errors available for use in NLP. |
| Outcome: | The proposed dataset contains more than 350k edits and 65M characters in more than 15 languages. |