Sakura: Large-scale Incorrect Example Retrieval System for Learners of Japanese as a Second Language (P19-3)
Copied to clipboard
| Challenge: | Existing example retrieval systems do not include grammatically incorrect examples . existing systems only provide a small number of examples, hence, learners cannot acquire sufficient information when they search . |
| Approach: | They propose an incorrect example retrieval system called Sakura using a large-scale dataset for Japanese language learners. |
| Outcome: | The proposed system is more useful than previous systems. |
Similar Papers
Scaling Data-Constrained Language Models with Synthetic Data (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) improve with more training data, but practical limitations on data collection constrain further scaling. |
| Approach: | They compare three strategies to generate Japanese text, repeat the limited Japanese Web text, and use English Web text to fill the data shortfall. |
| Outcome: | The proposed model outperforms baselines and achieves the performance achieved when the entire token budget is filled with additional organic Japanese Web text. |
Construction of an Evaluation Corpus for Grammatical Error Correction for Learners of Japanese as a Second Language (2020.lrec-1)
Copied to clipboard
| Challenge: | The Lang-8 corpus is suitable as a training dataset for machine translation-based grammatical error correction systems but it is not suitable as an evaluation dataset because corrected sentences sometimes include inappropriate sentences. |
| Approach: | They created an evaluation corpus for correcting grammatical errors made by Japanese as a second language learners using neural machine translation and statistical machine translation techniques. |
| Outcome: | The proposed corpus has less noise and its annotation scheme reflects the characteristics of the dataset, making it ideal for correcting grammatical errors in sentences written by learners of Japanese as a Second Language (JSL). |
Very Large-Scale Lexical Resources to Enhance Chinese and Japanese Machine Translation (L18-1)
Copied to clipboard
| Challenge: | A major issue in machine translation applications is the recognition and translation of named entities. |
| Approach: | They propose to integrate Very Large-Scale Lexical Resources (VLSLR) with lexicons to improve machine translation accuracy. |
| Outcome: | The proposed lexical resources can enhance the quality of MT in general and NMT systems, which currently don't use lexicons. |
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts. |
| Approach: | They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent . |
| Outcome: | The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora . |
Negative language transfer in learner English: A new dataset (2021.naacl-main)
Copied to clipboard
| Challenge: | This dataset contains annotated error causes for learner writing errors that tie learner mistakes to structures from their first language. |
| Approach: | They propose a learner English dataset enhanced with annotated error causes and concrete examples of learner errors that relate to their first languages. |
| Outcome: | The proposed dataset will be used to analyze learner errors related to language transfer from the learners’ first language. |
Building a Japanese Typo Dataset from Wikipedia’s Revision History (2020.acl-srw)
Copied to clipboard
| Challenge: | Typographical errors (typos) also occur in user generated content (UGC). |
| Approach: | They extract over half a million Japanese typo–correction pairs from Wikipedia’s revision history and combine character-based extraction rules, morphological analyzers to guess readings, and various filtering methods to address these challenges. |
| Outcome: | The proposed dataset extracts over half a million typo–correction pairs from Wikipedia’s revision history. |
New Dataset and Strong Baselines for the Grammatical Error Correction of Russian (2021.findings-acl)
Copied to clipboard
| Challenge: | a new resource is created to evaluate grammatical error correction models in English . a subset of the dataset is annotated in Russian, which is hard to come by and expensive to annotate . |
| Approach: | They develop an annotated learner corpus of Russian extracted from the Lang-8 website. |
| Outcome: | The proposed dataset is compared against two state-of-the-art grammatical error correction models . the results show that the created corpus is more diverse than the existing one . |
GitHub Typo Corpus: A Large-Scale Multilingual Dataset of Misspellings and Grammatical Errors (2020.lrec-1)
Copied to clipboard
| Challenge: | Lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction. |
| Approach: | They propose to make GitHub Typo Corpus a multilingual dataset of misspellings and grammatical errors available for use in NLP. |
| Outcome: | The proposed dataset contains more than 350k edits and 65M characters in more than 15 languages. |
Compositional Evaluation on Japanese Textual Entailment and Similarity (2022.tacl-1)
Copied to clipboard
| Challenge: | Despite growing interest in linguistic universals, most NLI/STS studies focus on English. |
| Approach: | They propose a Japanese NLI/STS dataset that was manually translated from the English dataset SICK. |
| Outcome: | The proposed datasets show that pre-trained language models are insensitive to word order and case particles. |
Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing: System Demonstrations (2022.aacl-demo)
Copied to clipboard
| Challenge: | the AACL-IJCNLP 2022 demonstrations track invited submissions ranging from early research prototypes to mature production-ready systems. |
| Approach: | the system demonstrations track invited submissions from early research prototypes to mature production-ready systems. |
| Outcome: | the AACL-IJCNLP 2022 demonstrations track received 17 submissions this year . the acceptance rate was 58.8% . |