Papers by Takenobu Tokunaga
Content-Equivalent Translated Parallel News Corpus and Extension of Domain Adaptation for NMT (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to train NMT systems with noisy data are not sufficient . a recent increase in foreigners visiting Japan has created a significant information gap . |
| Approach: | They propose a Japanese-English parallel news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus . |
| Outcome: | The proposed corpus improves translation quality and is more effective than existing methods. |
Analysis of Implicit Conditions in Database Search Dialogues (L18-1)
Copied to clipboard
| Challenge: | Annotators annotated 50 database search dialogues with database field tags . 10% of the utterances included non-database-field information, authors say . |
| Approach: | They propose to annotate database search dialogues on real estate and analyse their utterances for database queries. |
| Outcome: | The proposed method can extract the implicit conditions from user utterances and construct queries. |
Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019 (D19-52)
Copied to clipboard
| Challenge: | In addition to the JIJI Corpus, we developed a corpus of 0.22M sentence pairs by manually, translating Japanese news sentences into English content- equivalently. |
| Approach: | They propose to use JIJI Corpus and Equivalent-style sentences to translate Japanese news sentences into English content- equivalently. |
| Outcome: | The proposed translation models achieved the best human evaluation scores in the newswire translation tasks at WAT 2019 . they used the JIJI Corpus, which was provided by the task organizer, and the Equivalent-style translation model to translate Japanese news sentences into English content- equivalently. |
Interpretation of Implicit Conditions in Database Search Dialogues (C18-1)
Copied to clipboard
| Challenge: | Existing attempts to extract information from user utterances in database search dialogues have failed . |
| Approach: | They propose to utilise information in user utterances that do not directly mention database fields for constructing database queries. |
| Outcome: | The proposed model performs better than the existing model on a real estate agent-customer dialogue. |
Analyzing Interpretability of Summarization Model with Eye-gaze Information (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have provided saliency scores for neural summarization models . eye-gaze information is often used as a proxy for human attention in reading tasks . |
| Approach: | They propose to compare model saliency to human eye-gaze data to determine whether it conforms to human gaze during summarization. |
| Outcome: | The proposed framework compares the model behavior to human summarization performance. |
TIARA: A Tool for Annotating Discourse Relations and Sentence Reordering (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools for discourse relations and sentence reordering are difficult to use and clutter the display. |
| Approach: | They propose to use TIARA to simplify the annotation process by offering interactive visualisation, including coloured links, indentation, and dual-view. |
| Outcome: | The proposed tool simplifies the annotation process and offers visualisations including coloured links, indentation, and dual-view. |
Cross-domain Analysis on Japanese Legal Pretrained Language Models (2022.findings-aacl)
Copied to clipboard
| Challenge: | Existing studies do not care the performance of domain-adapted PLMs for a generic domain. |
| Approach: | They propose to use pretraining strategies to build pretrained language models specialised in the legal domain to improve their performance. |
| Outcome: | The pretrained language models can learn domain-specific and general word meanings simultaneously and can distinguish them. |
Annotation Study of Japanese Judgments on Tort for Legal Judgment Prediction with Rationales (2022.lrec-1)
Copied to clipboard
| Challenge: | An annotation scheme for Japanese judgment documents is proposed to provide a reliable dataset for Legal Judgment Prediction (LJP) the anticipated cost of LJP will be much lower than that of human legal professionals. |
| Approach: | They propose to build an annotation scheme for legal judgment prediction, especially for torts, which extracts decisions and rationales at character-level. |
| Outcome: | The proposed annotation scheme can produce a dataset of Japanese LJP at reasonable reliability. |
Gamification Platform for Collecting Task-oriented Dialogue Data (2020.lrec-1)
Copied to clipboard
| Challenge: | a crowd-sourced approach to gather dialogue data is still a challenge due to the complexity of human dialogue structure and diversity of dialogue topics. |
| Approach: | They propose a platform for collecting task-oriented situated dialogue data by using gamification. |
| Outcome: | The proposed platform collects task-oriented situated dialogue data by using gamification. |
Effective Use of Target-side Context for Neural Machine Translation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to train NMT systems with noisy data are not sufficient . et al., 2018) found that NMT models can learn with multiple types of corpora . |
| Approach: | They propose a Japanese-English news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus . |
| Outcome: | The proposed corpus improves translation quality and is more efficient than existing methods. |
Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | In second language learning, summaries are among the most popular type of student assignments. |
| Approach: | They propose to revise the annotation guidelines to allow machine implementation of the new annotation guidelines. |
| Outcome: | The proposed algorithm achieves 0.789 precision and 0.844 recall over the L2WS 2021 corpus. |