Papers by Takenobu Tokunaga

11 papers
Content-Equivalent Translated Parallel News Corpus and Extension of Domain Adaptation for NMT (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to train NMT systems with noisy data are not sufficient . a recent increase in foreigners visiting Japan has created a significant information gap .
Approach: They propose a Japanese-English parallel news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus .
Outcome: The proposed corpus improves translation quality and is more effective than existing methods.
Analysis of Implicit Conditions in Database Search Dialogues (L18-1)

Copied to clipboard

Challenge: Annotators annotated 50 database search dialogues with database field tags . 10% of the utterances included non-database-field information, authors say .
Approach: They propose to annotate database search dialogues on real estate and analyse their utterances for database queries.
Outcome: The proposed method can extract the implicit conditions from user utterances and construct queries.
Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019 (D19-52)

Copied to clipboard

Challenge: In addition to the JIJI Corpus, we developed a corpus of 0.22M sentence pairs by manually, translating Japanese news sentences into English content- equivalently.
Approach: They propose to use JIJI Corpus and Equivalent-style sentences to translate Japanese news sentences into English content- equivalently.
Outcome: The proposed translation models achieved the best human evaluation scores in the newswire translation tasks at WAT 2019 . they used the JIJI Corpus, which was provided by the task organizer, and the Equivalent-style translation model to translate Japanese news sentences into English content- equivalently.
Interpretation of Implicit Conditions in Database Search Dialogues (C18-1)

Copied to clipboard

Challenge: Existing attempts to extract information from user utterances in database search dialogues have failed .
Approach: They propose to utilise information in user utterances that do not directly mention database fields for constructing database queries.
Outcome: The proposed model performs better than the existing model on a real estate agent-customer dialogue.
Analyzing Interpretability of Summarization Model with Eye-gaze Information (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have provided saliency scores for neural summarization models . eye-gaze information is often used as a proxy for human attention in reading tasks .
Approach: They propose to compare model saliency to human eye-gaze data to determine whether it conforms to human gaze during summarization.
Outcome: The proposed framework compares the model behavior to human summarization performance.
TIARA: A Tool for Annotating Discourse Relations and Sentence Reordering (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools for discourse relations and sentence reordering are difficult to use and clutter the display.
Approach: They propose to use TIARA to simplify the annotation process by offering interactive visualisation, including coloured links, indentation, and dual-view.
Outcome: The proposed tool simplifies the annotation process and offers visualisations including coloured links, indentation, and dual-view.
Cross-domain Analysis on Japanese Legal Pretrained Language Models (2022.findings-aacl)

Copied to clipboard

Challenge: Existing studies do not care the performance of domain-adapted PLMs for a generic domain.
Approach: They propose to use pretraining strategies to build pretrained language models specialised in the legal domain to improve their performance.
Outcome: The pretrained language models can learn domain-specific and general word meanings simultaneously and can distinguish them.
Annotation Study of Japanese Judgments on Tort for Legal Judgment Prediction with Rationales (2022.lrec-1)

Copied to clipboard

Challenge: An annotation scheme for Japanese judgment documents is proposed to provide a reliable dataset for Legal Judgment Prediction (LJP) the anticipated cost of LJP will be much lower than that of human legal professionals.
Approach: They propose to build an annotation scheme for legal judgment prediction, especially for torts, which extracts decisions and rationales at character-level.
Outcome: The proposed annotation scheme can produce a dataset of Japanese LJP at reasonable reliability.
Gamification Platform for Collecting Task-oriented Dialogue Data (2020.lrec-1)

Copied to clipboard

Challenge: a crowd-sourced approach to gather dialogue data is still a challenge due to the complexity of human dialogue structure and diversity of dialogue topics.
Approach: They propose a platform for collecting task-oriented situated dialogue data by using gamification.
Outcome: The proposed platform collects task-oriented situated dialogue data by using gamification.
Effective Use of Target-side Context for Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Existing methods to train NMT systems with noisy data are not sufficient . et al., 2018) found that NMT models can learn with multiple types of corpora .
Approach: They propose a Japanese-English news corpus that is content-equivalent . they extend a domain-adaptation method to train NMT models with clean corpus .
Outcome: The proposed corpus improves translation quality and is more efficient than existing methods.
Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis (2022.lrec-1)

Copied to clipboard

Challenge: In second language learning, summaries are among the most popular type of student assignments.
Approach: They propose to revise the annotation guidelines to allow machine implementation of the new annotation guidelines.
Outcome: The proposed algorithm achieves 0.789 precision and 0.844 recall over the L2WS 2021 corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations