Papers by Tatsuya Kawahara
Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Ainu is an unwritten language spoken by Ainus, a minority of whom are critically endangered by UNESCO . a project of automatic speech recognition (ASR) for the Ainous language is being developed . |
| Approach: | They propose to use automatic speech recognition for the Ainu language to help preserve its language archives. |
| Outcome: | The proposed system improves word and phone recognition accuracy in speaker-open conditions. |
Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference (2025.coling-demos)
Copied to clipboard
| Challenge: | Qualitative interviews are foundational to social science research, offering deep insights through open-ended conversations. |
| Approach: | They introduce a human-like embodied AI interviewer which integrates android and humanoid robots equipped with advanced conversational capabilities. |
| Outcome: | The proposed system performs well in a real-world case study at SIGDIAL 2024 with 42 participants, of whom 69% reported positive experiences. |
Topic-relevant Response Generation using Optimal Transport for an Open-domain Dialog System (2020.coling-main)
Copied to clipboard
| Challenge: | Conventional neural generative models generate safe and generic responses which have little connection with previous utterances semantically and would disengage users in a dialog system. |
| Approach: | They propose a method that employs topical constraint and semantic constraint to generate relevant responses by regularizing the decoding objective function with semantic distance. |
| Outcome: | The proposed method generates more topic-relevant and content-rich responses than conventional models. |
Multi-Task Learning of Generation and Classification for Emotion-Aware Dialogue Response Generation (2021.naacl-srw)
Copied to clipboard
| Challenge: | Existing models for human-like interaction with humans are not expected to improve the accuracy of emotion recognition, but instead focus on generating emotion-aware responses. |
| Approach: | They propose a neural response generation model with multi-task learning of generation and classification, focusing on emotion. |
| Outcome: | The proposed model makes generated responses more emotionally aware. |
Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation (2021.naacl-main)
Copied to clipboard
| Challenge: | End-to-end speech translation models can be trained to leverage source text . however, since the input modalities are different, it is difficult to leverage the source text successfully. |
| Approach: | They propose to leverage source transcriptions via pre-training and joint training with ASR and NMT tasks. |
| Outcome: | The proposed model predicts paraphrased transcriptions as an auxiliary task with a single decoder. |
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)
Copied to clipboard
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
| Challenge: | Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed . |
| Approach: | They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities. |
| Outcome: | The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech . |
Designing Precise and Robust Dialogue Response Evaluators (2020.acl-main)
Copied to clipboard
| Challenge: | Existing automated dialogue response evaluators have only moderate correlation with human judgement and are not robust. |
| Approach: | They propose to build a reference-free dialogue response evaluator that exploits the power of semi-supervised training and pretrained (masked) language models. |
| Outcome: | The proposed model achieves strong correlation with human judgement and generalizes robustly to diverse responses and corpora. |
Multilingual Turn-taking Prediction Using Voice Activity Projection (2024.lrec-main)
Copied to clipboard
| Challenge: | a monolingual model does not make good predictions when applied to other languages, but a multilingual model is able to discern the language of the input signal. |
| Approach: | They propose to use a multilingual voice activity projection model to predict voice activities of spoken dialogue participants in English, Mandarin, and Japanese data. |
| Outcome: | The proposed model predicts the upcoming voice activities of participants in dyadic dialogue on multilingual data, encompassing English, Mandarin, and Japanese. |
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for backchannel prediction relied on turn-based or artificially balanced datasets. |
| Approach: | They propose a method for real-time, continuous backchannel prediction using a fine-tuned Voice Activity Projection model. |
| Outcome: | The proposed method outperforms baseline methods in timing and type prediction tasks in real-world environments. |
Building a Dialogue Corpus Annotated with Expressed and Experienced Emotions (2022.acl-srw)
Copied to clipboard
| Challenge: | a human would recognize the emotion of an interlocutor and respond with an appropriate emotion, such as empathy and comfort. |
| Approach: | They propose to build a dialogue corpus annotated with two kinds of emotions . they collect tweets and annotate them with the emotion they put into the utterance . |
| Outcome: | The proposed method shows that it is difficult to recognize experienced emotions and multitask learning is effective. |