Papers by Hiromi Wakaki
Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning (2024.findings-acl)
Copied to clipboard
Zhouhang Xie, Bodhisattwa Prasad Majumder, Mengjie Zhao, Yoshinori Maeda, Keiichi Yamada, Hiromi Wakaki, Julian McAuley
| Challenge: | Motivational Interviewing (MI) requires a system that can infer how to motivate users to adopt positive lifestyle changes. |
| Approach: | They propose a framework that can learn and apply conversation strategies from expert demonstrations by using natural language inductive rules. |
| Outcome: | The proposed framework outperforms in-context demonstrations that are over 50 times longer and can learn natural language strategies from demonstrations. |
DiffuCOMET: Contextual Commonsense Knowledge Diffusion (2024.acl-long)
Copied to clipboard
| Challenge: | Recent methods for identifying contextually relevant commonsense inferences are weak . knowledge models are trained to verbalize tuples from general commonsens knowledge graphs . |
| Approach: | They develop a series of knowledge models that leverage diffusion to reconstruct semantic connections between narrative contexts and relevant commonsense knowledge. |
| Outcome: | The proposed model improves on two benchmarks, ComFact and WebNLG+, to measure commonsense diversity and contextual relevance. |
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)
Copied to clipboard
Mengjie Zhao, Junya Ono, Zhi Zhong, Chieh-Hsin Lai, Yuhta Takida, Naoki Murata, Wei-Hsiang Liao, Takashi Shibuya, Hiromi Wakaki, Yuki Mitsufuji
| Challenge: | Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks. |
| Approach: | They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance. |
| Outcome: | The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large. |
ComFact: A Benchmark for Linking Contextual Commonsense Knowledge (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to retrieve facts from commonsense knowledge graphs are imprecise, requiring heuristics that ignore contexts and ambiguity . a novel benchmark, ComFact, contains 293k in-context relevance annotations for commonsensense triplets . |
| Approach: | They propose a task of commonsense fact linking where models are given contexts and trained to identify situationally-relevant commonsensical knowledge from KGs. |
| Outcome: | The proposed benchmark shows that heuristic fact linking approaches are imprecise . however, the models still significantly underperform humans in the commonsense augmentation task . |
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in music large language models have significantly improved music understanding tasks, but the potential of incorporating additional modalities such as images, videos and textual music features remains unexplored. |
| Approach: | They propose a multimodal music understanding LLM fine-tuned via multi-way instruction tuning with multi-ways aligned music, text, image, and video data. |
| Outcome: | The proposed model achieves state-of-the-art performance across six music understanding tasks and zero-shot scenarios. |
CARE: Multilingual Human Preference Learning for Cultural Awareness (2025.emnlp-main)
Copied to clipboard
| Challenge: | Language Models are tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. |
| Approach: | They introduce a multilingual resource that contains culturally specific questions and 31.7k responses with human judgments. |
| Outcome: | The proposed model outperforms models with stronger initial cultural performance . the proposed model has gaps in the literature on culturally relevant data . |
PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives (2023.acl-long)
Copied to clipboard
Silin Gao, Beatriz Borges, Soyoung Oh, Deniz Bayazit, Saya Kanno, Hiromi Wakaki, Yuki Mitsufuji, Antoine Bosselut
| Challenge: | a new knowledge graph for personas based on human-validated persona facts is constructed to model diverse persona attributes . a variety of persona characteristics are required to sustain coherent narratives . |
| Approach: | They construct a large-scale persona commonsense knowledge graph with 100K human-validated persona facts. |
| Outcome: | The proposed graph contains rich and precise world persona inferences that help systems generate more consistent and engaging narratives. |