Papers by Hayato Yamana
ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models (2024.emnlp-main)
Copied to clipboard
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, Tetsuya Sakai, Tian Feng, Hayato Yamana
| Challenge: | Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100. |
| Approach: | They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth. |
| Outcome: | The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth. |
HRCA+: Advanced Multiple-choice Machine Reading Comprehension Method (2022.lrec-1)
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) requires a model to understand natural languages and understand textual representations. |
| Approach: | They propose a model that uses human reading comprehension attention to increase accuracy for machine reading comprehension. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Semeval-2018 Task 11 dataset and on the DREAM dataset. |