Papers with FreshQA
FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation (2024.findings-acl)
Copied to clipboard
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, Thang Luong
| Challenge: | Modern large language models often "hallucinate" plausible but factually incorrect information, which reduces their trustworthiness especially in settings where accurate and up-to-date information is critical. |
| Approach: | They develop a human evaluation procedure to measure correctness and hallucination and use it to benchmark both closed and open-source LLMs. |
| Outcome: | The proposed method outperforms both competing search engine-augmented prompting methods and commercial systems on search-augmented QA. |