Assessing Web Search Credibility and Response Groundedness in Chat Assistants (2026.eacl-long)
Copied to clipboard
| Challenge: | Using 100 claims across five misinformation-prone topics, we assess GPT-4o, GPT-5, Perplexity, and Qwen Chat. |
| Approach: | They propose a method for evaluating assistants’ web search behavior focusing on source credibility and the groundedness of responses with respect to cited sources. |
| Outcome: | The proposed method focuses on source credibility and the groundedness of responses with respect to cited sources. |
Similar Papers
GPTEval: A Survey on Assessments of ChatGPT and GPT-4 (2024.lrec-main)
Copied to clipboard
| Challenge: | emergence of ChatGPT has generated speculation about its potential to disrupt social and economic systems. |
| Approach: | They analyze prior assessments of ChatGPT and GPT-4 to analyze their language and reasoning abilities, scientific knowledge, ethical considerations and existing evaluation methods. |
| Outcome: | The proposed model performs satisfactorily in science knowledge and can answer open questions. |
Evaluating Verifiability in Generative Search Engines (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing generative search engines are rapidly gaining users, according to a new study . existing systems are poorly cited and lack reliability, a study finds . |
| Approach: | They conduct human evaluations of four popular generative search engines . they find that existing generative engines are fluent and appear informative . |
| Outcome: | The results show that existing generative search engines are not reliable and often contain unsupported statements and inaccurate citations. |
Consistency Analysis of ChatGPT (2023.emnlp-main)
Copied to clipboard
| Challenge: | ChatGPT and GPT-4 have been reported to be more reliable and trustworthy, provided they behave similarly to humans. |
| Approach: | They propose to compare ChatGPT and GPT-4 in terms of logically consistent behaviour and the properties of negation, symmetric, and transitive consistency. |
| Outcome: | The proposed models show that they can be more reliable and trustworthy provided they behave similarly to humans. |
Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark (2022.tacl-1)
Copied to clipboard
| Challenge: | Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information. |
| Approach: | They propose to evaluate the validity of 12k dialogue turns generated by neural dialogue systems trained on three knowledge-grounded dialogue corpora and to use them to analyze eight evaluation metrics. |
| Outcome: | The proposed evaluation metrics rely on spurious correlations, do not reliably distinguish attributable abstractive responses from unattributable ones, and perform substantially worse when the knowledge source is longer. |
TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Document assistant chatbots are empowered with extensive capabilities by Large Language Models (LLMs) however, they suffer from hallucinations that are difficult to verify in the context of given documents. |
| Approach: | They propose a document assistant chatbot with reliable attribution that enables users to seek relevant information from given documents. |
| Outcome: | The proposed system generates answers with detailed inline citations, which can be attributed to the original document paragraphs, facilitating verification of factual consistency of the generated text. |
Evaluating Robustness of Generative Search Engine on Adversarial Factoid Questions (2024.findings-acl)
Copied to clipboard
Xuming Hu, Xiaochuan Li, Junzhe Chen, Yinghui Li, Yangning Li, Xiaoguang Li, Yasheng Wang, Qun Liu, Lijie Wen, Philip Yu, Zhijiang Guo
| Challenge: | Existing large language models (LLMs)-backed generative search engines may not always be accurate. |
| Approach: | They propose to evaluate the robustness of retrieval-augmented generation in a realistic and high-risk setting where adversaries have only black-box system access. |
| Outcome: | The proposed model exhibits higher susceptibility to factual errors compared to LLMs without retrieval. |
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems (2020.emnlp-main)
Copied to clipboard
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, Mark Cieliebak
| Challenge: | Lack of time efficient and reliable evalu-ation methods is hampering the development of conversational dialogue systems (chatbots). |
| Approach: | They propose a framework that replaces human-bot conversations with conversations between bots and an annotation tool that ranks chatbots based on their ability to mimic human behaviour. |
| Outcome: | The proposed evaluation framework replaces human-bot conversations with bot conversations and allows for frequent evaluations of chatbots during their evaluation cycle. |
Don’t Forget Your ABC’s: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems (2023.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods are biased because of their subjectivity and inconsistent evaluation can misinform the performance of a chat-oriented open-domain dialogue system. |
| Approach: | They propose to use a human evaluation method to estimate the rates of manypasted macro ‘LN’ dialogue system behaviors to compare them with existing evaluation methods. |
| Outcome: | The proposed method is more suitable than alternative Likert-style or comparative approaches for dimensional evaluation of open-domain dialogue systems. |
Challenges in Trustworthy Human Evaluation of Chatbots (2025.findings-naacl)
Copied to clipboard
| Challenge: | apathetic or adversarial annotators can corrupt the reliability of open leaderboard rankings . human annotation is widely accepted as the gold standard for open-ended text generation tasks . |
| Approach: | They show that bad annotations can corrupt the reliability of open leaderboard rankings . they argue that human annotation is widely accepted as the gold standard . |
| Outcome: | The proposed algorithm can corrupt the reliability of open leaderboard rankings by up to 5 places. |
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets (2023.findings-acl)
Copied to clipboard
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, Jimmy Huang
| Challenge: | Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth. |
| Approach: | They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets. |
| Outcome: | The proposed model performs well on 140 tasks and generates 255K responses in these datasets. |