Challenge: Existing studies on language models have focused on factual correctness and justification, but prior research has focused on the factual truth condition and justifier.
Approach: They analyze language models’ responses and confidence using verbalized confidence, token probability, and sampling to examine their knowledge of Bayesian epistemology.
Outcome: The language models that follow the Bayesian confirmation assumption with true evidence show varying performance depending on the degree of irrelevance, indicating they deviate from Bayes' assumptions.

Similar Papers

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in high-stakes areas such as healthcare, law, and education.
Approach: They propose a concept of Confidence-Probability Alignment that connects an LLM’s internal confidence to the confidence conveyed in the model’s response when explicitly asked about its certainty.
Outcome: The proposed model shows the strongest confidence-probability alignment across a wide range of tasks.
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering (2021.tacl-1)

Copied to clipboard

Challenge: Recent studies have shown that language models capture different types of knowledge regarding facts or commonsense knowledge.
Approach: They examine how language models can be calibrated to make their confidence scores correlate better with the likelihood of correctness.
Outcome: The proposed calibration methods improve confidence scores on QA tasks and improve accuracy.
What Evidence Do Language Models Find Convincing? (2024.acl-long)

Copied to clipboard

Challenge: Current retrieval-augmented language models are tasked with subjective, contentious, and conflicting queries.
Approach: They construct a dataset that pairs controversial queries with real-world evidence documents . they find current models rely heavily on relevance of a website to the query .
Outcome: The proposed dataset pairs controversial queries with real-world evidence documents that contain different facts, arguments, and answers.
Perceptions of Linguistic Uncertainty by Language Models and Humans (2024.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that humans are well-attuned to the use of uncertainty expressions, exhibiting population-level agreement in mapping these expressions to numerical responses.
Approach: They propose to map linguistic expressions of uncertainty to numerical responses by using a theory of mind approach to understand the uncertainty of another agent.
Outcome: The proposed model can map expressions to probabilistic responses in a human-like manner, but different behavior depending on whether a statement is actually true or false.
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
Credible without Credit: Domain Experts Assess Generative Language Models (2023.acl-short)

Copied to clipboard

Challenge: ChatGPT has been criticized for its lack of accuracy and coherence . authors argue that language models could replace search engines and make college essays obsolete .
Approach: a team of 10 domain experts conducts an initial assessment of language models using 100 expert-written questions.
Outcome: The results show that language models are mixed in their accuracy.
Deep Bayesian Learning and Understanding (C18-3)

Copied to clipboard

Challenge: COLING 2018 is a conference for researchers and practitioners working on machine learning and deep learning.
Approach: a tutorial on machine learning and deep learning will be presented at COLING 2018 . the tutorial will focus on statistical models, deep neural networks, sequential learning and natural language understanding .
Outcome: This tutorial will present the latest advances in deep Bayesian and sequential learning at COLING 2018 .
Can Large Language Models Discern Evidence for Scientific Hypotheses? Case Studies in the Social Sciences (2024.lrec-main)

Copied to clipboard

Challenge: scholarly databases fail to aggregate, compare, contrast, and contextualize existing studies in service to a targeted research question.
Approach: They propose to use large language models to discern evidence in support or refute of specific hypotheses based on abstracts.
Outcome: The proposed method outperforms state-of-the-art methods and highlights opportunities for future research.
Epistemology of Language Models: Do Language Models Have Holistic Knowledge? (2024.findings-acl)

Copied to clipboard

Challenge: et al., 2021) explores whether language models exhibit characteristics consistent with epistemological holism . authors examined the epistle of language models from the perspective of abduction, revision, and argument generation tasks.
Approach: They examine whether language models exhibit characteristics consistent with epistemological holism . they created a scientific reasoning dataset and examined the epistology of language models .
Outcome: The language models showed that they did not distinguish between core and peripheral knowledge, compared with other tasks.
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) tend to be unreliable on fact-based answers.
Approach: They propose a framework for comparing LLMs' confidence over fact-based answers with hidden-state probes that are more reliable than hidden-status probes.
Outcome: The proposed methods show that hidden-state probes provide the most reliable confidence estimates despite requiring access to weights and supervision data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations