Challenge: Language models often exhibit factual hallucination issue, exhibiting factual factual knowledge-grounded sentences.
Approach: They introduce a knowledge probing benchmark to evaluate the knowledge recall ability of pre-trained language models from diverse perspectives.
Outcome: The proposed benchmark evaluates the knowledge recall ability of encoder- and decoder-based pre-trained language models from diverse perspectives.

Similar Papers

Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models are trained on vast unlabeled data, rich in world knowledge.
Approach: They propose a categorization scheme for factual probing methods based on how inputs, outputs and probed PLMs are adapted . they synthesize insights about knowledge retention and prompt optimization in PLM models and analyze obstacles to adopting them as knowledge bases .
Outcome: The proposed method synthesizes insights about knowledge retention and prompt optimization in PLMs, analyzes obstacles to adopting them as knowledge bases and outline directions for future work.
Unveiling Factual Recall Behaviors of Large Language Models through Knowledge Neurons (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have underscored their exceptional reasoning prowess with natural language understanding across a broad spectrum of tasks.
Approach: They examine whether Large Language Models actively recall or retrieve their internal repositories of factual knowledge when faced with reasoning tasks.
Outcome: The proposed model improves reasoning performance while suppressing it leads to notable degradation.
X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Language models (LMs) capture factual knowledge by filling in the blanks of cloze-style prompts.
Approach: They propose a code-switching-based method to improve the ability of multilingual LMs to access knowledge and verify its effectiveness on several benchmark languages.
Outcome: The proposed method improves the ability of multilingual LMs to access knowledge and verify its effectiveness on several benchmark languages.
Calibrating Factual Knowledge in Pretrained Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that Pretrained Language Models can store factual knowledge, but facts stored in PLMs are not always correct.
Approach: They propose a lightweight method to calibrate factual knowledge in PLMs without re-training from scratch.
Outcome: The proposed method can be used to calibrate factual knowledge in PLMs without re-training from scratch.
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Language models (LMs) generate false or unverifiable content, often known as hallucination, despite ongoing efforts to enhance their factuality.
Approach: They propose a tool that measures LMs’ factuality in real-world user interactions by evaluating their factual accuracy and categorizing content units as Supported, Unsupported, or Undecidable based on Web-retrieved evidence.
Outcome: The proposed evaluation pipeline measures language models’ factuality in real-world user interactions.
Assessing Factual Reliability of Large Language Model Knowledge (2024.naacl-long)

Copied to clipboard

Challenge: Factual knowledge of LLMs is typically evaluated using accuracy, yet this metric does not capture the vulnerability of LRMs to hallucination-inducing factors like prompt and context variability.
Approach: They propose a metric designed to measure LLMs’ factual reliability by comparing the distance between the probability distributions of a valid output and its counterparts produced by the same LLM probing the same fact using different styles of prompts and contexts.
Outcome: The proposed metric measures the distance between the probability distributions of a valid output and its counterparts produced by the same LLM probing the same fact using different styles of prompts and contexts.
Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World Questions (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have shown promising ability to perform commonsense reasoning.
Approach: They propose a two-dimensional analysis framework that incorporates token back-tracing and token decoding to uncover how LLMs conduct factual knowledge recall.
Outcome: The proposed framework shows that LLMs lack relevant knowledge but struggle to select the most accurate information based on context during the retrieval and rerank phase.
How Do Multilingual Language Models Remember Facts? (2025.findings-acl)

Copied to clipboard

Challenge: Prior research has focused on English monolingual models, but how these mechanisms generalize to non-English languages remains unexplored.
Approach: They analyze three multilingual LLMs to find out how they can generalize recall mechanisms . they find that subject enrichment is language-independent, object extraction is language dependent .
Outcome: The proposed model performs better in multilingual contexts than in English models . the model is more efficient in multi-lingual context, but it is more complex in multilinguistic models compared to English models.
Tracing Multilingual Factual Knowledge Acquisition in Pretraining (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models are capable of recalling multilingual factual knowledge, but most studies evaluate only the final model, leaving the development of factual recall and crosslingual consistency unexplored.
Approach: They trace how factual recall and crosslingual consistency evolve during pretraining, focusing on OLMo-7B as a case study.
Outcome: The results show that fact frequency is the key to a better recall of multilingual facts, regardless of language, and some low-frequency facts in non-English languages can still be correctly recalled.
Factual Probing Is [MASK]: Learning vs. Learning to Recall (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for factual probing can interpret the model’s prediction accuracy as a lower bound on the amount of factual information it encodes.
Approach: They propose a method which directly optimizes in continuous embedding space and can predict an additional 6.4% of facts in the LAMA benchmark.
Outcome: The proposed method outperforms the best previous prompt method by 6.4% on the LAMA benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations