Investigating Memorization of Conspiracy Theories in Text Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies examine conspiracy theories in social media, but they have not evaluated their presence in generative language models. |
| Approach: | They examine the ability of generative language models to generate conspiracy theory text . they highlight the difficulties of this task and discuss the drawbacks . |
| Outcome: | The proposed model can generate conspiracy theories without access to training data. |
Similar Papers
Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language Models (2024.lrec-main)
Copied to clipboard
Boxi Cao, Qiaoyu Tang, Hongyu Lin, Shanshan Jiang, Bin Dong, Xianpei Han, Jiawei Chen, Tianshu Wang, Le Sun
| Challenge: | Pre-trained language models have shown remarkable memory formation, but vanilla networks without pre-training suffer catastrophic forgetting problem. |
| Approach: | They conduct experiments to investigate the retentive-forgetful contradiction between vanilla and pre-trained language models by controlling the target knowledge types, learning strategies and learning schedules. |
| Outcome: | The results show that pre-trained language models are forgetful and pre-training leads to retentive models . |
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them (2025.acl-long)
Copied to clipboard
| Challenge: | generative large language models produce hallucinations that are not aligned with world knowledge or input context. |
| Approach: | They propose a hallucination benchmark framework that measures hallucinism in large language models . they evaluate 150,000 generations from 14 language models and find they are riddled with hallucinos . |
| Outcome: | The proposed framework evaluates 150,000 generations from 14 language models. |
A Multi-Perspective Analysis of Memorization in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can generate the same sequences contained in the pre-train corpus, known as memorization. |
| Approach: | They analyze the relationship between memorization and outputs from Large Language Models (LLMs) they show a sudden drop and increase in the frequency of input tokens when generating memorized/unmemorized sequences . |
| Outcome: | The proposed model can generate the same sequences contained in the pre-train corpus, and it can predict unmemorized tokens. |
Pre-trained Language Models Return Distinguishable Probability Distributions to Unfaithfully Hallucinated Texts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | 88-98% of cases return distinguishable generation probability and uncertainty distributions to unfaithfully hallucinated texts, regardless of their size and structure. |
| Approach: | They examine 24 pre-trained language models on 6 data sets to examine their ability to distinguish unfaithfully hallucinated texts. |
| Outcome: | The proposed training algorithm outperforms baseline models while maintaining sound general text quality measures. |
Sources of Hallucination by Large Language Models on Inference Tasks (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are claimed to be capable of Natural Language Inference (NLI) |
| Approach: | They propose to use LLMs to probe their behavior using controlled experiments. |
| Outcome: | The proposed models perform significantly worse on NLI test samples which do not conform to these biases than those which do. |
Memorisation versus Generalisation in Pre-trained Language Models (2022.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art pre-trained language models have been shown to memorise facts and perform well with limited amounts of training data. |
| Approach: | They propose to extend pre-trained language models to generalise and memorise facts in noisy and low-resource scenarios. |
| Outcome: | The proposed extension improves performance in low-resource named entity recognition tasks. |
On the In-context Generation of Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have the ability of in-context generation (ICG) when given an in-text prompt, they can implicitly recognize the pattern of the examples and complete the prompt in the desired way. |
| Approach: | They propose a plausible latent variable model to model the distribution of pretrained corpora and formalize ICG as a problem of next topic prediction. |
| Outcome: | The proposed model can model the distribution of pretrained corpora and then formalize ICG as a problem of next topic prediction. |
Controlled Language Generation for Language Learning Items (2022.emnlp-industry)
Copied to clipboard
| Challenge: | Recent advances in pre-trained language models have resulted in success in generating fluent English text. |
| Approach: | They propose to employ natural language generation to rapidly generate English language items . they experiment with deep pretrained models and develop methods for controlling items for factors relevant in language learning . |
| Outcome: | The proposed framework shows high grammatically scores for all models and higher complexity over baseline models. |
Using Pre-Trained Language Models for Producing Counter Narratives Against Hate Speech: a Comparative Study (2022.findings-acl)
Copied to clipboard
| Challenge: | Autoregressive models combined with stochastic decodings are the most promising for generating CNs with regard to an unseen target of hate. |
| Approach: | They propose to use pre-trained language models to generate counter-narratives in English by adding an automatic post-editing step to refine generated CNs. |
| Outcome: | The proposed pipeline could be used to generate counter-narratives in English using pre-trained language models and stochastic decoding mechanisms. |
Do Language Models Perform Generalizable Commonsense Inference? (2021.findings-acl)
Copied to clipboard
| Challenge: | Recent work has applied pretrained language models to populate commonsense knowledge graphs (CKGs) but there is a lack of understanding on their generalization to multiple CKGs, unseen relations, and novel entities. |
| Approach: | They analyze the ability of pretrained language models to perform generalizable commonsense inference in terms of knowledge capacity, transferability and induction. |
| Outcome: | The proposed models can adapt to different schemas defined by multiple CKGs but fail to generalize to new relations. |