Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.

Similar Papers

Human Alignment: How Much Do We Adapt to LLMs? (2025.acl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming a common part of our lives, yet few studies have examined how they influence our behavior.
Approach: They propose a cooperative language game in which players aim to converge on a word and play a game in a group.
Outcome: The proposed game shows that humans notice and adapt to differences regardless of whether they are aware they are interacting with an LLM.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Can Large Language Models Be an Alternative to Human Evaluations? (2023.acl-long)

Copied to clipboard

Challenge: Human evaluation is indispensable for assessing the quality of texts generated by machine learning models or written by humans.
Approach: They propose to use large language models to evaluate unseen texts using the same instructions and samples . they also use LLMs to generate responses to questions that are used to conduct human evaluation .
Outcome: The proposed model can be used to evaluate texts in open-ended story generation and adversarial attacks.
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown capabilities close to human performance in various analytical tasks.
Approach: They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions .
Outcome: The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks.
Can Large Language Models Capture Dissenting Human Voices? (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive achievements in solving a broad range of tasks.
Approach: They evaluate the performance and alignment of large language models with humans using Monte Carlo Estimation and Log Probability Estimationic methods to estimate the multinomial distribution.
Outcome: The proposed models fail to capture human disagreement distribution and inference and human alignment performance plunge even further on data samples with high disagreement levels raising concerns about their natural language understanding ability and representativeness to a larger human population.
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the capabilities of AI systems.
Approach: This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction.
Outcome: This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape.
Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge.
Approach: They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data.
Outcome: The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages.
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models have enabled richer social simulations . however, the role of information asymmetry in these simulations has been overlooked .
Approach: They develop an evaluation framework to simulate social interactions with LLMs in different settings.
Outcome: The proposed framework performs better in unrealistic, omniscient simulation settings but struggles in those with information asymmetry.
Factuality of Large Language Models: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are factually incorrect, which limits their applicability in real-world scenarios.
Approach: They analyze existing work to identify major challenges and their associated causes . they propose to evaluate LLMs using a variety of measures to mitigate factual errors .
Outcome: The proposed methods are based on a variety of datasets and proposed strategies to mitigate factual errors.
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues (2024.naacl-srw)

Copied to clipboard

Challenge: Existing datasets with limited domain coverage and few challenging conversational phenomena are often unlabelled . Existing data is limited in quality and lacks a robust evaluation process .
Approach: They propose a high quality data generation system that generates high quality dialogues using 4,277 conversations across 100 intents.
Outcome: The proposed system produces high quality dialogue data with high quality labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations