Papers with realism

23 papers
Large Scale Multi-Actor Generative Dialog Modeling (2020.acl-main)

Copied to clipboard

Challenge: Non-goal oriented dialog agents typically exhibit inconsistent personality across conversations or the average personality of all users.
Approach: They propose a model that conditionally models past conversations to probabilistically model multi-turn conversations in the actor’s persona.
Outcome: The proposed model improves perplexity on 1.7M held out Reddit conversations by 0.47 on scaling from 117M to 8.3B parameters.
SocialForge: simulating the social internet to provide realistic training against influence operations (2025.acl-industry)

Copied to clipboard

Challenge: Social media platforms have enabled large-scale influence campaigns, impacting democratic processes.
Approach: They propose a system to enhance diversity and realism of the generated content while ensuring its adherence to the original scenario.
Outcome: The proposed system improves diversity and realism while ensuring its adherence to the original scenario.
Developing a Corpus of Indirect Speech Act Schemas (2020.lrec-1)

Copied to clipboard

Challenge: Indirect speech acts (ISAs) involve utterances whose literal meanings are not identical to their intended meanings.
Approach: They propose a formal representation of ISA Schemas required for such testing, including a measure of the difficulty of a particular schema.
Outcome: The proposed model minimizes the amount of expert authoring needed and maximizes realism.
Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have studied prompt sensitivity by altering formatting or generating paraphrases with automated techniques.
Approach: They propose a framework for generating controlled paraphrases grounded in user behaviors . they leverage linguistically informed rules and enforce quality through checks on instruction adherence .
Outcome: The proposed framework is able to detect weaknesses in large language models . it leverages linguistically informed rules and enforces quality through checks on instruction adherence, semantic similarity, and realism.
LLM-Based Multi-Agent Systems are Scalable Graph Generative Models (2025.findings-acl)

Copied to clipboard

Challenge: Social graphs are mathematical structures stem from pairwise interactions between entities through nodes and edges.
Approach: They propose a framework for dynamic, text-attributed social graph generation that simulates the temporal node and edge generation processes for zero-shot social graphs.
Outcome: The proposed framework improves macroscopic graph structure metrics by 11% . the proposed model can generate graphs with up to 100,000 nodes or 10 million edges .
BullyBench: Youth & Experts-in-the-loop Framework for Intrinsic and Extrinsic Cyberbullying NLP Benchmarking (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing youth-focused CB datasets lack conversational realism and ethical youth involvement with little or no evaluation of their social plausibility.
Approach: They propose a youth-in-the-loop dataset “BullyBench” that incorporates a structured intrinsic quality evaluation with experts-in the-looop (social scientists, psychologists, and content moderators) they perform extrinsic baseline evaluation by benchmarking encoder- and decoder-only language models for multi-class CB role classification.
Outcome: The proposed dataset is evaluated by a team of social scientists, psychologists, and content moderators to assess its quality, relevance, and coherence.
High-Quality Medical Dialogue Synthesis for Improving EMR Generation (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for generating EMRs from doctor-patient dialogues produce rigid and repetitive dialogues.
Approach: They propose a framework that integrates Intent Graph Planning, Dual-Agent Simulation and Rule-Reward Quality Control to generate realistic doctor-patient dialogues.
Outcome: The proposed framework significantly enhances realism, diversity and downstream EMR quality, reducing physician editing efforts.
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs (2025.findings-acl)

Copied to clipboard

Challenge: Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes.
Approach: They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning.
Outcome: The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are becoming increasingly popular in education, enabling researchers to simulate students' learning patterns and learning patterns.
Approach: They propose a training-free framework for student simulation that takes into account student cognitive diversity and realism.
Outcome: The proposed model outperforms baseline models and achieves 100% improvement in simulation accuracy and realism.
MMInA: Benchmarking Multihop Multimodal Internet Agents (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks fail to assess embodied agents in a realistic, evolving environment for compositional Internet tasks.
Approach: They propose a multihop and multimodal benchmark to evaluate embodied agents for compositional Internet tasks.
Outcome: The proposed protocol significantly improves the performance of both the single-hop and multihop web browsing abilities.
LSF-ANIMAL: A Motion Capture Corpus in French Sign Language Designed for the Animation of Signing Avatars (2020.lrec-1)

Copied to clipboard

Challenge: Signing avatars are often procedurally animated, resulting in robotic and unnatural movements, which are therefore rejected by the Deaf community.
Approach: They propose to use a French Sign Language corpus to create an avatar that can be edited from motion capture data to create new signs and utterances.
Outcome: The proposed corpus is based on a french Sign Language (LSF) corpus composed of captured signs and sentences.
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for embedding human personality traits into LLMs are limited by realism and validity issues.
Approach: They propose to use a large-scale dataset to embed human personality traits into LLMs . they use supervised fine-tuning and direct preference optimization to train LLM models .
Outcome: The proposed methods outperform prompting on personality assessments and IPIP-NEO, and show higher conscientiousness, agreeableness, lower extraversion, and lower neuroticism on reasoning tasks.
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks assess basic Theory of Mind abilities but neglect temporal evolution of mental states in real-world social contexts.
Approach: They propose a benchmark specifically designed to evaluate Large Language Models' ability to understand and track the temporal progression of mental states across interconnected scenarios.
Outcome: The proposed benchmarks underperform humans by 44.7% and show that they can model the dynamic nature of human mental states better than existing models.
CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular Tables (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on well-structured tables and fail to reflect irregular structures and complex reasoning commonly encountered in real-world scenarios.
Approach: They propose a benchmark to evaluate TableQA under complex reasoning and irregular table conditions.
Outcome: The proposed framework improves generalization and realism of large language models under complex and irregular table conditions.
Adapting Bias Evaluation to Domain Contexts using Generative Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to assess social bias in NLP systems face limitations in scalability and fidelity across domains.
Approach: They propose a domain-adaptive framework that uses prompting with Large Language Models to automatically transform template-based bias datasets into domain-specific variants.
Outcome: The proposed framework improves the accuracy and contextual relevance of bias evaluations in socially relevant datasets.
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question.
Approach: They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries.
Outcome: The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images.
A Multi-Agent Framework for High-Interaction Terminal Simulation (2026.acl-long)

Copied to clipboard

Challenge: Terminal simulation is a problem of symbolic language generation in dialogue and interactive systems.
Approach: They propose a terminal command-level Turing test framework that improves realism, consistency and robustness in command-language generation.
Outcome: The proposed framework outperforms state-of-the-art benchmarks by more than 9% on multi-turn terminal simulation.
Beyond Human Labels: A Multi-Linguistic Auto-Generated Benchmark for Evaluating Large Language Models on Resume Parsing (2025.emnlp-main)

Copied to clipboard

Challenge: Efficient resume parsing is critical for global hiring, yet the lack of dedicated benchmarks for evaluating large language models (LLMs) on multilingual, structure-rich resumes hinders progress.
Approach: They propose to use a human-in-the-loop pipeline to generate 2,500 synthetic resumes spanning 50 templates, 30 career fields, and 5 languages to evaluate large language models.
Outcome: The proposed benchmarks show that the models perform poorly on multilingual resumes and lack of standardized templates.
EMSDialog: Synthetic Multi-person Emergency Medical Service Dialogue Generation from Electronic Patient Care Reports via Multi-LLM Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing medical dialogue corpora are largely dyadic or lack multi-party workflow and annotations needed for this setting.
Approach: They propose an ePCR-grounded, topic-flow-based multi-agent generation pipeline that iteratively plans, generates, and self-refines dialogues with rule-based factual and topic flow checks.
Outcome: The proposed pipeline yields a dataset of 4,414 synthetic multi-speaker EMS conversations annotated with 43 diagnoses, speaker roles, and turn-level topics.
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating AFC systems are limited in terms of task scope, modalities, domain, language diversity, realism, or coverage of misinformation types.
Approach: They propose to use Verified Theses and Statements (VeriTaS) to evaluate AFC systems that are static and subject to data leakage as claims enter pretraining corpora.
Outcome: The proposed system is robust under large-scale pretraining of foundation models and can be updated in the future.
SciMDR: Advancing Scientific Multimodal Document Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents.
Approach: They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments.
Outcome: The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning.
SceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for synthesising 3D scenes from a single image are text-driven and lack precise metric understanding from images.
Approach: They propose a language-model-based framework that grounds 3D scene synthesis in visual evidence by recovering an executable metric 3D layout directly from a single image.
Outcome: The proposed framework recovers an executable metric 3D layout directly from an RGB image and instantiates, places, and edits objects for iterative refinement.
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing models for visual information extraction suffer from limitations in scale and realism . ReceiptBench is a large-scale, human-annotated benchmark for receipts .
Approach: They propose a large-scale, human-annotated benchmark for visual information extraction . the method organizes information extraction into four hierarchical sub-tasks .
Outcome: The proposed method surpasses proprietary models on complex reasoning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations