Papers with realism
Large Scale Multi-Actor Generative Dialog Modeling (2020.acl-main)
Copied to clipboard
| Challenge: | Non-goal oriented dialog agents typically exhibit inconsistent personality across conversations or the average personality of all users. |
| Approach: | They propose a model that conditionally models past conversations to probabilistically model multi-turn conversations in the actor’s persona. |
| Outcome: | The proposed model improves perplexity on 1.7M held out Reddit conversations by 0.47 on scaling from 117M to 8.3B parameters. |
SocialForge: simulating the social internet to provide realistic training against influence operations (2025.acl-industry)
Copied to clipboard
Ulysse Oliveri, Guillaume Gadek, Alexandre Dey, Benjamin Costé, Damien Lolive, Arnaud Delhay, Bruno Grilheres
| Challenge: | Social media platforms have enabled large-scale influence campaigns, impacting democratic processes. |
| Approach: | They propose a system to enhance diversity and realism of the generated content while ensuring its adherence to the original scenario. |
| Outcome: | The proposed system improves diversity and realism while ensuring its adherence to the original scenario. |
Developing a Corpus of Indirect Speech Act Schemas (2020.lrec-1)
Copied to clipboard
| Challenge: | Indirect speech acts (ISAs) involve utterances whose literal meanings are not identical to their intended meanings. |
| Approach: | They propose a formal representation of ISA Schemas required for such testing, including a measure of the difficulty of a particular schema. |
| Outcome: | The proposed model minimizes the amount of expert authoring needed and maximizes realism. |
Say It Another Way: Auditing LLMs with a User-Grounded Automated Paraphrasing Framework (2026.eacl-long)
Copied to clipboard
Clea Chataigner, Rebecca Ma, Prakhar Ganesh, Yuhao Chen, Afaf Taik, Elliot Creager, Golnoosh Farnadi
| Challenge: | Existing studies have studied prompt sensitivity by altering formatting or generating paraphrases with automated techniques. |
| Approach: | They propose a framework for generating controlled paraphrases grounded in user behaviors . they leverage linguistically informed rules and enforce quality through checks on instruction adherence . |
| Outcome: | The proposed framework is able to detect weaknesses in large language models . it leverages linguistically informed rules and enforces quality through checks on instruction adherence, semantic similarity, and realism. |
LLM-Based Multi-Agent Systems are Scalable Graph Generative Models (2025.findings-acl)
Copied to clipboard
Jiarui Ji, Runlin Lei, Jialing Bi, Zhewei Wei, Xu Chen, Yankai Lin, Xuchen Pan, Yaliang Li, Bolin Ding
| Challenge: | Social graphs are mathematical structures stem from pairwise interactions between entities through nodes and edges. |
| Approach: | They propose a framework for dynamic, text-attributed social graph generation that simulates the temporal node and edge generation processes for zero-shot social graphs. |
| Outcome: | The proposed framework improves macroscopic graph structure metrics by 11% . the proposed model can generate graphs with up to 100,000 nodes or 10 million edges . |
BullyBench: Youth & Experts-in-the-loop Framework for Intrinsic and Extrinsic Cyberbullying NLP Benchmarking (2025.emnlp-industry)
Copied to clipboard
Kanishk Verma, Sri Balaaji, Joachim Wagner, Arefeh Kazemi, Darragh Mccashin, Isobel Walsh@dcu, Sayani Basak, Sinan Asci, Yelena Cherkasova, Alexandros Poulis, James Ohiggins Norman, Rebecca Umbach Umbach, Tijana Milosevic, Brian Davis
| Challenge: | Existing youth-focused CB datasets lack conversational realism and ethical youth involvement with little or no evaluation of their social plausibility. |
| Approach: | They propose a youth-in-the-loop dataset “BullyBench” that incorporates a structured intrinsic quality evaluation with experts-in the-looop (social scientists, psychologists, and content moderators) they perform extrinsic baseline evaluation by benchmarking encoder- and decoder-only language models for multi-class CB role classification. |
| Outcome: | The proposed dataset is evaluated by a team of social scientists, psychologists, and content moderators to assess its quality, relevance, and coherence. |
High-Quality Medical Dialogue Synthesis for Improving EMR Generation (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing methods for generating EMRs from doctor-patient dialogues produce rigid and repetitive dialogues. |
| Approach: | They propose a framework that integrates Intent Graph Planning, Dual-Agent Simulation and Rule-Reward Quality Control to generate realistic doctor-patient dialogues. |
| Outcome: | The proposed framework significantly enhances realism, diversity and downstream EMR quality, reducing physician editing efforts. |
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes. |
| Approach: | They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning. |
| Outcome: | The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench. |
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are becoming increasingly popular in education, enabling researchers to simulate students' learning patterns and learning patterns. |
| Approach: | They propose a training-free framework for student simulation that takes into account student cognitive diversity and realism. |
| Outcome: | The proposed model outperforms baseline models and achieves 100% improvement in simulation accuracy and realism. |
MMInA: Benchmarking Multihop Multimodal Internet Agents (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks fail to assess embodied agents in a realistic, evolving environment for compositional Internet tasks. |
| Approach: | They propose a multihop and multimodal benchmark to evaluate embodied agents for compositional Internet tasks. |
| Outcome: | The proposed protocol significantly improves the performance of both the single-hop and multihop web browsing abilities. |
LSF-ANIMAL: A Motion Capture Corpus in French Sign Language Designed for the Animation of Signing Avatars (2020.lrec-1)
Copied to clipboard
| Challenge: | Signing avatars are often procedurally animated, resulting in robotic and unnatural movements, which are therefore rejected by the Deaf community. |
| Approach: | They propose to use a French Sign Language corpus to create an avatar that can be edited from motion capture data to create new signs and utterances. |
| Outcome: | The proposed corpus is based on a french Sign Language (LSF) corpus composed of captured signs and sentences. |
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for embedding human personality traits into LLMs are limited by realism and validity issues. |
| Approach: | They propose to use a large-scale dataset to embed human personality traits into LLMs . they use supervised fine-tuning and direct preference optimization to train LLM models . |
| Outcome: | The proposed methods outperform prompting on personality assessments and IPIP-NEO, and show higher conscientiousness, agreeableness, lower extraversion, and lower neuroticism on reasoning tasks. |
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks assess basic Theory of Mind abilities but neglect temporal evolution of mental states in real-world social contexts. |
| Approach: | They propose a benchmark specifically designed to evaluate Large Language Models' ability to understand and track the temporal progression of mental states across interconnected scenarios. |
| Outcome: | The proposed benchmarks underperform humans by 44.7% and show that they can model the dynamic nature of human mental states better than existing models. |
CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular Tables (2026.acl-long)
Copied to clipboard
Zhen Yang, Wei Du, Jie Wang, Wenze Zhou, Xiangfeng Meng, Zhengyang Wang, Suping Sun, Ziwei Du, Haodong Zou, Jie Chen, Yongbin Liu, Shicheng Tan, Jiahao Ying, Shu Zhao
| Challenge: | Existing benchmarks focus on well-structured tables and fail to reflect irregular structures and complex reasoning commonly encountered in real-world scenarios. |
| Approach: | They propose a benchmark to evaluate TableQA under complex reasoning and irregular table conditions. |
| Outcome: | The proposed framework improves generalization and realism of large language models under complex and irregular table conditions. |
Adapting Bias Evaluation to Domain Contexts using Generative Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to assess social bias in NLP systems face limitations in scalability and fidelity across domains. |
| Approach: | They propose a domain-adaptive framework that uses prompting with Large Language Models to automatically transform template-based bias datasets into domain-specific variants. |
| Outcome: | The proposed framework improves the accuracy and contextual relevance of bias evaluations in socially relevant datasets. |
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)
Copied to clipboard
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh
| Challenge: | Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question. |
| Approach: | They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries. |
| Outcome: | The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images. |
A Multi-Agent Framework for High-Interaction Terminal Simulation (2026.acl-long)
Copied to clipboard
| Challenge: | Terminal simulation is a problem of symbolic language generation in dialogue and interactive systems. |
| Approach: | They propose a terminal command-level Turing test framework that improves realism, consistency and robustness in command-language generation. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks by more than 9% on multi-turn terminal simulation. |
Beyond Human Labels: A Multi-Linguistic Auto-Generated Benchmark for Evaluating Large Language Models on Resume Parsing (2025.emnlp-main)
Copied to clipboard
| Challenge: | Efficient resume parsing is critical for global hiring, yet the lack of dedicated benchmarks for evaluating large language models (LLMs) on multilingual, structure-rich resumes hinders progress. |
| Approach: | They propose to use a human-in-the-loop pipeline to generate 2,500 synthetic resumes spanning 50 templates, 30 career fields, and 5 languages to evaluate large language models. |
| Outcome: | The proposed benchmarks show that the models perform poorly on multilingual resumes and lack of standardized templates. |
EMSDialog: Synthetic Multi-person Emergency Medical Service Dialogue Generation from Electronic Patient Care Reports via Multi-LLM Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing medical dialogue corpora are largely dyadic or lack multi-party workflow and annotations needed for this setting. |
| Approach: | They propose an ePCR-grounded, topic-flow-based multi-agent generation pipeline that iteratively plans, generates, and self-refines dialogues with rule-based factual and topic flow checks. |
| Outcome: | The proposed pipeline yields a dataset of 4,414 synthetic multi-speaker EMS conversations annotated with 43 diagnoses, speaker roles, and turn-level topics. |
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating AFC systems are limited in terms of task scope, modalities, domain, language diversity, realism, or coverage of misinformation types. |
| Approach: | They propose to use Verified Theses and Statements (VeriTaS) to evaluate AFC systems that are static and subject to data leakage as claims enter pretraining corpora. |
| Outcome: | The proposed system is robust under large-scale pretraining of foundation models and can be updated in the future. |
SciMDR: Advancing Scientific Multimodal Document Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents. |
| Approach: | They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments. |
| Outcome: | The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning. |
SceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for synthesising 3D scenes from a single image are text-driven and lack precise metric understanding from images. |
| Approach: | They propose a language-model-based framework that grounds 3D scene synthesis in visual evidence by recovering an executable metric 3D layout directly from a single image. |
| Outcome: | The proposed framework recovers an executable metric 3D layout directly from an RGB image and instantiates, places, and edits objects for iterative refinement. |
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models for visual information extraction suffer from limitations in scale and realism . ReceiptBench is a large-scale, human-annotated benchmark for receipts . |
| Approach: | They propose a large-scale, human-annotated benchmark for visual information extraction . the method organizes information extraction into four hierarchical sub-tasks . |
| Outcome: | The proposed method surpasses proprietary models on complex reasoning tasks. |