BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on the use of Large Language Models (LLMs) focus primarily on English, leaving the cross-lingual generalization of aligned behavior underexplored. |
| Approach: | They propose a structured generator-extractor pipeline and a multi-dimensional distributional analysis framework to examine how narrative attributes vary across languages, models, and social conditions. |
| Outcome: | The proposed model reveals substantial cross-lingual variability in narrative generation patterns, indicating that distributions observed in English do not always exhibit similar characteristics in other languages, particularly in lower-resource settings. |
Similar Papers
Biased Tales: Cultural and Topic Bias in Generating Children’s Stories (2025.emnlp-main)
Copied to clipboard
| Challenge: | Personalized stories are often preferred because they reflect a child's interests, experiences, and developmental needs. |
| Approach: | They analyze a dataset to examine how biases influence protagonists’ attributes and story elements in LLM-generated stories. |
| Outcome: | The proposed dataset shows that gender stereotypes influence protagonist attributes and story elements in LLM-generated stories. |
MirrorStories: Reflecting Diversity through Personalized Narrative Generation with Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used to create personalized “mirror stories” that reflect and resonate with individual readers’ identities. |
| Approach: | They propose to use Large Language Models to create personalized “mirror stories” that reflect and resonate with individual readers’ identities. |
| Outcome: | The proposed models outperform generic human-written and LLM-generated narratives on all metrics of engagement and textual diversity while preserving the intended moral. |
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains (2026.findings-acl)
Copied to clipboard
Md. Faiyaz Abdullah Sayeedi, Subhey Sadi Rahman, Md. Mahbub Alam, Md. Adnanul Islam, Jannatul Ferdous Deepti, Tasnim Mohiuddin, Md Mofijul Islam, Swakkhar Shatabda
| Challenge: | Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains. |
| Approach: | They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs. |
| Outcome: | The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs. |
Enriching children’s stories with LLMs: Delivering multilingual data enrichment for children’s books at scale and across markets (2025.acl-industry)
Copied to clipboard
| Challenge: | Using large language models and human-in-the-loop quality control, we enrich multilingual e-book and audio book data and make our product catalog easier to navigate for children. |
| Approach: | They propose a user-centered, empirically guided approach to multilingual metadata enrichment for children’s books using large language models and human-in-the-loop quality control. |
| Outcome: | The proposed approach delivers high-quality labels and improves user experience in real-world production environments. |
Are Large Language Models Capable of Generating Human-Level Narratives? (2024.emnlp-main)
Copied to clipboard
Yufei Tian, Tenghao Huang, Miri Liu, Derek Jiang, Alexander Spangher, Muhao Chen, Jonathan May, Nanyun Peng
| Challenge: | a recent HCI study has pointed to gaps in machine storytelling ability at the global level . authors show that LLMs have less suspense and less tension than human stories . |
| Approach: | They propose a computational framework to analyze narratives through three discourse-level aspects. |
| Outcome: | The proposed framework analyzes narratives through three discourse-level aspects . it shows that LLMs fall short of human abilities in discourse understanding . |
A LLM-based Ranking Method for the Evaluation of Automatic Counter-Narrative Generation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for evaluating CNs are expensive, time-consuming, and subjective, but lack a universal truth and the lack of a 'universal truth' . |
| Approach: | They propose a model ranking pipeline based on pairwise comparisons of generated CNs from different models organized in a tournament-style format to improve the evaluation process. |
| Outcome: | The proposed method achieves a high correlation with human preference, with a score of 0.88, and compares chat, instruct, and base models, exploring their strengths and limitations. |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation (2026.acl-long)
Copied to clipboard
Qianli Wang, Van Bach Nguyen, Yihong Liu, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schuetze, Sebastian Möller, Vera Schmitt
| Challenge: | Large language models excel at generating English counterfactuals but their effectiveness in generating multilingual counterfacts remains unclear. |
| Approach: | They conduct automatic evaluations on both directly generated and derived counterfactuals in six languages and find that cross-lingual perturbations follow common strategic principles. |
| Outcome: | The proposed models show that translation-based counterfactuals offer higher validity than their directly generated counterparts, but still fall short of matching the quality of the original English counterf actuals. |
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies focus on LLMs undertaking political questionnaires, which offers only limited insights into their biases and operational nuances. |
| Approach: | They propose to use a curated dataset to generate 56,700 synthetic articles using nine LLMs. |
| Outcome: | The proposed model can detect political biases using supervised models and LLMs. |
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)
Copied to clipboard
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth
| Challenge: | Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes. |
| Approach: | They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender. |
| Outcome: | The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning. |