Challenge: Using large language models and human-in-the-loop quality control, we enrich multilingual e-book and audio book data and make our product catalog easier to navigate for children.
Approach: They propose a user-centered, empirically guided approach to multilingual metadata enrichment for children’s books using large language models and human-in-the-loop quality control.
Outcome: The proposed approach delivers high-quality labels and improves user experience in real-world production environments.

Similar Papers

On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on the use of Large Language Models (LLMs) focus primarily on English, leaving the cross-lingual generalization of aligned behavior underexplored.
Approach: They propose a structured generator-extractor pipeline and a multi-dimensional distributional analysis framework to examine how narrative attributes vary across languages, models, and social conditions.
Outcome: The proposed model reveals substantial cross-lingual variability in narrative generation patterns, indicating that distributions observed in English do not always exhibit similar characteristics in other languages, particularly in lower-resource settings.
A Survey on LLMs for Story Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Methods for story generation with Large Language Models (LLMs) have come into the spotlight recently.
Approach: They propose a novel taxonomy of LLMs for story generation consisting of two major paradigms: independent story generation by an LLM, and author-assistance for story creation .
Outcome: The proposed taxonomy compares existing work on the topic with those of novel author-assistance models.
The Data Frontier for Large Language Models: Selection, Synthesis, and Tools (2026.acl-tutorials)

Copied to clipboard

Challenge: acquiring and curating high-quality training data remains a significant bottleneck . acquiring such high-quality data is a key challenge for researchers and practitioners .
Approach: This tutorial provides a comprehensive and practical guide to the state-of-the-art in data research directions for LLMs.
Outcome: The tutorial covers methods for curating the most valuable information from vast, noisy datasets and the synthetic data revolution.
On the Automatic Generation and Simplification of Children’s Stories (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have made it possible to generate children's educational texts with appropriate lexical and readability levels.
Approach: They first examine the ability of several popular LLMs to generate stories with properly adjusted lexical and readability levels.
Outcome: The proposed models can generalize to the domain of children's stories and create an efficient pipeline for their automatic generation.
Data Augmentation using LLMs: Data Perspectives, Learning Paradigms and Challenges (2024.findings-acl)

Copied to clipboard

Challenge: Data augmentation (DA) is a key technique for enhancing model performance by diversifying training examples without the need for additional data collection.
Approach: They examine various strategies that utilize LLMs for data augmentation, including a novel exploration of learning paradigms where LLM-generated data is used for diverse forms of further training.
Outcome: The proposed approach addresses the primary open challenges faced by LLMs in the field of large language models and aims to serve as a comprehensive guide for researchers and practitioners.
Making Large Language Models Better Data Creators (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have advanced the field of NLP significantly, but deploying them for downstream applications is still challenging due to cost, responsiveness, control, or concerns around privacy and security.
Approach: They propose a unified data creation pipeline that requires only a single formatting example.
Outcome: The proposed pipeline can generate data with a single formatting example.
LLM-Based Web Data Collection for Research Dataset Creation (2025.findings-emnlp)

Copied to clipboard

Challenge: researchers across many fields rely on web data to gain new insights and validate methods.
Approach: They propose a human-in-the-loop framework that automates web-scale data collection end-to-end using large language models (LLMs)
Outcome: The proposed framework outperforms existing methods in three different tasks and a user evaluation demonstrates its practical utility.
Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language.
Approach: They propose a model that integrates symbolic data into LLM training without loss of generality ability.
Outcome: The proposed model performs better on symbol- and NL-centric tasks.
Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions (2023.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can be used to generate text data for training and evaluating other models.
Approach: They propose to use logit suppression and temperature sampling to diversify text generation but at the cost of data accuracy.
Outcome: The proposed approach can increase diversity but at the cost of data accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations