Papers with Generation
Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (2023.eacl-main)
Copied to clipboard
| Challenge: | EACL 2023 submissions were divided into 21 areas . the areas were "Generation and Summarization", "Language Resources and Evaluation" and "Machine Learning in NLP". |
| Approach: | EACL 2023 submissions were divided into 21 areas . most popular areas were "Generation and Summarization", "Language Resources and Evaluation" and "Machine Learning in NLP" |
| Outcome: | EACL 2023 submissions were divided into 21 areas . the areas most popular with over 100 submissions included "Generation and Summarization", "Language Resources and Evaluation" |
Can Large Language Models Personalize Dialogues to Generational Styles? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a human evaluation reveals that annotators were able to most accurately identify the generation behind P-MultiWoZ dialogues, based only on a single query-reply pair. |
| Approach: | They create a personalized, generation-specific version of MultiWOZ 2.2 by prompting LLMs to generate personalized dialogue responses. |
| Outcome: | The proposed model is a personalized version of MultiWOZ 2.2 for Generation X, Y, and Z . it is validated by automatic and human evaluations to determine whether it reflects generational linguistic traits. |
NLP for Counterspeech against Hate and Misinformation (CSHAM) (2025.acl-tutorials)
Copied to clipboard
| Challenge: | tutorial aims to show how counterspeech is used to tackle abuse and misinformation by individuals, activists and organisations. |
| Approach: | tutorial aims to show how counterspeech is currently used to tackle abuse and misinformation . will also show how Natural Language Processing (NLP) and Generation (NLG) can be applied to automate its production. |
| Outcome: | The tutorial will bring diverse multidisciplinary perspectives to safety research . case studies from industry and public policy will be included . |
SPRING Goes Online: End-to-End AMR Parsing and Generation (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a formalism for representing the semantics of natural language in a readable and hierarchical way. |
| Approach: | They present SPRING Online Services, a Web interface and RESTful APIs for their AMR parsing and generation system, SPRING (Symmetric PaRsIng aNd Generation). |
| Outcome: | The proposed system provides a highly interactive visualization platform and feedback mechanism to obtain user suggestions for further improvements of the system’s output. |
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)
Copied to clipboard
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina Mcmillan-major, Anna Shvets, Ashish Upadhyay, Bernd Bohnet, Bingsheng Yao, Bryan Wilie, Chandra Bhagavatula, Chaobin You, Craig Thomson, Cristina Garbacea, Dakuo Wang, Daniel Deutsch, Deyi Xiong, Di Jin, Dimitra Gkatzia, Dragomir Radev, Elizabeth Clark, Esin Durmus, Faisal Ladhak, Filip Ginter, Genta Indra Winata, Hendrik Strobelt, Hiroaki Hayashi, Jekaterina Novikova, Jenna Kanerva, Jenny Chim, Jiawei Zhou, Jordan Clive, Joshua Maynez, João Sedoc, Juraj Juraska, Kaustubh Dhole, Khyathi Raghavi Chandu, Laura Perez Beltrachini, Leonardo F . R. Ribeiro, Lewis Tunstall, Li Zhang, Mahim Pushkarna, Mathias Creutz, Michael White, Mihir Sanjay Kale, Moussa Kamal Eddine, Nico Daheim, Nishant Subramani, Ondrej Dusek, Paul Pu Liang, Pawan Sasanka Ammanamanchi, Qi Zhu, Ratish Puduppully, Reno Kriz, Rifat Shahriyar, Ronald Cardenas, Saad Mahamood, Salomey Osei, Samuel Cahyawijaya, Sanja Štajner, Sebastien Montella, Shailza Jolly, Simon Mille, Tahmid Hasan, Tianhao Shen, Tosin Adewumi, Vikas Raunak, Vipul Raheja, Vitaly Nikolaev, Vivian Tsai, Yacine Jernite, Ying Xu, Yisi Sang, Yixin Liu, Yufang Hou
| Challenge: | Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work. |
| Approach: | They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations. |
| Outcome: | The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work. |
The Evolution of Gen Alpha Slang: Linguistic Patterns and AI Translation Challenges (2025.acl-srw)
Copied to clipboard
| Challenge: | Generation Alpha (born 2010-2024) exhibits unique linguistic behaviours influenced by rampant online communication and platform-specific cultures. |
| Approach: | They construct a comprehensive slang corpus from online platforms and evaluate four AI translation systems on over 100 sling terms. |
| Outcome: | The proposed translation systems outperform four existing translation models on over 100 slang terms. |
Language Generation with Multi-Hop Reasoning on Commonsense Knowledge Graph (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches that integrate commonsense knowledge into pre-trained language models simply transfer relational knowledge while ignoring rich connections within the knowledge graph. |
| Approach: | They propose a method that leverages structural and semantic information of the knowledge graph to generate commonsense-aware text. |
| Outcome: | The proposed method outperforms baseline models on three text generation tasks that require reasoning over commonsense knowledge. |
Annotate the Way You Think: An Incremental Note Generation Framework for the Summarization of Medical Conversations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for summarization of medical conversations are limited to conversation-summary pairs . a novel annotation framework is proposed to capture the summarizing process via an annotation task . |
| Approach: | They propose an incremental note generation framework that captures the human summarization process via an annotation task by instructing annotators to first incrementally create a draft note and polish it into a reference note. |
| Outcome: | The proposed framework shows that the human summarization process is much more efficient and accurate than the current method. |
FilBench: Can LLMs Understand and Generate Filipino? (2025.emnlp-main)
Copied to clipboard
Lester James Validad Miranda, Elyanah Aco, Conner G. Manuel, Jan Christian Blaise Cruz, Joseph Marvin Imperial
| Challenge: | Despite impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. |
| Approach: | They propose a benchmark to evaluate LLMs across a diverse set of tasks and capabilities in Filipino, Tagalog, and Cebuano. |
| Outcome: | The proposed benchmark reflects the priorities and trends of research in the Philippines . it finds that several LLMs suffer from reading comprehension and translation capabilities . |
Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models that ground retrieval on external evidence are limited in their ability to implement retrieval-augmented generation. |
| Approach: | They propose a retrieval-augmented generation model that embeds retrieval control directly into generation. |
| Outcome: | The proposed model surpasses strong RAG baselines and uses substantially fewer parameters. |
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)
Copied to clipboard
| Challenge: | a low-resource dataset is limited in training data, so generating task-specific data is challenging. |
| Approach: | They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations. |
| Outcome: | The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings. |
FUDGE: Controlled Text Generation With Future Discriminators (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent advances in large pretrained language models allow us to generate increasingly realistic text by modeling a distribution P (X) over natural language sequences X. |
| Approach: | They propose a flexible and modular method for controlled text generation that uses a Bayesian decomposition of the conditional distribution of G given an attribute predictor and can easily compose predictors for multiple desired attributes. |
| Outcome: | The proposed method can be easily composed and performs three tasks. |
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for assessing social bias in large language models (LLMs) do not capture nuanced and context-dependent nature of natural language generation. |
| Approach: | They propose a Bias Benchmark for Generation (BBG) that evaluates social bias in long-form generation by having LLMs generate continuations of story prompts. |
| Outcome: | The proposed benchmark is based on the English BBQ and Korean BBQ datasets and compares it with multiplechoice BBQ evaluation. |
AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on factual correctness, semantic grounding, visual reasoning, or multimodal large language models. |
| Approach: | They propose a benchmark to assess AICA, which integrates perception, reasoning, and generation into a unified framework. |
| Outcome: | The proposed framework corrects intensity errors and significantly enhances descriptive depth. |
Semantic Space Grounded Weighted Decoding for Multi-Attribute Controllable Dialogue Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Controlling chatbot utterance generation with multiple attributes is a useful but under-studied problem. |
| Approach: | They propose a framework that possesses strong controllability with a weighted decoding paradigm and improves generation quality with an attribute semantics space. |
| Outcome: | The proposed framework achieves high control accuracy with simultaneous control of 3 aspects while producing interesting and sensible responses even in an out-of-distribution robustness test. |
From Style to Story: A Curriculum Learning Approach for Imitative Novel Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Novels create rich, immersive worlds with intricate plots and distinct styles, captivating readers through complex storytelling. |
| Approach: | They propose a novel generation system that imitates novel elements by predicting plot developments and writing concrete details using vivid, expressive language. |
| Outcome: | The novel imitative novel generation system is trained through a curriculum learning paradigm, progressing from low-level stylistic mastery to high-level narrative coherence. |
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)
Copied to clipboard
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello
| Challenge: | Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks . |
| Approach: | They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. |
| Outcome: | The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English. |
Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World Questions (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have shown promising ability to perform commonsense reasoning. |
| Approach: | They propose a two-dimensional analysis framework that incorporates token back-tracing and token decoding to uncover how LLMs conduct factual knowledge recall. |
| Outcome: | The proposed framework shows that LLMs lack relevant knowledge but struggle to select the most accurate information based on context during the retrieval and rerank phase. |
Memory OS of AI Agent (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face a shortage of long-term memory capabilities and limited personalization due to fixed context windows. |
| Approach: | They propose a Memory Operating System to achieve efficient memory management for AI agents . MemoryOS enables hierarchical memory integration and dynamic updating . |
| Outcome: | The proposed architecture enables hierarchical memory integration and dynamic updating. |
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline (2026.acl-long)
Copied to clipboard
| Challenge: | Existing facial forgery detection methods focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. |
| Approach: | They propose a multimodal task that localizes forged regions and generates natural language explanations grounded in editing process. |
| Outcome: | The proposed task localizes forged regions and generates natural language explanations grounded in editing process. |
VISA: Retrieval Augmented Generation with Visual Source Attribution (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to retrieval-augmented generation primarily link generated content to document-level references, making it difficult for users to locate evidence among multiple content-rich retrieved documents. |
| Approach: | They propose a novel approach that combines answer generation with visual source attribution by leveraging large vision-language models to identify evidence and highlight exact regions that support the generated answers with bounding boxes in the retrieved document screenshots. |
| Outcome: | The proposed approach identifies evidence and highlights exact regions that support the generated answers with bounding boxes in the retrieved document screenshots. |