REVerSum: A Multi-staged Retrieval-Augmented Generation Method to Enhance Wikipedia Tail Biographies through Personal Narratives (2025.coling-industry)
Copied to clipboard
| Challenge: | Creating new articles and editing older ones is expensive and time-consuming. |
| Approach: | They propose a multi-staged retrieval-augmented generation technique to leverage personal narratives to enhance Wikipedia’s B and C category biography articles. |
| Outcome: | The proposed approach outperforms the best performing baseline by 17% in terms of integrability to the original Wikipedia article and 28.5% in terms informativeness. |
Similar Papers
Generating Biographies on Wikipedia: The Impact of Gender Bias on the Retrieval-Based Generation of Women Biographies (2022.acl-long)
Copied to clipboard
| Challenge: | Existing efforts to encourage article creation focus on reducing the gender gap in Wikipedia articles. |
| Approach: | They propose a model that retrieves web evidence and generates biographies section by section . they analyze available web evidence to determine the accuracy of the generated text . |
| Outcome: | The proposed model can generate biographies section by section, including citation information, using retrieval mechanisms and a cache-based pre-trained encoder-decoder. |
Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work has proposed to improve relevance modeling by having large language models actively involved in retrieval, i.e., to guide retrieval with generation. |
| Approach: | They propose to have large language models actively involved in retrieval to guide retrieval with generation. |
| Outcome: | The proposed method synergizes retrieval and generation in an iterative manner, and can generate better results in subsequent iterations. |
Enhancing Retrieval-Augmented Generation via Evidence Tree Search (2025.acl-long)
Copied to clipboard
| Challenge: | Evidence retrieval is used to enhance Large Language Models (LLMs) but in real-world applications, it often returns lengthy documents with redundant or irrelevant content, confusing downstream readers. |
| Approach: | They propose a framework that reformulates evidence retrieval as a dynamic tree expansion process. |
| Outcome: | The proposed framework outperforms existing methods on five datasets. |
Retrieval-augmented Generation across Heterogeneous Knowledge (2022.naacl-srw)
Copied to clipboard
| Challenge: | Existing methods for retrieving knowledge from a single source homogeneous corpus have been gaining increasing attention in the field of natural language processing (NLP) however, they still suffer from the following drawbacks: (i) They are usually trained offline, making the model agnostic to the latest information, e.g., asking a chat-bot about COVID-19. |
| Approach: | They propose to use a single-source homogeneous corpus to generate retrieval-augmented generation models that can learn from the pre-training corpus. |
| Outcome: | The proposed methods have been applied to various knowledge-intensive NLP tasks, but most of the work has focused on retrieving unstructured text documents from Wikipedia. |
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent research focuses on optimizing the use of Self-Docs with their inherent properties remaining underexplored. |
| Approach: | They develop a taxonomy to compare the effectiveness of different types of Self-Docs and explore strategies for combining them with external sources. |
| Outcome: | The proposed model can supplement retrieved content and provide a powerful way to improve knowledge-intensive question answering tasks. |
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)
Copied to clipboard
Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Xuan Long Do, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, Shafiq Joty
| Challenge: | Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities. |
| Approach: | They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs. |
| Outcome: | The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content. |
Targeted Augmentation for Low-Resource Event Extraction (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for low-resource information extraction struggle to strike a balance between weak augmentation and drastic augmentation. |
| Approach: | They propose a data augmentation paradigm that uses back validation and targeted augmentation to produce augmented examples with enhanced diversity, polarity, accuracy, and coherence. |
| Outcome: | The proposed paradigm produces augmented examples with enhanced diversity, polarity, accuracy, and coherence. |
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation (2025.findings-naacl)
Copied to clipboard
Yiruo Cheng, Kelong Mao, Ziliang Zhao, Guanting Dong, Hongjin Qian, Yongkang Wu, Tetsuya Sakai, Ji-Rong Wen, Zhicheng Dou
| Challenge: | Existing research focuses on single-turn RAG, leaving a gap in addressing multi-turn conversations . a new benchmark is designed to assess RAG systems in realistic multi-turned conversations based on Wikipedia . |
| Approach: | They propose a large-scale benchmark to assess RAG systems in multi-turn contexts . CORAL includes diverse information-seeking conversations automatically derived from Wikipedia . authors propose unified framework to standardize various conversational RAG methods . |
| Outcome: | The proposed framework supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling. |
Tell Me Again! a Large-Scale Dataset of Multiple Summaries for the Same Story (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to represent narratives on short-form texts are limited as narrative semantics are an open class. |
| Approach: | They propose to use Wikipedia summaries as a proxy for entire stories or for analysis of the summary itself. |
| Outcome: | The proposed dataset contains 96,831 individual summaries across 29,505 stories. |
OASum: Large-Scale Open Domain Aspect-based Summarization (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing generic summarization methods generate only one summary for all different requests which is not optimal for diverse demands. |
| Approach: | They use crowd-sourced knowledge on Wikipedia to create a large-scale open-domain aspect-based summarization dataset with 1 million different aspects on 2 million Wikipedia pages. |
| Outcome: | The proposed model can generate diverse aspect-based summarizations on Wikipedia with zero/few-shot and fine-tuning on seven downstream datasets. |