Challenge: Question-Answer Generation (QAG) is essential for domain-specific large language models post-training.
Approach: They propose a framework that balances semantic diversity and factual consistency . they propose entropy and consistency scores that harmonize the trade-off between diversity and correctness .
Outcome: The proposed framework outperforms baseline models in generating diverse QA pairs . the proposed framework harmonizes semantic entropy and consistency scores to quantify trade-off between diversity and correctness.

Similar Papers

Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to improve output quality without aggregating input tokens are limited by the complexity of aggregation of responses.
Approach: They propose to extract and integrate segment-level commonalities from candidate samples to enhance performance of LLMs in open-ended and reasoning tasks.
Outcome: The proposed method improves performance on reasoning, code generation and mathematical reasoning tasks without requiring additional models and overlooking the knowledge present among the candidates.
Explicit over Implict: Explicit Diversity Conditions for Effective Question Answer Generation (2024.lrec-main)

Copied to clipboard

Challenge: Recent pretrained and large language model-based QAG methods suffer from redundant generation of QA pairs, affecting downstream QA systems.
Approach: They propose to use explicit diversity conditions to generate diverse question-answer synthetic data by focusing on spatial aspects, question types, and entities.
Outcome: The proposed diversity conditions significantly increase diversity in QA generation over existing diversity techniques.
Exploring Precision and Recall to assess the quality and diversity of LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models are limited to specific tasks, but they are now widely available for a wide range of tasks.
Approach: They propose a framework for large language models such as Llama-2 and Mistral that imports precision and recall metrics from image generation to text generation.
Outcome: The proposed framework allows for a nuanced assessment of the quality and diversity of generated text without the need for aligned corpora.
Towards Diverse and Effective Question-Answer Pair Generation from Children Storybooks (2023.findings-acl)

Copied to clipboard

Challenge: Recent advances in QA pair generation (QAG) have raised interest in applying this technique to the educational field.
Approach: They propose a QAG framework that enhances QA type diversity by producing different interrogative sentences and implicit/explicit answers.
Outcome: The proposed framework outperforms state-of-the-art methods by significant margins, achieving improved diversity and quality.
Open-World Factually Consistent Question Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for question generation suffer from factual inconsistencies and incorrect entities and are not answerable from the input paragraph.
Approach: They propose a data processing technique based on de-lexicalization for consistent question generation across domains and a model that is generic across question-generation models.
Outcome: The proposed method produces entity-level factually consistent questions without significant impact on traditional metrics.
Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability.
Approach: They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.
Outcome: The proposed evaluation metric consistently shows higher correlation with human judgements.
Contextual Diversity Measure (CDM) for Controllable Story Generation in Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies on controllable text generation focus on controlling attributes such as sentiment, writing style, and writing style.
Approach: They introduce a metric that quantifies semantic diversity for scenario generation under fixed abstract semantic constraints and validate it through controlled experiments.
Outcome: The proposed metric achieves excellent discrimination accuracy (100% and 91.9%, respectively), with discriminative power up to 5.5 greater than the best baseline.
Adaptive Question Answering: Enhancing Language Model Proficiency for Addressing Knowledge Conflicts with Source Citations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing work on citation generation has focused on unambiguous settings with single answers, failing to address the complexity of real-world scenarios.
Approach: They propose a task of QA with source citation in ambiguous settings where multiple valid answers exist, where multiple sources exist.
Outcome: The proposed framework generates multiple answers and cites their sources, allowing users to verify the factuality of each answer and make informed decisions.
DiffQG: Generating Questions to Summarize Factual Changes (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods to identify factual changes between paired documents are limited . specialized entailment-like resources and models have been applied to fact verification .
Approach: They propose to represent factual changes between paired documents as question-answer pairs . they propose to generate a discriminating question given an answer span such that the question is answerable by one passage but not the other .
Outcome: The proposed model can flexibly and concisely capture the updated contents of paired documents.
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting (2025.findings-emnlp)

Copied to clipboard

Challenge: Fine-grained personas have been used for generating ‘diverse’ synthetic data for pre-training and supervised fine-tuning of Large Language Models (LLMs).
Approach: They measure the diversity of persona-driven synthetically generated prompts and responses with a suite of lexical diversity and redundancy metrics.
Outcome: The proposed model is based on human-written prompts and responses, but human-generated prompts are significantly less diverse than human-created ones.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations