Challenge: Existing approaches to generate Boolean queries for systematic reviews are limited by the lack of ground-truth best Boolesan queries.
Approach: They propose a reinforcement learning framework that trains large language models to generate effective Boolean queries for medical systematic reviews.
Outcome: The proposed framework outperforms zero-shot/few-shot prompting on 65 588 topics . it also matches or exceeds the effectiveness of larger GPT-based models using smaller backbones .

Similar Papers

SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in machine translation, but most MT-specific LLMs rely heavily on external supervision during training.
Approach: They propose a reinforcement learning framework for machine translation that is reference-free and relies solely on self-judging rewards.
Outcome: The proposed framework outperforms existing LLMs and larger general LLM models on English Chinese translation benchmarks and performs competitively with leading closed-source systems.
Reference-Free Schema Generation for Literature Review Tables via Multi-Faceted Rewards (2026.acl-srw)

Copied to clipboard

Challenge: Literature review systems generate literature review tables by inferring schemas and values from documents.
Approach: They propose to use schema generation as a reinforcement learning problem to determine which dimensions to compare a set of papers.
Outcome: The proposed model improves over the untuned model across intrinsic, reference-based, and LLM-judge metrics and remains competitive with supervised fine-tune models at 5 the parameter count on structural and diversity dimensions.
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient (2026.acl-long)

Copied to clipboard

Challenge: Using generic and efficient benchmark generators, human annotators are limited by inefficiency . current benchmark generator methods rely on seed signals, leading to long cycles and high costs .
Approach: They propose a framework to evaluate LLMs as generic benchmark generators and integrate them as BenchMaker.
Outcome: The proposed framework achieves comparable performance to human-annotated benchmarks on most metrics.
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large-scale reinforcement learning (RL) methods have proven effective in enhancing the reasoning abilities of large language models.
Approach: They propose an open-source adaptation of the R1-Zero RL framework for machine translation (MT) their code is available at https://github.com/fzp0424/MT-R1-zero.
Outcome: The proposed framework surpasses towerinstruct-7B-v0.2 on the english-chinese benchmark by 1.26 points.
Select High-quality Synthetic QA Pairs to Augment Training Data in MRC under the Reward Guidance of Generative Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches focus on downstream metrics to select QA pairs, which lack generalization across different datasets.
Approach: They propose a general selection method that uses a large pre-trained language model as a reward model in a Reinforcement Learning framework for the training of the selection agent.
Outcome: The proposed method improves performance on generative and extractive datasets.
Generation-driven Contrastive Self-training for Zero-shot Text Classification with Instruction-following LLM (2024.eacl-long)

Copied to clipboard

Challenge: a novel method to train a smaller model with LLMs for zero-shot text classification requires immense computational resources due to their substantial model size.
Approach: They propose a method which leverages the generative power of large language models to train a smaller model.
Outcome: The proposed method outperforms state-of-the-art methods when limited data is available.
ReviewRL: Towards Automated Scientific Review with RL (2025.emnlp-main)

Copied to clipboard

Challenge: Existing automated review systems struggle with factual accuracy, rating consistency, and analytical depth.
Approach: They propose a framework for generating comprehensive and factually grounded scientific paper reviews using supervised fine-tuning and reinforcement learning.
Outcome: The proposed framework outperforms existing methods on ICLR 2025 papers.
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are a promising solution to automate literature review writing tasks.
Approach: They propose a framework to automatically evaluate the performance of large language models in three key tasks of literature review writing: reference generation, abstract writing, and literature review composition.
Outcome: The proposed framework assesses the hallucination rates in generated references and measures the semantic coverage and factual consistency of the literature summaries and compositions against human-written counterparts.
Selecting Better Samples from Pre-trained LLMs: A Case Study on Question Generation (2023.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive prowess in natural language generation.
Approach: They propose a method to select high-quality questions from LLM-generated candidates using round-trip and prompt-based scoring.
Outcome: The proposed approach can select high-quality questions from a set of LLM-generated candidates without modification of the underlying model nor rely on human annotations.
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches typically assume access to ground-truth labeled data . Existing methods require a classifier to select models given an input .
Approach: They propose a routing setting where routers are trained exclusively on generated queries and answers from LLMs.
Outcome: The proposed router outperforms the best query-answer router by 4.6% absolute accuracy when trained on weak generator data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations