Challenge: Existing methods to improve language models require manual ranking and annotators.
Approach: They propose a self-supervised text ranking approach for applying Proximal-Policy-Optimization to fine-tune language models while eliminating the need for human annotators.
Outcome: The proposed method significantly outperforms baselines regarding BLEU, GLEU, and METEOR scores on three tasks and is consistent with humans.

Similar Papers

Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Instruction-fine-tuned large language models (LLMs) under 14B parameters underperform on NLU tasks . we explore a framework to improve the NLU capabilities of LLMs .
Approach: They propose to use Proximal Policy Optimization to improve NLU capabilities . they frame NLU as a reinforcement learning environment and optimize for reward signals .
Outcome: The proposed framework outperforms supervised fine-tuning on GLUE and superGLUE tasks.
P-TA: Using Proximal Policy Optimization to Enhance Tabular Data Augmentation via Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Contemporary approaches to generate tabular data are limited due to the lack of external knowledge.
Approach: They propose to use proximal policy optimization to apply GANs and fine-tune Large Language Models to enhance the probability distribution of tabular features.
Outcome: The proposed method improves accuracy of GANs and LLMs over state-of-the-art over three real-world datasets.
Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Reward hacking is a problem in reinforcement learning where the ability to specify the desired behavior of a reward function is difficult.
Approach: They propose to use feedback as a potential-based shaping function to solicit and apply feedback from large language models to improve convergence speed and policy returns.
Outcome: The proposed method improves convergence speed and policy returns over baselines even with significant ranking errors and eliminates the need for complex post-processing of reward functions.
RecGPT: Generative Pre-training for Text-based Recommendation (2024.acl-short)

Copied to clipboard

Challenge: Existing models for text-based recommendation lack data sparsity and flexibility to capture fluctuations in user preferences over time.
Approach: They present the first domain-adapted and fully-trained large language model for text-based recommendation.
Outcome: The proposed model outperforms baseline models on rating prediction and sequential recommendation tasks.
RLHF Algorithms Ranked: An Extensive Evaluation Across Diverse Tasks, Rewards, and Hyperparameters (2025.emnlp-industry)

Copied to clipboard

Challenge: Proximal Policy Optimization (PPO) has fallen out of favor for Large Language Models (LLMs), but its complexity and inefficiency have spurred the investigation of simpler alternatives.
Approach: They evaluate 17 RLHF algorithms on two benchmarks, OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness.
Outcome: The proposed methods are based on OpenAI’s TL;DR Summarization and Anthropic’s Helpfulness / Harmlessness benchmarks with two different reward models and a Rules based reward model.
LeLoRA: Learnable Low-Rank Adaptation of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to fine-tuning large language models (LLMs) rely on manually specified and fixed hyperparameters, resulting in suboptimal performance and low parameter efficiency.
Approach: They propose a framework that allows for dynamically learned adaptive adaptation strategies to be used to fine-tune large language models.
Outcome: The proposed framework outperforms baselines in adapting large language models.
Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs (2024.acl-long)

Copied to clipboard

Challenge: Proximal Policy Optimization (PPO) is used for RLHF but requires high computational cost and sensitive hyperparameter tuning.
Approach: They propose to use Proximal Policy Optimization to align large language models to human preferences.
Outcome: The proposed method preserves and even increases performance while preserving the motivational principles that led to the development of PPO.
Aligning Large Language Models via Fine-grained Supervision (2024.acl-short)

Copied to clipboard

Challenge: Pre-trained large-scale language models often generate biased or toxic text, misaligning with human intentions.
Approach: They propose to use human feedback to improve LLM alignment by fine-grained token supervision . they ask annotators to edit less preferred responses to make them more favorable .
Outcome: The proposed method improves LLM alignment by up to 5.1% in terms of win rate compared with the traditional model.
Aligning Large Language Models through Synthetic Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, alignment learning requires significant human demonstrations and feedback from proprietary LLMs such as ChatGPT.
Approach: They propose a framework that uses synthetic feedback to align large language models to human values without extensive human annotations and proprietary LLMs.
Outcome: The proposed model outperforms open-source models on human-annotated demonstrations in alignment benchmarks.
Language Models are Few-Shot Butlers (2021.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models demonstrate strong performance in most NLP tasks when fine-tuned on small task-specific datasets.
Approach: They propose a two-stage procedure to learn from a small set of demonstrations and a simple reinforcement learning algorithm to improve by interacting with an environment.
Outcome: The proposed method improves with only 1.2% of the demonstrations and a simple reinforcement learning algorithm over existing methods in the ALFWorld environment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations