GLaPE: Gold Label-agnostic Prompt Evaluation for Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have explored leveraging the LLM itself as an optimizer to identify optimal prompts that maximize task accuracy. |
| Approach: | They propose a gold label-agnostic prompt evaluation method to reduce dependence on gold labels. |
| Outcome: | The proposed method produces more effective prompts even without gold labels. |
Similar Papers
Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels (2025.coling-main)
Copied to clipboard
| Challenge: | Pre-trained language models have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. |
| Approach: | They propose a new paradigm termed zero-to-strong generalization that prompts LLMs to annotate unlabeled data and retain high-quality labels by filtering. |
| Outcome: | The proposed framework outperforms pre-trained language models on extensive classification and reasoning tasks on multiple model sizes. |
MAPO: Boosting Large Language Model Performance with Model-Adaptive Prompt Optimization (2023.findings-emnlp)
Copied to clipboard
Yuyan Chen, Zhihao Wen, Ge Fan, Zhengyu Chen, Wei Wu, Dayiheng Liu, Zhixu Li, Bang Liu, Yanghua Xiao
| Challenge: | Existing research emphasizes the importance of adapting prompts to specific tasks, rather than specific LLMs. |
| Approach: | They propose a model-adaptive prompt optimizer method that optimizes original prompts for each LLM in downstream tasks. |
| Outcome: | The proposed method can optimize prompts for an LLM in downstream tasks. |
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | a high prompt sensitivity has been widely accepted as a core limitation of large language models . a recent study suggests that prompt senescence may be an artifact of evaluation processes . |
| Approach: | They examine whether prompt sensitivity is an inherent weakness or an artifact of evaluation . they find that heuristic evaluation methods overlook semantically correct responses . large language models have achieved remarkable success across a wide range of tasks . |
| Outcome: | The proposed model is more robust to prompt templates than previously thought . the authors show that prompt sensitivity may be an artifact of evaluation rather than a flaw . |
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are useful for low-resource scenarios and time-restricted applications. |
| Approach: | They propose a large-scale evaluation tool for large language models that uses prompts . they evaluate 720 prompt templates for open-source LLM-based metrics on MT and summarization datasets a 6.6M evaluations. |
| Outcome: | The proposed model evaluates 720 prompt templates on machine translation and summarization datasets. |
PromISe: Releasing the Capabilities of LLMs with Prompt Introspective Search (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for large language models use uniform manual prompts, resulting in underestimation of performance. |
| Approach: | They propose a prompt introspective search framework that integrates self-introspect and self-refine to unlock the capabilities of LLMs. |
| Outcome: | The proposed framework significantly boosts the performance of 12 well-known LLMs compared to baseline methods. |
Rethinking Prompt Optimizers: From Prompt Merits to Optimization (2026.eacl-long)
Copied to clipboard
Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Tianjiao Li, Chua Jia Jim Deryl, Lee Onn Mak, Gee Wah Ng, Kezhi Mao
| Challenge: | Existing methods to optimize prompts rely on LLMs' self-generation ability but lack interpretability due to implicit optimization. |
| Approach: | They propose a model-agnostic prompt quality merits and a merit-guided, locally deployable prompt optimizer trained on a lightweight LLM to improve prompt quality. |
| Outcome: | The proposed model avoids online optimization, reduces privacy concerns, and generalizes effectively to both large-scale and lightweight inference models. |
Aligning Large Language Models for Controllable Recommendations (2024.acl-long)
Copied to clipboard
| Challenge: | Existing literature focuses on integrating domain-specific knowledge into LLMs to enhance accuracy using a fixed task template. |
| Approach: | They propose a collection of supervised learning tasks augmented with labels derived from a conventional recommender model to improve LLMs’ proficiency in adhering to recommendation-specific instructions. |
| Outcome: | The proposed approach significantly improves the capability of LLMs to respond to instructions within recommender systems, reducing formatting errors while maintaining a high level of accuracy. |
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs (2024.naacl-long)
Copied to clipboard
| Challenge: | Large language models exhibit undesirable preference toward predicting certain answers over others, despite their adaptability to diverse tasks. |
| Approach: | They propose a label bias calibration method that outperforms recent calibration approaches for improving performance and mitigating label bias. |
| Outcome: | The proposed method outperforms calibration approaches for improving performance and mitigating label bias. |
Gradient-Guided Multi-Judge Prompt Optimization (2026.acl-long)
Copied to clipboard
ChenZhuo Zhao, Xinda Wang, Pu Zhao, Yue Huang, Junting Lu, Ziqian Liu, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
| Challenge: | Existing approaches to prompt optimization trade off signal quality against computational cost. |
| Approach: | They propose a framework that uses a first-order gradient approximation to score segment importance in a continuous masking direction. |
| Outcome: | The proposed framework improves efficiency and robustness by using a first-order gradient approximation to score segment importance in a continuous masking direction. |
FIPO: Free-form Instruction-oriented Prompt Optimization with Preference Dataset and Modular Fine-tuning Schema (2025.coling-main)
Copied to clipboard
| Challenge: | naive prompts can enhance the task performance of large language models, but they are resource-intensive. |
| Approach: | They propose an automatic prompt optimization method that refines naive prompts according to task outputs from in-box testing models. |
| Outcome: | The proposed method is based on a large-scale dataset and performed fairly across multiple models. |