PRL: Prompts from Reinforcement Learning

About

Effective prompt engineering remains a central challenge in fully harnessing the capabilities of LLMs. While well-designed prompts can dramatically enhance performance, crafting them typically demands expert intuition and a nuanced understanding of the task. Moreover, the most impactful prompts often hinge on subtle semantic cues, ones that may elude human perception but are crucial for guiding LLM behavior. In this paper, we introduce PRL (Prompts from Reinforcement Learning), a novel RL-based approach for automatic prompt generation. Unlike previous methods, PRL can produce novel few-shot examples that were not seen during training. Our approach achieves state-of-the-art performance across a range of benchmarks, including text classification, simplification, and summarization. On the classification task, it surpasses prior methods by 2.58% over APE and 1.00% over EvoPrompt. Additionally, it improves the average ROUGE scores on the summarization task by 4.32 over APE and by 2.12 over EvoPrompt and the SARI score on simplification by 6.93 over APE and by 6.01 over EvoPrompt. Our code is available at https://github.com/Batorskq/prl .

Pawe{\l} Batorski, Adrian Kosmala, Paul Swoboda• 2025

Related benchmarks

Task	Dataset	Result
Mathematical Reasoning	GSM8K	Accuracy86.15	499
Mathematical Reasoning	MATH 500	Accuracy44.4	442
Text Classification	AG News (test)	Accuracy84.36	293
Text Classification	TREC	Accuracy77.07	281
Text Classification	SST-2 (test)	Accuracy96.32	185
Text Classification	MR	Accuracy91.27	174
Text Classification	MR (test)	Accuracy91.27	155
Medical Question Answering	MedQA	Accuracy53.34	153
Subjectivity Classification	Subj (test)	Accuracy76.9	152
Multi-hop QA	HotpotQA	Exact Match57.1	143

Showing 10 of 29 rows

Other info

Follow for update

@wizwand_team Discord