來源:arXiv · cs.AI查看原文 ↗
原文著作權歸來源方所有,本站僅作收錄、翻譯或格式整理。
事實脈絡
解讀與影響
TAPR的核心创新在于其训练方式。研究团队没有依赖昂贵的人工标注数据,而是采用了一种基于强化学习的自动化流程。具体来说,他们使用了组相对策略优化(GRPO)算法,并巧妙地利用大语言模型自身作为“裁判”[来源:arxiv.org]。这个裁判会同时评估重写后的提示词质量以及最终生成的任务结果,并将这些评估转化为奖励信号来训练TAPR。这意味着TAPR在学习如何改写提示时,直接以提升下游任务表现为目标,而非仅仅追求语法通顺。
从技术实现上看,TAPR的价值在于将复杂的“提示工程”能力模型化、自动化。对于普通用户而言,他们无需了解思维链、角色扮演等高级提示技巧,只需像日常对话一样提出需求,TAPR就能在后台将其重构为结构清晰、逻辑严谨的优化版本。这不仅能显著提升回答的准确性和相关性,也大幅降低了使用AI的门槛。原文未提供具体的性能提升数据,但其设计理念与当前业界追求AI易用性的趋势高度一致,例如,一些企业服务也在探索如何通过优化交互来提升生产力[来源:OpenAI Blog]。
该研究的另一个值得关注的启示是,它揭示了强化学习在优化语言模型交互层面的潜力。传统上,我们关注的是用强化学习微调模型本身的知识和推理能力,而TAPR则将优化目标前置到了输入阶段。这种“重写器+生成器”的分离式架构,或许能为未来更复杂、更可控的AI应用提供一种新的设计范式,让模型能力的释放不再依赖于用户的操作技巧。
參考來源
來源原文
\copyrightclause Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
\conference [PromptEng] Third International Workshop on Prompt Engineering for Pre-Trained Language Models co-located with the ACM WebConf, April.13, 2026, Dubai, United Arab Emirates
[email=oliver.savolainen@student.uva.nl, ]
[email=e.bastianelli@elsevier.com, ]
[email=h.azarbonyad@elsevier.com, ]
[email=a.lucic@uva.nl, ] TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
Oliver Savolainen University of Amsterdam, Amsterdam, The Netherlands
Emanuele Bastianelli Elsevier, Amsterdam, The Netherlands
Hosein Azarbonyad
Ana Lucic ( 2026 )
Abstract Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the challenge by introducing a Task-Aware Prompt Rewriter (TAPR), a model that reformulates user prompts into task-optimized prompts with the explicit goal of improving downstream LLM performance. We train TAPR using reinforcement learning with Group Relative Policy Optimization (GRPO), where rewards are derived from LLM-as-judge evaluations of both the reformulated prompt and the corresponding task output. Experimental results on diverse tasks, such as question answering, summarization, and arithmetic reasoning, show that our method yields consistent gains over base models in prompt rewriting ability. Fine-tuning Phi-4-mini-instruct (as the base model for TAPR) produces prompts that contain clearer and more instructive language, leading to higher accuracy on established benchmarks such as Natural Questions and GSM8K. Our code is available at: https://github.com/OliverSavolainen/task-specific-prompt-rewriter
keywords:
Automated Prompt Rewriting \sep Large Language Models (LLMs) \sep Reinforcement Learning (RL) \sep LLM-as-a-Judge
1 Introduction
Large Language Models (LLMs) have revolutionized the field of natural language processing, providing unprecedented capabilities in text generation, comprehension, and various other applications [ brown2020language , wei2022emergent ] . Despite their impressive performance, these models often require carefully crafted input prompts to produce the best possible responses [ brown2020language , srivastava2022beyond ] . Users, particularly those without expertise in prompt engineering, may find it challenging to formulate prompts that fully leverage the capabilities of these advanced models. This gap underscores the necessity for an automated system that can refine and optimise user prompts, making high-quality interaction with LLMs more accessible [ autop , prewrite ] .
Existing methods for prompt optimization rely largely on manual adjustments and iterative testing, which are not only time-consuming but also require a certain level of expertise. Automated prompt rewriting can bridge this gap by refining under‐specified queries into clear, robust instructions that ensure reliable performance for all users.
In this paper, we present a Task-Aware Prompt Rewriter (TAPR) , a lightweight model that automatically refines and adapts user prompts for LLMs, with the aim of enhancing downstream task performance. This system involves a small, specialized model dedicated to rewriting prompts, coupled with a larger LLM that executes the task based on these optimized prompts. To train TAPR, we build on the recent advancements on using reinforcement learning to improve LLM performance across a range of tasks. Therefore, the main research question of this work is:
How can we create a task-aware prompt rewriter such that it enhances the performance of LLMs in downstream tasks?
We use reinforcement learning to train a smaller LLM to optimize default prompts by rewriting them, using feedback derived from the performance of a frozen Task LLM responding to the rewritten prompts. Prior work has explored using reinforcement learning for this purpose. PRewrite [ prewrite ] applied PPO (Proximal Policy Optimization) [ ppo ] to fine-tune LLMs on individual datasets, yielding a single optimal prompt per dataset and demonstrating gains on several benchmarks.
Our approach extends this in two key ways. First, beyond the verifiable task-performance rewards used by prior work, we introduce using LLM-as-a-judge rewards for semantically more accurate evaluations, including a reward based on the quality of the prompts themselves, and we validate its impact through a dedicated ablation study. Second, we employ Group Relative Policy Optimization (GRPO) [ grpo ] in place of PPO and achieve great consistency with this algorithm. We also test generating multiple prompts per dataset and having TAPR select the best one, but we do not find this addition leading to further performance gains consistently.
To evaluate our method, we measure the performance of the Task LLM before and after rewriting on various tasks, including question answering, summarization, and arithmetic reasoning. Our framework leads to improvement in the prompt rewriting abilities of Phi-4-mini-instruct and LLaMA-3.2-3B-Instruct on multiple tasks. For Phi-4-mini-instruct, the highest metric score comes from a TAPR variant for each dataset.
In summary, our contributions are:
• We introduce TAPR , a task-aware prompt rewriter trained with reinforcement learning to optimize task-specific prompts using feedback derived from both rewritten prompts and responses from LLM performing the tasks.
• We leverage the LLM-as-a-judge paradigm to obtain semantically informed feedback and design a dedicated prompt quality reward , demonstrating that using LLM-based evaluations improves training stability and effectiveness.
• We perform various experiments across question answering, summarization, and arithmetic reasoning benchmarks, showing that prompts rewritten by TAPR enhance task performance compared to baseline prompts and rewrites from base models.
2 Related Work
In order to understand how to design the method for training a model to rewrite prompts, we look at previous work on LLM fine-tuning, prompt engineering, and existing prompt rewriting works. In addition, we look at the LLM-as-a-judge paradigm to explore alternative ways of evaluating and rewarding models.
Initial Prompt TAPR LLM (smaller) Rewritten Prompt Task LLM (frozen) Task Output Reward: Metrics / LLM-as-Judge Selection* GRPO Figure 1 : Overview of the training pipeline for TAPR. Orange nodes are prompts (inputs to models), blue nodes represent models (trainable or frozen), and green nodes are outputs —either the task answer or the scalar reward. Solid arrows show the forward data flow; the green reward is fed back via GRPO (curved arrow).
2.1 LLM Fine-tuning
Supervised fine-tuning (SFT) is a general technique for adapting any pretrained LM to a labeled dataset; in the context of generative models it’s often called instruction-tuning, since the labels take the form of (prompt, response) pairs [ wei2021finetuned ] . Although SFT provides strong initialization for instruction following, it does not explicitly optimize for long-term objectives, such as user satisfaction or safety. Reinforcement learning from human feedback (RLHF) addresses this by training a scalar reward model on human preference judgments between model outputs, and then using Proximal Policy Optimization (PPO) to maximize that learned reward [ rlhf , ppo ] . In the InstructGPT pipeline, a policy initialized by SFT is refined via RLHF, yielding substantial gains in helpfulness and coherence over purely supervised approaches [ instructgpt ] .
More recently, Group Relative Policy Optimization (GRPO) was introduced which generalizes PPO by updating the policy based on a group of sampled trajectories, thereby removing explicit value function learning and enabling multi-task or multi-agent training regimes, leading to impressive performance for the DeepSeek R1 model [ grpo , deepseek ] . Both GRPO and PPO have also been used to enhance other abilities, for example, Search-R1 [ jin2025searchr1trainingllmsreason ] showed the effectiveness of both algorithms to make LLMs better in the use of search engines. Our work is the first to explore GRPO for enhancing prompt rewriting abilities.
2.2 Prompt Engineering
Task framing has been proven to have a major impact on LLM performance. One of the first important breakthroughs for this was few-shot prompting, where providing just a handful of examples enables generalization without task-specific training [ brown2020language , boonstra2025prompt ] . Chain-of-thought prompting (i.e., “Let’s think step by step”) elicits intermediate reasoning and boosts accuracy on arithmetic and logical tasks [ kojima2022zeroshot , boonstra2025prompt ] . Few-shot chain-of-thought prompting remains essential for challenging reasoning benchmarks like GSM8K and StrategyQA, teaching models to break problems into steps [ wei2022chain ] .
Clear, structured instructions also matter. Specifying output formats, especially JSON, improves consistency and makes results easier to parse [ shen2017style , hu2017toward , long2024llmsbiasedoutputformats ] . Yet prompt design is still difficult, with no universally optimal formulation. Our approach tackles this by automatically rewriting prompts using prompt engineering principles.
2.3 Prompt Rewriting
Automated prompt engineering or prompt rewriting employs diverse optimization techniques to refine prompts that steer LLM behavior. AutoPrompt uses gradient‐based search over discrete token spaces to assemble sequences that elicit desired outputs, though often producing non‐interpretable prompts [ autop ] . Evolutionary methods such as PromptBreeder treat prompt design as a population‐based search, applying crossover and mutation to candidate prompts and selecting offspring by performance [ breed ] , but rely on carefully chosen initial prompts, limiting the search space.
Reinforcement learning offers a more sample‐efficient alternative by framing prompt generation as a sequential decision process. RLPrompt trains an agent to edit prompts token by token with rewards from task performance (e.g., classification accuracy, BLEU score), discovering prompts that generalize across datasets but remain hard to interpret [ rlprompt ] . TEMPERA improves stability by defining a custom action space over high‐level edits (e.g., paraphrasing, task description insertion) and using Q‐learning to maximize a composite reward of performance and linguistic diversity [ tempera ] , though its constrained actions may miss novel formulations.
PRewrite adopts PPO to fine‐tune LLMs (e.g., PaLM 2‐S/L [ palm ] ) to rewrite user prompts in a task‐aware manner [ prewrite ] . Unlike token‐level policies or fixed edit spaces, it processes the entire prompt and outputs a rewritten version optimized for metrics such as exact match rate on the Natural Questions dataset or solution rate for GSM8K. Empirical results show that PRewrite improves task performance and yields more interpretable and concise prompts than RLPrompt, without TEMPERA’s constraints. However, it only uses automated metrics such as exact match accuracy, which can lead to prompts that do not optimize for real-world utility. Our work leverages LLM-based judging to account for that limitation.
2.4 LLM-as-a-Judge
Traditional automatic evaluation metrics such as BLEU [ papineni2002bleu ] , ROUGE [ lin2004rouge ] , METEOR [ banerjee2005meteor ] , and BERT-Score [ zhang2020bertscore ] rely on n-gram overlap or embedding similarity with static reference texts, which can poorly reflect real‐world utility in open‐ended tasks such as summarization, dialogue, or long‐form question answering [ automatic ] . These metrics often penalize valid paraphrases or novel content, and cannot account for pragmatics, coherence, or factuality beyond surface overlap. To address these shortcomings, the LLM-as-a-judge paradigm repurposes LLMs to score or rank model outputs by simulating human evaluators [ turn0academia8 , gu2024survey ] . By prompting the judge model with evaluation criteria such as relevance, fluency, and informativeness, it can produce scalar scores or pairwise preferences that align more closely with human judgments and adapt dynamically to new tasks without costly reference annotation.
Beyond zero-shot scoring, LLM judges have been integrated directly into model training loops. For example, RLAIF demonstrates that synthetic preferences generated by an auxiliary LLM can approximate human feedback for RLHF, enabling large‐scale reinforcement learning at a lower cost with minimal quality drop-off [ rlaif , bai2022constit ] . This does not mean that using LLM judges leads to a perfect and faultless metric for every situation, but using LLMs to judge can often offer a better way to evaluate or calculate rewards for text generation, potentially leading to better LLM fine-tuning performance for various tasks. Our method is the first to integrate LLM-as-a-judge scores for rewards and evaluations to guide and improve automated prompt rewriting.
Dataset
Example
Initial Prompt
Natural Questions (NQ) [ kwiatkowski-etal-2019-natural ]
Who was the first president of the United States?
“Answer the question.” [ prewrite ]
HotpotQA [ yang-etal-2018-hotpotqa ]
Which novel did Mary Shelley write that was published in 1818?
“Answer the question based on the context.”
CNN/Daily Mail [ hermann2015teaching ]
Article excerpt: An Iraqi boy, badly burned, is flown to the U.S. for treatment funded by a charity.
“Summarize the text.”
SciTLDR [ cachola-etal-2020-tldr ]
Abstract excerpt: We propose a 2-simplicial Transformer for deep RL and show it improves reasoning.
“Summarize the text.”
GSM8K [ cobbe2021training ]
Katy mixes sugar and water in a 7:13 ratio. If she used 120 units total, how many units of sugar were used?
“SOLUTION.” [ prewrite ]
Table 1: All the datasets used in the experiments, example samples from them, and initial baseline prompt for rewriting we are using for them.
3 Task-Aware Prompt Rewriter
In this section, we describe the details of our method for creating the Task-Aware Prompt Rewriter (TAPR). The overall training method for TAPR can be seen in Figure 1 . The training procedure works as follows: for each sample, we rewrite the initial prompt using the TAPR LLM, then we give the rewritten prompt to the Task LLM, which performs the task, and finally, we calculate the rewards using both the rewritten prompt and the task. Based on the reward values, the model is trained with a reinforcement learning algorithm.
3.1 Problem Formulation
TAPR is a smaller LLM that reformulates queries for a bigger frozen Task LLM. We formulate the problem similarly to [ prewrite ] and [ Li2025Survey ] . Let 𝒫 \mathcal{P} denote the space of textual prompts. We consider the TAPR LLM π θ : 𝒫 → 𝒫 \pi{\theta}:\mathcal{P}\to\mathcal{P} that given an original prompt p ∈ 𝒫 p\in\mathcal{P} outputs a revised prompt p ~ = π θ ( p ) \tilde{p}=\pi{\theta}(p) . The revised prompt is then fed into a frozen Task LLM L L , yielding an output y = L ( p ~ ) y=L(\tilde{p}) . Let y ∗ y^{} be the desired (ground-truth) output corresponding to input p p . We define two reward functions: a task performance reward T ( y , y ∗ ) T(y,y^{}) (such as LLM-as-a-judge score or accuracy), measuring how well y y matches y ∗ y^{*} , and a prompt quality reward Q ( p ~ ) Q(\tilde{p}) that encourages the rewritten prompt to be accurate and follow common prompt engineering principles.
For reinforcement learning, we treat π θ \pi_{\theta} as a policy over the discrete prompt space. We define a combined reward
R ( y , p ~ ) = α T ( y , y ∗ ) + β Q ( p ~ ) , R(y,\tilde{p})=\alpha,T(y,y^{*})+\beta,Q(\tilde{p}),
where α , β > 0 \alpha,\beta>0 are weighting coefficients. The reinforcement learning objective is to maximize the expected reward under the data distribution 𝒟 \mathcal{D} of prompts:
𝒥 RL ( θ ) = 𝔼 p ∼ 𝒟 p ~ ∼ π θ ( ⋅ ∣ p ) [ R ( L ( p ~ ) , p ~ ) ] . \mathcal{J}{\mathrm{RL}}(\theta)=\mathbb{E}{\begin{subarray}{c}p\sim\mathcal{D}\ \tilde{p}\sim\pi_{\theta}(,\cdot\mid p)\end{subarray}}\bigl[R\bigl(L(\tilde{p}),,\tilde{p}\bigr)\bigr].
We optimize the expectation via policy gradient methods such as PPO [ ppo ] or GRPO [ grpo ] , treating prompts as policies. This formulation follows the optimization perspective of prompt tuning (treating prompts as differentiable policies and maximizing downstream task metrics).
3.2 Rewards
We train TAPR using the GRPO algorithm [ grpo ] , leveraging its recent success in advancing reasoning and other abilities within language models. It has also been used more commonly with verifiable rewards such as accuracy rather than rewards coming from a model trained on preference data, such as PPO and DPO [ dpo , ppo , grpo ] .
The TAPR model gets the initial prompts as input, and rewrites them. Then both the prompts and responses from the Task LLM are evaluated. The initial prompt can be a general instruction such as "Answer the question" , or it can also include the question and context if one exists. Example meta-prompt is given in Appendix A.1 . We use two types of rewards with equal weights: Task LLM performance rewards and the prompt quality reward. For the former, we use scores from an LLM-as-a-judge evaluator rather than traditional metrics such as exact match accuracy and ROUGE due to their limitations.
This approach enables the evaluation of semantic correctness, accounting for cases in which LLMs generate responses that are more verbose than the ground-truth labels but nonetheless accurate 1 1 1 We verified the effectiveness of this approach through running an internal test, and we evaluated that only 1 out of 100 of the judgments on the Natural Questions were incorrect. .
The prompt quality reward is a score on a scale of 0 to 5 coming from an LLM judge based on whether the rewritten prompt matches the meaning of the initial instruction and improves it according to common prompt engineering principles. This decreases the chances of generating prompts with drastic inaccuracies (for example, by completely changing the meaning of the initial instruction or even just essentially repeating the rewriting instruction). As LLM-as-a-judge usage can lead to unwanted bias, we try to mitigate this by using two different models as judges.
3.3 Prompt Selection Mechanism
To further improve rewriting performance, we introduce a Selection mechanism, where the trained TAPR model selects the most promising prompt from multiple generated options sampled with a higher temperature during inference. We hypothesize that, while this selection is not the main task of the TAPR LLM, the trained model has gathered enough knowledge about prompt engineering to evaluate the quality of prompts in addition to generating them. We compare this approach to an approach where prompts are generated by models using a temperature set to zero.
4 Experimental Setup
We evaluate the performance of TAPR across different tasks, including question answering, summarization, and arithmetic reasoning (GSM8K) [ cobbe2021training ] . Although we also report standard metrics for these datasets, our primary focus is on LLM-as-a-judge evaluations for question answering and summarization, as these better address the limitations of conventional evaluation methods as detailed in Section 2.4 .
In Table 1 , we show examples from each dataset we are using alongside the initial baseline prompts we are using for rewriting.
We evaluate prompt rewriting performance across a variety of models, tasks, and training methods. Each experiment involves two distinct LLMs: the TAPR LLM, which generates improved prompts, and the Task LLM, which solves the downstream task using these rewritten prompts. Our main experiments use LLaMA-3.1-8B-Instruct [ touvron2023llama ] as the Task LLM and compare two open-source TAPR LLMs: Phi-4-mini-instruct [ microsoft2024phi4 ] , LLaMA-3.2-3B-Instruct [ meta2024llama ] . Results with Qwen3-4B [ qwen ] are in Appendix C .
Reinforcement Learning Training Setup and Stopping Criterion: We train each TAPR LLM using reinforcement learning, precisely the GRPO [ grpo ] algorithm, until convergence, defined by a lack of improvement in the 25-step moving average of the reward for 100 consecutive steps. This stopping criterion prevents excessive training once the model’s performance plateaus. The rewards used are specified in Section 3.2 . To reduce bias and stabilize the reward signal, we average the LLM-as-a-judge scores from two different judge models (LLaMA-3.1-8B-Instruct and GPT-4o-mini).
Hyperparameters for training are listed in Appendix B . To maximize experimental coverage rather than repeat identical runs with different seeds, each training and evaluation is conducted only once.
We observed that the TAPR LLM might generate responses containing information other than the rewritten queries, such as explanations for the rewriting decisions made or responses to the initial task. This might lead to the Task LLM being unable to answer the query. Therefore, we instruct TAPR to use a JSON format. Apart from this change, we closely follow most templates and initial prompts used in PRewrite [ prewrite ] . This gives us reliable baselines and a starting point to improve any other prompts. A full example of a meta-prompt can be seen in Appendix A.1 .
Evaluation Protocol: Final evaluations are conducted on 1,000 samples from the validation or test split of each dataset. During evaluation, we use GPT-4o-mini as the sole judge model to score answers (avoiding any bias from using the same model for generation and judgment). In our internal tests (Section 3.2 ), verdicts from GPT-4o-mini were nearly perfectly accurate on the question-answering tasks.
Decoding Strategies: For consistency and reproducibility, we use a temperature of zero for the Task LLM’s answer generation and for all judge evaluations. For TAPR’s generation, we compare two strategies: (a) TAPR , where the rewriter uses temperature zero and returns a single fixed rewrite; and (b) TAPR + Selection , where the rewriter samples five candidate prompts using a moderate temperature (0.5) and then selects the best candidate based on its judgment. We chose five samples to balance diversity and computational efficiency.
Ablation Studies: To isolate the contribution of the GRPO algorithm, prompt quality reward, and using SFT before TAPR training, we also conduct ablation studies using the Phi-4-mini-instruct as TAPR LLM with LLaMA-3.1-8B-Instruct as the Task LLM. We consider Phi-4-mini-instruct to be a capable model for its size, which responds well to reinforcement learning training in our experiments. LLaMA-3.1-8B-Instruct was chosen as the Task LLM as it is a slightly bigger LLM than Phi-4-mini-instruct, while not being from the same model family.
We also perform experiments using full-prompt rewriting, and alternative Task LLMs, and report these results in the Appendices D and E . To assess how task-aware and generalizable our method is, we also report experimental results across tasks in Appendix F .
5 Experimental Results
We compare four conditions for each base model in each experiment:
• Baseline Prompt : The unmodified initial instruction.
• Base Model : Rewriting with the base version of the LLM before training.
• TAPR : Rewriting after training the rewriter LLM with our method.
• TAPR + Selection : Generating multiple rewritten prompts after training and selecting one of them.
This allows us to answer our research question and analyze whether our method is effective in enhancing the prompt rewriting abilities of LLMs. All results are reported based on the quality of the Task LLM responses.
5.1 Impact of Prompt Rewriting on Different Tasks
Here, we evaluate whether TAPR enhances the performance of LLMs in downstream tasks.
Question Answering
Table 2 reports accuracy based on LLM judgments on the Natural Questions (NQ) dataset. In Appendix Table 14 , we demonstrate that standard exact-match metrics often yield low scores not aligned with the actual answer quality, suggesting that LLM-as-a-judge evaluation can capture practical performance better. For Phi-4-mini-instruct and LLaMA-3.2-3B-Instruct, both deterministic rewrites and TAPR + Selection variants outperform their Base Model versions and surpass the Baseline prompt’s accuracy.
!
TAPR LLM Rewriting Variant Accuracy ↑ \uparrow Baseline prompt – 56.30 [.5pt/2pt] Phi-4-mini-instruct Base Model 48.80 Phi-4-mini-instruct TAPR 59.20 Phi-4-mini-instruct TAPR + Selection 57.30 [.5pt/2pt] LLaMA-3.2-3B-Instruct Base Model 58.30 LLaMA-3.2-3B-Instruct TAPR 62.20 LLaMA-3.2-3B-Instruct TAPR + Selection 59.10 Table 2 : NQ dataset results, each value is an accuracy percentage (%), measured by GPT-4o-mini, reflecting answer correctness for each prompt variant.
TAPR LLM Rewriting Variant Accuracy ↑ \uparrow
Baseline prompt – 53.20
[.5pt/2pt] Phi-4-mini-instruct Base Model 53.20
Phi-4-mini-instruct TAPR 56.70
Phi-4-mini-instruct TAPR + Selection 53.16
[.5pt/2pt] LLaMA-3.2-3B-Instruct Base Model 53.15
LLaMA-3.2-3B-Instruct TAPR 56.00
LLaMA-3.2-3B-Instruct TAPR + Selection 55.00 Table 3 : HotpotQA dataset results, each value is an accuracy percentage (%), measured by GPT-4o-mini, reflecting answer correctness for each prompt variant.
Example prompts illustrating the rewriting for Natural Questions are provided in Table 4 . The results show that the Phi-4-mini-instruct’s base model rewrite is focused on describing the process, while RL fine-tuning yields more targeted and concise instructions. The best-performing prompt, generated by Llama 3.2-3B-Instruct TAPR, combines accuracy requirements with explicit formatting and clarity constraints.
TAPR LLM Rewriting Variant
Prompt
– Initial Prompt
Answer the question
Phi-4-mini-instruct Base Model
Provide a detailed explanation of the process you used to answer the question, including any relevant examples or illustrations.
Phi-4-mini-instruct TAPR
Answer the question and provide a brief explanation for your answer.
LLaMA-3.2-3B-Instruct TAPR
Provide a well-supported answer to the question in a concise paragraph (approximately 50–75 words) or a list of 3–5 bullet points. Ensure your response is accurate and relevant, and include evidence or examples to support your answer. Use clear, concise language and avoid ambiguity.
Table 4 : Example prompts for the Natural Questions dataset. Rows compare the baseline, the pre‑RL rewrites, and the TAPR outputs.
Summarization
For summarization, we report LLM-as-a-judge scores out of 5 in Tables 5 and 6 on the CNN/Daily Mail and SciTLDR datasets.
The baseline summarization instruction performs noticeably worse than several rewritten prompt variants. On both datasets, Phi-4-mini-instruct trained variants yield the highest scores.
The best overall scores come from TAPR and TAPR+Selection variants of Phi-4-mini-instruct on the datasets, respectively. These prompts are shown in Table 8 . Both prompts are well-matched to their respective dataset styles and provide strong length constraints, leading to more focused summaries. In contrast, over-specified prompts (such as LLaMA-3.2-3B-Instruct’s multi-sentence summary format) can degrade performance by instructing the use of formats that do not match the reference summaries. LLaMA results overall are weaker here, which could be due to the hyperparameters being optimized based on experiments with Phi-4-mini.
TAPR LLM Rewriting Variant Judge Score (1-5) ↑ \uparrow
Baseline prompt – 3.373
[.5pt/2pt] Phi-4-mini-instruct Base Model 3.772
Phi-4-mini-instruct TAPR 3.782
Phi-4-mini-instruct TAPR + Selection 3.774
[.5pt/2pt] LLaMA-3.2-3B-Instruct Base Model 3.794
LLaMA-3.2-3B-Instruct TAPR 3.435
LLaMA-3.2-3B-Instruct TAPR + Selection 3.411 Table 5 : CNN/DM dataset summarization results, each value is an LLM-as-a-judge score out of 5, as rated by GPT-4o-mini, reflecting summary quality for each prompt variant.
TAPR LLM Rewriting Variant Judge Score (1-5) ↑ \uparrow
Baseline prompt – 3.182
[.5pt/2pt] Phi-4-mini-instruct Base Model 3.802
Phi-4-mini-instruct TAPR 3.710
Phi-4-mini-instruct TAPR + Selection 3.869
[.5pt/2pt] LLaMA-3.2-3B-Instruct Base Model 3.749
LLaMA-3.2-3B-Instruct TAPR 3.107
LLaMA-3.2-3B-Instruct TAPR + Selection 3.041 Table 6 : SciTLDR dataset summarization results, each value is an LLM-as-a-judge score out of 5, as rated by GPT-4o-mini, reflecting summary quality for each prompt variant.
Pairwise Win-Rate Results for Summarization
To complement the numeric scores, we also report win-rates comparing summaries generated from the baseline prompts with the ones generated with the rewritten prompts in Table 7 . These results confirm that TAPR variants are often preferred over both the baseline and base model outputs. None of the TAPR models had a win-rate lower than 40%; notably, Phi-4-mini-instruct achieves almost 85% win-rates on SciTLDR, while LLaMA-3.2-3B achieves between 47%–63% win-rates despite the numeric score showing worse performance. We do observe a strong bias of the LLM-as-a-judge towards the second shown option, which is why we shuffled outputs randomly for each sample at comparison time.
Dataset TAPR LLM vs. Baseline vs. Base
Overall ↑ \uparrow Placed 1st/2nd ↑ \uparrow Overall ↑ \uparrow Placed 1st/2nd ↑ \uparrow
CNN/Daily Mail Phi-4-mini 46.11 18.13/72.99 40.66 20.41/59.69
LLaMA-3.2-3B 62.89 37.45/89.09 67.03 39.07/94.28
SciTLDR Phi-4-mini 84.58 75.23/94.88 84.42 86.09/82.80
LLaMA-3.2-3B 47.97 37.11/59.53 50.73 39.88/62.89 Table 7 : Win–rate comparison on CNN/Daily Mail and SciTLDR datasets. “Overall” is the percentage of cases where TAPR was preferred by GPT-4o-mini; “Placed 1st/2nd” reports preferences when shown first vs. second, compared to both baseline and base model prompts.
Dataset TAPR LLM Rewriting Variant
Prompt
CNN/DM, SciTLDR – Initial Prompt
Summarize the text
CNN/DM Phi-4-mini-instruct TAPR
Please provide a concise summary of the provided text, aiming for a succinct overview suitable for a general audience. The summary should be no longer than one-third of the original text length, maintaining a neutral and informative tone. Ensure that the key points and main ideas are preserved while eliminating any redundant or less critical information.
SciTLDR LLaMA-3.2-3B-Instruct TAPR
Condense the given text into a 3- to 5-sentence summary, highlighting key points and main ideas. Present the summary in a neutral, third-person voice, without using the author’s name or any personal pronouns. For example, A recent study found that [key finding]…
Table 8 : Example prompts for the CNN/DM and SciTLDR datasets. Rows compare the baseline, the pre‑RL rewrites, and the TAPR outputs.
Arithmetic Reasoning (GSM8K)
For the arithmetic reasoning task with the GSM8K dataset, we report accuracy based on the last integer found in the model’s response, matching the reward used for training. While this approach may occasionally misattribute the answer if multiple numbers are present, it also serves as a strong indicator of whether prompts guide models to output formats suited to the evaluation method. Results can be seen in Table 9 .
The baseline “SOLUTION” prompt is surprisingly robust and yields high accuracy relatively to many of the more detailed rewritten prompts. Nevertheless, the Phi-4-mini-instruct TAPR variant surpasses it, while also exhibiting more than a 70% improvement over the base model.
Table 10 includes example prompt rewrites for GSM8K. The Phi-4-mini-instruct’s base model rewrite asks for a Python function to sum integers, likely leading to confusion for the Task LLM and inability to produce the final correct answer with the limited maximum response tokens. This leads to poor performance, which all other rewrites clearly surpass.
In contrast, the TAPR variant produces a highly relevant, task-specific chain-of-thought prompt that encourages explicit step-by-step reasoning, justifications for each step, and a clear final answer. This style directly supports the kind of solutions needed for GSM8K, resulting in substantial performance improvements.
Despite improvements, neither the summarization nor the arithmetic reasoning results show substantial improvement over the base model in the former or the short general prompt in the latter task, with some trainings and variants leading to worse task performance. In addition to the potential issue with hyperparameters, this highlights how RL training for a task that is not measured directly based on the output of the model being trained is challenging, and not always reliable. Furthermore, despite significant research, prompt engineering is still an open problem where we can not exactly know which prompt leads to the optimal performance for the Task LLM. Using Selection leads to improvement sometimes, but overall, we find no strong evidence that training for prompt rewriting consistently enhances a model’s ability to judge and select among candidate prompts.
TAPR LLM Rewriting Variant Accuracy ↑ \uparrow
Baseline – “SOLUTION” – 82.38
[.5pt/2pt] Phi-4-mini-instruct Base Model 11.91
Phi-4-mini-instruct TAPR 83.60
Phi-4-mini-instruct TAPR + Selection 83.60
[.5pt/2pt] LLaMA-3.2-3B-Instruct Base Model 74.80
LLaMA-3.2-3B-Instruct TAPR 82.80
LLaMA-3.2-3B-Instruct TAPR + Selection 81.80 Table 9 : GSM8K dataset results comparing the Baseline against the base models and TAPR variants. Each value is an accuracy percentage (%), reflecting answer correctness for each prompt variant.
TAPR LLM Rewriting Variant
Prompt
– Initial Prompt
SOLUTION
Phi-4-mini-instruct Base Model
Please provide a detailed solution to the following problem: Given a list of integers, write a function in Python that returns the sum of all even numbers in the list…
Phi-4-mini-instruct TAPR
Provide a step-by-step solution to the problem, including intermediate steps, justifications for each step, and a clear explanation of the final answer.
Table 10 : Example prompts for the GSM8K dataset. Rows compare the baseline, the pre‑RL rewrites, and the TAPR outputs.
5.2 GRPO vs PPO
Here we show results comparing using the GRPO algorithm for our method to the PPO algorithm used in PRewrite [ prewrite ] . Table 11 summarizes the results on both the GSM8K and Natural Questions datasets. We observe that GRPO reliably improves the training reward and downstream task accuracy for both datasets. Depending on the hyperparameters, PPO training either results in stagnant rewards or leads to model collapse, with degenerate outputs such as repetitive symbols or nonsense token.
Although prior work in PRewrite [ prewrite ] showed that PPO can be effective for prompt rewriting training, differences in model architecture, hyperparameter choices, and using a critic head instead of a full critic model may contribute to the discrepancies observed here. Due to these challenges and the lack of access to the exact implementation in PRewrite [ prewrite ] , we opted to use GRPO for all other experiments, and couldn’t include a PRewrite baseline to compare to.
TAPR LLM
Rewriting Variant
NQ ↑ \uparrow
GSM8K ↑ \uparrow
Baseline
–
56.30 82.38
[.5pt/2pt] Phi-4-mini-instruct
Base Model
48.80 11.91
Phi-4-mini-instruct
TAPR with GRPO
59.20 83.60
Phi-4-mini-instruct
TAPR with PPO — Con.
35.90 71.10
Phi-4-mini-instruct
TAPR with PPO — Mod.
4.00 75.20
Phi-4-mini-instruct
TAPR with PPO — Agg.
52.10 78.40 Table 11: Comparison of performance using GRPO and three PPO variants for training the TAPR on NQ and GSM8K (Accuracy %). "Con." (“Conservative”), "Mod" (“Moderate”), and "Agg." (“Aggressive”) are PPO hyperparameter settings differing mainly in learning rate and KL-divergence control values.
5.3 SFT impact on TAPR
In this section, we see how effective is using Supervised Fine-Tuning (SFT) prior to reinforcement learning.
We conduct SFT for 5 epochs, with data coming from various sources showing rewrites of prompts, chain-of-thought prompts and responses, and ratings to prompts given by humans [ 10kpromptsranked , promptoptimizationdataset ] . In addition, we instructed GPT-4o-mini [ openai2024gpt4omini ] to generate prompts using common prompt engineering techniques for queries coming from the datasets used in the RL phase. We compare training on our internally generated dataset with training on publicly available datasets combined with our internally generated dataset. As seen in Table 12 , SFT improves the base performance of Phi-4-mini-instruct, particularly when conducted with both internally generated and external data. However, this improvement does not translate into enhanced RL training outcomes.
TAPR LLM
Rewriting Variant
NQ ↑ \uparrow
GSM8K ↑ \uparrow
Baseline
–
56.30 82.38
[.5pt/2pt] Phi-4-mini-instruct
Base Model
48.80 11.91
Phi-4-mini-instruct
SFT with generated data
52.50 35.34
Phi-4-mini-instruct
Full SFT
54.00 71.00
Phi-4-mini-instruct
TAPR
59.20 83.60
Phi-4-mini-instruct
TAPR after SFT
55.50 70.50 Table 12: Comparison of TAPR performance on the NQ and GSM8K datasets, evaluating the impact of supervised fine-tuning (SFT) with internally generated data, externally sourced data (“Full SFT”), and reinforcement learning (RL), both with and without prior SFT. Results are reported as accuracy percentages (%).
5.4 Impact of the Prompt Quality Reward
As an additional ablation study, we evaluate the effect of using the prompt quality reward.
As shown in Table 13 and Figures 2 and 3 , integrating the prompt quality reward improves both the accuracy of the downstream task and the convergence of training. In contrast, omitting this reward severely limits training progress. Without it, the model rewrites remain essentially unchanged from the initial prompts (“Answer the question." for Natural Questions and “Provide the solution." for GSM8K), indicating negligible learning.
Figure 2 : Training rewards comparison for the TAPR on the NQ dataset. The upper panel shows the raw per‐step reward (blue) and its 25‐step moving average or running mean (orange) with our method, and the bottom panel shows the training without the prompt quality reward.
Figure 3 : Training rewards comparison for the TAPR on the GSM8K dataset. The upper panel shows the raw per‐step reward (blue) and its 25‐step moving average or running mean (orange) with our method, and the bottom panel shows the training without the prompt quality reward.
TAPR LLM Rewriting Variant NQ ↑ \uparrow
GSM8K ↑ \uparrow
Baseline – 56.30 82.38
[.5pt/2pt] Phi-4-mini Base Model 48.80 11.91
Phi-4-mini TAPR 59.20 83.60
Phi-4-mini TAPR without PQ Reward 57.20 82.80 Table 13 : Comparison of performance with and without prompt quality reward on the NQ and GSM8K datasets using Phi-4-mini-instruct as the base model (Accuracy %).
6 Discussion
6.1 Performance Discussion
Our results indicate that, although the TAPR method can produce improved prompt rewrites and boost performance, these gains are not always consistent. In several cases, even simple or generic prompts achieve competitive or better task results than more highly optimized rewrites. This underscores the difficulty of reinforcement learning for prompt engineering: learning a rewriting policy based on rewards that do not directly reference the Task LLM’s outputs is inherently challenging.
While the prompt quality reward correlates with better training rewards and often improved task scores, as seen in Section 5.4 , much of the reward improvement may reflect the model’s ability to satisfy prompt quality criteria rather than the end-task metrics. Thus, increases in training reward do not always translate directly into downstream accuracy gains.
Prompt engineering is, by nature, a challenging problem. The relationship between prompt structure and LLM performance is complex and sometimes unpredictable. During RL training, models can receive mixed signals: for some samples, even a suboptimal or unrelated prompt may elicit a correct answer, while a carefully crafted prompt might not.
Further complicating matters, RL for text generation is inherently unstable. Since the reward is based on the entire output rather than single decisions, training is highly sensitive to hyperparameter choices. As a result, identifying robust configurations can require significant trial and error. In our experiments, this issue was apparent for LLaMA-3.2-3B-Instruct, likely due to the hyperparameters being optimized based on training Phi-4-mini-instruct. In summary, while the Task-Aware Prompt Rewriter approach shows promise, reliably optimizing prompts through RL remains an open challenge.
6.2 Prompt Selection Mechanism
The Selection mechanism sometimes improves over basic TAPR outputs, but these gains are inconsistent. As a result, we find no strong evidence that training for prompt rewriting consistently enhances a model’s ability to judge and select among candidate prompts. This is likely due to the fact that our training only improves performance with a specific prompt without intending to improve any other abilities. Given the additional computational cost, 0-temperature generation may be preferable for simplicity, despite the risk of a single failed prompt affecting all outputs. For instance, in one Phi-4-mini-instruct training run on CNN/Daily Mail, test-time selection modestly outperformed the base model, yet 0-temperature generation produced a prompt such as “Could you tell me a light-hearted, family-friendly joke, preferably related to technology or everyday life?”, a complete mismatch that caused the Task LLM to generate jokes rather than summaries for all inputs.
6.3 LLM-as-a-Judge Evaluation
Although LLM-as-a-judge evaluation is central to this work, we acknowledge its limitations and potential biases. For example, as observed with GPT-4o-mini in Section 5.1 , the model exhibited a strong preference for whichever response was presented second. Still, we argue that this approach offers a more realistic assessment of open-ended generation tasks like question answering and summarization, compared to overlap-based metrics. Appendix Table 14 shows that exact-match accuracy can severely underestimate the quality of model outputs, while manual inspection confirms many responses deemed incorrect by traditional metrics are actually valid and possibly preferred due to natural phrasing or richer explanations. Despite the higher cost, we therefore recommend LLM-as-a-judge scoring as a valuable complement to traditional metrics and human evaluation for future work on LLM fine-tuning and benchmarking.
6.4 Future Work
The results and analysis in this work only partially capture the true effectiveness of the TAPR method. Both the datasets and prompts considered represent a narrow look at how people interact with LLMs in real-world scenarios. In practice, many user queries are more diverse, nuanced, and less well-formed than those found in common benchmarks. As such, small changes to questions or instructions on familiar datasets may yield limited improvements, while automatic prompt optimization could provide even greater benefits when adapting to users’ unique, everyday needs. This is partially supported by our finding that rewritten prompts tended to be clearer, more detailed, and more consistent with prompt engineering best practices than the baseline instructions, suggesting that our approach may have more potential not fully reflected by our evaluation metrics. In the future, a more robust test would be to deploy this method in real agentic workflows, where crafting effective prompts is a central challenge, and directly compare the outcomes achieved by human-written prompts versus automatically optimized prompts.
Our study was limited to smaller, non-reasoning models. Although scaling up to larger models is computationally expensive, it is also plausible that larger LLMs, with their broader knowledge and enhanced capabilities, could be more effective at both generating and benefiting from optimized prompts, particularly in the context of full-prompt rewriting where information loss is a concern. However, using advanced reasoning models that already employ chain-of-thought or step-by-step prompting may require specialized evaluation protocols and different prompt objectives to achieve maximum benefit.
Finally, one motivation for this approach was to help close the gap between small and large LLMs by maximizing the effectiveness of less capable models through better prompt engineering. If successful, this would enable cost savings by achieving similar downstream performance with less expensive models. However, our results show that the gains from prompt rewriting are not yet consistent enough to offset the additional training and inference costs. More detailed cost–benefit analyses, as well as further method refinements, will be valuable directions for future work.
7 Conclusion
In this work, we set out to address how to create a task-aware prompt rewriter model that effectively enhances LLM performance. Our findings demonstrate that a smaller model can be successfully trained through reinforcement learning to rewrite prompts that consistently improve downstream LLM outputs. Specifically, we showed that training with Group Relative Policy Optimization (GRPO) yields stable learning. Applying this trained TAPR model resulted in performance gains across various benchmarks. Notably, Phi-4-mini-instruct-based TAPR variants consistently outperformed baseline and base model prompts, delivering improved accuracy alongside qualitatively clearer, more detailed prompts adhering closely to established prompt engineering principles. Our ablation study revealed how the incorporation of LLM-as-a-judge evaluations, particularly the prompt quality reward, enhanced the training process, guiding the model towards clearer and more effective rewrites. We also observed LLM-as-a-judge evaluations to be more accurate than many traditional metrics.
Despite these promising results, our approach does have limitations, including variability in improvements across tasks and models, potential biases inherent to LLM-based reward evaluations, and increased computational demands. Future research should focus on refining the RL training and scaling to more diverse tasks and realistic user scenarios.
Overall, this work demonstrates the viability of automated prompt rewriting as a practical and effective technique for unlocking the full potential of LLMs, reducing the reliance on manual prompt engineering.