
Turning a Prompt That Worked Once Into One That Works Every Time
Most prompts start as a lucky result. You typed something, the answer was good, and you saved it. Then you ran it on ten real inputs and three came back wrong in ways that were hard to describe. The gap between a prompt that worked once and one you can put inside a process is not about finding better wording. It is about removing the assumptions the prompt was carrying without stating them.
Why the same prompt gives different answers
Two reasons, and they need different fixes. The first is sampling: a language model picks each next token from a probability distribution, so two runs of an identical prompt can differ. Lowering the temperature narrows that variation and is worth doing for extraction and classification work, where you want the same answer every time. It does not make output identical, and it does nothing about the second reason.
The second reason is that the inputs differ more than you think. Your test document had headings and the real ones do not. One customer email is three lines, another is a forwarded thread with four quoted replies. The prompt that worked was relying on the shape of the example you happened to test it on, and nothing in the wording said what to do when that shape is missing.
Say what the task is, not how impressive it should be
Adjectives are the least useful part of most prompts. "Write a professional, engaging summary" gives the model almost nothing it can act on. Replace each adjective with the decision it implies:
- Who the reader is and what they already know.
- What the output is for: a decision, a file, a reply someone will send.
- How long it should be, in a unit you can check.
- What to leave out, which is the instruction people forget.
Compare "summarise this support ticket professionally" with "summarise this support ticket in three sentences for an engineer who has not seen it: what the customer tried, what happened instead, and which product area it touches. Do not suggest a fix." The second can be graded by someone else, which is the real test of whether an instruction is specific.
Give it the material
A model asked about your refund policy will answer from what it learned in training plus whatever it can infer, which is a good way to get a plausible policy that is not yours. If a fact matters, put it in the prompt. Paste the policy, the product list, the previous email. This is the single biggest improvement available in most prompts, and it costs nothing but tokens.
When you supply material, mark where it starts and ends. Delimiters such as XML-style tags do two jobs: they tell the model which part is data rather than instruction, and they give you something to refer back to later in the prompt.
Here is the policy:
<policy>
...text...
</policy>
Answer the question using only the policy above.
If the policy does not cover it, say so and stop.
That last line matters more than it looks. Without it, a model asked a question the material does not answer will usually produce something anyway. With it, you get a signal you can route to a human.
One example beats five adjectives
If you want a particular tone or structure, show it. A single worked example of an input and its ideal output communicates more than a paragraph of description, and it is easier to write. Two or three examples help when the task has distinct cases, such as a classifier that needs to see one instance of each label.
Choose examples that are typical rather than impressive. An example showing the model handle an unusually rich input teaches it to expect rich inputs.
Fix the output shape
Anything downstream of the model needs a predictable shape. Ask for one, in the form the consumer wants: a JSON object with named fields, a fixed set of headings, one label from a list you supply. Several providers support a structured output mode where you give a schema and the response is guaranteed to match it. That is more reliable than asking politely for JSON and parsing whatever comes back.
Whether or not you use a schema, decide what the output looks like when the model cannot do the task. A confidence field, or a permitted label such as unclear, gives it somewhere to put uncertainty other than a confident guess.
Test with cases, not vibes
This is the step that separates a prompt you trust from one you hope about. Collect ten to twenty real inputs, including the awkward ones: the empty field, the wrong language, the input twice as long as usual, the one where the answer genuinely is "not covered here". Write down what a correct output looks like for each.
Then run the whole set after every change to the prompt. You will find that half your improvements fix one case and break another, which you would never notice by trying a single input after each edit. The set needs no tooling to begin with; a spreadsheet and a script that prints inputs next to outputs is enough to start.
Treat prompts as code
A prompt that runs in production has the same needs as the code around it. Keep it in version control rather than in a text box in a dashboard, so you can see what changed and revert it. Note which model and which settings it was tuned against, because a prompt tuned on one model is not automatically right on the next one. When you do switch models, rerun the test set before assuming the newer model is better at your specific task.
Things that do not help
Magic phrases have a short shelf life. Telling a model to think step by step was a genuine finding on older models; current reasoning models already do that internally, and the instruction mostly adds noise. Offering a tip, threatening consequences, or insisting the task is critically important adds tokens and little else. Stacking six instructions that all mean "be accurate" makes the prompt harder to maintain without making the output more accurate.
What does help is boring: state the task precisely, supply the material, fix the output shape, and test against cases you can inspect. A prompt built that way survives a new model, a new colleague and a Monday morning. A lucky one does not.
Comments
No comments yet. Be the first to share your thoughts.


