PRODUCTION NOTES · #01PROMPT & AGENTS
Lucky ≠ Tuned
Published 1 min read
A prompt that gives you one great answer isn’t finished. It’s just lucky.
Here’s how I tell the difference.
Think of a senior designer. Every deck looks different, but every one is right and on-brand. You don’t check it slide by slide. You trust how they work.
I want the same from prompts and agents. My two-step test:
- 01TunedSame model, real inputs, many runs (e.g. 20 inputs × 5). If about 98% of the results meet my spec, it’s tuned. Not the same words. The same structure, decisions and quality.
- 02RobustThen I run it on a smaller, cheaper model. If it still hits 85–92%, it’s robust. If it falls apart, the big model was quietly filling gaps I never wrote down.
It’s not just me:
- [1]Sierra’s τ-bench checks if an agent succeeds on every try, not just once. GPT-4o fell below 25% when it had to pass all 8.
- [2]An ICLR 2024 paper: small formatting changes moved accuracy by up to 76 points, and what works on one model may fail on another.
It saves money too: moving steps to smaller models was one of three reasons we cut costs 82% in the AI pipeline I own.
What I write in every prompt
- Role: who the model is
- Thinking rules: how it should think
- Limits: what it must never do
- Output format: exactly how the answer should look
- Examples: powerful but risky (models copy them). More in #03.
Judge a prompt by its 100th run, not its first.