Alaa Al-bdewi

PRODUCTION NOTES · #01PROMPT & AGENTS

Lucky ≠ Tuned

Published 1 min read

Read in ArabicRead on LinkedIn

A prompt that gives you one great answer isn’t finished. It’s just lucky.

Here’s how I tell the difference.

Think of a senior designer. Every deck looks different, but every one is right and on-brand. You don’t check it slide by slide. You trust how they work.

I want the same from prompts and agents. My two-step test:

  1. 01TunedSame model, real inputs, many runs (e.g. 20 inputs × 5). If about 98% of the results meet my spec, it’s tuned. Not the same words. The same structure, decisions and quality.
  2. 02RobustThen I run it on a smaller, cheaper model. If it still hits 85–92%, it’s robust. If it falls apart, the big model was quietly filling gaps I never wrote down.

It’s not just me:

  1. [1]Sierra’s τ-bench checks if an agent succeeds on every try, not just once. GPT-4o fell below 25% when it had to pass all 8.
  2. [2]An ICLR 2024 paper: small formatting changes moved accuracy by up to 76 points, and what works on one model may fail on another.

It saves money too: moving steps to smaller models was one of three reasons we cut costs 82% in the AI pipeline I own.

What I write in every prompt

  • Role: who the model is
  • Thinking rules: how it should think
  • Limits: what it must never do
  • Output format: exactly how the answer should look
  • Examples: powerful but risky (models copy them). More in #03.
The rule I work by:
Judge a prompt by its 100th run, not its first.

Your turn

How do you know when a prompt is “done”?

Read on LinkedIn