Skip to content
Manoj Deshmukh
All English essays

The Practical Technologist · 14 Jul 2026 · 3 min read

Common Sense Gets You a Draft. Evals Get You to Production.

By Manoj Deshmukh
Common Sense Gets You a Draft. Evals Get You to Production.

You wouldn't push code to production without tests.

Yet every day, teams push prompts to production on nothing but a good feeling. "It worked when I tried it" and out it goes.

I've spent 26 years in IT, and over the last year+ I've been building AI-ready products hands-on. Here's the hard lesson that surprised even me: a prompt is code but code that behaves differently every time you run it.

We forget that. We treat prompts like common sense. Write a few clever lines, eyeball three outputs, ship. In a deterministic world, that's careless. In a non-deterministic world, it's disastrous.

Think of the difference between cooking one dinner and catering a wedding.

At home, you taste once and trust your instinct. For a wedding of 500 guests, instinct isn't enough you standardise the recipe, cook test batches, and put tasters in front of them to score consistency. Not because you're a worse cook. Because the stakes and the variability just went up.

Your prompt in production is the wedding. Not the dinner.

So before "prompt engineering," there has to be "prompt evaluation." Two disciplines, in that order. Most teams skip straight to the clever part and wonder why quality is a rollercoaster.

Evaluation rests on three things:

👉 A test dataset — a real spread of inputs your prompt will face in the wild, including the messy edge cases, not just the three happy paths you demoed to your boss.

👉 A grader — a defined way to score each output. Sometimes another model as judge, sometimes a rule, sometimes a human. But the standard is written down, not living in your head.

👉 An iteration loop — you run the prompt across the dataset, grade it, read the failures, improve, and re-run. You're chasing a rising score, not a single lucky output.

Only once you can measure quality does prompt engineering actually pay off. Now the craft techniques stop being folklore and start being levers you can pull with evidence:

✅ Be clear and direct — say exactly what you want. The model is a brilliant new colleague on day one: capable, willing, and with zero context on your intent.

✅ Be specific — "summarise this" and "summarise this in 3 bullets for a CFO who has 30 seconds" are not the same instruction.

✅ Provide structure — separate the role, the task, the rules, and the input. Structure removes ambiguity, and ambiguity is where non-determinism does its worst damage.

✅ Give examples — multi-shot. Two or three worked examples of input → ideal output teach more than a paragraph of description ever will. Show, don't just tell.

The pattern underneath all of it is old and familiar to any engineer: you can't improve what you don't measure. Evals are your test suite. Prompt engineering is your refactor. One without the other is either flying blind or polishing something you never validated.

Clarity is born from constraints and in AI, evals are the constraint that makes clarity possible.

The teams winning with GenAI right now aren't the ones with the cleverest single prompt. They're the ones who built the boring measurement scaffolding around it and kept improving the rank.

Common sense gets you the first draft. Engineering discipline gets you to production.

If you're shipping LLM features, here's my question: do you have a test dataset and a grader today or are you still tasting one dinner and calling it catering? Tell me where you are in the comments.


I write The Practical Technologist every week — practical takes on AI, engineering, and business from 26+ years of building things. Subscribe if that's your kind of thing.

First published on LinkedIn.

Read next