Every LLM project I pick up now starts the same way: someone saw a demo, the demo was magic, and then reality arrived. The model answered differently on Tuesday. A customer typed something hostile. The invoice doubled overnight.
In my experience the fix is rarely a better prompt. It's a golden set of fifty real cases, a regression run in CI, a cost ceiling with an alarm, and a human handoff path. None of it demos well. All of it decides whether the thing survives the quarter.
So when I scope AI work these days, I quote the boring half first. It's the half that holds.
