There's a moment in most AI projects where somebody says "this is working really well," everyone in the room agrees, and the thing goes live. No test was run. Nobody wrote down what "working really well" means. The evidence was a demo where the person driving knew which questions to ask.
That's the moment most AI systems ship, and it explains a lot about why so many of them quietly stop being useful. Below is the gate we run instead. It's published in full so you can hold us to it, and so you have something to compare against when a vendor tells you their agent is ready.
If a vendor can't tell you what would have stopped them from shipping, nothing would have.
Gate 1 — There is a dataset
Twenty examples minimum for any prompt going into production, in a file, in version control. Each one labeled by where it came from: a real production sample, or something we made up. Real beats invented every time, because the examples you write yourself are always tidier than the ones your customers send.
Three of those examples have to be adversarial, and they each have a job. One with personal data in the input, to check the agent doesn't repeat it back. One containing instructions aimed at the agent, to check it treats documents as information rather than orders. And one that's empty, contradictory, or plainly out of scope, where the right behavior is a clean handoff rather than a confident guess.
Gate 2 — There is a suite, and some of it blocks
At least four categories of check running against that dataset. Does it sound right, is the output structurally valid, is it grounded in the source material, is it safe, is it fast and cheap enough, has it got worse since last time.
Two of those have to be blocking, meaning a failure stops the release rather than producing a note somebody reads later. Blocking checks pass at 95% or above, warnings at 80%.
One rule that surprises people: if everything scores 100%, we suspect the test set rather than the agent. A suite that has never failed isn't measuring anything.
Gate 3 — Every call is traced
All model calls go through a wrapper that records them, with a consistent set of tags: which client, which agent, which version of the prompt, which model, whether it was a live run. Personal data gets stripped before anything is stored.
This is the gate people skip, and it costs you in month four when somebody says "it gave a weird answer to a customer last Tuesday" and the only honest reply is that we have no way of finding that conversation.
Gate 4 — There is a baseline to compare against
Before a change ships, we run it against the stored results of the version currently live. A drop of more than 5% in the mean score blocks the release. More than 2% and somebody reads the five most-changed outputs by hand before deciding.
Prompt edits are famously local. You fix the thing a client complained about on Monday and break two things nobody will notice until October. Champion versus challenger is the only honest way to ship a change to a prompt.
Gate 5 — The check runs itself
The suite runs automatically when the prompt file changes, and blocks the merge when it fails. Under five minutes for a twenty-example set.
Manual quality checks get skipped. Not through negligence, just through having a busy week twice in a row. Automation is the only enforcement that survives August.
For client work, three more
- A tenant configuration bundle. Every client-specific rule lives in config rather than in somebody's memory of a call.
- A scheduled monthly quality report. The answer to "how do we know it's still working?" should be a document that arrives whether or not anyone asks for it.
- A signed rubric and handoff. The client agrees in writing what good output looks like, before launch rather than during the first argument about it.
The gate has to actually stop things
A standard nobody has ever failed isn't a standard.
The most recent thing ours blocked was our own sales agent. The run came back BLOCK on a structural check across the whole dataset, which looked alarming for about ten minutes until we traced it to a broken validator in the harness rather than anything the agent had produced. Slightly embarrassing — I was the one who wrote the broken validator — and exactly the point: we knew within the hour because something ran and failed loudly.
The alternative version of that week is that we shipped it, felt good about it, and found out from a client.
Exceptions do happen. When we bend a threshold, the number, the reason, and the name of whoever approved it get written into the agent's baseline document. An undocumented exception becomes the new default within about a month.
Five questions to ask your vendor
- What's in your test set, and how many of the cases are real?
- Which checks block a release, and what's the passing score?
- How would I find a specific bad answer from three weeks ago?
- When you change the prompt, what do you compare against?
- Show me something this stopped you from shipping.
Good answers are specific and a bit boring. Bad answers talk about how carefully the team reviews everything, which is not a measurement. The fifth question is the one that separates them, and it's the one almost nobody asks.
If you're earlier than this and still working out which system is worth building at all, start with the two-week audit instead. When you're ready to talk about how this applies to your own stack, book a call.