Gate B4Writing
How I gate LLM releases without slowing the team down
An eval-gated release loop that catches regressions before rollout, without turning every deploy into a research project.
Most teams shipping LLM features don't have a release gate. They have a vibe check — someone reads a few outputs, nods, and ships. That works until it doesn't, usually the week you swap models or edit the system prompt.
The gate I've been running looks like this.
One eval suite, two tiers
Tier 1: smoke. Twenty to thirty prompts covering the top intents. Runs on every PR. Deterministic scoring where possible (JSON schema match, keyword presence, refusal detection), LLM-as-judge for the rest. If tier 1 fails, the PR is red.
Tier 2: red-team. Adversarial prompts and known-bad inputs. Runs weekly. Any regression here escalates to the AI safety review, not the deploy queue.
What actually breaks
The failure I have seen kill releases most often:
- The judge drifts. If the LLM-as-judge is a different model than production, an update to the judge model looks like a product regression. I now version the judge and treat judge changes as their own release.
What I don't do
- I don't chase 100% on any tier. The point is to catch drops, not to prove perfection.
- I don't gate on token counts, cost, or latency at the same level as quality. Those are their own dashboards.
- I don't let anyone "override" a gate failure without a written note in the release doc explaining why.
If you're standing up something like this and want to compare notes, contact me.
