Guardrails are not eval suites
A common confusion: the difference between offline model evaluation and online runtime enforcement, and why you need both.
We see teams ship a beautiful evaluation harness - golden datasets, regression tests, leaderboards - and then go to production with no runtime guardrails at all. The eval suite tells you what the model can do on a fixed corpus. It tells you very little about what the model will do at 2 a.m. against an adversarial prompt your harness has never seen.
Guardrails are the production runtime. They evaluate every single request, in line, against current policy. They don't replace evals - they complement them. Evals catch capability regressions. Guardrails catch policy violations.
If you only have one, you should have guardrails. A model that scores well on your eval suite and leaks customer data in production has not been governed.
Further reading