Introducing run-assert-eval: Find the risk, fix it, prove it
Today, we’re introducing run-assert-eval, a skill that discovers the risks that matter for a given agent, measures how often the agent fails, generates runtime policy directly from those findings, and reruns the eval to prove whether the fix worked.