Guardrails & Evals: How to Build Hard Constraints for AI — Technical Workshop | Maneesh Maddala Skip to main content

Guardrails & Evals: How to Build Hard Constraints for AI

How to test whether constraints hold under pressure

Ministry of Testing · Technical Workshop · September 23, 2026

LLM Evals · Guardrails · Risk Engineering · Red Teaming · CI/CD

I ran a technical workshop for the Ministry of Testing community on building hard constraints for AI systems and testing whether they hold under pressure. It separates the guardrail from the eval that tries to break it, then uses a case study of 22 agent runs to show why checking every constraint with an LLM does not scale. The answer is a tiered suite: deterministic schema checks at the base, mechanical checks on each pull request, and a few LLM-judged evals kept for the cases that need them. A live demo runs a 14-eval suite against contract-drift detection.

  1. 01

    Why assertions fail on generative output and how semantic evaluation replaces them.

  2. 02

    Using input filters, schema checks, and output validation together.

  3. 03

    Crash-testing LLM agents against prompt injection and model drift before reaching users.