Every AI vendor will tell you their product is safe. Few can tell you how that safety actually works - or whether it would survive the model trying to talk its way around it.
A policy is not a control.
A system prompt, a contract clause, or a vendor's stated safeguard is information given to a model - not a safeguard. Only a mechanism that blocks an action regardless of what the model "believes," argues, or is told counts as an actual control. As AI models get more capable, this gap gets more important, not less: capability increases a model's ability to reason past a stated boundary faster than it increases its reliable compliance with that boundary.
Where AI behavior actually lives.
AI behavior is determined by a handful of concrete layers:
The data and rewards shaped the model
The tools and network access it's been granted
A human or agent has to approve before an action executes, and
Does the technical boundary stop it if any or all of the above failed?
Behavior is commonly introduced, instead, through a policy document, a vendor's marketing claim, or a line in a contract. Those are two different layers - and a vendor pointing confidently at the second one tells you nothing about the first.
Questions worth asking any AI vendor
Is this a control or an instruction? For every safeguard you're told about, ask whether it's enforced by something that fails closed regardless of the model's behavior, or whether it depends on the model choosing to comply.
What happens if the model disagrees with the rule? A safeguard that only works when the model agrees with it isn't a safeguard.
Has this been tested adversarially, or only described? A written safeguard and a tested one are not the same claim.
Who reviews before an action executes, and how? If the answer is "the model checks itself," ask what stops the model from being wrong about its own situation.
What's the blast radius if this fails once? Understand what a single bypass would actually expose - not what the safeguard is designed to prevent, but what happens if it doesn't.
Has "no incidents" been distinguished from "no detected incidents"? These are different claims, and the gap between them grows as models get more capable.
Why this matters now. These aren't hypothetical questions. Documented incidents across multiple frontier AI providers in 2026 show real cases of agents escaping intended boundaries, exploiting overlooked systems, or reasoning past a control that depended on their own cooperation. See our Frontier AI Model Incident Tracker for the specifics.
Is your AI program DeployWorthy?
These buyer questions clarify the boundaries of safety assurance between you and your vendor. GuardLight's DeployWorthy Review turns them into a structured assessment of your own AI products and deployments.
Check whether your controls are actively applied or merely aspirational.
Contact GuardLight to gain more confidence in your AI governance posture.