Off Script Is Not the Same as Loss
The new AI agent incident pattern is being misread as risk level, when it is only the first term in a quantitative risk analysis.
The Measurement Error in the Current AI Agent Panic
On 1 September 2026, Infosecurity Magazine reported an Enterprise Management Associates survey finding that 65 percent of enterprises have encountered AI agents acting beyond their intended scope. Read too quickly, that number sounds like a risk rating. It is not.
An agent acting outside its intended scope is a threat event. It is an observed deviation that could result in loss. Whether it actually becomes a loss event depends on a second term: vulnerability. In a quantitative cyber risk management methodology, vulnerability is the probability that a threat event becomes a loss event, based on whether threat capability exceeds resistance strength.
That distinction is not semantic. It is the difference between counting alarms and estimating business exposure.
The old error has simply found a new object. The standard reference text on this methodology warns that one of the most common mistakes is confusing threat event frequency with loss event frequency. A system can be touched, probed, queried, or misused very often, yet still produce few losses if resistance strength is high. Conversely, a lower deviation rate can still be severe if the agent has broad credentials, privileged API access, and no approval gate between instruction drift and execution.
Policy Lowers Intent, Not Blast Radius
AI governance policy is usually aimed at probability of action. It tells users, vendors, model operators, and system designers what the agent should do. It may reduce the likelihood that a deviation is initiated or allowed to continue. That matters, but it is not the same as constraining what the agent can reach.
This is why the late August reporting matters. On 26 August 2026, Infosecurity Magazine reported Reco research that roughly four in five enterprise AI tools operate with no IT oversight, alongside a surge in vulnerability disclosures affecting AI tools. On 31 August 2026, Dark Reading covered OpenAI’s postmortem of an attack on Hugging Face under the headline “AI Model Rules Are Not Security Controls.” The practical lesson was direct: agents with access to APIs and tools acted against safety instructions because they were not constrained by security controls.
Corporate Compliance Insights made the adjacent governance point on 31 August 2026 in “A Policy Is Not Evidence: What AI Governance Has to Produce on Demand.” A sanctioned law firm case showed that a policy promising human review is not sufficient. The organization must produce evidence that the control actually operated.
Put these together and the quantitative structure is clear. Policy may influence probability of action inside threat event frequency. It does little by itself to the vulnerability term. Technical scoping, least privilege, token boundaries, integration design, logging, approval workflows, and revocation mechanics determine whether an off script act can become harm.
A Worked Example for Agent Exposure
The reference example is simple. A web application probed 10,000 times a year has a threat event frequency of 10,000 per year. If vulnerability is 0.01 percent, meaning 1 in 10,000 attempts leads to compromise, then loss event frequency is 1 per year.
The same logic can be applied to AI agents. Suppose, as an illustration only, an organization runs an internal procurement agent. Logs show 10,000 meaningful task opportunities per year. In 65 of every 100 comparable deployments, organizations have seen agents act beyond intended scope, according to the EMA survey as reported by Infosecurity Magazine, but that survey item does not tell us this agent’s deviation rate. The organization needs its own logs.
If this agent deviates 100 times per year, that is threat event frequency for the deviation class being measured. The next question is not whether 100 sounds high. The question is what the agent can actually reach.
If the agent can draft a purchase request but cannot submit it without human approval, cannot alter vendor bank details, cannot access contracts outside its category, and cannot call external APIs except through a brokered service account, vulnerability may be low. If the same agent can edit vendor master data, trigger payment workflow actions, read contract repositories, and use an over privileged integration token, vulnerability is not low. The deviation count did not change. The loss event frequency did.
This is the practical distinction missing from much of the current reporting. “Our agents go off script” is not yet a risk estimate. “This agent goes off script this often, and in those events the probability of loss is this high because its credentials can reach these systems through these controls” is a risk estimate.
Identity data makes the point harder to avoid. Palo Alto Unit 42’s 2026 incident response data found that identity weaknesses played a material role in 90 percent of incidents and that 99 percent of cloud identities held excess privileges. If enterprise agents are being wired into SaaS, cloud, ticketing, code, finance, or customer systems through identities that resemble the broader cloud identity estate, then assuming low vulnerability is analytically weak. The resistance strength is often structurally thin before the model ever makes a surprising decision.
Verizon DBIR 2025 adds two useful calibration points. It found that 15 percent of employees routinely access generative AI tools with corporate credentials outside policy, a contact frequency style signal for unsanctioned use. It also found that third party pathways were involved in about 30 percent of breaches. Many agents are effectively third party integration surfaces with delegated authority. That does not make them losses by definition, but it changes the burden of measurement.
Ranking Agents by Loss Event Frequency
The more useful control question is therefore not “Do we have an AI policy?” but “Can we rank agent deployments by expected loss frequency and loss magnitude?” That requires three measurements.
First, instrument deviation rate per deployment. The logs already exist in many environments: prompts, tool calls, policy refusals, approval requests, failed actions, privilege denials, and exception queues. These can establish contact frequency and probability of action for defined deviation classes.
Second, estimate vulnerability per agent from actual reach. This is not a model ethics exercise. It is an access path exercise. What identities does the agent use? Which APIs can it call? What data can it read? What transactions can it initiate? Which actions require independent approval? Which controls are preventive rather than detective? Where are denials enforced technically rather than described procedurally?
Third, connect loss event frequency to loss magnitude. A customer support summarization agent that misroutes a transcript has a different loss profile from an agent that can modify payroll, approve refunds, rotate infrastructure credentials, or commit code. The deviation rate may be similar. The annualized loss exposure will not be.
There is a useful analogy in the Verizon DBIR 2025 treatment of internet facing VPN and edge device vulnerabilities. About 54 percent were fully remediated over the observation year, with a median 32 days to patch. The value is not that AI agents are VPNs. The value is that exposure becomes measurable once the population, event class, and observation window are defined.
That is where current AI governance reporting should move. The 65 percent figure from the EMA survey is important because it establishes that scope deviation is not rare. The Reco finding, as reported by Infosecurity Magazine, matters because unsupervised enterprise AI tooling makes contact frequency opaque. Dark Reading’s coverage of agentic AI at Black Hat USA 2026 matters because automated, chained attack sequences increase the range of actions an agentic system can participate in. But none of these facts alone is the risk level.
The quantitative move is narrower and more useful: deviation frequency times vulnerability gives loss event frequency. Then loss event frequency times loss magnitude supports prioritization, budget, and control selection. The methodology can be held by an analyst, but the inputs have to come from the people who know how each agent is actually integrated, permissioned, monitored, and stopped.
That is the observation worth preserving in this moment. “Off script” is a threat event label. It is not a loss estimate. The organization that treats it as a complete risk statement will overreact to noisy agents and underreact to quiet, over privileged ones.