When Breach Statistics Lose Their Uncertainty

The DBIR’s intervals are not presentation detail, they are the part a quantitative risk model should preserve.

The Number Was Never Just 20 Percent

On 26 August 2026, as reported by Infosecurity Magazine, CISA added six already known vulnerabilities to its Known Exploited Vulnerabilities catalog, covering Microsoft, Linux, Red Hat and Citrix products, with active exploitation confirmed in the wild. Around those same days CISA also issued a binding directive requiring US federal agencies to patch an actively exploited Citrix NetScaler remote code execution flaw on a compressed deadline.

That timing matters because Citrix NetScaler is not just another product name in a vulnerability bulletin. It is an edge and remote access appliance, the exact class of technology that has been moving fastest in breach data. In the Verizon Data Breach Investigations Report 2025, exploitation of vulnerabilities as an initial access vector appears in 20 percent of breaches, up from about 15 percent, which the report frames as a 34 percent year over year increase. Within that vulnerability exploitation subvector, VPN and edge devices were 22 percent, up from 3 percent the prior year, roughly an 8x increase.

A risk register still carrying last year’s midpoint for “vulnerability exploitation as initial access” is therefore wrong twice. It is wrong to use a point at all, and it is wrong to use a stale one.

The first error is statistical. The DBIR does not present its figures as hard constants. Its charts use slanted bars to represent a 95 percent confidence interval. The authors explicitly forbid pie charts. They also tell readers not to treat the percentages as point estimates. That warning exists because the data is contributed, sampled and classified under real constraints.

The second error is operational. Even if the midpoint were acceptable, which it is not, the movement in VPN and edge device exploitation means the older assumption no longer describes the threat surface with enough fidelity. A number copied into a slide looks stable because it has no visible uncertainty around it. The environment it describes is not stable.

The Interval Is The Input

A quantitative cyber risk management methodology needs frequency inputs as distributions. Threat event frequency describes how often a threat community acts. Resistance strength and vulnerability determine how often those actions become loss events. Loss event frequency then feeds loss magnitude to estimate annualized loss exposure, often through Monte Carlo simulation.

In that workflow, a published benchmark such as the DBIR is not best understood as a source of precise answers. It is a calibration source. It helps shape assumptions about loss event frequency and vulnerability, especially when an organization has limited internal data. But the useful object is the range, not the midpoint.

If the DBIR says a category sits around 20 percent and shows a 95 percent confidence interval, the interval should travel into the model. It can become a calibrated range, a parameter for a distribution, or a reason to widen uncertainty when evidence is thin. What should not happen is the common transcription error: strip the interval, keep the visible label, and paste the midpoint into a risk register as if it were a measured constant.

That habit creates false precision. It also distorts prioritization. Stolen or abused credentials remained the most common initial access vector in the DBIR 2025 at 22 percent, down from 31 percent. Phishing was about 15 percent and roughly stable. Vulnerability exploitation was 20 percent, but moving materially. A control conversation that treats 22 percent, 20 percent and about 15 percent as rigid ranks misses the point. The better question is how much uncertainty sits around each estimate, whether the organization’s own exposure resembles the report population, and whether the direction of movement changes the decision.

The same issue appears in remediation data. For internet facing VPN and edge device vulnerabilities, organizations had fully remediated about 54 percent, with a median remediation time of 32 days. A leaked secret found in a public GitHub repository had a median remediation time of 94 days. These are not interchangeable “patching is slow” anecdotes. They describe different control processes, different discovery paths and different exposure windows. A quantitative model can use those medians, but only if it keeps the definition of the population attached.

This is where many registers collapse categories. Third party risk is a frequent example. The DBIR 2025 reports third party involvement in 30 percent of breaches, doubled from 15 percent the prior year, and calls it the most significant year over year shift. But that 30 percent is an “any involvement” measure. It is not an initial access vector share. It cannot be dropped into the same model slot as the 20 percent vulnerability exploitation figure without changing the meaning of the variable, and with it the frequency pathway being modeled.

Caveats Are Model Data Too

The interval is not the only part that gets lost. Written caveats often disappear even faster.

The clearest example in the DBIR 2025 is espionage. Espionage motivated breaches were 17 percent, up from about 6 percent, a 163 percent increase. In espionage breaches specifically, vulnerability exploitation was the entry point 70 percent of the time. Those are important figures for any enterprise with exposure to state aligned activity, sensitive intellectual property, regulated infrastructure, or geopolitical relevance.

But the report explicitly notes that the rise is partly attributable to a change in which organizations contributed data this year, not purely to a real world shift. That caveat should change how the statistic is used. It does not make the finding useless. It makes it conditional.

In a quantitative risk model, that condition might widen the distribution. It might reduce confidence in year over year trend extrapolation. It might trigger segmentation by threat community or business unit exposure. It might lead to scenario analysis rather than a single blended frequency estimate. What it should not do is vanish between the report PDF and the board deck.

The same discipline applies to ransomware. The DBIR 2025 reports ransomware present in 44 percent of breaches, up from 32 percent. It also reports a median ransom paid of 115,000 dollars, down from 150,000 dollars, and victims refusing to pay at 64 percent, up from 50 percent two years earlier. A qualitative summary might say ransomware is rising while payments are falling. A quantitative analysis has to ask what distribution of loss magnitude is implied when payment behavior changes, what portion of the loss is independent of ransom payment, and whether refusal to pay shifts cost into restoration, interruption, legal response or customer impact. The human element, present in 60 percent of breaches and down slightly from 61 percent, deserves the same care. It is a common condition in breach pathways, not a single loss event path, and it should not receive a flat multiplier across every scenario.

The practical fix is modest. When a DBIR statistic enters a risk register, threat model, control business case, or board risk narrative, carry three things with it: the definition, the interval and the caveat. Define what the figure actually measures. Preserve the uncertainty as a range or distribution. Record why the report authors think the number moved.

That practice works best when the analyst holding the quantitative discipline sits with control owners who understand the local environment. The analyst asks what the interval and the caveat were. The control owner answers whether the estate actually looks like that population. For edge and remote access appliances, that means knowing whether internet facing services are inventoried, whether emergency patching can move faster than median behavior, and whether compensating controls meaningfully change resistance strength while remediation is incomplete.

This is not a call for more external data. The DBIR already gives more statistical humility than many readers preserve. The problem is that uncertainty is treated as a cosmetic feature of the chart rather than as an input to the model. When CISA adds actively exploited vulnerabilities affecting edge infrastructure and federal deadlines compress almost immediately, the cost of that transcription error becomes concrete. The risk question is not whether vulnerability exploitation is exactly 20 percent of breaches. It is what range of loss event frequency is credible for this environment now, given the observed movement in edge device exploitation, the organization’s own remediation performance, and the caveats attached to the source.

Quantifying what others leave qualitative does not mean pretending uncertainty has disappeared. It means carrying uncertainty forward with enough structure that decisions can absorb it.