The Number One Attack Vector Is Not Your Number One Attack Vector

A ranking from a credible corpus can sharpen cyber risk analysis, but only after it is converted from an ordinal headline into conditional rates with visible denominators.

The Seduction of First Place

Infosecurity Magazine reported on 9 September 2026 on research by SpyCloud finding that non-human identities, including service accounts, API keys and tokens, had become the most common entry point for corporate cyberattacks, surpassing human credentials as the initial attack vector. That is a useful finding. It is also the kind of finding that becomes dangerous when it travels too far from the corpus that produced it.

The problem is not that the vendor is wrong. The problem is that “number one” is not a risk estimate. It is an ordinal statement about somebody else’s collection of observations. A rank cannot be multiplied by an exposure count, split across asset classes, compared against a control budget, or fed into a frequency estimate. The practical damage is not that the ranking is false, it is that ranks get read as rates, and the two are not interchangeable.

The same critique applies to every ranked list, including the ones security teams already trust. Top attack vectors, leading causes, most common initial access paths and dominant breach patterns all describe the ordering inside a collection mechanism, not the expected frequency of loss for a particular organization. A ranking may be a good prior. It is not a denominator.

This distinction matters because quantitative risk analysis is built from quantities that can compose. Threat event frequency decomposes into contact frequency multiplied by probability of action. Loss event frequency is threat event frequency multiplied by the vulnerability term, meaning the proportion of attempts that actually result in loss. A global ranking of entry points is, at best, a statement about contact frequency inside the reporting corpus. It says nothing about an individual organization’s probability of action or about its resistance strength. Confusing an attempt rate with a loss rate is the single most expensive error in the whole estimate.

The Corpus Changes the Ordering

The VERIS Community Database, the open, community-contributed incident corpus maintained by the Verizon VTRAC team, uses the same VERIS schema as the annual Verizon Data Breach Investigations Report but holds row-level incidents rather than aggregate percentages. It holds 10,591 incidents in total, of which 3,769 belong to healthcare, finance and manufacturing victims, classified by NAICS industry code.

Within that subset, healthcare accounts for 2,563 incidents, finance for 980 and manufacturing for 226. The action types are counted per incident, and an incident can carry more than one action. In healthcare, Physical appears in 698 incidents, Error in 685, Misuse in 647, Hacking in 358, Malware in 227 and Social in 145; in finance, Hacking in 307, Physical in 270, Error in 210, Malware in 132 and Misuse in 131; in manufacturing, Hacking in 98, Misuse in 44, Malware in 40, Error in 30 and Physical in 28.

That is not a minor reshuffling of labels. It is a sector-conditioned reversal. Hacking appears in 358 of 2,563 healthcare incidents, about 14 percent, in 307 of 980 finance incidents, about 31 percent, and in 98 of 226 manufacturing incidents, about 43 percent. The same action type ranks fourth in one sector and first in the other two, and its share of incidents roughly triples across them.

In healthcare, Physical plus Error is 1,383 against 358 for Hacking, close to four times as many. In the VERIS schema, Physical covers lost and stolen devices and paper records, and Error covers misdelivery and improper disposal. Both are non-adversarial, or at least not network intrusions. That matters because the controls implied by an entry-point headline are usually identity controls, secrets rotation, privileged access hygiene and detection around authentication misuse. Those controls may be necessary. They do not stop a misdelivered discharge summary or a stolen unencrypted laptop.

This is where ordinal thinking does real operational harm. A portfolio steered by the global first-ranked entry point can raise resistance strength against that vector while leaving the two action types that generate most of the sector’s recorded loss events untouched. The problem is not overinvestment in identity hygiene, it is treating identity hygiene as the answer to an ordering problem whose local denominator has not been established.

Data variety reinforces the same point. In healthcare, Medical appears in 1,955 incidents and Personal in 688. In manufacturing, Personal appears in 86, Payment in 37 and Secrets, meaning trade secrets and intellectual property, in 34. Manufacturing is the only one of the three sectors where Secrets appears near the top, and that matters because loss from trade-secret compromise belongs to a competitive-advantage category rather than a regulated-personal-data category. The action ordering and the loss category have to be read together, because a control that reduces one route into systems may have little bearing on the loss magnitude distribution that matters most.

Caveats Are Part of the Measurement

The open corpus should not be treated as better than the vendor corpus, only as visible enough to inspect. That visibility shows why every ranked list needs its collection mechanism stated out loud.

Of the 3,769 matched incidents, 1,811, or 48 percent, carry no year at all, so no trend claim can be made from this subset without first checking the denominator per year. In healthcare, 1,757 of the 2,563 incidents, about 69 percent, have an unknown organization size. Classification is done by whoever submitted each incident, so coding consistency varies more than in a single-analyst dataset. Mandatory breach-notification regimes in healthcare plausibly surface lost-laptop and misdelivery events that would never become public in a sector without them, which is itself a selection effect in the corpus and a reason the healthcare ordering may overstate the non-adversarial share.

Those caveats do not cancel the analysis. They define the analysis. A ranked list is a statement about a collection mechanism before it is a statement about the world. A survey of security leaders, a public incident corpus, a breach notification dataset and a managed detection provider’s telemetry each observe a different part of the threat landscape, with their own blind spots, incentives and paths by which an event becomes visible.

The honest move is therefore not to pick one corpus and crown it, but to name the corpus, state the selection mechanism, convert its ordinal claim into a share with the denominator visible, and carry that share as a range rather than as a point. The ranking can then become one prior among several, rather than an input masquerading as a measurement.

From Ranking to Frequency

The translation from headline to risk estimate starts by refusing to multiply a rank. “Most common” has to become “this share of this denominator, under this collection mechanism.” Only then can the finding be compared with local incident history, asset exposure, control coverage and expected loss magnitude.

For non-human identities, the practical question is not whether they deserve attention. They do. Service accounts, API keys and tokens often carry broad access, weak ownership and poor rotation discipline, and they persist after the application, integration or administrator that created them has changed. A credible finding that they are now the leading entry point in a reported corpus should increase their weight as an external prior.

But a prior is not a replacement for local evidence. For each action type, the control question has to be explicit: which budgeted control reduces contact frequency, which reduces probability of action, which increases resistance strength, and which reduces loss magnitude after failure? Identity hygiene and secrets rotation reduce some forms of contact and some forms of successful action. They do not reduce misdelivery or improper disposal, they do not encrypt an endpoint that is already gone, and they do not change whether a paper record can leave a building.

The quantitative risk model forces that separation, because contact frequency, probability of action, vulnerability and loss magnitude are answered by different evidence and moved by different controls. A headline ranking can inform one of those terms only after it has been converted into a rate, and even then only with uncertainty carried forward.

The final estimate belongs at the intersection of outside and inside evidence. External corpora provide base rates and reveal how collection mechanisms shape what becomes visible. Internal records provide the denominator no outsider can see: which assets exist, which are exposed, which identities are privileged, which devices leave controlled environments, and which incidents were recorded internally without ever becoming public. The number one attack vector in a headline may be the right starting prior. It becomes a risk estimate only after it survives contact with the organization’s own denominator.