Application security tooling encodes a belief about the attacker: that their time is expensive. Scanners fire known payloads at known shapes because exhaustively reasoning about an unfamiliar application used to cost a skilled human several days. Everything downstream inherits that belief: severity models, scan windows, the idea of a quarterly pen test.
What changed
An attacker with an agent reads your documentation, your public API surface, your job postings and your changelog before touching anything. They form hypotheses about how the system is put together and test the cheap ones first. They notice that an endpoint returns one field more than the UI displays. They give up on dead ends in seconds rather than hours, which means they can afford to be wrong constantly.
None of that produces traffic a signature matches. The requests are individually well-formed and individually authorised. The attack is in the sequence.
Why a findings list does not help
Three findings rated Low, in the right order, are one finding rated Critical.
A list of independently-scored issues is the wrong artefact for this. The severity of a behaviour depends entirely on what else is reachable from it, and no tool that scores findings in isolation can see that.
What we do instead
- Model the application as a set of actors, human and agent, with what each is implicitly trusted to do.
- Propose chains across trust boundaries, including the ones that belong to no single team.
- Run them in a shadow environment and keep the evidence, so a finding is a demonstration rather than an assertion.
- Re-run on every change, because the chain that does not exist today is created by the pull request that ships on Thursday.
The test is not whether a tool can find a known vulnerability class in your code. It is whether it can find the thing that only becomes dangerous in combination with the three other things you already shipped.