Skip to content
The Cyber Security Place

Defence

Application security in 2026: finding flaws was never the constraint

Two decades of tooling made detection nearly free and left repair exactly as expensive as it always was. Everything uncomfortable about the discipline today follows from that one asymmetry.

Last reviewed August 25, 2026

Security debt — a known flaw left unrepaired beyond a year — now affects about 82% of organisations, up roughly 11% in twelve months. The median repairs near 10% of its backlog monthly, which is a fixed point rather than an effort level: at that rate a queue settles at ten times the monthly inflow and the average finding waits about ten months. Third-party components carry 66% of critical debt and clear far slower, a half-life of 358 days against 243 overall. High-risk flaws rose 36% year on year. Roughly 45% of generated code samples failed security testing, and repositories using assistants report about 1.57× the findings and 2.74× the cross-site scripting. The lever is reachability, which shortens the queue without lowering the bar.

Where the backlog settles

Suppose a hundred findings arrive each month and the team clears a tenth of whatever is outstanding. Advance the clock and watch what the queue does. It does not run away, and it does not empty — it climbs and then flattens at a level neither the scanner nor the team chose.

The two controls are the only two variables anybody actually has. One is how much arrives, which is set by how much code gets written and how hard the tooling looks for problems in it. The other is what share gets repaired, which is set by how much engineering time is genuinely available for repair rather than nominally allocated to it.

Worked exampleWhere a backlog settles

718open findings after 12 months
1,000where it settles and stops growing
10.0months the average finding waits

Twelve months in, the queue holds 718 findings and is still climbing towards 1,000.

Where the queue settles, and how long the average finding waits
Fixed per month50 arriving100 arriving250 arrivingAverage wait
5%1,0002,0005,00020.0 months
10%5001,0002,50010.0 months
20%2505001,2505.0 months
30%1673338333.3 months

A proportional fix rate has a fixed point: the queue stops growing when the share repaired equals the volume arriving, which is inflow divided by rate. Nothing in that result depends on effort, and no amount of scanning changes it. Reported figures put the median organisation near ten per cent a month.

The level is inflow divided by rate, and the average wait is the reciprocal of the rate. Ten per cent a month puts the resting level at ten times the monthly arrival and the mean age near ten months, which is uncomfortably close to the reported half-life across all scan types. That correspondence is not a coincidence; it is the same quantity approached from two directions.

Two consequences follow, and both are unpopular. Buying a scanner that finds twice as much raises the arrival rate, so the resting level doubles while every dashboard reports improved coverage. And a team that doubles its throughput halves both the level and the wait, which is the only intervention on the page that touches the number a regulator will ask about.

It also explains a familiar and demoralising experience: a quarter of hard work that visibly closes hundreds of items and leaves the total roughly where it started. Nobody was idle. The system was at its fixed point, and the fixed point does not care how tired anybody is.

Why did finding more flaws not help?

The industry solved detection with real success. Static analysis, composition analysis, dynamic testing, secret scanning and fuzzing all became cheap, fast and available in the editor. One published dataset covers 1.6 million applications and over 141 million raw findings, which is a remarkable engineering achievement and an equally remarkable operational problem.

Repair did not follow, because it is a different kind of work. Detection is a machine reading text. Repair means understanding intent, changing behaviour, proving nothing else broke, and shipping — human work that scales with headcount rather than with compute. Making the first side a hundred times cheaper while the second stayed flat did not shrink the gap. It widened it, and the widening is what the debt statistics describe.

This is why the field's dominant metric quietly stopped being useful. A count of findings measures how hard the tooling looked, and it rises when detection improves and when code volume grows. The number that describes the state of an application is what fraction of what you found is still open, and how old that fraction is.

The vocabulary for closing that gap arrived early enough: Secrets of 'shift left' success ran in August 29, 2018. What the vocabulary underestimated was that moving a finding earlier does not make repairing it cheaper — it only makes it cheaper to decide whether to bother.

Two-thirds of the critical debt is not your code

Half-life is the honest measure here: the days until half of a batch of findings has been repaired. Split by where the code came from, it separates cleanly.

165dYour own code, scanned statically243dAll findings, every scan type358dThird-party components365 days — a finding becomes debt

The bar that matters crosses the line. A flaw in a third-party component reaches its half-life at around 358 days, which is to say the median one becomes security debt before anybody repairs it. Those components also supply roughly two-thirds of the critical debt, so the slowest category is also the heaviest.

The reason is structural rather than cultural. Repairing your own code means editing a file somebody on the team understands. Repairing a dependency means upgrading a version you do not control, in a project you did not choose, with a change log written for somebody else, and then establishing that nothing you built on top of it has quietly changed behaviour. That work is coordination, and coordination does not accelerate when you hire another reviewer.

The dependence itself was visible early — Trojanized, info-stealing PuTTY version lurking online, May 20, 2015 — but the framing at the time was about licence compliance and provenance. The maintenance liability took another several years to be counted, and it is now the larger half of the problem.

Who owns a flaw in code nobody wrote?

Every workflow for repairing a defect assumes there is somebody whose defect it is. For inherited components that assumption quietly fails, and the failure explains more of the half-life gap than any technical difficulty does.

Follow one of these tickets and the shape is familiar. It is raised against a service, so it lands with the team that maintains the service. They did not add the library; it arrived four levels down, pulled in by a framework chosen before two of them joined. They can upgrade the framework, which means a migration guide, a behaviour change in a component they use heavily, and a fortnight they were not given. So the ticket is real, the fix exists, and there is no plausible week in which anybody does it.

Compare that with a flaw in a file the team wrote. Same severity, same rating, completely different economics: an afternoon, one reviewer, done. The two items sit adjacent in a queue sorted by severity, looking identical, and one of them is thirty times more expensive than the other. Any prioritisation that ignores repair cost will keep scheduling the cheap ones and keep deferring the expensive ones, and the expensive ones are the two-thirds.

The organisations that get out of this stop treating upgrades as interruptions to product work and give them a standing allocation — a fixed share of every cycle, defended when the roadmap gets tight. It is unglamorous and it is the intervention that moves the fix rate, which is the only variable in the model that changes where the queue rests.

The alternative, which is more common than anybody admits, is to let the ownership question go unanswered and allow the age of the queue to answer it instead. That is a decision too. It is simply one that gets made by default, and reviewed only when something in it is exploited.

The code arriving faster than anyone can review it

Assistants changed the arrival rate, which is the one variable in the model above that nobody deliberately chose to move.

The measurements are consistent enough to take seriously. In one large evaluation roughly forty-five per cent of generated samples failed security testing outright, with particular classes doing much worse: cross-site scripting failed in the great majority of attempts, and log injection similarly. Repository-level studies report around 1.57 times the security findings and about 2.74 times the rate of cross-site scripting compared with human-written equivalents, alongside a tenfold rise in monthly findings across six months in the codebases studied.

The subtler result is about the people rather than the output. Developers using assistants have been observed writing less secure code while rating their own solutions as more secure than they were. Fluent, well-formatted, plausible code reads as reviewed code, and a reviewer who is nodding along is not reviewing.

None of this is an argument against the tooling, and pretending otherwise would be both futile and wrong: the productivity is real and it is not going to be given back. The argument is narrower and follows from the arithmetic. If arrivals rise and the fix rate holds, the resting level rises in proportion. An organisation that adopted assistants and left its repair capacity untouched has already decided to carry a larger backlog; it has simply not written that decision down.

Does shift left still make sense?

Half of it does, and the half that failed took the reputation of the other half with it.

The durable insight was always about feedback latency. Telling somebody about a problem while the code is still in their head costs a few minutes; telling them four months later costs a day of re-reading plus the risk of breaking whatever was built on top. Nothing has undermined that.

What failed was the operational reading: that moving detection earlier would move the work earlier too. It moved the alerts earlier, to people who were already at capacity and had no authority to decide which ones mattered. The documented consequence is fatigue — a queue full of findings that are theoretical, duplicated or unreachable, presented with the same urgency as the two that are genuinely exploitable. People stop reading a channel that has cried wolf, and then the real one arrives in a channel nobody reads.

The correction now being argued for is not a retreat to a gate at the end. It is that specialists, not developers, should own triage: decide what is real, what is reachable and what is urgent, and hand developers a short list they can trust. Developers keep the fast feedback; they stop being made responsible for judging a firehose they never asked for.

What does reachability actually remove?

Composition analysis traditionally answers a question of inventory: is a vulnerable version of this library present. That question has a high yield of true-but-useless answers, because presence is not the same as exposure.

Reachability asks a harder one: can execution actually get from your code to the affected function. Most of the time it cannot. A library may be pulled in for one helper, and the vulnerable path may sit in a subsystem nothing in your application touches. Answering that turns a list of components into a much shorter list of exploitable paths — and unlike every other way of shortening a queue, it does so without lowering the standard.

That distinction is worth insisting on. Raising the severity threshold shortens the list by ignoring real problems. Reachability shortens it by removing things that were never problems in your particular build. The first buys quiet, the second buys accuracy, and only the second survives contact with an auditor.

The practical caveat is that reachability is not free and not perfect. Dynamic dispatch, reflection and configuration-driven behaviour defeat static call graphs, so the analysis is conservative and errs towards reporting. Used as a ranking signal it is excellent; used as permission to ignore everything it cannot reach, it eventually goes wrong.

There is also a question of who is allowed to act on the result. If the analysis lives inside a security team's console and developers see the unfiltered list, the noise reaches them anyway and the investment buys nothing. The value is realised at the moment the shortened list becomes the only list anybody is asked to work through, which is an organisational change wearing a technical disguise.

A decade of the same names

This section carried 1,350 reports across seven years, and the year-by-year volume tracks the industry's attention rather than the state of the code.

Reports in this section, by year
2015201620172018201920202021
15717615226834013887

Read the headlines across that stretch and the striking thing is how few of the flaw classes changed. Injection in its several forms, broken access control, unsafe deserialisation and mishandled untrusted input recur year after year, which is why the perennial top-ten lists are revised so gently between editions. The techniques for finding them improved enormously; the frequency with which they are written did not.

Only one genuinely new surface appeared in that period at scale, and it appeared quietly. Interfaces designed for machines multiplied, each one an entry point with its own authorisation logic, frequently undocumented and rarely inventoried.

Going Beyond Usernames and Passwords ran in January 8, 2016, at a point when most organisations could not have listed the interfaces they exposed. A decade on, the inventory problem is largely unchanged and the count is an order of magnitude larger.

The stability of that list is easy to read as failure, and it is worth resisting the reading. Injection persists not because the industry forgot how to prevent it but because new code keeps being written, by new people, against new frameworks, and each generation meets the same trap for the first time. A defect class disappears only when the platform makes it unrepresentable — which is what parameterised queries did for one variant of injection and what memory-safe languages are now doing for an entire family of others.

That points at the intervention with the longest reach and the slowest payback. Every other control on this page reduces the number of flaws that survive; a platform change reduces the number that can be written. It is the only lever that lowers the arrival rate rather than raising the repair rate, and the two are interchangeable in the arithmetic.

Debt with a deadline attached

An ageing backlog used to be an internal embarrassment with no external consequence. In Europe that is ending, and the mechanism is worth understanding because it changes what the backlog is.

The Cyber Resilience Act obliges manufacturers placing products on the European market to handle vulnerabilities across the supported lifetime and to report actively exploited ones within twenty-four hours, with those duties beginning in September 2026. Attach that to the half-life figures above and the collision is immediate: a category of flaw whose median repair time exceeds a year now sits inside a regime that expects disclosure within a day of exploitation.

The practical effect is to convert a queue into an inventory of dated liabilities, each one older than the last review that failed to notice it. A finding that is eleven months old and one that is thirteen months old are indistinguishable on a dashboard sorted by severity, and completely different under a regime that asks how long you have known.

Which suggests the reporting change that most organisations have still not made. Stop leading with a count of open findings, which rises when the tooling improves, and lead with the share older than a year, plotted over time. It cannot be improved by scanning less, it moves only when repair actually happens, and it is the number the regime is built around.

How should a backlog be prioritised?

Severity alone is a poor sort order, because it describes what a flaw could do in the abstract rather than what it can do in your deployment. Four questions produce a better ranking, and all four are answerable.

Is it reachable? Can execution get there from code you actually run. This removes more of the queue than every other filter combined.

Is it exposed? An internet-facing endpoint and an internal batch job with the same rating are not the same finding.

Is it being exploited? Public exploitation moves an item from theoretical to scheduled, and high-risk flaws — severe and likely exploited — rose about thirty-six per cent year on year, so this filter is doing more work than it used to.

What does it guard? A moderate flaw in front of customer records outranks a critical one in a demo service, and no automated severity score knows the difference.

Applied honestly, those four turn an unreadable queue into a list short enough to finish. That is worth stating plainly: the goal is a list somebody can complete, because a queue nobody can finish is a queue nobody starts.

Where to start on a Monday

Three measurements, none of which needs a purchase, and all of which change the conversation more than another scanner would.

Measure your actual fix rate. Count what was open at the start of last month and what closed during it. The ratio is the number that sets everything else, and most teams have never calculated it.

Age the backlog and plot the share over a year old. One chart. It will be worse than expected and it is the only figure on the page a regulator would recognise.

Expect the first version of that chart to be wrong, and publish it anyway. Ageing a backlog accurately requires a first-seen date that many tools record badly: rescans reset timestamps, re-imported findings arrive as new, and a queue migrated between platforms usually lost its history in the move. The first chart therefore understates the problem, sometimes badly. It is still the most informative artefact the team will produce that quarter, and the arguments it triggers about which dates can be trusted are themselves worth having.

Take a sample of twenty findings and ask whether each is reachable. Not the whole queue — twenty. The proportion that turns out to be unreachable tells you how much of your team's attention is currently spent on nothing.

After those three the harder work is genuinely prioritisable: raising throughput with dedicated repair time rather than goodwill, moving triage to people whose job it is, and negotiating upgrade paths for the dependencies that supply most of the debt. An organisation that has done the first three knows whether its queue is shrinking or merely being restated, and that distinction is what the entire discipline currently turns on.

Common questions

What is security debt?

A known flaw that has gone unfixed for more than a year. It is a useful category because it separates the work in progress from the work that has quietly become permanent, and it is now the normal condition: roughly four in five organisations carry some, an increase of about eleven per cent in a single year.

Why does the backlog stop growing at a particular size?

Because a proportional fix rate has a fixed point. If a fixed number of findings arrive each month and a constant share of the queue is repaired, the queue settles where the share repaired equals the volume arriving — inflow divided by rate. At the reported median of ten per cent a month, the average finding waits about ten months regardless of how large the team is.

How much of the debt is our own code?

A minority of the critical portion. Third-party components and open-source libraries account for around two-thirds of critical security debt, and they take substantially longer to clear: a half-life near 358 days against 243 across all scan types.

Why do third-party flaws take longer to fix?

Because fixing one is usually a version upgrade you do not control, in a dependency you did not choose, that may break something you did not write. The work is coordination rather than editing, and coordination does not respond to hiring more reviewers.

Is generated code less secure than hand-written code?

Measurably, on current evidence. Around 45% of generated samples failed security tests in one large evaluation, with cross-site scripting and log injection failing far more often than that. Separate repository studies report roughly 1.57 times more security findings and about 2.74 times the rate of cross-site scripting compared with human-written code.

Does that mean assistants should not be used?

No, and the argument does not follow. The tooling raises output volume; the security consequence is that findings arrive faster while repair capacity stays flat. That is a scheduling problem with a known shape, and it is addressed by raising the fix rate rather than by refusing the productivity.

Has shift left failed?

The narrow version has. Moving detection earlier without moving repair capacity produced more findings, sooner, for people already at capacity — and alert fatigue is the documented result. What survives is the useful half: fast feedback on code somebody is still holding in their head.

What is reachability analysis?

A check on whether a vulnerable function is actually callable from your application, rather than merely present in a dependency. It converts a list of components into a much shorter list of exploitable paths, and it is the most reliable way to cut noise without lowering standards.

Which flaw classes recur most?

The same ones as a decade ago. Injection in its several forms, broken access control, and unsafe handling of untrusted input dominate, which is why the perennial lists change so little between editions.

How should a backlog be prioritised?

By exploitability and exposure rather than by severity score alone. A critical rating on a component that no request path reaches is worth less attention than a moderate one on an internet-facing endpoint holding customer records.

What does the Cyber Resilience Act change here?

It attaches deadlines to something that previously had none. Manufacturers placing products on the European market must handle vulnerabilities and report actively exploited ones within 24 hours from September 2026, which turns an ageing backlog into a compliance exposure.

What single metric is worth reporting to a board?

The share of the backlog older than a year, plotted over time. It is harder to game than a count of findings, it moves only when repair actually happens, and its direction answers the question a board is really asking.

Application security coverage

1,350 reports, newest first. Most cited sources: helpnetsecurity.com (248), infosecurity-magazine.com (111), itproportal.com (68), csoonline.com (44), informationsecuritybuzz.com (44), information-age.com (37).