Skip to content
The Cyber Security Place

Resilience

The record attack lasted thirty-five seconds

The largest denial of service ever measured was over before most people finished a sentence. The outage that mattered that year ran for more than a day.

Last reviewed August 30, 2026

The biggest recorded attack reached 31.4 terabits per second and lasted about 35 seconds. A single cloud region has been unavailable for roughly 28 hours — some 2,880 times longer. Attacks above a terabit have become ordinary, with about 935 blocked in one half-year by a single provider, and they cause little harm because absorbing volume is a solved commercial problem. What is not solved is concentration: three suppliers hold near 66% of the cloud market, one network fronts around 20% of the web, and in two of three recent major outages the cause was a configuration change rather than any adversary. Some 81% say a 7-day supplier outage would be severe or critical, and almost none have priced the alternative. The cheap remedy is not a second provider but a degraded mode — read-only, cached, capture-now-authorise-later — decided before the day it is needed.

What does it actually rest on?

Start at any of these 4 and keep asking the same question. The starting points are unrelated and the answers are not: authoritative dns is reached by all 4 of them.

Your public website

It runs on servers you administer. And those depend on?

  1. Your application servers. The part you actually operate and can restart.
  2. A content delivery network. Caches and fronts every request before it reaches you, and terminates the encryption.
  3. Authoritative DNS. Turns your name into an address. Nobody reaches you without it, cache or no cache.
  4. A certificate authority. Vouches for that encryption. When renewal fails, browsers refuse the site outright.
  5. One cloud region. Where most of the above is hosted, frequently the same one.

Taking card payments

It goes through your payment provider. And that depends on?

  1. Your checkout code. Yours, and the only link in this chain you can change today.
  2. A payment gateway. Holds the card details so you do not have to, and speaks to the banks.
  3. An acquiring bank's interface. Where the money actually moves, on a schedule nobody in your building sets.
  4. Authoritative DNS. Every one of those calls begins by resolving a name.
  5. One cloud region. Most payment infrastructure is hosted, and hosted in remarkably few places.

Staff signing in

They authenticate against your identity provider. And that depends on?

  1. Your identity platform. Decides who everybody is, for every application you run.
  2. A second-factor service. Delivers or verifies the proof. Often a different company again.
  3. Authoritative DNS. The login page has to be found before it can be typed into.
  4. One cloud region. Identity platforms are hosted like everything else.

Your status page

It exists to tell customers when things break. And it depends on?

  1. A hosted status service. Deliberately kept outside your own systems, which is correct as far as it goes.
  2. A content delivery network. Usually the same one fronting your website, because both chose the market leader.
  3. Authoritative DNS. The status page has a name too.
  4. One cloud region. And here everybody arrives, including the page whose job was to survive this.

Authoritative DNS is reached by 4 of the 4; One cloud region is reached by 4 of the 4.

The exercise is unfair in one direction only, and it is worth naming. Real estates are messier than four tidy chains, and some organisations genuinely have independent paths where this shows convergence. What almost none of them have is a written record of which, and the reason the descent feels uncomfortable is that most people discover the shared floor during the outage rather than before it.

How did the attack stop mattering?

For roughly a decade this subject was the attack. A botnet assembled from domestic cameras and recorders demonstrated that consumer equipment could generate more traffic than large parts of the internet could absorb, and organisations discovered that their answer to it was to telephone somebody and hope. The threat was real, the remedy did not exist at a price ordinary companies could pay, and the coverage reflected that.

What followed is one of the few unambiguous wins in this field. Mitigation became a commodity. Networks with capacity measured in tens of terabits per second now filter traffic hundreds of miles from the target, the arrangement costs less per month than a junior engineer, and the effect is that an organisation behind one experiences a record-breaking attack as a line on a report. The attacks did not stop growing — about 935 above a terabit in a single half-year, and roughly 23.2 million network-layer attacks in the same period — but the harm curve flattened underneath them.

The duration figure is what makes the point unarguable. A peak of 31.4 terabits per second sounds like a siege and describes a burst lasting 35 seconds. Records of this kind are set by botnets firing everything they have at once, which is spectacular, brief, and precisely the shape of thing a large filtering network is built to swallow. Reporting them as escalation is technically accurate and practically misleading.

Two populations remain genuinely exposed, and they are identifiable. Organisations with no mitigation arrangement at all, who can still be removed from the internet by a modest attack and are the ones extortion demands are aimed at. And services where the expensive part is not bandwidth but computation — a search, a login, a checkout — which can be overwhelmed by a volume of requests far too small to register on anybody's terabit chart.

The outages nobody caused

Meanwhile the actual unavailability moved somewhere else entirely. A single cloud region has produced an outage of about 28 hours; another incident generated something like 17 million public reports over 15-odd hours. In two of the three most disruptive recent events the cause was a configuration or metadata change rather than hardware, capacity or anybody hostile. Somebody deployed something, and a fifth of the visible internet stopped answering.

This is not a story about incompetence at large providers, whose engineering is generally better than that of their customers. It is a story about correlation. The same three suppliers hold roughly 66% of the cloud market; one front-line network sits ahead of something like 20% of the web. When the failure domain is that large, an ordinary mistake — the kind every engineering organisation makes monthly — produces a national event instead of an incident report.

The exposure is understood and unresolved. Around 81% of organisations say a 7-day outage at a single supplier would cause severe or critical disruption, which is a remarkable thing for four fifths of a market to agree on and then do nothing about. The reason they do nothing is that the alternatives are worse: running elsewhere costs more, running in two places costs much more, and the outages, however dramatic, remain rarer than the savings.

That calculation is defensible, and it stops being defensible the moment nobody has done it explicitly. An organisation that has weighed a rare multi-hour outage against permanently higher costs and chosen the outage is managing a risk. One that has never asked is carrying the same risk and calling it an architecture.

When the story changed underneath the subject

514 entries here concern availability, denial of service or outages. 322 name denial of service, 90 the botnets behind it, and 32 outages or downtime.

5201410320157120161052017762018392019562020502021

The movement worth watching is inside the columns rather than across them. Denial of service is about 74% of this subject across the first half of the period and around 49% across the second. Early on the attack is the whole story, because that is when a botnet of household devices removed a chunk of the consumer internet and nobody had an answer. The earliest piece here on those botnets, Mirai malware is infecting sierra wireless cellular network equipment help net security, is writing about a problem that was about to be solved commercially.

That the share falls by roughly a third is not a sign the attacks stopped; they grew by every measure available. It is a sign that the coverage found something else to be about, because the answer arrived, was bought, and worked, while the things that kept going down did so for reasons nobody had attacked. A subject can be superseded without being solved, and this one was solved and superseded in the same decade.

There is a caution worth attaching to any archive of this kind, and it applies to the column heights rather than the proportions inside them. Coverage measures what was written about, which tracks what was novel as much as what was harmful. The years where the columns are tallest are the years when this was surprising, and the later years are quieter partly because a solved problem stops generating articles. The proportion is the more honest signal, which is why it is the one marked.

Cheaper than resilience

The instinctive response to concentration is duplication, and for most organisations it is the wrong one. Running a genuinely independent copy of a service at a second provider means two of everything, including two sets of operational knowledge, and introduces its own failures: divergent configuration, data that disagrees between the copies, and a failover procedure exercised once a year by people who are nervous. The expense is certain and the benefit arrives only in the rare event.

The cheaper move is a degraded mode. Decide in advance what the service does when a dependency is missing, and build that instead of building redundancy. Read-only when the database cannot be written. Cached prices when pricing does not answer. Card details captured now and authorised later when authorisation is unreachable. Each is a modest piece of engineering, each converts an outage into an inconvenience, and each works against causes nobody predicted, including the ones inside your own building.

The second cheap move is knowing which services genuinely cannot stop. Most organisations, asked to name them, produce a list of three or four and are surprised by how short it is — and by how much of the estate turns out to tolerate a day offline without anybody minding. That distinction is the entire basis for spending proportionately, and it is the one supervisors in financial services now require firms to write down precisely because so few did it voluntarily.

Third, rehearse the restoration rather than the backup. Almost every organisation takes copies and a startling number have never timed a full recovery, which is the only figure that matters when the question arrives. Restoring a large database from cold storage across a slow link can take days, and discovering that during the incident converts a bad afternoon into a bad week. The number to obtain is simple: how many hours from the decision to restore until customers can transact again, measured once, honestly, on real volumes.

Fourth, and easily forgotten in a subject dominated by digital failures: much of what stops an organisation is physical and prosaic. A power distribution board, a cooling failure in a comms room nobody has entered in two years, a single fibre route between two buildings that the map shows as diverse and the ground does not. These have no attacker, no advisory and no vendor, and they cause a steady trickle of days lost that never appears in any security report. The organisations that handle them well are usually the ones with somebody old enough to remember when this was the whole job, and the ones that handle them badly have generally never walked the route their diagram claims exists. Walking it is a morning of somebody's time, needs no budget approval, and has an unusually high rate of producing a finding: the two supposedly separate feeds entering the same duct, the generator that has not been run under load since it was installed, the door held open by a wedge because the card reader is slow. None of those appear in a scan, and each of them has taken an organisation offline at some point this year.

Last, and least glamorous: keep the thing that reports the outage away from the thing that has the outage. Status pages behind the same front-line network as the service they describe are common, and they fail at the exact moment their purpose begins. A page hosted somewhere unrelated, with a name resolved somewhere unrelated, costs nothing and is the difference between an outage and an outage nobody can get any information about.

What happens during an outage you cannot fix?

The technical answer is nothing, and that is precisely the difficulty. When the failure belongs to a supplier, an organisation has no lever, no diagnostic access and no estimate beyond whatever appears on a public page written by people who are busy. Meanwhile its own customers are telephoning, its own staff are asking, and every hour of silence is interpreted as either incompetence or concealment.

This is why the most valuable preparation is not technical at all. Somebody must be named in advance to speak, and given permission to say the true thing — that the cause lies with a supplier, that there is no estimate, and that the next update comes at a stated time whether or not anything has changed. Regular updates containing no news are far better received than sporadic updates containing news, because the audience is measuring whether anybody is paying attention rather than whether progress exists.

The dangerous impulse in the same hour is to act. Failing over to a secondary arrangement under pressure, with an incomplete picture and a team that last rehearsed it a year ago, has ended more days badly than the original outage would have. If the failover is genuinely automatic and genuinely tested, it has already happened without anybody deciding. If it is neither, the middle of somebody else's incident is the worst moment to find out.

Afterwards there is a review, and the useful version of it asks an unpopular question. Not what the supplier should have done differently, which is unknowable and irrelevant, but what the organisation would need in order to be unbothered next time — and then whether that is worth its cost. Most reviews of supplier outages produce a stern letter and no change, which is a legitimate outcome only if somebody chose it deliberately.

What is an availability guarantee worth?

Less than its arithmetic suggests, and far less than its reputation. Three nines of promised availability permit roughly nine hours of absence a year, four permit under an hour, and both are measured by the supplier, over a window the supplier defines, excluding whatever the contract lists as maintenance. A month containing a multi-hour outage can comfortably remain inside a headline figure that sounds like a guarantee of continuity.

What the agreement actually provides is a refund. Miss the target and the customer receives credits calculated as a percentage of that month's fee, capped, claimable within a window, and payable in further service rather than money. For an organisation whose revenue depends on being open, the compensation for a day offline is a partial discount on a subscription — a sum with no relationship whatever to what the day cost.

None of which makes the document worthless; it makes it a different document than people assume. It is a statement of intent, a definition of what counts as being down, and a mechanism for noticing. Read it for the definitions rather than the numbers: what the supplier counts as an outage, whether partial degradation counts at all, who declares it, and how you would prove your own case if the two of you disagreed about whether the service was working.

The practical consequence is that availability has to be bought with engineering rather than with contract terms. No clause has ever kept a service running. What clauses do is tell you which failures the supplier considers its problem, and everything outside that boundary remains yours regardless of what you paid.

The dependency you bought without noticing

Concentration is usually discussed in terms of infrastructure, and most organisations acquired their heaviest dependencies through purchasing rather than architecture. The customer system, the payroll, the ticketing, the electronic signature, the video meetings, the tool that grants everybody access to all of the above: each was bought separately, by different departments, in different years, and each is now a component the organisation cannot function without and does not operate.

Nobody designed this and the incentives make it inevitable. Buying a service is cheaper and faster than running one, and correctly so. The consequence is that the real availability map of a modern organisation is a purchasing history, and the person who could draw it — someone who has seen every contract and every integration — usually does not exist.

The compounding problem is that these services depend on each other in ways invisible from outside. Several will authenticate through the same identity platform. Several will be hosted in the same region of the same cloud. One will call another for document storage, which the buyer of the first never knew. When the shared component fails, an organisation experiences six unrelated products breaking simultaneously and concludes, reasonably and wrongly, that something has gone catastrophically wrong internally.

The remedy is dull and takes a fortnight: list what has been bought, ask each supplier which region and which providers they depend on, and record the answers. Some will refuse, which is itself informative and worth recording as a finding rather than an inconvenience. The output is not a resilience programme. It is the first honest picture of what the organisation is standing on, and almost nobody has one.

Common questions

How big is the largest recorded attack?

Around 31.4 terabits per second, generated by a large botnet of compromised consumer devices. The detail worth carrying is the duration: it lasted about 35 seconds. These records are bursts, not sieges, and they are absorbed by mitigation networks rather than survived by the target.

Are enormous attacks becoming routine?

Yes, in the sense that they stopped being remarkable. One mitigation provider blocked some 935 network-layer attacks above a terabit per second in a single half-year, 805 of them in the second quarter alone, a rise of roughly 519% between quarters. What has not risen correspondingly is the amount of damage they do.

Why do they do so little damage now?

Because absorbing volume is a solved commercial problem. Mitigation networks have vastly more capacity than any botnet can generate, and traffic is filtered far from the target. An organisation behind one experiences a record-breaking attack as a line on a report, which is why the attacks keep growing and the outage statistics do not follow them.

So what does take services down?

Configuration changes, mostly, at very large providers. A single cloud region has produced an outage lasting about 28 hours, and in two of the three most disruptive recent incidents the root cause was a configuration or metadata problem rather than hardware, capacity or any attacker.

How concentrated is the risk?

Three providers hold roughly 66% of the cloud market between them, and a single front-line network sits in front of something like 20% of the web. Around 81% of organisations say a 7-day outage at one supplier would cause severe or critical disruption.

Is multi-cloud the answer?

Rarely, and it is expensive enough that the question deserves a straight answer. Running genuinely independent copies of a service across two providers roughly doubles operational complexity and introduces failure modes of its own. Most organisations get more availability from removing single points inside their own design than from duplicating the whole thing.

What is worth doing instead?

Knowing which of your services actually cannot stop, for how long, and what each of them rests on. That list is short in most organisations and almost never written down. Everything useful — degraded modes, cached fallbacks, a second provider for one specific component — follows from having it.

What is a degraded mode?

A deliberately reduced version of a service that keeps working when a dependency is gone: read-only when the database is unavailable, cached prices when the pricing service is not answering, offline card capture when authorisation cannot be reached. Designing one is cheaper than resilience and helps more often.

Does a status page help?

Only if it does not share infrastructure with the thing it reports on, which is a mistake made regularly. A status page hosted behind the same front-line network as the service will go down with it, at exactly the moment its entire purpose begins.

Are extortion threats around denial of service still common?

They persist and they mostly fail. Demands arrive claiming an attack will follow non-payment, and organisations behind competent mitigation can ignore them safely. The ones who cannot ignore them are those with no mitigation arrangement at all, which is now a small and identifiable population.

Where do the botnets come from?

Consumer equipment, overwhelmingly: cameras, recorders, routers and other devices shipped with weak defaults and no update path. The economics are unchanged from a decade ago because the incentives are unchanged — the owner of the device suffers nothing, and the manufacturer sold it years ago.

Is regulation addressing concentration?

Slowly, and mainly in financial services, where supervisors now require firms to identify important business services, set tolerances for how long they may be unavailable, and demonstrate they can stay inside them. The approach is sound and the concentration itself is largely untouched by it.

Availability in the archive

514 entries, peaking in 2017 with 105.