2019 in the archive: the best-recorded year, and what it shows
A year with no defining event and an unusually complete record. The second of those is what makes it worth reading closely.
Last reviewed September 3, 2026
- Opinion outweighs incident. 18.2% against 17.2% on the same measure.
- The busiest month is a calendar effect. Dec peaks on forecasting, not on anything that happened.
- Predictions are written in December. In 4 of 6 years, more than in the January they describe.
- A better record is not a truer one. It makes this year harder to compare with thinner ones, not easier.
The shape of the year, and the season inside it
Each column is a month, with the forecasting portion marked at the base. No month is lifted by an event. The two that rise are the ones at either end of the calendar.
Red: forecasting pieces. Peak Dec at 166, trough Oct at 110, monthly average about 141.
Why does this year get read differently?
Because there is more text in it. The share of entries carrying a written summary, across the archive:
- 201521%643 of 2,997
- 201628%632 of 2,277
- 201738%907 of 2,368
- 201866%1,156 of 1,758
- 201977%1,300 of 1,687
- 202050%947 of 1,891
- 202152%850 of 1,628
77% here against 21% in 2015 — better than three times the coverage. Every other page in this series has had to work around that limitation, stating each subject count as a floor because the text it searches is partial.
With three quarters of a year described rather than merely listed, a different question becomes answerable. Not what an entry is about, which a headline supplies, but what kind of thing it is: an account of something that happened, an argument about what should happen, a set of findings, a piece of instruction, an announcement.
That question is worth asking once, on the year that can support it, because the answer turns out to apply to all of them.
What is security writing actually made of?
Matching the 1,300 summarised entries against indicators for five kinds of piece:
- 18.2%236opinion and prediction
- 17.2%223a specific incident
- 14.7%191a survey or report
- 10.1%131how to do something
- 7.3%95a product or company
Opinion and prediction lead, marginally but consistently, over accounts of specific incidents. Add surveys and instructional pieces and the material that argues, measures or advises is comfortably larger than the material that reports.
That is not a criticism, and it is close to inevitable. Incidents are finite and arguments are not: an organisation is breached once and the case for what it should have done can be made every week for a decade. A publication with a weekly obligation and a finite supply of events fills the difference with the only thing available.
It does have a consequence for anybody treating a security archive as a record of events. Roughly one entry in six here describes something that happened. The rest describes what people thought about it, what a survey found, or what a reader ought to do — and all of it counts identically in any total.
What those categories are not
They are not a classification, and treating them as one would overstate what a list of words can do.
They overlap by design. A piece reporting a survey about a breach matches two of the five, and it should, because it is both. The shares do not add to a hundred and are not meant to. 47% of summarised entries match none of them at all, which is roughly what a set of five crude indicators should be expected to miss.
What the numbers support is a comparison between two of them measured the same way. That opinion exceeds incident is a statement about two expressions applied to the same corpus, and it survives the imprecision because the imprecision falls on both.
What they do not support is a claim about the composition of the field. A different set of words would produce different shares, and anybody wanting a proper taxonomy of security writing would have to read the entries rather than match them. This is the cheap version, and its value is that it can be run over thirteen hundred items and checked by anybody who disagrees with the word list.
Does the field have a season?
It has a strong one. Forecasting reaches 14.5% of December and 15.2% of January, against 5.9% across the five months in the middle of the year. December is also the busiest month of 2019 overall, which makes this the only year in the archive whose highest month is a calendar effect rather than a response to something.
The pattern is not particular to this year. Counting forecasting pieces in December against the January that follows it:
- 201536in December ·19in the January after
- 201627in December ·17in the January after
- 201729in December ·36in the January after
- 201818in December ·27in the January after
- 201924in December ·22in the January after
- 202029in December ·23in the January after
December carries more in 4 of the 6 years. The predictions for a year are largely written before it starts, which is obvious once stated and changes how the genre should be read: they are not assessments of a situation, they are produced to a deadline set by the calendar.
An earlier page in this series counted the forecasts filed in January 2017 and asked what the year expected. This is the fuller answer: the expecting happens in December, and January is where it gets published again.
Where is the actual reporting?
In the middle of the year, and it is quieter there. The five months from Jun to Oct hold 666 entries with forecasting at 5.9% — less than half its share at either end.
Oct, at 110, is the thinnest month of 2019. Nothing about the autumn is quiet in the world; it is quiet in the record because it sits furthest from both ends of the publishing calendar.
There is a practical reading for anybody sampling this material. A month drawn from the middle of a year is a more representative sample of what the field is doing than a month drawn from either end, and January in particular is the least representative month available — it is a review of a year that has not happened.
The effect is not confined to forecasting. August, at 128, is the second thinnest month of the year, and the reason is that a good part of the industry stops publishing for a fortnight. Anybody comparing summer coverage of a subject against autumn coverage of the same subject is partly measuring annual leave.
It also suggests where to look for anything unusual. If a subject appears in July, it is there because somebody had a reason to write about it that week. If it appears in December, it may be there because it belongs on a list of ten things to watch.
Two hundred and twenty-eight publications
1,186 entries name where they came from, across 228 distinct titles. The six most frequent:
- 376helpnetsecurity.com
- 144infosecurity-magazine.com
- 62itproportal.com
- 55forbes.com
- 41securityboulevard.com
- 38information-age.com
That figure matters beyond this page, because the year before it reads as an outlier at fifty-one. The page for 2018 originally treated that number as the mechanism behind a decline in volume, and this year is the measurement that removed the conclusion: the count returns to the ordinary range while the volume stays low.
What the return restores is the tail. A year drawn from more than two hundred titles has a long list of publications contributing one or two items each, and that list is where anything unusual sits — the specialist outlet covering a sector nobody else writes about, the regional paper with a local incident, the researcher's own blog.
The concentrated middle is stable across every year here. Roughly the same handful of trade titles supplies the bulk of each one, and they overlap heavily on the week's obvious material. Whether a year feels varied depends almost entirely on how much of the tail survived into the record.
Artificial intelligence arrives as a subject
63 entries in 2019 concern machine learning, artificial intelligence or synthetic media, which makes it the largest of the year's emerging subjects.
What is absent is any account of it being used. The entries argue about what it will make possible, describe products built on it, and warn about fabricated audio and video, and none of them reports an attack that used any of it. The subject exists here entirely in the future tense.
That is worth marking because it is the ordinary way a subject enters a record. It arrives as speculation, accumulates for several years, and is already familiar by the time anything actually happens — at which point the writing about it looks like a continuation rather than a response.
The synthetic media entries are the clearest case. They describe fabricated audio and video as an approaching problem for identity verification, for evidence and for executives whose voices are publicly available. Six years later this site was describing an attempted intrusion at several large funds that used exactly that, against firms whose controls were built for a world where recognising a voice meant something — the argument set out in the guide to scams and the persuasion behind them.
Being early is therefore not the same as being useful. The warning was correct, it was published years in advance, it was repeated, and the control it argued against was still in place when the thing arrived.
The consequence is a genuine measurement difficulty. A category that has been growing steadily for years on the strength of anticipation cannot easily be distinguished, from inside a count, from one growing because of events. Deciding which is happening requires reading, and reading does not scale.
The data nobody had to steal
42 entries concern material that was reachable because a system had been configured to allow it — a storage bucket left public, a database with no password in front of it, a backup on an open port.
These are the quietest disclosures in security. There is no intrusion, no attacker to attribute, no technique to explain and frequently no evidence that anybody other than the researcher who reported it ever looked. They are also, across this archive, among the most common ways that large quantities of personal information become available.
The uncomfortable part is what they imply about the rest of the record. An exposure of this kind is discovered by somebody scanning for it, so the ones that appear here are the ones a researcher happened to find and chose to publish. There is no reason to think that set is representative, and every reason to think it is small.
The mechanics are dull enough to be worth spelling out, because dullness is the whole problem. A developer needs colleagues to reach a test dataset, so the permission is widened to everybody for an afternoon. A backup routine writes to a bucket whose default was public before the provider changed it. An analytics database is stood up for a project, the project ends, the machine keeps running with no authentication because nobody remembers it exists. None of those is a mistake anybody would defend, and each is a decision that was reasonable for about a week.
The remedy is unglamorous and known: knowing what you run, and checking what it is showing the internet. It appears in this archive continuously for seven years, and 42 entries in the best-documented of them suggests the advice was not the binding constraint.
Eight entries for the cities
Cities, counties and local government appear 8 times across the year, in a period when ransomware against municipalities was becoming routine enough to have a recognisable pattern: services offline for weeks, a decision about whether to pay taken in public, and a recovery bill several times the demand.
The subjects it competed against give the proportion. Staffing shortages appear 42 times, regulators and fines 25, and artificial intelligence 63. A category of incident that was actually happening to real institutions, repeatedly, occupied a fraction of the space taken by a category that had not happened at all.
The reason is the one this series keeps arriving at. A municipal ransomware attack is the same story each time, and the second telling of a story is worth a fraction of the first regardless of how many times the thing recurs. Recurrence is precisely what coverage handles worst.
The shape of a municipal attack is worth stating once for the same reason. Payroll, permits, court scheduling, water billing and the telephone system share an infrastructure bought at different times and maintained by whoever is available, and there is rarely anybody whose job is only security. Recovery is measured in weeks because the systems are old enough that restoring them means rebuilding them, and the bill lands on a body whose budget was set the previous year.
Which means the categories most under-represented in a record like this one are not the rare and dramatic events. They are the ordinary, repeating failures that keep happening to organisations without the budget to prevent them.
The shortage that never resolves
Staffing appears 42 times in 2019: the difficulty of hiring, the size of the gap, the burnout of the people already in post. It is a modest number and an unusual subject, because in seven years of this archive it never produces an event and never goes away.
Every other recurring category has moments. Ransomware has the year it became famous and the year it stalled. Regulation has a deadline. The cloud has migrations that begin and finish. The shortage has none of that. It is present in the first year of this collection and present in the last, at a similar modest share, with the same three arguments made each time — the estimates of unfilled posts, the observation that entry-level roles ask for five years of experience, and the complaint that the training pipeline produces people the industry then declines to hire.
That constancy is itself the finding. A problem discussed continuously for seven years without visible movement is either not a problem or not tractable by discussion, and the archive cannot distinguish between those. What it can show is that writing about it had no discernible effect on how often it was written about.
The subject also sits oddly against the rest of the record. Nearly everything else here concerns systems, and this concerns people — how many there are, what they are paid, whether they stay. Those are the constraints that decide whether any of the advice elsewhere in the archive gets carried out, and they occupy a fraction of the space given to the advice itself.
An organisation reading this archive for guidance would find several thousand entries on what to do and a few dozen on whether there is anybody available to do it. The proportions are almost certainly the wrong way round.
The labels at their halfway point
The eight most-used tags of 2019:
- 351Software Security
- 249Network Security
- 221Cyber Security Strategy
- 209Cyber Security
- 206Data Security
- 172CISO
- 165Cyber Attack
- 159Cyber Threats
This list sits between two states the archive shows at either end of it. In 2015 the leading labels were software, identity, mobile, network and hardware — divisions of a product catalogue, each sorting material into a different pile. By 2020 the top of the list has collapsed into near-synonyms, several of them variants of the same word, applied to a quarter of everything.
Here both are present at once. Software and network security still lead, doing the work they were designed for. Alongside them sit strategy and risk, which are not categories of subject matter but categories of reader — and one entirely generic term, which is the beginning of the collapse.
A classification decays in a recognisable order. Specific labels are joined by audience-shaped ones, audience-shaped ones are joined by generic ones, and the generic term eventually wins because it is never wrong to apply. Nobody decides this; it happens because adding a label is easy and merging two is work that nobody is assigned.
Watching it happen across seven years is the clearest argument available for why the counts on these pages come from titles and summaries instead. A decayed index still looks like an index, and it will return numbers for anything asked of it.
Does a better record make a truer one?
More usable, and in one respect harder to use. The extra text makes subject matching inside 2019 considerably more reliable, and it makes comparison between 2019 and a thinner year less reliable, because the difference in available text is itself a source of apparent change.
Every cross-year figure in this series is calculated only among entries carrying a summary for exactly that reason. It is a restriction that costs a great deal of data —2015 contributes about a fifth of itself to any such comparison — and there is no alternative that does not smuggle the record's own unevenness into the result.
The deeper point is that a record improving in quality partway through produces trends that belong to the record. A subject can appear to rise simply because more of the text that mentions it now exists to be searched, and the rise will be perfectly visible, perfectly measurable and entirely artificial.
The only defence available is the one used throughout these pages: prefer proportions within a year, insist that cross-year comparisons run on a common basis, and treat any finding that depends on absolute totals as provisional until something independent supports it.
What 2019 settled
That regulators would actually impose penalties. Fines and enforcement appear throughout the year, and the argument that data protection law was a paperwork exercise stops being available at this point — which is the reasoning behind this site's guide to breaches and the deadlines attached to them.
That configuration is a bigger source of disclosure than intrusion. A year in which exposed storage outnumbers most categories of attack settles an argument about where effort belongs, and it has not been unsettled since.
That the useful unit of a security programme is an inventory. Both of the year's larger quiet findings — exposed storage and unmanaged systems — reduce to not knowing what you have, and every remedy proposed for either begins with an enumeration somebody has to maintain by hand. The instrument is a spreadsheet more often than a product, and it is stale within a fortnight of anybody finishing it. Which is why the entries arguing for one keep appearing: the task is never finished, so the advice is never spent, and it recurs in every year of this archive without once being resolved, closed or superseded.
And that the shortage of people is structural rather than temporary. It appears continuously and never resolves, which over seven years is the signature of a condition rather than a problem.
What 2019 did not settle is what the following year would do to all of it. The reallocation described on the page for 2020 begins ten weeks after this year ends, and most of the subjects listed above spend the next twelve months competing for a smaller share of the same space.
Common questions
How many entries does this archive hold for 2019?
1,687, of which 1,300 carry a written summary. At 77% that is the highest proportion of any year here, against 21% in 2015.
Why does 2019 get read differently from the other years?
Because there is enough text in it to ask what kind of piece each entry is, rather than only what subject it names. In years where a fifth of entries carry a summary, that question measures how much summary exists rather than what was written.
What is security writing actually made of?
On these indicators, mostly not news. Opinion and prediction appear in 18.2% of summarised entries and specific incidents in 17.2%. Surveys, reports and how-to advice add a good deal more. The field writes about itself at least as much as about what happened to it.
Are those categories exact?
No, and they are not presented as one. They are overlapping indicators from text matching: a piece can be a survey about an incident and count in both, and 47% of entries match none of them. The comparison between two of them is meaningful; the shares do not add to a hundred and are not meant to.
Does the field have a season?
Yes, and it is December. Forecasting pieces reach 14.5% of that month and 15.2% of January, against 5.9% across the middle of the year. In 4 of the 6 years measured, December carries more of them than the January that follows.
Which month is busiest?
Dec, at 166. It is the only year in this archive that ends on its own highest month, and the reason is the forecasting season rather than anything that happened.
What was new in 2019?
Artificial intelligence, at 63 entries — including synthetic media, which arrives here as a subject before it arrives as an incident. Nothing in this year describes an attack that used it.
What about data left exposed rather than stolen?
42 entries concern material that was reachable because a system had been configured to allow it rather than because anybody broke in. It is one of the quieter categories in security writing and one of the more common causes of disclosure.
Did the attacks on local government register?
Barely. Cities and local government appear 8 times across the year, in a period when municipal ransomware was becoming routine. The pattern is the one every year in this archive shows: events are small and conditions are large.
Is a well-documented year a more accurate one?
More usable rather than more accurate. Better summaries make subject matching more reliable within the year and make comparison with thinner years harder, because an apparent rise between a sparse year and a rich one may be nothing but the difference in text available to search.
Was 2019 an eventful year?
Not by the standards of this archive. It has no month lifted by a single event, its busiest month is December, and its named incidents are small. What distinguishes it is the quality of the record rather than the drama of the year.
Can these figures be checked?
Yes. Every count is computed from the archive when the page is built. The type indicators are stated as overlapping matches rather than a classification, and the expression behind each one is a plain list of words.
2019 in the archive
The 1,687 entries behind this page run from January 1, 2019 to December 31, 2019, drawn from 228 publications.
Browse the archive by month, read the same treatment of 2021, 2020, 2018, 2017 or 2015, or search across every year.