2014 in the archive: where the record begins
Five months, one of which is three quarters of the year. The beginning of a collection, and the two modes it turns out to run in.
Last reviewed September 4, 2026
- The beginning is unmarked. An ordinary Thursday's filing, and then there is a record.
- 74% of the year is one month. Any total for 2014 describes Dec.
- The slow months describe everything. 100% against 43% once volume arrives.
- The first thing it does at full volume is predict. 33 of 257 entries forecast the year ahead.
The shape of five months
Twelve columns, seven of them empty because the record does not exist before August. The solid portion of each column is the part carrying a written summary; the hollow part is entries that exist as a headline and a date.
Solid: entries carrying a written summary. Outline: the rest. Aug 24 at 100%, Sep 17 at 100%, Oct 7 at 100%, Nov 40 at 100%, Dec 257 at 43%.
Where does a record begin?
On 2014-08-07, with nothing to distinguish it from the day before, which is not in the collection, or the day after, which is.
The page for 2021 asks the same question at the other end and reaches the same answer. A record stops on an ordinary Tuesday because nobody knew it was the last day, and it starts on an ordinary Thursday for the corresponding reason: a beginning is only recognisable once there is something to compare it against.
What the first three entries say
- New Research to Show Aircraft, Ships and Traffic Systems at Risk From Hackers
- NSS Launches Cyber Resiliency Center
- Russian Gang Steals 1.2 Billion User Credentials in Biggest Ever Hack
Research indicating that aircraft, ships and traffic systems are exposed. A company opening a facility. A claim that a criminal group holds well over a billion stolen credentials.
As an opening trio it is unintentionally representative of everything that follows. One warning about infrastructure, published years before the incidents that would make the argument concrete. One announcement, of the kind that occupies a steady share of every year in this collection. And one very large number, of the kind that arrives without a method attached and is repeated for years afterwards by people who never saw one.
The billion-credential claim repays a moment of scepticism, and the reasons are generic enough to apply to any figure of that kind. A hoard assembled by scraping many separate incidents contains duplicates, and deduplicating across sources is difficult when the same address appears with different capitalisation. Credentials accumulated over years include a large share already rotated, already abandoned, or belonging to accounts that no longer exist. And the party announcing the total is frequently the party selling monitoring against it, which is not disqualifying but is worth putting on the record beside the number.
None of that means the figure was invented. It means the figure needs a method attached before it can be compared with anything, and the method is what did not travel.
That third entry is worth pausing on, because the figure was contested at the time and the contest is not in the archive. What is recorded is the claim; what is missing is the argument about whether the count was meaningful, how the sample was drawn, and whether the credentials were current or an accumulation of old ones. A record of headlines preserves assertions considerably better than it preserves the doubts about them.
What changes in December?
August through November hold 88 entries between them, at a rate that never exceeds 40 a month and once falls to 7. Then Dec holds 257.
That is not a ramp. A collection finding its feet would show growth across several months, and this shows four months of one thing followed abruptly by a different one. The transition happens between two calendar months with nothing in between.
Whatever changed is not visible in the entries. They are drawn from the same kinds of publication, cover the same kinds of subject, and read identically on either side of the boundary. Only the counting changes, and the counting is exactly what a record cannot explain about itself.
It is worth noticing what the switch does not do. The subjects do not change: retail card data, surveys, vendor announcements and infrastructure warnings appear on both sides of it in similar proportions. Nor does the sourcing change character. What arrives in December is more of the same kind of material, described less fully — which is the signature of a throughput change rather than an editorial one.
The practical consequence is that 345 is not a measurement of 2014. It is a measurement of Dec with four months of something else attached, and the two parts should not be added together for any purpose that treats them as comparable.
Who supplied the first months?
197 of the 345 entries name where they came from, and they name 132 different publications. That is roughly three attributed entries per title across the whole period, and it makes 2014 by some distance the most dispersed stretch in this archive.
The six appearing most often:
- 11tech.einnews.com
- 6net-security.org
- 6information-age.com
- 5prnewswire.com
- 5finance.yahoo.com
- 5computerworld.com.sg
Nobody supplies more than a handful. In later years a single trade title routinely accounts for a fifth or a quarter of everything attributed, and here the leader is in double figures only just. The material arrives from wire aggregators, national business pages, regional technology sites and specialist outlets in roughly equal measure.
Dispersion of that degree has a cost and a benefit, and they are the same property seen from two directions. Nothing dominates, so no single editorial judgement shapes the record — but nothing recurs either, so a subject cannot build across weeks the way it does when a handful of outlets are following it together. The early archive is broad and discontinuous, which is what a collection looks like before anybody has decided who is worth reading regularly.
Two modes, and what each is good for
The difference between the two halves is not only in volume. It is in how much of each entry exists.
- Aug24entries ·100%described
- Sep17entries ·100%described
- Oct7entries ·100%described
- Nov40entries ·100%described
- Dec257entries ·43%described
Four months at a hundred percent, then 43%. Every entry taken in the slow period has a written summary underneath it; well over half of what arrives in Dec is a headline, a date and a link.
This archive shows the same trade exactly once more, and it is the reason this year is worth a page. In 2016, three months running at a fifth of the ordinary rate carry summaries on 82% of their entries against 24% for the rest of that year. Different year, different circumstances, same relationship: less captured, better described.
Two instances is a pattern worth naming and not a law, and no mechanism is offered because none is available. Breadth and depth are not independent here, and anything combining periods of both kinds is measuring their mixture.
What the first full month chose
With 257 entries available for the first time, the record spends them like this:
- 33forecasts for the year ahead
- 32retailers and payment cards
- 29surveys and reports
- 13how-to advice
- 9the studio breach
The largest single category is forecasting. 33 entries of 257 concern the year ahead: predictions, trend lists, things to watch. The archive reaches full volume in December and immediately starts describing a year that has not happened.
That is the earliest confirmation of something the page for 2019 measures properly across six years — that forecasting is a December activity, published before the year it describes, produced to a deadline set by the calendar rather than by any assessment. The very first month this collection runs at strength is a forecasting month, and it was always going to be, because December is when that work is done.
Surveys and reports account for another 29. Between the two, roughly a quarter of the first full month is material about the field rather than about anything that happened in it — a proportion that stays remarkably stable for the next seven years.
Why it begins on retailers
Retailers and payment cards account for 32 entries in Dec, second only to the forecasting. The period this archive opens in was defined by card data taken from shop tills at a scale nobody had seen before — tens of millions of cards from single chains, extracted from the machines customers tap.
The mechanism was a supplier. Access came through the systems of a contractor with a legitimate connection, and from there into the network that carried card data between the till and the bank. That is the argument this site now makes constantly about the route in belonging to somebody else, and it is on the archive's first page.
The remediation described in those weeks is worth listing, because it is unglamorous and because almost none of it involves security products. Segmenting a flat network so a thermostat contract cannot reach a payment terminal. Cataloguing which outside firms hold a live connection and revoking the ones nobody can justify. Replacing shared vendor logins with named accounts. Instrumenting tills so that unexpected software running on one raises something. Rehearsing who telephones the acquiring bank, and when. Each is a week of somebody's time, none produces a demonstration, and collectively they are the whole of the answer.
It is also the earliest form of the other argument that recurs every year here. The customer chose the shop. The customer did not choose the payment processor, the contractor, the till manufacturer or the network between them, and the failure occurred in a relationship they had no visibility of and no ability to decline.
Seven years later the same structure produces an attack reaching companies through the firms that administer their systems. The shape does not change. Only the layer it operates at moves further from the person eventually affected.
What else was happening
The months this archive opens in were not quiet ones. They contained a flaw in a shell interpreter present on an enormous number of web servers, exploitable by planting a command inside an environment variable. They contained private photographs pulled from consumer cloud accounts through guessed credentials and unlimited login attempts. They contained the disclosure of espionage software elaborate enough that its authorship was argued about for years. And they contained a film studio losing its unreleased work, its payroll and its executives' correspondence, with the material published in instalments.
Their combined presence in the record:
- 9the studio breach
- 4espionage malware disclosed in November
- 3the shell interpreter flaw
- 1photographs taken from cloud accounts
One of the decade's most consequential cryptographic disclosures had happened four months before the collection existed and is therefore absent altogether. The others arrived while it was running at two dozen entries a month, and are represented by a handful apiece.
The lesson is the same one the gap two years later teaches, arriving from a different direction. A record running slowly is not a record of a quiet period; it is a thin sample of a normal one. Anybody using these months to judge how eventful the autumn of 2014 was would conclude that almost nothing occurred, and would be describing the collection rather than the season.
Three mechanisms, briefly
Since so little of the autumn survives here, the three flaws behind it are worth describing rather than named, because each illustrates a different failure and none has stopped occurring.
The shell interpreter. A command interpreter reads environment variables when it starts. A defect in how it parsed function definitions meant a variable holding a particular sequence would execute whatever trailed it. Web servers routinely copy visitor-supplied values — the browser's identifying string, the referring page — into environment variables before invoking a script. A request carrying the right header therefore ran commands on the machine receiving it, without authentication, without memory corruption, and against software that had behaved that way for two decades.
The photographs. Private images retrieved from consumer storage accounts, not through any defect in the storage, but by guessing passwords against an interface that neither limited attempts nor alerted the owner. The material was personal in a way that made the harm immediate and reputational rather than financial, and the accounts belonged to people who had never chosen to store anything anywhere — synchronisation was the default and the backup was automatic.
The espionage implant. Software assembled in stages, each stage decrypting the next, resident on telecommunications equipment and research institutions across several countries for years before anybody described it. Its sophistication became the story, which is a recurring distraction: attribution arguments absorb enormous attention and change almost nothing a defender does on a Tuesday.
Guessed passwords, unbounded retries, a parser accepting more than it should, and patience. Nothing in that list has been retired since, and the reason the mechanisms deserve setting down is that they remain the ordinary ones a decade later.
How the till data was taken
Worth setting down, because the mechanism explains why the subject occupied the archive's opening months and why the eventual remedy was a change to the cards rather than to the shops.
A magnetic stripe carries the account number, an expiry and a short verification value. When a customer swipes, the terminal reads that stripe and passes it toward the payment processor. Encryption protects it in transit, but for a fraction of a second — while the till assembles the transaction — the stripe contents sit unencrypted in the memory of an ordinary computer running an ordinary operating system.
Software placed on that computer can watch the memory and copy the contents as they appear. It requires no interception, no defeat of any cryptography and no access to the processor's systems. The technique became known as scraping, and its defining property is that every control protecting the data elsewhere in its journey is irrelevant to it.
Reaching the tills was the other half. In the largest case the route ran through a contractor maintaining heating and ventilation equipment, which held a remote connection for monitoring and billing and had no business being able to reach a payment network at all. The network was flat enough that it could.
The remedy was structural and slow: embedded chips that authorise each transaction with a one-time cryptogram, making a copied stripe worth very little. Deployment took years, and the liability rules that forced it arrived before most terminals could accept it.
What followed was displacement rather than prevention. As counterfeiting a physical card became unprofitable, fraud moved to transactions where no card is presented at all — the web checkout, where a number and an expiry are still sufficient. The category did not shrink. It relocated to the one place the new mechanism does not reach.
Is a small record a worse one?
For the four slow months, no — and the finding is more useful than it first appears. Those 88 entries are the most completely described material in this archive. Every one carries a summary, which is true of no other period in eight years.
The share of entries carrying a summary, year by year:
- 201458%345 entries
- 201521%2,997 entries
- 201628%2,277 entries
- 201738%2,368 entries
- 201866%1,758 entries
- 201977%1,687 entries
- 202050%1,891 entries
- 202152%1,628 entries
2014 sits at 58% overall, which is itself a blend of a hundred and 43. The two years immediately following it, at their much larger volumes, record 21% and 28%.
So a smaller record is not automatically a poorer one, and a larger one is not automatically better. What a reader needs from an archive depends on the question: a census needs breadth and an argument needs depth, and this collection supplies them in different proportions at different times without ever announcing which it is doing.
What can be said about five months?
Less than about any other year here, and the limits are worth stating precisely because this page is the one most likely to be quoted carelessly.
Nothing about the year. 345 entries covering 5 months is not a year and cannot be compared with one.
Nothing seasonal. The months present are consecutive and end in December, so every question about how a subject behaves across a year is unanswerable here.
Proportions inside a single month, carefully. Dec is large enough to support them, and the four slow months are large enough only for the crudest.
And the two structural findings. Neither depends on the count being complete.
What does a beginning tell you about the rest?
Three things, and each of them recurs in every later year of this collection.
That the record is an instrument with settings. Volume and description are not fixed properties of an archive; they are choices somebody made, changed here between November and December without explanation. Every count taken from this collection is therefore partly a measurement of how it was being operated that month, which is the argument the page for 2018 arrives at from the opposite direction after a conclusion had to be withdrawn.
That the calendar shapes the contents. Begin the collection in June and it opens on conference season; begin it in April and it opens on quarterly filings and budget cycles. The starting month decides the first subject and decides nothing about the field.
That the arguments arrive fully formed. Suppliers as an attack path, very large numbers travelling without their methods, warnings about infrastructure published years early — all three are present in the first weeks and none of them is resolved by the last. Seven further years of writing did not settle questions that were already legible in the opening days.
What 2014 settled
That a supplier relationship is an attack path. The card breaches this archive opens on established it with a clarity that took the rest of the industry years to absorb, and every subsequent argument here about third parties descends from them.
That a very large number travels further than the method behind it. The archive's third entry is a credential figure that was disputed within days and repeated for years, which is the pattern this site was built to interrupt.
And, for a reader of records: that beginnings are invented afterwards. Nothing about 2014-08-07 distinguishes it. It became the first day of this archive because it is the earliest one in it, which is a fact about the collection and not about the field it describes.
The eight pages in this series end where they began, on that observation. A record of what was written is a record of what somebody chose to write down, starting when they started and stopping when they stopped — and the most useful thing it can be asked is not what happened, but what looked worth keeping at the time.
Common questions
How many entries does this archive hold for 2014?
345, across 5 months. The record begins on 2014-08-07; there is nothing before that date, and 74% of the year sits in Dec alone.
Where does the record begin?
On 2014-08-07, with no marker of any kind. The first entries are an ordinary day's filing — a piece of research about transport systems, a vendor opening a facility, and a claim about a very large number of stolen credentials.
Why is Dec so much larger than the rest?
Because the collection changes mode. The four months before it hold 88 entries between them; Dec holds 257. Nothing in the entries explains the switch and this archive contains nothing that would.
Are the early months worse recorded?
The opposite, and it is the most useful thing here. 100% of the entries in the slow months carry a written summary — every single one — against 43% in Dec. Running slowly, the record describes everything it takes.
Has that pattern appeared elsewhere?
Once more, in 2016, where three months running at a fifth of the normal rate carry summaries on 82% of entries against 24% for the rest of that year. Two instances is a pattern worth naming and not a law; both show the same trade between how much is captured and how well.
What did the first full month write about?
Mostly the year ahead. Forecasts for the following year account for 33 entries of 257, alongside 32 about retailers and payment cards. The record reaches full volume and immediately starts predicting.
Why retailers?
Because the period this archive opens in was defined by card data taken from shop tills at very large scale. It is the subject the collection begins on, and it is the earliest form of an argument that recurs throughout: the organisation losing the data is rarely the one the customer chose to trust with it.
Is a short year usable?
For proportions inside it, yes. For anything else, no: 345 entries across 5 months cannot be compared with a full year, and 74% of it comes from a single month whose recording mode differs from the rest.
Does this affect the other pages?
Only where 2014 would be used as a baseline, and no page in this series does that. It is excluded from every cross-year comparison because it is neither a complete year nor a consistent one.
Why include it at all?
Because where a record starts is as informative as what it contains. The beginning is unmarked, arbitrary and invisible from inside — the same properties the last year of the archive shows at the other end, and worth seeing at both.
What is the earliest entry about?
Research indicating that aircraft, ships and traffic systems were exposed. It is a fair opening for a collection of this kind: a warning about infrastructure, published years before the incidents that would make the argument concrete.
Can these figures be checked?
Yes. Every count is computed from the archive when the page is built, the months present are named, and every proportion states which months it was calculated over.