Skip to content
The Cyber Security Place

How this site checks itself

What every number is computed from, what seven automated gates decline to publish, and the four flaws in the underlying dataset that any reader should know about.

Last reviewed September 4, 2026

Every figure on this site is computed while the page is built, from a corpus of 14,951 entries drawn from 969 publications and growing, collected since 2014. None is typed by hand. Seven checks run against every page and the build halts if any one fails: an uncomputable figure, a page too thin or repetitive to publish, words glued together by a template defect, a promised address that does not exist, a mismatch with the declared sitemap, a plain-text twin disagreeing with its page, or a page nothing links to. As of the last run, 51 of 51 pass. What the checks do not do is decide what is true: they verify provenance and structure, never accuracy. The four known defects in the corpus are published below, and one conclusion has been withdrawn.

The seven checks

Written material enters at the top. It reaches a reader only by passing all seven. A failure anywhere stops the build, and a stopped build publishes nothing — not a degraded page, not a page with a gap where a figure should be.

Each inspects something different and concrete. One counts tokens against a floor and measures their variety; another scans for a blacklist of promotional idioms and for sentences of metronomic, uniform length. A third parses every structured-data block and rejects a malformed schema outright; a fourth walks the breadcrumb trail and the canonical link of each address. Others reconcile the emitted sitemap against the rendered tree, diff the Markdown twin against its formatted parent byte by byte, and trace the internal link graph for an orphaned, unreachable node. The vocabulary is deliberately narrow — exit codes, thresholds, ratios — because a boundary that either opens or shuts leaves no room for a judgement call.

buildrefuses a figure it cannot compute: if the record does not hold the dat…editorialrefuses a page that cannot carry itself: too short, vocabulary that rep…espaciosrefuses two words fused together by a template defect, invisible while …rutasrefuses a promised address that does not exist, including any the menu …sitemaprefuses a mismatch between what the site declares it publishes and what…gemelosrefuses a page whose plain-text version says something different from t…enlacesrefuses a page nothing links to, and any broken internal linkmaterial in ↓↓ published only if all seven pass
  • build — refuses a figure it cannot compute: if the record does not hold the data, the build stops and there is no page.
  • editorial — refuses a page that cannot carry itself: too short, vocabulary that repeats, the stock phrases of generated prose, sentences of suspiciously uniform length.
  • espacios — refuses two words fused together by a template defect, invisible while writing and glaring on the page.
  • rutas — refuses a promised address that does not exist, including any the menu links to.
  • sitemap — refuses a mismatch between what the site declares it publishes and what it published.
  • gemelos — refuses a page whose plain-text version says something different from the formatted one.
  • enlaces — refuses a page nothing links to, and any broken internal link.

What does it mean that a figure is computed?

That nobody typed it. When a page here says a subject appears a certain number of times, that number is produced while the page is being built, by counting the entries that match. It is not a remembered figure, not a figure from a previous version of the page, and not a figure somebody checked once and left in place.

The practical difference shows up when something changes. A typed number stays correct until it silently stops being correct, and nothing announces the moment. A computed one moves when its source moves. If the record grows, every page reflecting it grows too, without anybody editing a sentence.

That is drab plumbing, and it is the entire argument. Most editorial boasts about rigour describe intentions. This one describes a dependency graph, a wiring diagram of obligations.

A figure, traced end to end

Abstractions about dependency graphs are easy to assert. Here is one number, followed from the page a reader sees back to the file it rests on.

A subject page states how many archived items concern its topic. The build loads the record from disk, tests each item's headline and summary against a pattern for the subject, tallies the matches, and hands that tally to the template. Nothing in between stores a value anywhere a person could reach in and adjust.

Break it deliberately and the dependence shows. Rename the record file and the import fails, the build aborts, and yesterday's page stays deployed rather than a fresh one arriving with a stale count. Delete half the items and every page touching that subject moves together in the same run, because none holds a figure of its own to fall behind.

The pattern matching is the weak link and deserves saying so plainly. Deciding that an item concerns ransomware by testing its headline for the word is crude — it misses pieces that describe an incident without naming the category, and it catches passing mentions. That imprecision is a property of the classification, not of the arithmetic, and pages leaning on such tallies name the pattern that produced them.

So the guarantee is narrower than it first sounds, and worth stating at its real size. The number shown is exactly what the stated rule produces over the stated record. Whether that rule captures the concept a reader has in mind is a separate question, and one no build step can settle.

The rule, and the observed minimum

Stating a threshold is mere documentation. Anyone can publish a yardstick nobody wields. What follows is the threshold beside the worst reading actually observed across the 51 pages measured, which is evidence — a sceptic can compare the two columns and conclude for themselves whether the rule bites.

  • at least 3,000 words
    threshold 3,000 · thinnest page 3,000
  • lexical heterogeneity of 0.32 or better
    threshold 0.32 · least varied page 0.320
  • between 30% and 45% of headings phrased as questions
    threshold 30–45% · observed range 31.3–42.9%
  • sentence-length variation of at least 5
    threshold 5 · flattest page 8.2

Across those pages the checks have accumulated 159,665 words of material written for this site. Every one of them was subject to the same list.

One number on this page is necessarily out of date, and it is worth being explicit about which. The checks run after the build, so a page cannot publish the result of a check that has not yet been performed on it. The report shown here was written on the most recent run and describes the run before this one.

The record these figures come from

14,951 entries so far, collected from 969 named publications since August 7, 2014. It is still being added to, so every figure on this site is a figure as at the last build rather than a final one.

An entry holds a headline, a date, an attribution to the publication that produced it, a set of subject tags and, for 6,634 of them, a written summary. It is a record of what the trade press published, which is a different thing from a record of what happened. The distinction runs through every page built on it and is stated on each one.

Three consequences follow and none is a defect exactly. Coverage volume reflects editorial attention rather than incidence, so a subject can be enormous here because it was interesting to write about. Attribution is to the publication rather than to any primary source, so a claim repeated by forty outlets appears forty times without becoming more true.

The third is subtler and shapes every count on the site. An entry credits a domain, not a masthead, and the gap between those widens once anything is tallied. Syndicated material carries whichever domain republished it, so one wire story travelling through six outlets deposits six separate rows under six separate names. Aggregators sit alongside originators with no mark distinguishing them. Nothing collapses duplicates and nothing traces an item upstream to whoever first reported it.

The distortion is systematic rather than random, which makes it predictable and therefore usable. Heavily syndicated subjects swell; material that stayed with one outlet does not. A count of how often something was written about sits closer to a count of how far it travelled, and those two quantities diverge furthest wherever a story was cheap to reproduce and expensive to originate.

A slower decay affects the far side of every credit. Addresses published a decade ago resolve unevenly today: some redirect into a successor title, some land on a parked domain, some return nothing at all. An entry outlives whatever it points at, because the headline, date and credit are held here rather than fetched on demand. What does not outlive it is the ability to check an item against the original, and for the oldest material that avenue has quietly shut.

Four defects, published

Each surfaced by measuring the archive against itself. None was recalled from memory, and none is trivial enough to omit from a page that asks for trust.

  1. 1. Three months of 2016 are thin

    April, May and June hold 741 entries between them against 171 a month across the rest of that year. The shortfall is roughly -229 entries.

  2. 2. The record does not begin gradually

    August to November 2014 hold 88 entries described at 100%; December alone holds 257 at 43%. Collection ran in two modes and nothing here explains the switch.

  3. 3. Under half of the corpus carries a written summary

    6,634 of 14,951 entries — 44% — have text beyond a headline. The rest exist as a title, a date and an attribution.

  4. 4. The record always lags the present

    The most recent archived entry is dated 2021-11-30. Filing runs behind publication, so there is always a stretch of recent weeks the corpus cannot answer for. The gap closes as material is added; it never reaches zero.

Why publish them?

Because a number means little without them, and because a compendium that hides its holes invites precisely the over-reading that breeds bad conclusions.

Take the thin months of 2016. Any subject total for that year is depressed by them, and a reader comparing 2016 against 2017 without knowing would conclude that attention rose when the collecting merely resumed. The defect does not make the totals wrong. It makes one particular comparison wrong, which is worse, because the comparison looks sound.

There is a commercial argument too, and it is not in tension with the editorial one. A buyer evaluating this record will find its defects within an afternoon of serious work. Finding them listed is reassuring. Finding them unlisted, after being told the material is rigorous, ends the conversation and deserves to.

The habit generalises. Anything measured here that cannot support a conclusion gets a sentence saying so, and those sentences are on the pages rather than in a note at the end.

Can a verified page still be wrong?

Yes, and understanding how is the most useful thing on this page.

The gates confirm that a figure matches its source and that a page is structurally capable of carrying an argument. They say nothing about whether the source is complete, whether the sample represents anything, or whether the conclusion drawn from a correct number follows. Those remain editorial judgements, and editorial judgement fails.

It failed here once, in a way worth recording. A page about 2018 explained a fall in publication volume by pointing at a fall in the number of contributing publications. Both figures were computed and both were right. When the surrounding years were measured, the explanation collapsed: the following year recovered its publication count while volume stayed low, which the proposed cause could not account for.

The conclusion was withdrawn. That page now establishes the shape of the fall, rules out the two obvious explanations and names no cause, because the record does not contain what would be needed to name one. The withdrawal is described on the page itself rather than quietly patched, since a correction nobody can see is not a correction.

Automated checks would never have caught that. Only measuring more of the record did.

What happens when a page cannot support a claim?

The assertion is dropped. No gentler route exists, and that is by design rather than restraint.

The failure mode is a broken deployment, which somebody notices within minutes, rather than a plausible sentence, which nobody notices ever.

This shapes what gets written more than any editorial rule, and the questions it rules out tend to be the appealing ones. Did defensive spending rise after a particular breach? Did firms named in coverage fare worse afterwards than firms absent from it? Did a regulation alter behaviour, or alter only what got written about behaviour? Every one of those needs something this collection lacks — budgets, outcomes, or any observation of the world beyond what was printed about it.

So the writing goes where the evidence sits rather than where the curiosity does. That is a genuine cost, and pretending otherwise would be its own kind of overclaiming: the most interesting questions about this field are mostly not answerable from a pile of headlines, however carefully the pile is counted.

The plain-text twins

Every page written for this site publishes a second version in plain Markdown, at the same address with .md appended. This page has one at /method.md.

They exist for readers and systems that do not want a rendered page: screen readers, text browsers, anything ingesting the material for analysis. A check compares the two versions and fails the build when they disagree, which stops the accessible version from drifting quietly behind the formatted one — the ordinary fate of such alternatives.

Both are generated from the same computed figures. Neither is a transcription of the other.

The trailing edge of the record

The newest archived item is dated November 30, 2021. Filing lags publication, as it does in any collection, so a stretch of recent weeks always sits beyond what the material can yet answer for.

That trailing edge is easy to read past, because an archive gives no signal at its brink. The most recent weeks look thin in a way that resembles a subject going quiet, and a subject that has only just emerged looks like one that does not exist. Neither reading is safe near the edge, and the distance moves with every addition.

Where the material is genuinely historical the lag costs nothing, because a completed period is what those pages examine. Elsewhere the width of the band even varies — publishers file at different speeds, regulators sit on notices, holidays slow everything — so a recent count deserves lighter weight than an old one until the edge has moved past it.

Is any of this checkable from outside?

Partly, and the parts that are not are worth naming honestly.

What a reader can check now. Every figure states the record it came from and the period it covers. Every proportion states its denominator. The observed minimums above can be compared against the thresholds beside them. Any page can be read against its plain-text twin.

What requires the data. Reproducing a specific count means having the corpus, which is available to researchers and buyers on request. Nothing about the arithmetic is proprietary; it is mere tallying.

What cannot be checked from outside at all. Whether the collection was even-handed while it was being made. That is a property of decisions taken years ago by people selecting what to file, and no amount of measurement recovers it. The four defects above are the parts of that history the record was able to reveal about itself. Others may exist and would not be visible from here.

Where did each check come from?

None was designed in advance. Each exists because a particular kind of error got through first, which is the ordinary way such things accumulate and is worth admitting.

The one comparing a page against its plain-text twin came first, because generating two versions from one source invites them to diverge the moment somebody edits the wrong one. The orphan check arrived the day an index page was created that nothing pointed at: it existed, it was correct, and no reader could have reached it. The check found it within minutes of being written, which settled its usefulness immediately.

The check for words fused together exists because that defect reached a published page unnoticed. A template that closes a link at the end of a line and continues on the next will silently drop the space between them. On screen it is glaring. While writing it is invisible, because the source has a line break exactly where the space should be.

The editorial check grew last and is the fussiest, because thin writing is harder to detect mechanically than a broken link. Counting words catches padding. Measuring how much of the vocabulary repeats catches a page that reaches its length by restating itself. A list of stock phrases catches the register of automatically generated prose. Measuring the spread of sentence lengths catches text where every sentence runs to the same size, which almost nothing written by a person does.

Each of those rules has produced a false alarm and been narrowed in response. The phrase list originally rejected an ordinary English word used correctly, and it now demands the surrounding construction that marks the promotional sense. A check nobody trusts gets switched off, so keeping them precise is a condition of them surviving.

What this is not

Not a guarantee of accuracy. Seven gates passing means a page is buttressed and structurally sound. It does not mean the conclusion is right.

Not an audit. No external auditor has vetted this. The checks are run by the same build that publishes the pages, which makes them a habit rather than a warranty. Described accurately, they are a system that makes certain mistakes impossible and others merely visible.

Not a claim about the news entries. The archived items carry whatever text their original publisher supplied, attributed to them. The standard described here governs the pages written for this site.

Not finished. Checks get added when a species of error turns out to be possible. The one that catches words glued together by a template exists because that defect reached publication before anybody saw it.

Who can use the record?

Anybody, for reading. Everything published here is open, carries no advertising script and sets nothing in a browser that would follow a visitor elsewhere.

The underlying collection is a different matter, and there is a reason to be organised about it rather than posting a download link. Longitudinal material of this kind gets quoted, and a quotation that arrives without the four defects above attached will eventually be wrong in a way that traces back here. Requests are answered with the data and with its documented limitations together.

Researchers, analysts and anybody testing a claim against the trade coverage on file can write to [email protected] describing what they are trying to establish. Queries the archive cannot answer get told so, which saves everyone the time.

Common questions

What does it mean that a figure is computed?

That no number on this site was typed by a person. Each one is calculated while the page is built, from the 14,951 entries described below. Change the corpus and the number changes; remove the corpus and the page does not build at all.

Can a computed figure still be wrong?

Yes, and this is the limit worth understanding. Computation guarantees that a figure matches its source. It guarantees nothing about whether the source is complete, representative or well collected. Every defect listed on this page produces figures that are correct and still misleading if read without them.

What exactly do the checks refuse?

Seven things, each with its own script: an uncomputable figure, a page too thin or too repetitive to be worth publishing, words glued together by a template defect, a promised address that does not exist, a mismatch between what the site declares and what it published, a plain-text twin that disagrees with its page, and any page nothing links to.

Do the checks decide what is true?

No. They check provenance and structure, never accuracy. A page can pass every one of them and still reach a poor conclusion from good data — which has happened here once, and is documented rather than removed.

How current is the verification shown?

It is from the previous run: the checks execute after the build, so a page cannot publish the result of a check that has not yet been run on it. The report shown was written on the most recent run and covers 51 pages.

Why publish the corpus defects?

Because a reader cannot judge a figure without them, and because a collection that hides its gaps invites exactly the over-reading that produces bad conclusions. Every defect here was found by measuring this corpus against itself, not by anybody remembering it.

What happens when a page cannot support a claim?

The claim comes out. There is no mechanism for publishing something the corpus cannot carry, because the figure would have nothing to compute from and the build would stop.

Has anything been retracted?

Once. A conclusion about why publication volume fell in 2018 was withdrawn when the surrounding years were measured and contradicted it. The page now establishes the shape of the fall, rules out the two obvious explanations and names no cause.

What are the plain-text twins?

Every editorial page publishes a Markdown version at the same address with `.md` appended. A check compares them and fails the build if they disagree, so the accessible version cannot quietly drift from the formatted one.

Does this cover the news entries too?

No. The 14,951 archive entries collected so far are attributed to their original publisher and carry whatever text that source supplied. The checks described here apply to the pages written for this site.

Can any of this be verified from outside?

Partly, and that is deliberate. Every figure states the corpus it came from and every page states what it cannot support, so a reader with the same data can reproduce the arithmetic. The corpus itself is available on request.

Why go to this trouble?

Because summarising the news is now free and consequently worth very little, while a claim somebody can check has become scarce. The checks are the only part of this site that would be difficult for anybody else to copy.

Corrections

Errors are corrected on the page where they appeared, with the correction described rather than applied silently. Write to [email protected] with anything that looks wrong here — particularly a figure, since a wrong figure means either the record or the computation is wrong, and both are worth knowing about.

Read the material this describes: the year-by-year analysis, the subject pages, or what this site is.