Skip to content
The Cyber Security Place

Governance

Security awareness in 2026: even a 4% click rate loses over twelve campaigns

Two decades of budget have gone into persuading staff to be careful — against people who persuade for a living — and the best published results still leave hundreds of people in a thousand opening something they should not have. That is not an argument against teaching people. It is an argument about where you put the wall.

Last reviewed August 25, 2026

A randomised trial across roughly 19,500 employees found annual training had no significant effect on phishing susceptibility, and that the module which fires straight after somebody clicks reduced later clicking by about 2% — while around a third closed it without reading a line. Vendor cohort data reports the opposite: roughly 86% improvement over twelve months, from 33% down to 4%. Both can be honest, and neither rescues the arithmetic. At 4% per campaign across 12 campaigns a year, 387 of 1,000 people still follow a link at least once. Generated lures now draw about 4.5 times the engagement of human-written ones, retiring the advice about spotting mistakes. NIS2 Article 20 makes this a board-level duty, and phishing-resistant sign-in is the control that makes a click survivable.

How much does training actually move the number?

Ask two credible sources and you get answers that differ by more than an order of magnitude, which ought to be a scandal in a field that spends this much money.

On one side sits the largest controlled study anybody has run on the question. A health system with roughly nineteen and a half thousand staff was studied over eight months, with mandatory annual instruction on one axis and monthly simulated lures on the other. The finding on the annual module was flat: no significant relationship between having recently completed it and whether somebody fell for a subsequent message. The finding on embedded remediation — the page that appears the instant you click, widely regarded as the part that works — was a reduction of roughly two percentage points in the relative likelihood of clicking again. About a third of those who landed on that page shut it immediately without engaging.

On the other side sit the cohort numbers published by the platforms themselves, drawn from very large populations: a baseline around a third of staff failing an initial exercise, falling to the low single digits after twelve months of continual drilling. Eighty-six per cent better, on the face of it. Those datasets are far larger than any academic study and it would be lazy to wave them away.

The honest reading is that both are measuring something real and the two things are not the same thing. What follows is an attempt to say which is which — and then to show why, in the specific sense that matters to whoever has to explain an incident, the argument is somewhat beside the point.

The arithmetic that survives any click rate

Set aside whose figure is right and take the most flattering one seriously. Suppose a programme genuinely gets a workforce down to four per cent per campaign. Then ask the question an intruder asks, which is not what fraction of messages fail but how many doors open somewhere across a year.

Worked exampleOne thousand people, one year of attempts

387 of 1,000 people take the bait at least once over the year

The figure vendor cohort data reports after twelve months of continuous drills.

People per thousand who follow a link at least once
Click rate1 a year4 a year12 a year24 a year52 a year
30% — No training3007609861,0001,000
4% — A year of training40151387625880
2% — Best published result2078215384650

Expected values, not a simulation: one minus the chance of getting through every attempt untouched. It assumes attempts are independent and that everyone is equally exposed, which is generous to the training. In practice susceptibility is uneven and attackers aim repeatedly at the same finance and executive inboxes.

The grid is the whole argument. Four per cent sounds like a solved problem until it is applied twelve times, at which point it describes nearly four hundred people in every thousand. Push the cadence to weekly, which is closer to what a large organisation actually receives, and even the best result anybody has published leaves two in three colleagues having followed something at least once.

This is not a rhetorical trick. It is the same reason a component with three nines of availability is unacceptable in a chain of forty: small individual failure probabilities compose badly, and defences built on them inherit the composition. An intruder does not need a majority. Repetition is free for them and expensive for you, which is the asymmetry that no amount of instruction reverses.

There is a further generosity built into those figures. The model assumes everybody is equally likely to be targeted, and nobody works that way. Finance inboxes, executive assistants and anyone whose address appears on a public filing receive attention that is both more frequent and better researched than the average. The real distribution concentrates the exposure exactly where the consequences are worst, which is the case for measuring risk per person rather than per company.

Why do two credible sources disagree by a factor of forty?

Four differences account for most of the gap, and none of them requires anybody to be dishonest.

Who is left in the sample. Cohort figures describe organisations that kept renewing. Those that saw nothing useful in month four stopped buying and stopped contributing data. The surviving population is therefore selected for the conditions under which the product performs, which is a well-understood effect and an unavoidable one for any vendor reporting on its own installed base.

What the exercise looks like by month nine. A simulated lure has to be deliverable, has to survive the filters, and has to be defensible if somebody complains. Those constraints push the difficulty down over time. Staff also learn the tells of the exercise itself — the sending domain, the tracking link format, the day of the month it tends to arrive. Improvement against the exercise is real and does not necessarily transfer to a message written by somebody who wants your payroll.

What else changed in twelve months. Filtering improved, sign-in got harder, banners appeared on external mail. Attributing the whole delta to the human intervention credits it with work done by the infrastructure alongside it.

What the comparison is against. This is the decisive one. A before-and-after measurement compares a population with its own past; a randomised design compares it with people who were not taught, over the same weeks, facing the same messages. The second answers the question a chief information security officer is actually asking, and it is the design that produced the small number.

The reasonable position is therefore neither triumph nor dismissal. Programmes produce a genuine improvement against the exercise, a smaller and less certain one against the real thing, and no improvement whatsoever in the composition arithmetic above. Buying one is defensible. Treating it as the control that keeps intruders out is not.

Attention peaked in 2019; the attacks did not

The human side of security was covered here through the whole of the last decade, and the volume of that coverage has a shape worth looking at.

14201511201635201719420181302019332020162021reports on the human side of security, by year

The rise through 2018 and 2019 tracks a period when every vendor had a culture product and every conference had a track about it. The fall afterwards tracks nothing in the threat landscape at all. Messages remained the most common way in, the proportion of incidents that began with somebody opening one held roughly steady, and the losses attributed to fraudulent payment requests kept climbing.

What changed was the novelty. A subject stops generating reports long before it stops generating incidents, and the gap between those two moments is where complacency lives. The pattern is not unique to this topic; it is what the whole decade of coverage looks like for anything that becomes a checkbox.

The scepticism now supported by controlled evidence was already being voiced inside the industry well before it was: Outdated cybersecurity training erodes trust, hurts more than it helps ran in July 12, 2021, arguing that a badly designed programme costs more in trust than it returns in vigilance.

The language of a security culture arrived earlier still — Practical IT: How to create a culture of cybersecurity at work, October 9, 2015 — and it was a genuine improvement on treating people as a compliance problem, even where it was mostly used to sell the same course under a warmer name.

Is the report rate a better metric?

It is the metric most programmes should be managed by, and the one most of them report second if at all.

Consider two colleagues. One receives a convincing message, senses something is wrong, and deletes it. The other opens it, gets as far as the sign-in page, hesitates, and presses the report button ninety seconds later. Scored on clicks the first is a success and the second a failure. Scored on what happens next, the second is worth considerably more: the response team now has the sending infrastructure, the lure text and the knowledge that a campaign is live against this organisation, in time to pull the rest of the copies out of everybody else's inbox.

That is why the number to watch is elapsed time from first delivery to first report. It is measurable without a vendor, it improves quickly when reporting is made frictionless and consequence-free, and it correlates with how an incident actually goes in a way that a click percentage never has. Programmes that shift their emphasis this way tend to see reporting volumes rise by multiples rather than percentages, which is the sort of movement worth having.

The corollary is uncomfortable for anyone who has built a scoreboard. If you publish a league table of departments by click rate, you have created a strong private incentive to keep quiet after a mistake. Every hour of silence is an hour of undisturbed access. A programme that suppresses reporting has made the organisation less safe while improving its own headline figure, and it will be renewed on the strength of it.

What actually removes the failure mode

Most of what training was aimed at was one specific loss: somebody types a password and a one-time code into a page that is not the page they think it is. That loss no longer requires a human solution, because the protocol can refuse.

A passkey is bound cryptographically to the domain that issued it. Presented with a convincing replica on a lookalike address, the authenticator does not offer the credential — not because it is suspicious, but because the name does not match and there is nothing to offer. The secret never leaves the device and there is no code to relay, which also closes the real-time proxy technique that made earlier multi-factor methods so much weaker than their reputation.

The practical caveat is the one people skip. A deployment that keeps a code-based fallback for the days somebody forgets their device has kept the old failure mode and added a new one, because the fallback is exactly what an attacker will steer the target towards. The security of the arrangement is the security of its weakest enrolled method, which makes recovery design the hard part rather than the rollout.

Alongside that sit the unglamorous mechanics: a report button that takes one press and works on a phone, the ability to remove a message from every mailbox at once, and an out-of-band confirmation step for payment changes that no email can satisfy. None of these ask anybody to be more vigilant. They assume the mistake and make it cheap, which is the only design assumption the arithmetic supports.

Does telling people to look for typos still work?

It stopped working, and the advice is now worse than saying nothing, because it teaches confidence in precisely the wrong signal.

The old heuristics were artefacts of the economics of the trade. Bulk lures were written quickly by people working in a second language with no editor, so awkward phrasing and misplaced apostrophes really were correlated with fraud. Generated text removed that correlation at zero cost. Messages arrive fluent, correctly formatted, matched to the recipient's role, and referencing a supplier relationship that genuinely exists because the details were scraped from a public filing. Measured against human-written equivalents, generated lures draw several times the engagement.

Voice adds a second front. A short sample is enough to reproduce a recognisable voice, and a request that arrives as a call from a familiar-sounding executive defeats a control most organisations never built. The response is procedural rather than perceptual: certain classes of instruction — a change of bank details, an urgent transfer, a request to bypass a step — require confirmation through a channel the requester did not choose, no matter who appears to be asking.

What replaces the old advice is narrower and more honest. Stop asking people to judge authenticity from the surface of a message, which they can no longer do, and start naming the small set of actions that require a second channel regardless of how convincing the request looks. That is a rule people can follow under pressure, which is when it will be needed.

From awareness to human risk management

The label changed in 2024, when Forrester retired the awareness-and-training category from its coverage and replaced it with human risk management. Renaming a market is usually a marketing event, and much of what is sold under the new heading is the old course with a dashboard attached. The underlying shift is nevertheless the right one.

The distinction that matters is what gets counted. The old model counted completions: how many staff sat the module, what they scored, whether the audit evidence existed. The new model counts behaviour and exposure per individual — who reports, how quickly, who holds access that would make a mistake expensive, whose role puts them in front of the messages that matter. Those are different questions, and only the second can be acted on.

Acted on how, in practice? By spending controls where the consequence is concentrated rather than spreading them evenly for fairness. Hardware-backed sign-in for the forty people whose accounts would be worth the most. A mandatory second channel on the finance workflows. Tighter data handling around the teams who touch customer records. Instruction remains part of that, but as one intervention among several rather than the programme itself.

The failure mode of the new category is the same as the old one wearing better clothes: a risk score per employee, published internally, which recreates the league table and with it the incentive to stay quiet. If a number can be used against somebody, it will be, and the reporting you needed will evaporate.

What regulation now requires

For two decades this subject was governed by best practice, which is another way of saying by nothing. That has changed in Europe, and the change is more interesting than another mandatory module.

Article 20 of NIS2 places the obligation on the management body itself: directors must undergo regular training, must approve the risk measures, and are personally accountable for them. Member state audit deadlines fall through the middle of 2026. The significance is not the syllabus, which will be unremarkable, but the seating plan. Every previous scheme aimed at staff and exempted the people who decide the budget, and those are the same people whose approval an invoice-fraud campaign is trying to obtain.

DORA reaches the same place for financial entities by a different route, folding awareness into the operational resilience framework rather than naming it separately. Both share a premise worth stating plainly: this is a governance obligation, not an annual formality that can be delegated downwards and evidenced with a completion report.

Whether that produces better outcomes or merely better records depends on something no directive can legislate. If board sessions become a scheduled hour of slides, nothing follows. If they become the moment a director learns that a message from their own address is trivial to fabricate, and that the finance team has no procedure for refusing it, the hour will have earned its place.

Whose failure is a click?

The phrase about the weakest link has done real damage, because it locates the problem in a person and therefore locates the remedy in persuading that person to be better.

Look at what is actually being asked. A member of staff receives a hundred or more messages a day, most of them legitimate requests to open an attachment, follow a link or sign in somewhere. They are evaluated on responsiveness. The one message that matters is designed to be indistinguishable from the ninety-nine that do not, by somebody with time, tooling and a financial motive. A system that requires a perfect record from every participant, every day, under time pressure, is not a system anybody should have built.

The comparison other fields reached long ago is worth borrowing. Aviation stopped attributing accidents to inattentive individuals and started designing controls that tolerate a lapse, along with reporting schemes that guarantee no consequences for admitting one. The result was not more careful pilots. It was aircraft where a single mistake no longer determines the outcome — and, crucially, a workforce that volunteers its errors because doing so costs them nothing.

Applied here, that means judging a programme by whether it makes people willing to say I think I just did something stupid within a minute of doing it. Every design decision that makes that sentence more expensive to utter has traded an hour of detection time for a tidier report, and it is a bad trade every time.

What a programme worth running looks like

Everything above argues for a smaller, sharper effort rather than for abandoning the idea. Five properties separate the programmes that repay their cost from the ones that generate evidence for auditors.

Short and frequent, not long and annual. Measured effects decay towards baseline within roughly six months, so a forty-minute session every March is a record of attendance rather than an intervention. Two minutes attached to something that just happened does more.

Specific to the role. The finance team needs a payment verification procedure. Developers need to know how a token ends up in a public repository. Directors need to see their own voice reproduced. Generic content is content nobody recognises as being about their own working day.

Reporting is the goal, clicking is the diagnostic. Publish the report rate and the time to first report. Keep click data for planning and never for scoring individuals.

Exercises that resemble the real thing. If the simulation is easier than what arrives on a Tuesday, the improvement it measures is improvement against the simulation. Difficulty should be calibrated against current campaigns.

Consequence-free by design and by statement. Say explicitly that reporting a mistake carries no penalty, then behave that way the first time somebody senior makes one. That single episode determines whether anybody believes the policy.

Where to start on a Monday

Three of these can be finished this week and none of them needs a purchase order.

Time your own report path. Send a harmless message to a handful of colleagues and measure how long it takes to reach whoever would act on it. Most organisations discover the button files into a mailbox nobody watches after six in the evening.

List the accounts worth the most and check how they sign in. Not the whole directory — the few dozen whose compromise would be an incident. If any of them can still authenticate with a code that can be typed into a fake page, that is the finding.

Ask finance what would happen if a supplier emailed new bank details today. Follow the answer to the person who would make the change and ask them the same thing. The gap between the written procedure and the answer you get is the exposure.

After those, the harder work is genuinely prioritisable: retiring code-based fallbacks without stranding anybody, rebuilding exercises to resemble live campaigns, and getting a board session that is about their own accounts rather than about the workforce. An organisation that has done the first three already knows whether a click on Tuesday afternoon becomes an incident or a note in a log, and that is what the entire subject reduces to.

Common questions

Does security awareness training work?

It depends entirely on what you expect from it. A randomised trial across roughly 19,500 employees found no significant relationship between recent completion of annual training and susceptibility to phishing, and measured about a 2% improvement from the training that fires immediately after somebody clicks. Vendor cohort data reports far larger reductions. What nobody disputes is that no achievable rate is low enough to make people the last line of defence.

Why do published figures disagree so widely?

They measure different things on different populations. Cohort figures track organisations that kept paying for a programme, using simulated lures that staff learn the shape of, over a period in which other controls also improved. A randomised trial isolates the intervention and compares against people who did not receive it. Both numbers can be honest and still describe different worlds.

What click rate should we expect?

Untrained populations tend to land around 30% on a well-built lure. Programmes report reaching single digits within a year. The more useful question is what happens across a year of attempts: at 4% per campaign and twelve campaigns, roughly 387 people in a thousand follow a link at least once.

Is the report rate a better metric than the click rate?

Yes, and it is also the one that improves fastest. Somebody who opens a lure and reports it within two minutes has given the response team a head start, which is a better outcome than a colleague who ignored it silently. Time from first delivery to first report is the number that predicts how an incident goes.

Do phishing simulations do harm?

They can, when they are designed as traps and scored as failures. Programmes that punish clicking suppress reporting, which is the opposite of what you want: staff who fear consequences go quiet, and silence is what lets an intrusion mature. Reporting from 2021 was already describing programmes that eroded trust and cost more than they returned.

Should people still be told to look for spelling mistakes?

No. That advice was calibrated for lures written by somebody working in a second language with no editor. Generated messages are fluent, contextual and personalised, and are measured drawing several times the engagement of human-written ones. Advice that teaches people to trust a well-written message is now actively harmful.

What is human risk management?

The category that replaced awareness training in analyst coverage, after Forrester retired the old label in 2024. In practice it means measuring behaviour and exposure per person rather than course completions, and directing controls at the people whose role and access make a mistake expensive.

What does NIS2 require on training?

Article 20 obliges the management body itself to undergo regular cyber security training, alongside offering it to staff, and makes senior management personally accountable for the risk measures adopted. It is the first mainstream regulation that puts the board in the room rather than exempting it.

Does DORA require the same thing?

DORA covers awareness within its risk management framework for financial entities, with less prescription about who sits in the session. The direction is identical: training is treated as part of governance rather than as an annual formality delegated downwards.

What single control removes most of the exposure?

Phishing-resistant authentication. A passkey is bound to the site that issued it and cannot be replayed against a lookalike domain, so a person who follows a link and tries to sign in simply fails to authenticate. It converts a human decision into a protocol outcome.

If authentication is fixed, is training pointless?

No, but its job changes. Once credentials stop being stealable the remaining human-mediated losses are payment fraud, data sent to the wrong place, and approval of fraudulent requests. Those need judgement and a procedure, which is a different curriculum from spotting a suspicious link.

How often should refreshers run?

Field research puts the decay of any measured effect at around six months, so anything annual is mostly a compliance record. Short, frequent and specific beats long, rare and general — and a two-minute prompt tied to something that just happened outperforms a forty-minute module in March.

Security awareness coverage

435 reports on training, culture and the human side of security, newest first.