Governance
Security awareness in 2026: even a 4% click rate loses over twelve campaigns
Two decades of budget have gone into persuading staff to be careful — against people who persuade for a living — and the best published results still leave hundreds of people in a thousand opening something they should not have. That is not an argument against teaching people. It is an argument about where you put the wall.
Last reviewed August 25, 2026
- No achievable rate makes people the last line. Repetition beats vigilance: a small per-message probability compounds across a year of attempts.
- The report rate is the metric worth managing. Somebody who opens a lure and says so within two minutes is a good outcome, not a failure.
- Punitive programmes buy silence. Staff who expect consequences stop reporting, and quiet is what lets an intrusion mature.
- Passkeys remove the failure mode entirely for credential theft, which is the one category that training was mostly aimed at.
How much does training actually move the number?
Ask two credible sources and you get answers that differ by more than an order of magnitude, which ought to be a scandal in a field that spends this much money.
On one side sits the largest controlled study anybody has run on the question. A health system with roughly nineteen and a half thousand staff was studied over eight months, with mandatory annual instruction on one axis and monthly simulated lures on the other. The finding on the annual module was flat: no significant relationship between having recently completed it and whether somebody fell for a subsequent message. The finding on embedded remediation — the page that appears the instant you click, widely regarded as the part that works — was a reduction of roughly two percentage points in the relative likelihood of clicking again. About a third of those who landed on that page shut it immediately without engaging.
On the other side sit the cohort numbers published by the platforms themselves, drawn from very large populations: a baseline around a third of staff failing an initial exercise, falling to the low single digits after twelve months of continual drilling. Eighty-six per cent better, on the face of it. Those datasets are far larger than any academic study and it would be lazy to wave them away.
The honest reading is that both are measuring something real and the two things are not the same thing. What follows is an attempt to say which is which — and then to show why, in the specific sense that matters to whoever has to explain an incident, the argument is somewhat beside the point.
The arithmetic that survives any click rate
Set aside whose figure is right and take the most flattering one seriously. Suppose a programme genuinely gets a workforce down to four per cent per campaign. Then ask the question an intruder asks, which is not what fraction of messages fail but how many doors open somewhere across a year.
The figure vendor cohort data reports after twelve months of continuous drills.
| Click rate | 1 a year | 4 a year | 12 a year | 24 a year | 52 a year |
|---|---|---|---|---|---|
| 30% — No training | 300 | 760 | 986 | 1,000 | 1,000 |
| 4% — A year of training | 40 | 151 | 387 | 625 | 880 |
| 2% — Best published result | 20 | 78 | 215 | 384 | 650 |
Expected values, not a simulation: one minus the chance of getting through every attempt untouched. It assumes attempts are independent and that everyone is equally exposed, which is generous to the training. In practice susceptibility is uneven and attackers aim repeatedly at the same finance and executive inboxes.
The grid is the whole argument. Four per cent sounds like a solved problem until it is applied twelve times, at which point it describes nearly four hundred people in every thousand. Push the cadence to weekly, which is closer to what a large organisation actually receives, and even the best result anybody has published leaves two in three colleagues having followed something at least once.
This is not a rhetorical trick. It is the same reason a component with three nines of availability is unacceptable in a chain of forty: small individual failure probabilities compose badly, and defences built on them inherit the composition. An intruder does not need a majority. Repetition is free for them and expensive for you, which is the asymmetry that no amount of instruction reverses.
There is a further generosity built into those figures. The model assumes everybody is equally likely to be targeted, and nobody works that way. Finance inboxes, executive assistants and anyone whose address appears on a public filing receive attention that is both more frequent and better researched than the average. The real distribution concentrates the exposure exactly where the consequences are worst, which is the case for measuring risk per person rather than per company.
Why do two credible sources disagree by a factor of forty?
Four differences account for most of the gap, and none of them requires anybody to be dishonest.
Who is left in the sample. Cohort figures describe organisations that kept renewing. Those that saw nothing useful in month four stopped buying and stopped contributing data. The surviving population is therefore selected for the conditions under which the product performs, which is a well-understood effect and an unavoidable one for any vendor reporting on its own installed base.
What the exercise looks like by month nine. A simulated lure has to be deliverable, has to survive the filters, and has to be defensible if somebody complains. Those constraints push the difficulty down over time. Staff also learn the tells of the exercise itself — the sending domain, the tracking link format, the day of the month it tends to arrive. Improvement against the exercise is real and does not necessarily transfer to a message written by somebody who wants your payroll.
What else changed in twelve months. Filtering improved, sign-in got harder, banners appeared on external mail. Attributing the whole delta to the human intervention credits it with work done by the infrastructure alongside it.
What the comparison is against. This is the decisive one. A before-and-after measurement compares a population with its own past; a randomised design compares it with people who were not taught, over the same weeks, facing the same messages. The second answers the question a chief information security officer is actually asking, and it is the design that produced the small number.
The reasonable position is therefore neither triumph nor dismissal. Programmes produce a genuine improvement against the exercise, a smaller and less certain one against the real thing, and no improvement whatsoever in the composition arithmetic above. Buying one is defensible. Treating it as the control that keeps intruders out is not.
Attention peaked in 2019; the attacks did not
The human side of security was covered here through the whole of the last decade, and the volume of that coverage has a shape worth looking at.
The rise through 2018 and 2019 tracks a period when every vendor had a culture product and every conference had a track about it. The fall afterwards tracks nothing in the threat landscape at all. Messages remained the most common way in, the proportion of incidents that began with somebody opening one held roughly steady, and the losses attributed to fraudulent payment requests kept climbing.
What changed was the novelty. A subject stops generating reports long before it stops generating incidents, and the gap between those two moments is where complacency lives. The pattern is not unique to this topic; it is what the whole decade of coverage looks like for anything that becomes a checkbox.
The scepticism now supported by controlled evidence was already being voiced inside the industry well before it was: Outdated cybersecurity training erodes trust, hurts more than it helps ran in July 12, 2021, arguing that a badly designed programme costs more in trust than it returns in vigilance.
The language of a security culture arrived earlier still — Practical IT: How to create a culture of cybersecurity at work, October 9, 2015 — and it was a genuine improvement on treating people as a compliance problem, even where it was mostly used to sell the same course under a warmer name.
Is the report rate a better metric?
It is the metric most programmes should be managed by, and the one most of them report second if at all.
Consider two colleagues. One receives a convincing message, senses something is wrong, and deletes it. The other opens it, gets as far as the sign-in page, hesitates, and presses the report button ninety seconds later. Scored on clicks the first is a success and the second a failure. Scored on what happens next, the second is worth considerably more: the response team now has the sending infrastructure, the lure text and the knowledge that a campaign is live against this organisation, in time to pull the rest of the copies out of everybody else's inbox.
That is why the number to watch is elapsed time from first delivery to first report. It is measurable without a vendor, it improves quickly when reporting is made frictionless and consequence-free, and it correlates with how an incident actually goes in a way that a click percentage never has. Programmes that shift their emphasis this way tend to see reporting volumes rise by multiples rather than percentages, which is the sort of movement worth having.
The corollary is uncomfortable for anyone who has built a scoreboard. If you publish a league table of departments by click rate, you have created a strong private incentive to keep quiet after a mistake. Every hour of silence is an hour of undisturbed access. A programme that suppresses reporting has made the organisation less safe while improving its own headline figure, and it will be renewed on the strength of it.
What actually removes the failure mode
Most of what training was aimed at was one specific loss: somebody types a password and a one-time code into a page that is not the page they think it is. That loss no longer requires a human solution, because the protocol can refuse.
A passkey is bound cryptographically to the domain that issued it. Presented with a convincing replica on a lookalike address, the authenticator does not offer the credential — not because it is suspicious, but because the name does not match and there is nothing to offer. The secret never leaves the device and there is no code to relay, which also closes the real-time proxy technique that made earlier multi-factor methods so much weaker than their reputation.
The practical caveat is the one people skip. A deployment that keeps a code-based fallback for the days somebody forgets their device has kept the old failure mode and added a new one, because the fallback is exactly what an attacker will steer the target towards. The security of the arrangement is the security of its weakest enrolled method, which makes recovery design the hard part rather than the rollout.
Alongside that sit the unglamorous mechanics: a report button that takes one press and works on a phone, the ability to remove a message from every mailbox at once, and an out-of-band confirmation step for payment changes that no email can satisfy. None of these ask anybody to be more vigilant. They assume the mistake and make it cheap, which is the only design assumption the arithmetic supports.
Does telling people to look for typos still work?
It stopped working, and the advice is now worse than saying nothing, because it teaches confidence in precisely the wrong signal.
The old heuristics were artefacts of the economics of the trade. Bulk lures were written quickly by people working in a second language with no editor, so awkward phrasing and misplaced apostrophes really were correlated with fraud. Generated text removed that correlation at zero cost. Messages arrive fluent, correctly formatted, matched to the recipient's role, and referencing a supplier relationship that genuinely exists because the details were scraped from a public filing. Measured against human-written equivalents, generated lures draw several times the engagement.
Voice adds a second front. A short sample is enough to reproduce a recognisable voice, and a request that arrives as a call from a familiar-sounding executive defeats a control most organisations never built. The response is procedural rather than perceptual: certain classes of instruction — a change of bank details, an urgent transfer, a request to bypass a step — require confirmation through a channel the requester did not choose, no matter who appears to be asking.
What replaces the old advice is narrower and more honest. Stop asking people to judge authenticity from the surface of a message, which they can no longer do, and start naming the small set of actions that require a second channel regardless of how convincing the request looks. That is a rule people can follow under pressure, which is when it will be needed.
From awareness to human risk management
The label changed in 2024, when Forrester retired the awareness-and-training category from its coverage and replaced it with human risk management. Renaming a market is usually a marketing event, and much of what is sold under the new heading is the old course with a dashboard attached. The underlying shift is nevertheless the right one.
The distinction that matters is what gets counted. The old model counted completions: how many staff sat the module, what they scored, whether the audit evidence existed. The new model counts behaviour and exposure per individual — who reports, how quickly, who holds access that would make a mistake expensive, whose role puts them in front of the messages that matter. Those are different questions, and only the second can be acted on.
Acted on how, in practice? By spending controls where the consequence is concentrated rather than spreading them evenly for fairness. Hardware-backed sign-in for the forty people whose accounts would be worth the most. A mandatory second channel on the finance workflows. Tighter data handling around the teams who touch customer records. Instruction remains part of that, but as one intervention among several rather than the programme itself.
The failure mode of the new category is the same as the old one wearing better clothes: a risk score per employee, published internally, which recreates the league table and with it the incentive to stay quiet. If a number can be used against somebody, it will be, and the reporting you needed will evaporate.
What regulation now requires
For two decades this subject was governed by best practice, which is another way of saying by nothing. That has changed in Europe, and the change is more interesting than another mandatory module.
Article 20 of NIS2 places the obligation on the management body itself: directors must undergo regular training, must approve the risk measures, and are personally accountable for them. Member state audit deadlines fall through the middle of 2026. The significance is not the syllabus, which will be unremarkable, but the seating plan. Every previous scheme aimed at staff and exempted the people who decide the budget, and those are the same people whose approval an invoice-fraud campaign is trying to obtain.
DORA reaches the same place for financial entities by a different route, folding awareness into the operational resilience framework rather than naming it separately. Both share a premise worth stating plainly: this is a governance obligation, not an annual formality that can be delegated downwards and evidenced with a completion report.
Whether that produces better outcomes or merely better records depends on something no directive can legislate. If board sessions become a scheduled hour of slides, nothing follows. If they become the moment a director learns that a message from their own address is trivial to fabricate, and that the finance team has no procedure for refusing it, the hour will have earned its place.
Whose failure is a click?
The phrase about the weakest link has done real damage, because it locates the problem in a person and therefore locates the remedy in persuading that person to be better.
Look at what is actually being asked. A member of staff receives a hundred or more messages a day, most of them legitimate requests to open an attachment, follow a link or sign in somewhere. They are evaluated on responsiveness. The one message that matters is designed to be indistinguishable from the ninety-nine that do not, by somebody with time, tooling and a financial motive. A system that requires a perfect record from every participant, every day, under time pressure, is not a system anybody should have built.
The comparison other fields reached long ago is worth borrowing. Aviation stopped attributing accidents to inattentive individuals and started designing controls that tolerate a lapse, along with reporting schemes that guarantee no consequences for admitting one. The result was not more careful pilots. It was aircraft where a single mistake no longer determines the outcome — and, crucially, a workforce that volunteers its errors because doing so costs them nothing.
Applied here, that means judging a programme by whether it makes people willing to say I think I just did something stupid within a minute of doing it. Every design decision that makes that sentence more expensive to utter has traded an hour of detection time for a tidier report, and it is a bad trade every time.
What a programme worth running looks like
Everything above argues for a smaller, sharper effort rather than for abandoning the idea. Five properties separate the programmes that repay their cost from the ones that generate evidence for auditors.
Short and frequent, not long and annual. Measured effects decay towards baseline within roughly six months, so a forty-minute session every March is a record of attendance rather than an intervention. Two minutes attached to something that just happened does more.
Specific to the role. The finance team needs a payment verification procedure. Developers need to know how a token ends up in a public repository. Directors need to see their own voice reproduced. Generic content is content nobody recognises as being about their own working day.
Reporting is the goal, clicking is the diagnostic. Publish the report rate and the time to first report. Keep click data for planning and never for scoring individuals.
Exercises that resemble the real thing. If the simulation is easier than what arrives on a Tuesday, the improvement it measures is improvement against the simulation. Difficulty should be calibrated against current campaigns.
Consequence-free by design and by statement. Say explicitly that reporting a mistake carries no penalty, then behave that way the first time somebody senior makes one. That single episode determines whether anybody believes the policy.
Where to start on a Monday
Three of these can be finished this week and none of them needs a purchase order.
Time your own report path. Send a harmless message to a handful of colleagues and measure how long it takes to reach whoever would act on it. Most organisations discover the button files into a mailbox nobody watches after six in the evening.
List the accounts worth the most and check how they sign in. Not the whole directory — the few dozen whose compromise would be an incident. If any of them can still authenticate with a code that can be typed into a fake page, that is the finding.
Ask finance what would happen if a supplier emailed new bank details today. Follow the answer to the person who would make the change and ask them the same thing. The gap between the written procedure and the answer you get is the exposure.
After those, the harder work is genuinely prioritisable: retiring code-based fallbacks without stranding anybody, rebuilding exercises to resemble live campaigns, and getting a board session that is about their own accounts rather than about the workforce. An organisation that has done the first three already knows whether a click on Tuesday afternoon becomes an incident or a note in a log, and that is what the entire subject reduces to.
Common questions
Does security awareness training work?
It depends entirely on what you expect from it. A randomised trial across roughly 19,500 employees found no significant relationship between recent completion of annual training and susceptibility to phishing, and measured about a 2% improvement from the training that fires immediately after somebody clicks. Vendor cohort data reports far larger reductions. What nobody disputes is that no achievable rate is low enough to make people the last line of defence.
Why do published figures disagree so widely?
They measure different things on different populations. Cohort figures track organisations that kept paying for a programme, using simulated lures that staff learn the shape of, over a period in which other controls also improved. A randomised trial isolates the intervention and compares against people who did not receive it. Both numbers can be honest and still describe different worlds.
What click rate should we expect?
Untrained populations tend to land around 30% on a well-built lure. Programmes report reaching single digits within a year. The more useful question is what happens across a year of attempts: at 4% per campaign and twelve campaigns, roughly 387 people in a thousand follow a link at least once.
Is the report rate a better metric than the click rate?
Yes, and it is also the one that improves fastest. Somebody who opens a lure and reports it within two minutes has given the response team a head start, which is a better outcome than a colleague who ignored it silently. Time from first delivery to first report is the number that predicts how an incident goes.
Do phishing simulations do harm?
They can, when they are designed as traps and scored as failures. Programmes that punish clicking suppress reporting, which is the opposite of what you want: staff who fear consequences go quiet, and silence is what lets an intrusion mature. Reporting from 2021 was already describing programmes that eroded trust and cost more than they returned.
Should people still be told to look for spelling mistakes?
No. That advice was calibrated for lures written by somebody working in a second language with no editor. Generated messages are fluent, contextual and personalised, and are measured drawing several times the engagement of human-written ones. Advice that teaches people to trust a well-written message is now actively harmful.
What is human risk management?
The category that replaced awareness training in analyst coverage, after Forrester retired the old label in 2024. In practice it means measuring behaviour and exposure per person rather than course completions, and directing controls at the people whose role and access make a mistake expensive.
What does NIS2 require on training?
Article 20 obliges the management body itself to undergo regular cyber security training, alongside offering it to staff, and makes senior management personally accountable for the risk measures adopted. It is the first mainstream regulation that puts the board in the room rather than exempting it.
Does DORA require the same thing?
DORA covers awareness within its risk management framework for financial entities, with less prescription about who sits in the session. The direction is identical: training is treated as part of governance rather than as an annual formality delegated downwards.
What single control removes most of the exposure?
Phishing-resistant authentication. A passkey is bound to the site that issued it and cannot be replayed against a lookalike domain, so a person who follows a link and tries to sign in simply fails to authenticate. It converts a human decision into a protocol outcome.
If authentication is fixed, is training pointless?
No, but its job changes. Once credentials stop being stealable the remaining human-mediated losses are payment fraud, data sent to the wrong place, and approval of fraudulent requests. Those need judgement and a procedure, which is a different curriculum from spotting a suspicious link.
How often should refreshers run?
Field research puts the decay of any measured effect at around six months, so anything annual is mostly a compliance record. Short, frequent and specific beats long, rare and general — and a two-minute prompt tied to something that just happened outperforms a forty-minute module in March.
Security awareness coverage
435 reports on training, culture and the human side of security, newest first.
- Navigating the cybersecurity implications of remote work
November 1, 2021 · crainscleveland.com
- Applying The Power Of Deep Learning To Cybersecurity
October 14, 2021 · forbes.com
- Cybersecurity Is A Journey, Not A Destination
October 12, 2021 · forbes.com
- Cybersecurity Awareness Month: Time for your safety check
October 8, 2021 · cnet.com
- Cybersecurity tips from a reformed hacker
October 8, 2021 · abc7ny.com
- Cybersecurity best practices lagging, despite people being aware of the risks
October 7, 2021 · helpnetsecurity.com
- 3 ways any company can guard against insider threats this October
September 27, 2021 · helpnetsecurity.com
- Data Backup – More Important Than Ever
August 26, 2021 · which-50.com
- Users Can Be Just As Dangerous As Hackers
August 10, 2021 · thehackernews.com
- 48 million malware messages: Proofpoint reveals the reality of today's threat landscape
August 5, 2021 · securitybrief.asia
- Secure your IoT: Why smart attack and insider threat detection is key
July 20, 2021 · techbeacon.com
- Outdated cybersecurity training erodes trust, hurts more than it helps
July 12, 2021 · securitymagazine.com
- Combating Risk Negligence Using Cybersecurity Culture
March 11, 2021 · tripwire.com
- DTX Tech Predictions Mini Summit: How to Build a Strong Cybersecurity Culture
February 18, 2021 · infosecurity-magazine.com
- Phishing email attacks targeting remote workers on the rise
January 25, 2021 · securitybrief.eu
- Cybersecurity strategies must involve every part of the organisation - study
January 7, 2021 · securitybrief.eu
- How to Mitigate the Risk of Social Engineering and BEC Attacks
December 22, 2020 · channelfutures.com
- Insider Threats: A Byproduct of the New Normal
December 16, 2020 · cisomag.eccouncil.org
- Moving from Human Error to Human Firewall
December 15, 2020 · cisomag.eccouncil.org
- Don’t get hooked by GDPR compliance phishing scams
December 8, 2020 · itproportal.com
- Every employee has a cybersecurity blind spot
November 9, 2020 · helpnetsecurity.com
- Cyber risk literacy should be part of every defensive strategy
October 27, 2020 · helpnetsecurity.com
- Cybersecurity: Security awareness to boost trade in digital money
October 26, 2020 · vanguardngr.com
- Insider threat report reveals deception in the workforce
October 21, 2020 · securitybrief.eu
- Experiencing ransomware significantly impacts cybersecurity approach
October 16, 2020
- Malicious Apps Pose as Contact Tracing to Infect Android Devices
June 12, 2020 · infosecurity-magazine.com
- Dtex, a specialist in insider threat cybersecurity, raises $17.5M
May 8, 2020 · techcrunch.com
- COVID-19, Cyber Security and the “New Normal”
May 6, 2020
- Most UK Remote Workers Haven’t Had Cybersecurity Training in Past Year
April 27, 2020
- CyberNite is bringing Cyber Security Awareness to those working from home
April 22, 2020 · wyomingnewsnow.tv