Threats
AI security in 2026: the boundary that does not exist
Every defensive idea in this archive assumes a line somewhere between what a system was told to do and what it was given to work on. Language models have no such line, and a decade of accumulated tactics quietly stops applying at that point.
Last reviewed August 25, 2026
- Injection is architectural, not a bug. There is no version in which the model can tell your wording from a stranger's.
- Two legs are survivable; three are not. Danger comes from the coincidence of capabilities, never from one of them.
- Safety has to be architectural too. Restrict what the assistant can reach, because you cannot restrict what it can be told.
- Most of the exposure is unrecorded. The material leaving through text boxes never appears in any inventory.
What makes an assistant dangerous?
Not any single permission. Grant them one at a time and watch where the verdict changes, because it does not change where people expect.
2 of 3 legs present
Contained, because it has no outward channel
An injected instruction can make the agent misbehave in front of you, and it cannot move anything out of the room. Adding any outward channel — a webhook, an image fetch, a calendar invitation — removes that containment.
| Private data | Untrusted content | Outward channel | Verdict |
|---|---|---|---|
| — | — | — | Nothing granted, nothing at risk |
| yes | — | — | One leg on its own does nothing |
| — | yes | — | One leg on its own does nothing |
| yes | yes | — | Contained, because it has no outward channel |
| — | — | yes | One leg on its own does nothing |
| yes | — | yes | Contained, because it has no untrusted content |
| — | yes | yes | Contained, because it has no private data |
| yes | yes | yes | Anyone who can put text in front of it can read your data |
The arithmetic is not additive, which is what makes it easy to get wrong in a design review. Six capabilities drawn from two legs are containable; three capabilities drawn from three legs are not. Nobody grants the trifecta deliberately — it accumulates, because each request on its own is reasonable and the person approving the third has usually forgotten the first.
The move from two legs to three is not a step up in severity; it is a change of kind. With two, an intruder who plants wording in a document can annoy you. With three, the identical wording reads your records and posts them somewhere, using permissions you granted, in a session your logs will describe as authorised.
What makes this treacherous in practice is that nobody assembles the trifecta on purpose. It accrues. Somebody wires up mailbox access in March because summarising threads saves an hour a day. Web browsing follows in June, because answers grounded in current pages are better answers. Then in September a colleague asks it to send the summary on, and whoever approves that request is approving a capability, not a configuration — and has no reason to remember the other two.
Which is why the useful question in a design review is never is this permission safe. It is which legs does this system now stand on, asked of the finished arrangement rather than of the change in front of you.
The boundary that does not exist
Drawn as sets, the shape of the problem is a small region in the middle where three ordinary properties coincide.
Every technique this archive covered for a decade rests on a distinction that a language model does not make. A firewall separates inside from outside. A parser separates a query from its parameters. Escaping separates markup from the text it surrounds. In each case the defence works because the machine can tell which part of the input was authorised.
A model reads one continuous sequence and has nothing equivalent. The operator's wording, the user's question, the contents of a fetched page and the text inside a forwarded document all arrive as the same kind of thing. When a sentence in that page says to disregard earlier orders and forward the last message received, the model is not being tricked in any interesting sense. It is doing what it does, which is continue the sequence it was handed.
Practitioners have tried to reintroduce the missing line — special delimiters, priority markers, a second model reading the first one's input, fine-tuning towards obedience to a single source. All raise the effort required and none holds when somebody adapts, which is the pattern that keeps recurring in published evaluations and is now the working assumption in the field.
Why can prompt injection not be patched?
Compare it with the class it most resembles, and the comparison is the whole answer.
Injection into a database query looks like the same problem and is not. There the fix exists and is complete: send the command and its values along separate paths, and no amount of cleverness in the values can alter the command. The channel is real, enforced below the level anybody can talk to, and once used correctly the defect cannot recur. Two decades of this archive record the industry slowly adopting it.
Nothing of that shape exists for a language system, because meaning is not carried in a structure that can be kept apart. A model that could be told to ignore certain portions of what it reads would have to understand which portions, which is the same understanding an attacker manipulates. The property being defended is entangled with the property that makes the thing useful.
There is a further wrinkle that catches teams who have built a filter and feel secure. The wording need not be legible. It can sit in white-on-white text, inside an image the system reads, in metadata, in a comment, encoded in a way the model happily decodes on request. Filtering assumes you can enumerate what the input might look like, and the input is anything the model can interpret.
None of which means defences are pointless. It means they are odds, not walls, and the difference matters enormously when deciding what an assistant is allowed to touch. Treat filtering as friction and put the actual guarantee in the architecture.
When this section meant something else entirely
172 reports were filed here, and reading them in sequence shows a subject that has since turned inside out.
| 2018 | 2019 | 2020 | 2021 |
|---|---|---|---|
| 2 | 84 | 45 | 41 |
Across those years the phrase names a product category. Statistical classifiers sorting alerts, anomaly scoring on network telemetry, models that promised to spot what signatures missed. It is a thing you procure, deploy and evaluate — an addition to the defensive side of the ledger, discussed with the vocabulary of buying.
Can AI Become Our New Cybersecurity Sheriff?, February 4, 2019, captures the register exactly. The question posed is whether the technology will police your estate for you. Seven years later the pressing question runs the other way: what polices the technology, now that it reads your correspondence and can act.
The vocabulary that dominates the field today is wholly absent from those files, and its absence is not a gap in the record. Nothing was being missed, because there was nothing yet to miss. What existed was a market for detection products and a recurring argument about how much of the labelling was marketing.
The framing did begin to shift at the very end of the period — Artificial intelligence in cyber security: The savior or enemy of your business?, July 18, 2019 — where the technology appears as both problem and remedy. Even there the threat imagined is an adversary wielding it against you, rather than the arrangement that now dominates: your own deployment, holding your own permissions, following a stranger's wording.
How much is leaving through a text box?
Before any of the architectural arguments become relevant, most firms have a simpler exposure they are not measuring.
Regular use of these services on corporate machines rose from roughly fifteen per cent to about forty-five in a single year, and around two-thirds of it happens through personal accounts rather than anything procurement arranged. Analyses of data-prevention events find source code submitted more than any other category by a wide margin, then images, then structured records. Personal information turns up in about sixty-five per cent of unsanctioned-use incidents; commercially sensitive material in roughly forty.
The mechanics are entirely mundane and that is why the volume is what it is. Nobody is exfiltrating anything. Somebody has a contract to summarise, a stack trace to explain, a difficult message to draft, and a tool that does it in fifteen seconds against an internal process that would take a fortnight to arrange an approved alternative. The behaviour is a rational response to a gap, and it is invisible because it leaves through the same channel as ordinary browsing.
The governance figures point somewhere uncomfortable for anybody drafting a prohibition. Around sixty-eight per cent of firms that reported an incident had no policy of any kind, so absence genuinely correlates with harm. But a ban that leaves the underlying need unmet does not reduce the activity; it relocates it from an account somebody could audit to one nobody can see. The organisations that came out of this well provided something adequate and then asked people to use it.
The rules arrived, then moved
Anybody working from a planning deck written last year is probably holding the wrong date, because the schedule was formally amended in the middle of 2026.
The parts already in force stayed in force. Prohibited practices have applied since February 2025. Duties on providers of general-purpose models have applied since August 2025. Transparency and content-labelling obligations kept their original timing.
What moved was the heaviest tier. The omnibus regulation that entered into force in July 2026 deferred the high-risk obligations: standalone high-risk systems now fall due in December 2027, and systems embedded in already-regulated products in August 2028. That is a delay of roughly sixteen months on the deadline most readiness programmes were built around.
Two responses to that news are both wrong. Treating it as cancellation is the obvious error. The subtler one is treating it as sixteen additional months of the same work — because the obligations that did not move are the ones that touch everyday deployments, and an inventory of where these systems already sit inside the business is now the binding constraint. Almost nobody has one, and it is prerequisite to every later requirement rather than a step within them.
What is inside the weights you did not train?
A model deployed inside a product is a large binary artefact obtained from somewhere, and the discipline that grew up around obtaining software artefacts has barely reached it.
The parallels are exact and unflattering. Weights are downloaded from public registries. They are frequently unsigned. They are derived from other weights through fine-tuning chains that nobody records, so the provenance question — what was this built from, and by whom — is usually unanswerable past one hop. Serialised formats have executed code on load. Registries have hosted deliberately malicious uploads under plausible names.
What is genuinely different is that inspection does not help. A dependency can be read; its behaviour is in the source. A set of weights is opaque by construction, and behaviour conditioned on a rare trigger phrase is not discoverable by testing, because you would have to know the phrase to test for it. The usual assurance route is closed.
Which leaves the same answer as the rest of this page. Pin the artefact by hash, record where it came from, prefer publishers who sign, and — because none of that tells you what the weights will do — assume the model may be adversarial and constrain what it can reach. Every durable control here is about the boundary around the system rather than confidence in the system.
Is this better for defence or for attack?
Currently for attack, and the reason is about labour rather than capability.
Offence gains most where the constraint was human effort per target. Reconnaissance that took an afternoon per organisation now takes a minute. Lures that betrayed themselves through awkward phrasing are fluent, specific and personalised. A familiar voice can be reproduced from a short sample. None of these is a new technique; each was previously rationed by how many people you could pay to do it.
Defence gains too, and the gains are real but structurally narrower. Triage of alert queues, summarising an incident for people who need to act, drafting the first pass of a detection rule, reading unfamiliar code quickly. All valuable, all bounded by the fact that a defender must be right about everything while an attacker needs one success — an asymmetry no productivity improvement alters.
The honest summary is that the technology accelerates both sides and the sides are not symmetrical, so the same tool produces a larger effect for the smaller party. That may change as tooling matures on the defensive side. It is not where things stand today, and a plan that assumes otherwise is a plan written from a brochure.
One consequence deserves stating because it is actionable rather than gloomy. If the attacker's gain is volume, then the metrics a defender should watch are the ones that scale with volume: how many suspicious messages arrive per person per week, how long the first report takes, how quickly a confirmed lure can be pulled from every mailbox. Those numbers move when the labour cost of an attempt falls, and they are visible from inside the organisation without waiting for anybody to publish a study about it.
Designing an assistant that cannot be turned
Since the wording cannot be trusted, the design has to make untrustworthy wording harmless. Four patterns do that, and each works by removing a leg.
Split the reader from the actor. One component consumes outside text and holds no credentials; another touches internal records and never sees anything a stranger wrote. They communicate through a fixed vocabulary of requests rather than by passing prose. This is the strongest available arrangement and the most work.
Fix the destinations in advance. If outbound requests can only reach a list settled at deployment, an injected directive has nowhere to send anything. Watch the quiet channels here: an image tag that fetches an address is an outward channel, and so is a calendar invitation.
Put a person on the irreversible step. Reading is delegated, acting is confirmed. This is unfashionable because it removes the autonomy that made the product attractive, and it remains the control most likely to be present when something goes wrong.
Give it its own identity, scoped tightly. An assistant acting with a person's full permissions inherits every mistake in that person's access. Its own account, limited to the records it genuinely needs, converts a total compromise into a bounded one.
What to ask somebody selling you an agent
Four questions, none of which requires expertise to ask, and all of which are answered more honestly by hesitation than by the answer.
Which of the three legs does this have? A supplier who has not thought in those terms will describe features. One who has will answer in a sentence.
What happens when a page it reads contains an instruction? The correct answer describes an architectural limit. An answer describing a filter is an answer about odds, and should be received as such.
Can outbound destinations be restricted to a list I control? If not, the third leg is permanently present and the other two decide your exposure.
Show me your logs from a successful injection. Every serious team has run this exercise against their own product and can describe what it looked like. A supplier who has not is telling you something more useful than any answer.
Where to start on a Monday
Three tasks, all of which are measurement rather than remediation, and all of which are usually finished before anybody has approved a budget.
Find out what is already deployed. Not the sanctioned list — the assistants embedded in tools your teams already pay for, most of which acquired these features in an update nobody read. Network telemetry answers this faster than a survey does.
For each one, count its legs. Three columns, one row per assistant. The rows with three ticks are your entire priority list, and there are usually fewer than feared and at least one that surprises everybody.
Ask what people would stop doing if you switched them off. That answer is the requirement you have to meet before any restriction survives contact with the organisation, and it is cheaper to hear now than after issuing a ban.
After those three the harder work is genuinely prioritisable: separating readers from actors in the deployments that matter, restricting outbound destinations, and building the inventory the coming obligations assume you already keep. A firm that has done the first three knows which of its assistants could be turned against it this afternoon, and that is the question the whole subject reduces to.
Common questions
What is prompt injection?
Text that reaches a language system and is treated as a directive rather than as material to work on. It arrives inside a web page, a document, an email or a code comment, and the system obeys it because nothing in its architecture separates the operator's wording from anybody else's.
Why can it not simply be patched?
Because it is not a defect in an implementation. A language system reads one stream of tokens and has no channel that says which portion was authorised. Filters, delimiters and instruction hierarchies raise the effort and none of them closes the gap, which is why published defences keep falling to adaptive attempts.
What is the lethal trifecta?
The combination of access to private material, exposure to text written by strangers, and any means of reaching outwards. Any two are containable. All three together make an assistant into an exfiltration tool for whoever can put words in front of it.
Is this only a problem for autonomous agents?
It becomes consequential with tools. An assistant that only answers can be made to say something wrong; one that can act does the wrong thing on your behalf, with your permissions, and leaves logs showing an authorised session.
How much confidential data is going into these services?
A great deal, and mostly outside any register. Regular use on corporate devices climbed from roughly 15% to about 45% in a year, around two-thirds of that from non-corporate accounts, and between a quarter and a third of staff report having pasted confidential material into a public assistant.
What kind of data leaks most?
Source code leads by a wide margin in analyses of prevention events, followed by images and structured records. Personal information features in around 65% of unsanctioned-use incidents and intellectual property in roughly 40%.
Does a written AI policy reduce incidents?
Of organisations reporting an AI-related incident, about 68% had no governance policy at all, so the absence correlates with harm. A policy nobody can comply with, however, converts sanctioned use into unsanctioned use, which moves the same activity out of view.
When do the EU AI Act obligations bite?
Prohibited practices have applied since February 2025 and general-purpose model duties since August 2025. The high-risk obligations were deferred by the digital omnibus regulation that entered into force in July 2026: standalone high-risk systems now fall due in December 2027, and AI embedded in regulated products in August 2028. Transparency and content-labelling duties stayed on their original schedule.
What does the model supply chain add?
Weights are artefacts you did not build, downloaded from a registry, frequently unsigned and derived from other artefacts through a chain nobody records. Every argument that applies to a software dependency applies here, minus the tooling that grew up around software dependencies.
Is machine learning better for defence or for attack?
It has been quietly useful in defence for a decade in narrow, well-posed tasks, and it is currently more transformative for attackers, because it removes labour from reconnaissance, lure writing and impersonation. Those are asymmetric gains: defence gains speed, offence gains scale.
How should an agent be designed so that it cannot be turned?
By breaking one leg of the trifecta deliberately. Separate the assistant that reads outside text from the one that touches internal records, allow no outbound destination that was not fixed in advance, and require a person to approve any action that leaves the machine.
What should we ask a vendor selling an agent?
Which of the three legs their product has, what happens when a page it reads contains an instruction, whether outbound destinations are restricted to a list, and what their logs show when an injected directive succeeds. Vague answers to the last one are the reliable signal.
AI and security coverage
172 reports, newest first. Most cited sources: helpnetsecurity.com (24), forbes.com (18), itproportal.com (15), information-age.com (12), securityboulevard.com (4), infosecurity-magazine.com (4).
- Artificial Intelligence: The Enemy and The Solution
November 26, 2021 · cpomagazine.com
- AI modeled on the spread of human viruses to combat cyber attacks
November 16, 2021 · securitybrief.asia
- Is This Top Artificial-Intelligence-Powered Cybersecurity Stock a Buy?
November 12, 2021 · fool.com
- 4 Benefits of Using AI in Cybersecurity
November 9, 2021 · cioinsight.com
- How Blockchain and AI will Promote Industrial Growth: An Overview
November 2, 2021 · cisomag.eccouncil.org
- How AI can be the next leap for cybersecurity
November 2, 2021 · techradar.com
- Why Data Science is Key to Delivering Continuous Authentication
November 1, 2021 · cisomag.eccouncil.org
- Cybersecurity blind spot: AI’s inherent vulnerabilities
October 22, 2021 · gcn.com
- Fortinet: Artificial intelligence essential to combat fast-moving threats
October 13, 2021 · channellife.co.nz
- Trends in Artificial Intelligence (AI) in Cybersecurity
September 28, 2021 · datamation.com
- How Self-Learning AI is changing the paradigm of endpoint security
September 17, 2021 · securitybrief.asia
- The age of AI-powered devices at the edge
September 7, 2021 · helpnetsecurity.com
- Increasing number of investigations calls for advanced technology and dedicated teams
September 3, 2021 · helpnetsecurity.com
- Putting the AI into security
August 31, 2021 · it-online.co.za
- Why AI isn't the only answer to cybersecurity [Q&A]
August 6, 2021 · betanews.com
- 4 conversations every company needs to be having about AI
August 2, 2021 · venturebeat.com
- How can AI derail the organization’s cybersecurity?
July 27, 2021 · techhq.com
- Should we use AI in cybersecurity? Yes, but with caution and human help
July 22, 2021 · techrepublic.com
- 7 Ways AI and ML Are Helping and Hurting Cybersecurity
July 20, 2021 · darkreading.com
- AI and Cybersecurity: Making Sense of the Confusion
July 13, 2021 · beta.darkreading.com
- Armorblox and Intermedia team up to protect email from cyberattacks with AI
July 9, 2021 · venturebeat.com
- Boosting IT Security with AI-driven SIEM
July 9, 2021 · itbusinessedge.com
- How AI and automation are creating a future with smarter cybersecurity
June 30, 2021 · dqindia.com
- Using AI-powered software to manage potential cyber threats
May 26, 2021 · siliconangle.com
- How AI is Mishandled to Become a Cybersecurity Risk
April 30, 2021 · eweek.com
- Cybersecurity challenges in AI age
April 29, 2021 · miragenews.com
- Preparing for AI-enabled cyberattacks
April 9, 2021 · technologyreview.com
- How To Make Autonomous Cars Trustworthy and Free from Cybersecurity Threats
April 8, 2021 · spectrum.ieee.org
- Defending AI With AI: The AI-Enabled Solutions to Next-Gen Cyberthreats
April 6, 2021 · readwrite.com
- The superpowered SOC: How AI can drive agencies to the next level of cyber defense
March 30, 2021 · gcn.com