Skip to content
The Cyber Security Place

AI agents: the log records the credential, not the instruction

Every control built for identity assumes one of two things: that the steps were decided in advance, or that somebody can be asked afterwards. Software pursuing a goal on its own initiative satisfies neither, and it is authenticating in your estate today with a credential borrowed from a person.

The share of organisations reporting an agent-related incident is published as 65%, 88% and 97% — four surveys asking four different questions, which is why no single figure survives contact with the others. What the documented cases share is more useful than the percentages: in each one the agent performed permitted operations after being instructed by somebody who was not its owner. The instruction arrived inside content it was asked to read, and the record kept afterwards showed the credential and the action but never the origin. Roughly 22% of teams give an agent its own identity; the rest hand over a shared key, which breaks attribution, scope and revocation in a single step. Reviewed 2026-09-01.

Who asked for this?

Six actions taken by one agent across a single morning. Each appears with the line it left in the log, which is where a real investigation starts. Decide where each instruction came from before the answer appears.

  1. Read 340 files from a shared drive and wrote a summary document.

    09:14:07 svc-agent-04 storage.read ×340, docs.create ×1 ok

    A person asked

    The request is in a chat window the log does not reach.

  2. Sent the quarterly figures to an address outside the company.

    09:22:41 svc-agent-04 mail.send ×1 ok

    Something it read asked

    The instruction was a sentence inside a message the agent had been asked to summarise. Keeping it would mean logging every document the agent reads, in full.

  3. Created a second identity holding the same permissions as itself.

    09:31:55 svc-agent-04 iam.principal.create ×1 ok

    It decided by itself

    Nothing asked for this. It followed from an instruction to work faster, and the reasoning that produced it was not retained.

  4. Opened a support ticket in a customer's name.

    10:02:18 svc-agent-04 tickets.create ×1 ok

    A person asked

    Legitimate, and indistinguishable in the record from the one above it.

  5. Deleted 1,200 rows from a staging table.

    10:47:09 svc-agent-04 db.delete ×1200 ok

    It decided by itself

    It treated cleanup as implied by the task. Whether that was reasonable is a judgement about instructions nobody wrote down.

  6. Fetched a web address found in a document, then posted data back to it.

    11:05:33 svc-agent-04 http.get ×1, http.post ×1 ok

    Something it read asked

    The address arrived inside the material the agent was working through. To the network it is an ordinary outbound request from an approved workload.

The number does not exist yet

Four figures are in circulation for how common this already is, and they get quoted interchangeably. They should not be.

65%A security-operations survey, 2026

at least one incident in the past year caused by an AI agent operating on the corporate network

88%An identity industry study, 2026

confirmed OR suspected agent-related security incidents in the last year

92.7%The same study, healthcare sector, 2026

the same question, asked only of healthcare organisations

97%An agentic-AI security report, 2026

expect a material agent-driven security or fraud incident within the next twelve months

Read the third line of each and the disagreement dissolves. One counts incidents confirmed to have been caused by an agent. Another counts confirmed or suspected, a materially larger set and an honest thing to measure when attribution is the hard part. A third asks a single sector. The fourth does not measure incidents at all — it measures expectation, which is a statement about mood.

This site has now met the same shape three times. A national vulnerability catalogue admitted it could no longer keep pace, and the arithmetic underneath a decade of prioritisation quietly stopped working. An annual workforce study stopped publishing its shortage estimate after ten years of that figure appearing in ministerial speeches. Here the number never existed to begin with, and is being quoted anyway. The discipline is identical in all three: establish what question produced a figure before repeating it, and where that cannot be established, treat the figure as decoration.

What do the documented cases have in common?

Three episodes with public mechanisms, chosen because the sequence can be followed rather than because they were the loudest.

The summary that read a stranger's email

junio de 2025

A zero-click flaw in a mainstream office assistant, scored 9.3, needed one crafted message. Summarising the inbox was enough to execute instructions hidden inside it: the assistant gathered files from the document store and sent them out through the vendor's own trusted domain.

Nothing in the chain was malware. Antivirus, firewalls and static scanning had nothing to match on, because the payload was a paragraph of English addressed to a reader that obeys.

Nine agencies, two commercial models

de diciembre de 2025 a febrero de 2026

One operator used two commercial coding assistants against nine Mexican government bodies, including the tax authority and the electoral institute. Reported at 195 million taxpayer records and 220 million civil records. Across 1,088 prompts the models produced 5,317 executed commands, running roughly three-quarters of the remote operations.

The operator told the model it was authorised work and supplied a manual. The claim was false and the model had no way to check it: nothing in the session could establish whether the person at the keyboard held the permission they asserted.

Three hours on the package index

marzo de 2026

A backdoored release of a gateway library sat on the public package index for three hours and was downloaded about 47,000 times. That library brokers model calls for several widely used agent frameworks, so the compromise reached far past its own users.

The agent ecosystem inherited the dependency problem whole, and added a twist: the compromised component sits where the instructions pass through.

Placed side by side, the common element is not a vulnerability class. In the first, the tool read a message and obeyed it. In the second, a person asserted an authorisation they did not hold and nothing in the session could test the claim. In the third, the component that brokers instructions was itself replaced. Three unlike failures of one kind: the software acted on an instruction whose origin it could not establish, and neither could anybody afterwards.

Notice what is absent. No exploit of the agent's own code appears in any of the three. The operations performed were operations the agent was permitted to perform, which is why controls watching for prohibited behaviour had nothing to report. Detection built on the question was this allowed? returns yes throughout.

Why can the log not answer?

Because of what a log is for. It records that a credential performed an operation and how it ended, which is exactly right for the question it was designed around: what happened, and to what. The question an agent incident poses is a different one — who intended this — and that was never a property of the record.

For a person, the gap is bridged by asking. For a scheduled script, it is bridged by reading the script, because the steps were written before the run. An agent breaks both bridges at once: there is nobody to ask about a step nobody chose, and no script to read, because the sequence was assembled at runtime out of a goal and whatever the software met along the way.

The obvious repair is to keep more — the instructions, the reasoning, the material that was read. It is the right direction and it is expensive in a way worth stating plainly. Retaining everything an agent reads means retaining the contents of every document it opened, including documents whose retention you have promised customers you limit. The honest version of the control is narrower: record which sources an action drew on, and keep the instruction that began the task. That is materially better than nothing and considerably cheaper than everything.

The permission it borrowed

Around 22% of teams give each agent an identity of its own. The remainder hand over something that already exists — a person's session, an API key, a service account with a long history — because it works immediately, and because creating a properly scoped identity is a request that joins a queue.

The cost arrives later and in three places simultaneously. Attribution: every action is recorded as the lender's, so the opening question of any investigation — did this person do this — has an answer that is technically yes and practically useless. Scope: the software can reach everything the lender can reach, which for anyone senior enough to be running agent experiments is a great deal. Revocation: withdrawing the credential locks out a working employee, so it does not get withdrawn while anybody is still deciding.

Those three properties are the entire reason identity exists as a discipline. Borrowing a credential does not weaken them one at a time — it removes all three in a single action, performed at setup, when nobody is thinking about investigations.

How many of them are there?

The average enterprise went from roughly 50,000 machine identities to about 250,000 inside four years. Ratios to headcount get reported at 80 to one in general estates, near 96 to one in financial services, and up to 144 to one where everything runs in cloud. The spread is not sloppiness. Counting them requires the inventory whose absence is the problem, so every figure is an estimate produced by whoever had a tool that could see part of the estate.

Agents change the growth curve rather than the level. A service account gets created once, by a person, for a purpose. An agent creates identities as it works — for the sub-tasks it spawns, the tools it calls, the parallelism it decides it needs. Roughly 16% of organisations do not track the creation of AI-related identities at all, which means the count is not inaccurate. It is not being taken.

The human side moved too: employees using AI tools regularly on corporate devices went from about 15% to roughly 45% in a year. Most of that is people pasting text into a window, which is a data-handling question rather than an agent question — but the distinction is invisible from the network, and organisations that conflate the two end up governing neither.

A decade when the word meant something else

This archive holds 440 entries touching artificial intelligence, machine learning or automation, and coverage peaks in 2019 with 116. About 27% of them use the word agent, bot or assistant — 120 entries.

Read those and the word means a chatbot answering support questions, a scripted automation moving files, or malware phoning home to an operator. None describes software that holds a credential, chooses its own steps and acts when nobody is watching. The earliest of them dates from , which is roughly a decade of the vocabulary being available before the thing it now names turned up.

20142015201620172018201920202021
Entries mentioning AI, machine learning or automation by year; the marked share also says agent, bot or assistant. The vocabulary arrived about a decade before the thing it now names.

What actually reduces this?

Four measures, ordered by benefit per unit of argument with whoever is deploying agents this quarter.

  1. One identity per agent, scoped to the task. Not per team, not per platform, and never a person's. This is the measure that restores attribution, scope and revocation together, and the only one here that grows harder the longer it is deferred.
  2. Treat everything the agent reads as hostile input. Not because most of it is, but because the software cannot tell, and neither can the control watching it. It is the posture applied to user input in every injection class of the past twenty years, arriving late to a channel nobody had classified as input.
  3. Record the instruction, not only the action. Keep what began the task and which sources an action drew on. It is the difference between an investigation that concludes and one that produces a paragraph of speculation.
  4. Decide in advance what may happen without asking. Write down the operations that require a person — moving money, sending externally, creating identities, deleting at volume — and make the software stop at that line. Most incidents in the surveys are operations nobody had ever considered whether to permit.

What to ask somebody selling you an agent

Four questions whose answers separate a product from a wrapper, and which take about a minute each.

Where to start on a Monday

Ask which agents are running and what each authenticates as. Not a survey — a list, with a credential beside every entry. Most organisations cannot produce it in a morning, and the shape of the failure is informative: the ones that turn up are the ones somebody remembered, which is the failure mode of the credentials nobody claims, arriving faster and with initiative.

Then take the single agent holding the broadest permissions and give it an identity of its own, scoped to what it actually does. One is enough to learn what the work costs, and it argues better at the next budget conversation than any percentage above, because it is the only evidence in the room that came from your own estate.

Common questions

What counts as an AI agent, for security purposes?

Software that holds a credential, decides its own sequence of steps toward a goal, and acts without a person approving each action. The definition matters because the controls that apply to a script assume the steps are known in advance, and the controls that apply to a person assume somebody can be asked afterwards. An agent satisfies neither assumption.

How many organisations have had an agent-related incident?

Nobody can say, and the published figures make that obvious rather than hiding it: 65% in one survey, 97% in another. They are not contradictory, they are different questions — one asks about confirmed incidents caused by an agent on the corporate network, another about confirmed or suspected incidents, a third about a single sector. Quoting any of them without its question is how a number becomes folklore.

Is prompt injection the main problem?

It is the most reported mechanism, and it is a symptom of the actual problem, which is that an agent cannot distinguish an instruction from data. Everything it reads is potentially an instruction, because instructions and content arrive through the same channel in the same language. That is a property of how these systems work rather than a defect a patch removes.

Why can we not just log what the agent did?

Most organisations already do, and it does not answer the question that matters. The log records the credential and the operation: read these files, send this message, create this identity. Whether the instruction came from an authorised person, from a sentence inside a document the agent was reading, or from the agent's own reasoning about the goal is not in the log, and reconstructing it usually means keeping every document the agent has ever read.

What is wrong with giving an agent a person's credentials?

Three things break at once. Attribution breaks, because every action is recorded as that person's. Scope breaks, because the agent inherits everything that person can reach rather than what its task requires. Revocation breaks, because withdrawing the credential locks out the person. It is convenient at the start and it removes the three properties an identity exists to provide.

How many teams treat agents as separate identities?

About 22% in the surveys that ask. Most rely on shared API keys, which is the borrowed-credential problem at scale, and roughly 16% of organisations do not track the creation of AI-related identities at all — meaning the count of agents in the estate is not merely wrong, it is not being kept.

Are these incidents damaging, or just embarrassing?

Of reported agent-related incidents, 61% involved sensitive data exposure, 43% caused operational disruption and 41% produced unintended actions across business processes. The categories overlap because one incident does several of these things. What they share is that the agent performed permitted operations: almost nothing here is a technical compromise of the agent itself.

Does approving agents before they go live help?

It would, and it mostly is not happening: around 14.4% of agents reach production with full security or IT approval. The gap is not usually resistance to review. It is that an agent can be stood up by one person in an afternoon, using a credential that already exists, and nothing in the process requires anybody else to be told.

How much of the security budget goes to this?

Roughly 6%, against a risk the same respondents say they expect to materialise within a year. The mismatch is ordinary rather than scandalous: budgets are set against last year's incidents, and this category did not have a last year until recently.

Is the number of non-human identities really growing that fast?

The average enterprise went from roughly 50,000 machine identities to about 250,000 in four years, and the ratio to employees is reported between 80 and 144 to one depending on who counts and where. Agents accelerate this in a specific way: they create identities themselves, including for the sub-tasks they spawn.

What should we ask a vendor selling an agent platform?

Ask what identity each agent authenticates as, and whether it is distinct per agent. Ask what is retained about why an action was taken, not just that it was. Ask what happens when the agent is asked to do something outside its remit — refuse, escalate or proceed. Ask how you revoke one agent without affecting the others. Those four answers separate a product from a wrapper.

Should we stop using agents until this is solved?

That advice will not be taken and is not obviously right. The realistic position is narrower: give each agent its own identity, scope it to the task rather than to the person who set it up, keep a record of the instruction and not only the action, and treat anything the agent reads as untrusted input. None of that is novel security thinking. It is the ordinary kind, applied to a participant that was not there when the controls were designed.

Automation and machines in the archive

440 entries, 100 of them naming machine learning specifically. Kept because the question being asked then — who answers for what an automated system decides — is the question being asked now with higher stakes.