Skip to content
The Cyber Security Place

Devices and data

BYOD was the rehearsal

A decade of arguing about personal devices produced one durable lesson, and the industry is currently relearning it from the beginning with a different tool.

Last reviewed August 29, 2026

Personal devices went from policy question to ordinary condition without the security problem being solved — the vocabulary simply stopped. The same shape is running again with unsanctioned tools: something useful arrives faster than approval can consider it, prohibition produces concealment, and only defining what the tool may touch holds. Reported figures on shadow AI breaches range from about 20% of incidents adding some $670k to as high as 43% averaging $5.39m, and 68% of breached organisations had no governance policy at all. Meanwhile the data those years accumulated remains: more than half of enterprise data is dark — unclassified and unmonitored — and unmonitored stores carry roughly $900k of extra cost per breach. The cheapest remedy remains the one with no supplier attached: hold less of it, for less time, with a named reason written beside each period.

A category that stopped being written about

This section carried 101 entries, 47 of them about personal devices and 36 about large data estates. Coverage peaks in 2017 and then stops.

202014282015102016302017112018220190202002021

Nothing happened in 2019 to make personal devices safe. What happened is that they stopped being remarkable, and a topic that is no longer remarkable stops generating articles regardless of whether it has been dealt with. The same is true of the other half of this section: large data estates did not shrink, they became the normal condition of running any business, and the special vocabulary retired. That is worth pausing on, because it is a general hazard of reading any archive of trade coverage. Volume tracks novelty. A subject that has been settled and a subject that has merely become boring produce the same falling line, and telling them apart requires knowing something about the world rather than something about the coverage. The earliest piece here, What a BYOD disaster looks like - and how to prevent it, sets out concerns that read as current.

What exactly was rehearsed?

Strip the technology out of the BYOD decade and a sequence remains that has since repeated almost exactly. Staff acquire something that makes their work easier. It arrives without passing through any approval process, because the approval process exists for purchases and this was not one. Security discovers the practice already widespread. A prohibition is issued. The practice continues, now unobserved, because the thing being forbidden solves a real problem for the person doing it and the prohibition solves a real problem for somebody else.

What eventually worked was neither the ban nor surrender. It was narrowing the question from whether the device may exist to what the device may reach — a session rather than a machine, a container rather than a phone, access checked at each request rather than assumed from an enrolment. That reframing took most of a decade to become conventional, and it is the single transferable finding of the period.

The current version substitutes a service for a device. Somebody pastes a document into a tool that summarises it well. The tool was not procured, no contract governs it, and the material has left. Reported measurements of the consequence differ enough to be worth stating carefully: one widely cited figure puts incidents involving unsanctioned AI at roughly a fifth of breaches with some $670,000 added to the average, while another reports the share at 43% and an average near $5.39m. The methodologies differ; the direction does not. What is not in dispute is that 68% of breached organisations had no governance policy covering it at all, which is the 2014 position with the nouns changed.

Anyone drafting that policy now has a decade of evidence about which approaches fail, and it is available at the price of noticing the resemblance. The arguments in What a BYOD disaster looks like - and how to prevent it can be read almost unaltered against the current question.

What those years left behind

The other half of this section was about collecting more. It succeeded. Analyst estimates now put more than half of enterprise data in the category of dark data — held but not classified, monitored or used — and unmonitored stores have been associated with something like $900,000 of additional cost when a breach reaches them. Almost none of it was collected deliberately in bulk. It accumulated one copy at a time.

Cameras and badge readers are the purest illustration. They write continuously, their recorder was configured once by whoever installed it, and the setting that governs how long the footage survives is a default nobody has revisited since. Retention is where this becomes tractable, because the arithmetic is a multiplication that few teams perform. A store receiving a given number of records a month, kept for a given number of months, holds their product. Set the period once in a policy document and the volume it produces is never looked at again.

Customer records

Nobody has been asked to delete them, and the marketing team would rather keep the option open. Seven years is the number people reach for when there is no number.

Held now:
336,000 records (84 months)
If kept 24 months:
96,000
Modelled exposure:
$2.5m

Cutting to 24 months divides what a breach could expose by 3.5, and removes 240,000 records from the estate.

Access and system logs

Kept too briefly rather than too long, because storage is billed monthly and the value only appears during an investigation that has not happened yet.

Held now:
2,700,000 records (3 months)
If kept 12 months:
10,800,000
Modelled exposure:
$20.3m

Kept for less than the 12 months an investigation would need. This one runs the other way: the fix is to keep it longer.

Backups and snapshots

Snapshots taken for a migration in 2019 that nobody deleted afterwards. They hold the data as it was, including the fields since removed from production.

Held now:
1,440,000 records (120 months)
If kept 12 months:
144,000
Modelled exposure:
$10.8m

Cutting to 12 months divides what a breach could expose by 10, and removes 1,296,000 records from the estate.

Video and door records

Cameras and badge readers write continuously, and the recorder was configured by the installer to fill the disk it shipped with. Almost nobody revisits that setting, and few obligations anywhere require three years of footage of a corridor.

Held now:
9,360,000 records (36 months)
If kept 1 months:
260,000
Modelled exposure:
$70.2m

Cutting to 1 months divides what a breach could expose by 36, and removes 9,100,000 records from the estate.

Analytics extracts

Copies made for a question somebody asked once. They sit outside the systems that enforce retention, which is usually why nobody counts them.

Held now:
1,800,000 records (60 months)
If kept 6 months:
180,000
Modelled exposure:
$13.5m

Cutting to 6 months divides what a breach could expose by 10, and removes 1,620,000 records from the estate.

Exposure is a straight-line anchor of $7.50 per record, taken from published incident figures. Real cost per record falls as volume rises and climbs with sensitivity, so this compares two retention periods against each other. It does not budget for anything.

Across the four profiles above, the excess adds to roughly 12,256,000 records held beyond any stated obligation. One of the four runs the other way — access and system logs are kept for less time than an investigation would need, which is why retention cannot be fixed by shortening everything. It is fixed by attaching a named reason to each period, and some of those reasons argue for longer.

Why is deletion the control nobody buys?

Because there is nothing to buy. Every other measure adds something around the data — a control, a monitor, a layer of encryption — and each arrives with a supplier, a budget line and somebody whose job is to explain its value. Deletion removes the thing being protected. It has no vendor, no dashboard, and no way to demonstrate that it worked, since its entire result is an absence.

It is also the only measure whose benefit is unconditional. A control can be misconfigured, bypassed or switched off during a migration. Data that no longer exists cannot be exposed by any of those failures, and it costs nothing to maintain for the rest of time. Given that breach cost scales with the number of records involved — the largest incidents, in the tens of millions of records, run to hundreds of millions of dollars — halving what is held reduces the largest term in that calculation directly.

The obstacles are organisational rather than technical. Nobody is rewarded for deleting something, and everybody can imagine a future question the data might have answered. The counter is to require a named reason and a named owner for each retention period, which converts an open-ended default into a decision somebody has to defend. Stores that survive that test are kept with confidence. Stores that cannot be attributed to any obligation are usually the majority, and they can go at no cost and no risk.

Start small. Pick one store. Ask three questions about it — who asked for this, what breaks if it disappears, and when did anyone last read it — then act on the answers. Most teams find the third question does the work: a table nobody has queried in four years is not an asset waiting for its moment. It is a liability with a storage bill. Do that once a quarter, on one store at a time, and the estate shrinks without any project, any budget line, or any meeting about data governance frameworks.

How do you find data nobody remembers creating?

Discovery tooling scans what you point it at, which makes it excellent for the stores you already know about and structurally blind to the ones you do not. The useful complement is to approach the problem from the direction of money. Somebody had to pay for storage, so cloud billing lines, expense claims and supplier invoices name systems that no inventory contains, and they name them with a date and an owner attached.

Leavers are the second seam. When somebody departs, the question of what they were running is usually reduced to disabling their accounts, and the analytics extract they built for a project in 2021 continues to exist without anyone able to say what it holds. A leaver process that asks what information the person created, rather than only what access they had, finds material that no scan will.

The third is the one nobody enjoys: asking teams directly what copies they hold, with an explicit assurance that the answer will not be used against them. Copies made outside the sanctioned path exist because the sanctioned path was too slow, which is a finding about the path. Treating the disclosure as an incident guarantees the next one stays hidden, and the same dynamic governed device policy for a decade.

Every copy is a second estate

Retention periods are written against systems, and information does not stay in systems. It gets extracted for a report, loaded into a test environment so a developer can reproduce a fault, attached to an email, pasted into a spreadsheet that becomes the way a team actually works. Each of those is a full copy of some subset, and each sits outside whatever enforces the policy on the original.

Test environments deserve particular attention because the incentive runs so strongly the wrong way. Realistic data makes testing better, production data is the most realistic data available, and the environment holding it is by design less protected than production — fewer controls, broader access, and often a copy taken once and never refreshed, which means it also preserves records the live system has since deleted. It is the same material with the protections removed, and it is rarely in scope for the audit that examined the system it came from.

The spreadsheet version is more mundane and at least as common. Somebody needs a view the system does not provide, exports what they need, and the export becomes the artefact the team refers to. It is emailed, it is copied to a shared drive, a version of it goes home on a personal device — which is where the two halves of this section meet — and none of that is visible from the system of record.

The realistic response is not to forbid copies, which fails for the reasons the device decade established. It is to make the sanctioned route fast enough that the copy is not worth making, and to treat a proliferation of extracts as a report about a missing capability instead of as a discipline problem. Where copies must exist, the cheapest control is masking at the point of extraction, because it removes the sensitive content before the copy escapes rather than trying to follow it afterwards.

What happens when somebody asks for their data back

Data protection regimes across many jurisdictions give individuals the right to ask what an organisation holds about them and, in defined circumstances, to have it deleted. Both rights have deadlines measured in weeks. Both are answered from an inventory, and an organisation that cannot enumerate its own stores cannot answer either accurately — it can only answer for the stores it remembers.

This is where dark data stops being an abstraction about cost per breach and becomes a concrete exposure. An erasure request honoured in the customer database and not in the analytics extract, the 2019 snapshot and the departed employee's spreadsheet has not been honoured, and the response given to the individual was incorrect. The organisation will usually not know that until something surfaces the copy, at which point the original failure and the inaccurate response are both on the record.

The practical protection is the same inventory that everything else in this section depends on, which is a reasonable argument for building one even where no obligation forces it. The secondary protection is to reduce the number of places an answer has to be checked, and that is retention again: a store that has aged out cannot hold a record that a request should have reached.

Worth separating from the compliance framing is that this is also simply a service question. Somebody asking what a company knows about them is entitled to a correct answer, and the organisations that give one quickly tend to be the ones that already knew where their data was for their own reasons.

What happens to company data when a device's owner leaves?

Offboarding is designed around accounts. Access is revoked, mail is forwarded, the laptop comes back, and the process records itself as complete. The personal phone does not come back, because it was never company property, and whatever it holds leaves with it.

Under an arrangement where the device only ever displayed data, this is a non-event: revoking the session ends the access and nothing local survives. Under any arrangement where data was allowed to land — a synced mail store, downloaded attachments, an offline copy of a document library — the leaver walks out with a copy that no revocation reaches. Which of those two situations an organisation is in was decided years earlier by a configuration choice, and it is discovered at the moment somebody resigns.

The remote wipe that policies rely on has a narrower reach than it is credited with. It requires the device to be enrolled, connected and not already reset; it is frequently limited to a managed container rather than the whole handset, which is the correct design and also means anything the user copied out of the container is untouched; and it depends on somebody triggering it promptly, which competes with every other task in a departure week. As a control it is worth having and it is not a substitute for the data never having been there.

A departure is also the moment to ask the question the copy problem raises: what did this person create. The accounts they held are recorded. The extract they built, the shared spreadsheet the team still uses, the analytics job that runs on a schedule nobody else understands — none of that is in the offboarding checklist, and all of it outlives them. Adding one line to that checklist finds more unmanaged data than most discovery exercises.

A workable position on tools people bring

The decade of device policy narrowed to a small set of things that held up, and they transfer directly to services. The first is that the unit of control is the data, not the tool. A rule saying which categories of information may go to which classes of destination survives the arrival of a product nobody has heard of; a list of approved products does not, and it is out of date the week it is published.

The second is that the sanctioned route has to be genuinely usable. Every unsanctioned copy and every unapproved service is evidence about the approved path — that it was slower, or absent, or required an approval nobody could obtain in the time available. Treating each instance as a discipline matter addresses the symptom and guarantees the next one is hidden, which removes the only signal available about where the official process fails.

The third is that visibility beats permission. Knowing that forty people use a particular service is worth considerably more than a policy stating that nobody may, because the first can be acted on and the second is a claim about a world that does not exist. That is an argument for making disclosure safe and for looking at what billing and network egress already show, rather than for accumulating more rules.

None of this is permissive. A rule that certain categories of data never leave, enforced where the data lives, is stricter in effect than a prohibition everybody works around — and it has the advantage of describing something that can actually be checked.

Common questions

Is BYOD still a thing in 2026?

The practice is universal and the term is not. Personal devices reach work systems almost everywhere, including at organisations whose written policy forbids it. What disappeared was the idea that this needed a dedicated policy category rather than being the ordinary condition that access control has to assume.

What replaced BYOD as the unmanaged thing staff bring to work?

Tools they signed up for themselves, and increasingly AI services. The pattern is identical: something genuinely useful arrives faster than the approval process can consider it, prohibition produces concealment rather than compliance, and the workable answer is to define what the tool may touch instead of whether it may exist.

How costly are breaches involving unsanctioned AI tools?

Published figures vary considerably by methodology. One widely cited set puts incidents involving shadow AI at around 20% of breaches, adding roughly $670,000 to the average cost; another reports the share climbing to 43% with an average near $5.39m. The direction is consistent across sources even where the levels are not.

What is dark data?

Data an organisation holds but does not classify, monitor or use. Analyst estimates put it at more than half of enterprise data, and unmonitored stores of it have been associated with roughly $900,000 of additional cost per breach. Most of it accumulated one copy at a time for questions somebody asked once.

Does deleting data actually reduce risk?

It is the only control that removes the thing being protected instead of wrapping a layer around it. Breach cost scales with the number of records exposed, so halving what you hold halves the largest term in that calculation, and it costs nothing to run afterwards.

How long should we keep customer data?

As long as a named obligation requires, and no longer, with the obligation written next to the period. The common failure is not a period that is too long but a period nobody can attribute to anything — seven years is what people reach for when there is no number, and it is rarely the number any rule actually specifies.

Are logs an exception to data minimisation?

Largely, yes, and they are usually the one category kept too briefly rather than too long. With a mean of roughly six months between compromise and discovery, retaining authentication and administrative logs for a few months guarantees the investigation begins after the evidence has rotated away.

Why do old backups and snapshots matter?

They hold the records as they were, including fields since removed from production and accounts since closed. A deletion request honoured in the live system and not in a 2019 snapshot has not been honoured, and the snapshot is usually outside whatever enforces retention elsewhere.

Should we ban staff from using AI tools?

A prohibition that people need to break in order to do their work produces concealment, which removes the visibility that made the prohibition enforceable. The BYOD decade tested this thoroughly. Defining which data may go to which category of service is harder to write and considerably more likely to hold.

How do you find data stores nobody remembers creating?

Start with what pays for storage rather than with what stores data. Cloud billing, expense claims and supplier invoices name systems that no inventory contains, because the bill had to be paid by someone. It is a less elegant method than discovery tooling and it finds a different set.

Who should own data retention?

Whoever can say no to keeping something. In practice retention fails when it is owned by a team that can write a policy but cannot delete anything, which produces an accurate document describing a state of affairs that has never existed.

What is the cheapest improvement available here?

Writing down, for each significant data store, why it exists and what would have to be true to delete it. The exercise is unglamorous and it routinely finds stores nobody can justify, which are the ones that can go immediately at no cost and no risk.

Devices and data in the archive

101 entries, newest first, from 2014 to 2019.