Shadow AI Statistics 2026: Adoption, Cost, and Real Risk
43% of breached organisations hit a shadow AI incident in 2026, up from 20% a year earlier (IBM). Those breaches averaged $5.39M against a $4.99M global average. Meanwhile, among the organisations breached through AI, the share holding any AI governance policy fell from 37% to 32%.
Here is the catch, and it shapes everything below. Almost every shadow AI statistic in circulation comes from a company selling shadow AI detection, and a lot of them cannot be traced back to anyone at all. We know, because we audited our own database first and found three figures citing content farms instead of the researchers who produced them. Every number on this page names and links the organisation that published it. Where a figure carries a limit that changes what it means — a narrow sample, an overlapping cohort, a measurement of conversations rather than people — we say so at the point you read it. And where we tested something ourselves, we show you what broke.
TL;DR — the numbers that matter
- 43% of breached organisations experienced a shadow AI incident, up from 20% (IBM 2026).
- $5.39M average cost, roughly 8% above the $4.99M global average.
- 66% of office workers used AI they believed was banned; 88% put work data into public AI tools (PagerDuty, n=1,250).
- 47.11% of enterprise AI conversations run on personal identities (LayerX telemetry).
- 18,033 TB of enterprise data went to AI tools in 2025, up 93% (Zscaler).
- Grammarly received 3,615 TB — about 1.8x ChatGPT's 2,021 TB. The biggest destination is not the chatbot your policy names.
- 32% of AI-breached orgs have an AI governance policy, down from 37%.
- We tested the free tools: Presidio caught 4 of 12 sensitive items out of the box, 10 of 12 after tuning. The best free blocklist misses chatgpt.com.
- No government or statistics office measures shadow AI. Academia has one small study. Every figure with real scale behind it is vendor research.
Last updated: August 2026
📊 Key Shadow AI Statistics at a Glance
Shadow AI means staff using AI tools the business has not approved — a personal ChatGPT account for a work report, an AI notetaker in a client call, a browser extension nobody vetted. These are the headline figures, each traced to the organisation that produced it.
| Finding | Value | Source |
|---|---|---|
| Breached orgs with a shadow AI incident | 43% | IBM Cost of a Data Breach Report 2026 |
| Same figure, one year earlier | 20% | IBM Cost of a Data Breach Report 2025 |
| Breaches of an org's own AI models or apps | 21% | IBM Cost of a Data Breach Report 2026 |
| Workers using personal GenAI accounts for work | 57% | Gartner Top Cybersecurity Trends 2026 |
| Workers entering sensitive data into unapproved GenAI | 33% | Gartner Top Cybersecurity Trends 2026 |
| Average cost of a shadow AI breach | $5.39M | IBM Cost of a Data Breach Report 2026 |
| Global average cost of any breach | $4.99M | IBM Cost of a Data Breach Report 2026 |
| Average cost of an AI-enabled malicious breach | $6M | IBM Cost of a Data Breach Report 2026 |
🔍 We Audited Our Own Statistics First. Three of Nine Cited the Wrong Source.
Before writing this page we audited every shadow AI figure already in our database. Nine entries. Three cited SEO content farms rather than the organisations that did the research, and two more duplicated the same IBM figure under different names.
Most of the underlying numbers were defensible. The citations were not. That is the ordinary failure mode in this subject: a figure gets published once, then repeated for years while the trail back to its source rots away. We re-sourced every one before publishing this article.
Nathan House's Analysis — the provenance problem, measured on ourselves
When we audited our own database in August 2026, 3 of the 9 shadow AI statistics we held cited SEO aggregator sites instead of the organizations that produced the research, and 2 more duplicated the same IBM figure under different names. Most of the underlying numbers were defensible — the citations were not. That is the ordinary failure mode in this subject: figures survive while their provenance rots. Method: Manual provenance audit of every shadow AI entry in this store, August 2026: 9 entries checked; 3 cited aggregators (secondtalent.com, electroiq.com, gitprotect.io) rather than the primary publisher; 2 duplicated the IBM $670K cohort figure. Finding is about CITATION quality, not numerical falsity. Re-auditing now returns 100% because the entries were corrected — this is a fixed record of the August 2026 state.
One correction is worth stating plainly, because we got it wrong in the other direction too. During earlier research we blacklisted the "43%" figure as an unverifiable inflation of IBM's 2025 number. It is IBM's genuine 2026 statistic, published on 29 July 2026. We were right about the evidence available at the time and wrong about the number. Both errors — trusting a bad citation and rejecting a good figure — come from the same place: nobody checks.
📈 How Common Shadow AI Actually Is
43% of breached organisations reported a shadow AI incident in IBM's 2026 Cost of a Data Breach Report. A year earlier it was 20%.
Nathan House's Analysis — the growth rate
The share of breached organizations reporting a shadow AI incident was 2.15x higher year over year (20% to 43%, a 23-point jump). This is a change in reported prevalence between two annual samples, not a count of incidents. Method: 43% / 20% = 2.15x. Separate annual samples (IBM 2025 and 2026), so this is not a longitudinal measurement of the same organizations.
Read that figure carefully. It is 43% of organisations that suffered a breach — not 43% of all breaches, and not 43% of all companies. IBM surveyed 602 organisations across 17 industries and 16 countries via the Ponemon Institute, with fieldwork running March 2025 to February 2026. Breached organisations are not a random sample of all organisations, so this tells you about companies that had a bad year, not about the market as a whole.
| Finding | Value | Source |
|---|---|---|
| Breached orgs with a shadow AI incident | 43% | IBM Cost of a Data Breach Report 2026 |
| Same figure, one year earlier | 20% | IBM Cost of a Data Breach Report 2025 |
| Breaches of an org's own AI models or apps | 21% | IBM Cost of a Data Breach Report 2026 |
| Workers using personal GenAI accounts for work | 57% | Gartner Top Cybersecurity Trends 2026 |
| Workers entering sensitive data into unapproved GenAI | 33% | Gartner Top Cybersecurity Trends 2026 |
💰 What a Shadow AI Breach Costs
$5.39M is the average cost of a breach involving shadow AI (IBM 2026). The global average for all breaches is $4.99M, itself a record high and up 12% year over year.
Nathan House's Analysis — the cost gap
Breaches involving shadow AI averaged 8% more than the global average — $400K higher ($5.39M vs $4.99M). This is a comparison between cohorts in the same IBM study, not evidence that shadow AI caused the difference. Method: ($5.39M - $4.99M) / $4.99M — both from IBM Cost of a Data Breach 2026. Observational, not causal.
Two caveats. First, those three cohorts overlap — a breach can be both AI-enabled and involve shadow AI, so they are not exclusive categories you can stack. Second, this is an observational comparison inside one study. Organisations with shadow AI incidents likely differ from those without in ways IBM does not control for. The data shows an association. It does not prove shadow AI caused the extra $400,000.
| Finding | Value | Source |
|---|---|---|
| Average cost of a shadow AI breach | $5.39M | IBM Cost of a Data Breach Report 2026 |
| Global average cost of any breach | $4.99M | IBM Cost of a Data Breach Report 2026 |
| Average cost of an AI-enabled malicious breach | $6M | IBM Cost of a Data Breach Report 2026 |
| Shadow AI incidents that drew a regulatory fine | 21% | IBM Cost of a Data Breach Report 2026 |
| High-shadow-AI vs low-shadow-AI cohort gap | $670,000 | IBM Cost of a Data Breach Report 2025 |
⚠️ The $670,000 figure is real, and almost always misquoted
IBM compared two groups: organisations with high levels of shadow AI, and organisations with little or none. The gap between those cohorts was $670,000. It is not a surcharge added to a breach, which is how you will see it written almost everywhere. If you quote it, quote it as a cohort comparison.
👥 What Employees Are Actually Doing
66% of office professionals have used AI at work believing it was not permitted under company policy. 88% have shared work information with public AI tools like ChatGPT, Claude or Gemini (PagerDuty, 2026).
The methodology matters here. Wakefield Research surveyed 1,250 office professionals at organisations with $500M+ annual revenue, deliberately excluding IT and technology roles, across the US, UK, Australia and Japan. So this measures ordinary office staff, not engineers.
One nuance worth preserving: the 66% figure describes what staff believed was not permitted. It is a measure of knowing-and-doing-anyway, which is a different and more useful thing than measuring policy breaches directly.
| Finding | Value | Source |
|---|---|---|
| Used AI at work believing it was not permitted | 66% | PagerDuty Shadow AI Survey 2026 |
| Shared work information with public AI tools | 88% | PagerDuty Shadow AI Survey 2026 |
| Entered customer data into public AI tools | 34% | PagerDuty Shadow AI Survey 2026 |
| First encountered the AI tool in their personal life | 89% | PagerDuty Shadow AI Survey 2026 |
| Believe they know AI better than their own tech team | 72% | PagerDuty Shadow AI Survey 2026 |
| Believe leadership plays by different AI rules | 81% | PagerDuty Shadow AI Survey 2026 |
Nathan House's Analysis — awareness does not produce compliance
Among the same 1,250 office professionals, 66% used AI at work believing it was against policy and 88% shared work information with public AI tools. Both figures share one denominator, but PagerDuty did not publish a cross-tabulation, so these are two separate measures of the same population — not a measured overlap. Method: PagerDuty Shadow AI Survey 2026 (Wakefield Research), n=1,250, US/UK/Australia/Japan. Same denominator; no respondent-level cross-tab published.
89% of those using AI for work first met the tool in their personal life. That is the mechanism: people bring a habit from home, find it works, and keep going. 72% believe they understand AI better than their own tech team, and 81% think leadership plays by different rules. Whatever you think of those beliefs, they explain why a policy memo does not change behaviour.
🕵️ The Visibility Gap
47.11% of enterprise AI conversations happen through personal identities rather than corporate-managed accounts (LayerX telemetry, 2026). That traffic sits outside your identity controls by design: no corporate login, so no account-level policy, no retention control and nothing to revoke when someone leaves. Your network, endpoint or browser tooling may still see it — LayerX measured this precisely because its own browser telemetry does — but the provider-side account is not yours.
There is a second layer underneath that. Of the conversations that do use a corporate identity, 14.39% run on a personal AI licence. The login looks corporate, so it passes a casual audit, but the organisation still has no control over how that data is stored, retained or used for training.
Nathan House's Analysis — the real ungoverned share
Adding the personal-identity share (47.11%) to the personal-licence share hidden inside corporate logins (14.39% of the 52.89% corporate slice) gives roughly 54.7% of AI conversations outside governance. LayerX reports the two shares separately; this sum is StationX's calculation, not a figure LayerX published. Method: 47.11% + (52.89% × 14.39%) = 54.7%. The two LayerX categories are disjoint by construction (personal identity vs corporate identity), so they do not double-count.
| Finding | Value | Source |
|---|---|---|
| AI conversations on personal identities | 47.11% | LayerX State of AI Usage Report 2026 |
| Corporate logins running on a personal AI licence | 14.39% | LayerX State of AI Usage Report 2026 |
| Workers using personal GenAI accounts for work | 57% | Gartner Top Cybersecurity Trends 2026 |
| Unmanaged BYOD devices mixing work and personal credentials | 46% | Verizon DBIR 2025 |
| Insider incidents involving cloud or SaaS | 78% | CrowdStrike / Cybersecurity Insiders |
Gartner reaches a similar place from a different direction: 57% of workers use personal GenAI accounts for work. And the problem is not confined to AI — Verizon found that 46% of compromised systems holding corporate logins were unmanaged devices carrying personal credentials alongside them, while 78% of insider incidents already involve cloud or SaaS platforms. Unsanctioned AI is the newest layer on an access-control problem that predates it.
📤 How Much Data Is Actually Leaving
18,033 terabytes of enterprise data went to AI and machine learning applications in 2025 — a 93% increase year over year. Zscaler measured this across roughly 989 billion AI/ML transactions on its Zero Trust Exchange between January and December 2025.
⚠️ 83% and 93% are different measurements
Zscaler reports both. 93% is the growth in data volume; 83% is the growth in AI/ML activity across 3,400+ applications. Zscaler's own press release headline says 83% while the data-transfer section says 93%. Quoting either as “AI growth” without the qualifier is wrong.
| Finding | Value | Source |
|---|---|---|
| Enterprise data sent to AI/ML apps in 2025 | 18,033 TB | Zscaler ThreatLabz 2026 AI Security Report |
| Growth in that data volume, year over year | 93% | Zscaler ThreatLabz 2026 AI Security Report |
| Growth in AI/ML activity, year over year | 83% | Zscaler ThreatLabz 2026 AI Security Report |
| Apps driving AI/ML transactions (quadrupled YoY) | 3,400+ | Zscaler ThreatLabz 2026 AI Security Report |
| DLP policy violations tied to ChatGPT alone | 410 million | Zscaler ThreatLabz 2026 AI Security Report |
410 million data-loss-prevention violations were tied to ChatGPT alone, including attempts to share Social Security numbers, source code and medical records. The number of applications driving AI traffic quadrupled to more than 3,400 — which is the practical problem with any blocklist approach, as we will show with measurements further down.
✍️ The Biggest AI Data Destination Is Not ChatGPT
Grammarly received 3,615 TB of enterprise data in 2025. ChatGPT received 2,021 TB. The writing assistant took in roughly 1.8 times as much company data as the chatbot every AI policy names first.
Nathan House's Analysis — the destination nobody blocks
Grammarly received 3,615 TB of enterprise data in 2025 — 1.8x what went to ChatGPT (2,021 TB). The largest AI data destination is not the chatbot most policies name. Method: 3,615 TB / 2,021 TB (Zscaler ThreatLabz 2026, 989.3bn AI/ML transactions)
This is not an argument about Grammarly specifically. It is an argument about where your data actually goes versus where you assume it goes. A writing assistant sees every document it helps with. If your controls are built around the tools that make headlines, the largest flow is the one you are not watching.
| Finding | Value | Source |
|---|---|---|
| Data sent to Grammarly in 2025 | 3,615 TB | Zscaler ThreatLabz 2026 AI Security Report |
| DLP violations tied to ChatGPT alone | 410 million | Zscaler ThreatLabz 2026 AI Security Report |
| Applications driving AI/ML transactions | 3,400+ | Zscaler ThreatLabz 2026 AI Security Report |
| Time to compromise most enterprise AI systems | 16 minutes | Zscaler ThreatLabz 2026 AI Security Report |
The 16-minute figure is worth sitting with. Zscaler found critical flaws in 100% of the enterprise AI systems it analysed, and judged that most could be compromised in about a quarter of an hour. That is the sanctioned side of the estate — the AI you deployed on purpose — which makes the 3,400-application sprawl on the unsanctioned side harder to wave away.
📉 Governance Is Going Backwards
Of the organisations that suffered an AI-related incident, 68% had no AI governance policy — leaving just 32% that did. The year before, 63% lacked one. The gap widened while shadow AI incidents more than doubled.
Nathan House's Analysis — the divergence
Shadow AI incidents rose 23 points among breached organizations, while the share of AI-breached organizations holding an AI governance policy FELL 5 points (37% to 32%). The problem grew; the response shrank. Note the two figures use different sub-samples — breached orgs for incidence, AI-breached orgs for governance. Method: Incidence 20%→43% (breached orgs) against governance 37%→32% (AI-breached orgs). IBM Cost of a Data Breach 2025 and 2026.
An important limit on this claim. Both figures describe breached organisations only. A fall from 37% to 32% within that group does not prove governance is declining across all businesses — the composition of who got breached can shift year to year. What it does show is that among companies that suffered a breach, governance was rarer this year than last.
| Finding | Value | Source |
|---|---|---|
| AI-breached orgs with an AI governance policy | 32% | IBM Cost of a Data Breach Report 2026 |
| Orgs coordinating AI governance with security | 19% | IBM Cost of a Data Breach Report 2026 |
| Orgs with a mature AI governance committee | 12% | Cisco 2026 Data and Privacy Benchmark Study |
| Orgs using access controls on AI models and data | 40% | IBM Cost of a Data Breach Report 2026 |
| AI-breached orgs that lacked those controls | 92% | IBM Cost of a Data Breach Report 2026 |
Nathan House's Analysis — three sources, one conclusion
Two studies, two different questions, one direction of travel: 12% of organizations have a mature AI governance committee (Cisco), and 32% of organizations breached through AI had any AI governance policy at all (IBM). These measure different populations against different thresholds, so they are not points on one scale — but the lowest bar and the highest both land in the minority. Method: Cisco 12% (all organizations, mature committee) · IBM 32% (AI-breached organizations, any policy). Different constructs — presented as corroborating direction, NOT as a range.
⚖️ Shadow AI vs Your Own AI
Among breached organisations, shadow AI incidents were reported roughly twice as often as breaches of the organisation's own sanctioned AI models and applications — 43% against 21%.
Nathan House's Analysis — where the exposure sits
Among breached organizations, shadow AI incidents were reported 2.0x as often (43%) as breaches of the organization's own AI models or applications (21%). These are not mutually exclusive categories — an organization can appear in both. Method: 43% / 21% = 2.0x — both from IBM Cost of a Data Breach 2026, same 602-organization sample. Overlapping cohorts.
These categories are not mutually exclusive — an organisation can appear in both, and IBM is counting organisations rather than incidents. Even so: most AI security work goes into hardening the AI you deployed on purpose, while the unsanctioned side turns up in twice as many organisations. IBM does not measure where the budget went, so treat that as a prompt to check your own split rather than a finding about the market.
📋 Policy Exists. Compliance Does Not.
86% of PagerDuty's respondents work somewhere they believe has an AI policy. 66% used AI they believed that policy prohibited. Having the document is not the same as changing what people do.
What organisations have
- 86% believe their employer has an AI policy
- 32% of AI-breached orgs have an AI governance policy
- 40% use access controls on AI models and data
- 19% coordinate AI governance with security
What staff actually do
- 66% used AI they believed was not permitted
- 88% shared work information with public AI
- 34% entered customer data
- 31% shared financial or confidential documents
Both figures come from the same PagerDuty survey, but treat the comparison carefully. PagerDuty published no respondent-level cross-tabulation, so these are not demonstrably the same individuals, and the press release does not state the base for each question. Read them as two measurements of one surveyed population, not as a proven overlap — and if you need to cite the gap precisely, go to the full report for the bases rather than relying on this page.
🎯 What Security Teams Are Actually Worried About
Data leakage from GenAI is now the top AI-related concern in the World Economic Forum's 2026 outlook at 34%, ahead of adversarial AI capabilities on 29%. That is a reversal — the fear used to be attackers with better tools, not staff with a chatbot.
| Finding | Value | Source |
|---|---|---|
| Top AI concern: data leaks from GenAI | 34% | WEF Global Cybersecurity Outlook 2026 |
| Orgs highly concerned about AI misuse by insiders | 60% | Proofpoint |
| Orgs reporting GenAI-related security issues or breaches | 97% | Capgemini Research Institute |
| Workers entering sensitive data into unapproved GenAI | 33% | Gartner Top Cybersecurity Trends 2026 |
| Annual breach costs avoided with a formal insider risk programme | $8.2M | Ponemon Institute / DTEX 2026 Cost of Insider Risks Report |
The concern is broadly correct, but note how it sits against the spending. 60% of organisations report high concern about AI misuse by insiders, yet only 19% coordinate AI governance with their security team. Concern is cheap. Ponemon puts a number on the alternative: organisations running a formal insider risk programme avoided $8.2M in annual breach costs.
🏢 Embedded AI: The Part You Cannot Block
The applications driving AI traffic quadrupled to more than 3,400 in a year. A large share of that growth is AI features built into software you already bought and approved.
Zscaler's finding is blunt: embedded AI features "are often active by default and escape detection by legacy security filters". Atlassian was among the leading sources of embedded AI activity, reflecting AI inside Jira and Confluence. You cannot block those domains — they are the tools your business runs on.
Why this breaks the blocking approach
Domain blocking assumes the AI lives somewhere you can cut off. When AI is a feature inside Microsoft 365, Slack, Zoom, Jira or Confluence, there is no separate domain to add to a list — the traffic goes to the SaaS you already sanctioned. That is a hard ceiling on domain-level blocking specifically. It is not a ceiling on control in general: tenant-level admin settings, feature toggles, inspection-based DLP and API controls can all still reach these features. But none of those are a blocklist, and if your AI strategy is a blocklist, embedded AI is where it stops working.
🔐 Access Controls: 92% Did Not Have Them
Of organisations that suffered an AI-related breach, 92% lacked proper AI access controls (IBM 2026). Across IBM's surveyed organisations, only 40% use access controls on AI models and data at all.
IBM's X-Force analysis independently confirms both figures. The 2025 equivalent was 97%, so there is movement in the right direction — but from a very low base, and only 19% of organisations coordinate AI governance with their security team, which is the rarest control measured anywhere in this data.
🧠 Shadow AI Is an Insider Problem With a New Tool
75% of insider incidents are non-malicious, and 55% come from negligent or mistaken employees rather than bad actors (Ponemon/DTEX 2025). Ponemon does not separate shadow AI out, so we cannot say what share of those incidents involve AI. But the profile fits the behaviour every survey on this page describes: people trying to do their job faster, not people trying to cause harm.
Nathan House's Analysis — the reframe
75% of insider incidents are non-malicious and 55% come from negligent or mistaken employees rather than attackers. Ponemon does not break shadow AI out separately, so this is the shape of insider risk overall — but it is the category shadow AI most plausibly belongs to, and it argues for treating unsanctioned AI as an insider-risk problem rather than a new threat class. Method: Ponemon/DTEX Cost of Insider Risks 2025 — 55% negligent, 75% non-malicious overall
This matters because insider risk has been measured for years, and the trend was already ugly before AI arrived.
Nathan House's Analysis — the curve shadow AI joined
Annual insider-risk cost per organization rose 123% between 2018 and 2026, from $8.76M to $19.5M. Shadow AI arrives on top of a curve that was already climbing — it did not start it. Method: ($19.5M - $8.76M) / $8.76M = +123% over 8 years (Ponemon/DTEX Cost of Insider Risks series)
| Finding | Value | Source |
|---|---|---|
| Insider incidents that are non-malicious | 75% | Ponemon / DTEX 2025 |
| Insider incidents from negligent or mistaken staff | 55% | Ponemon Institute / DTEX 2025 Cost of Insider Risks Report |
| Average annual insider-risk cost per organisation | $19.5M | Ponemon Institute / DTEX |
| Average cost per insider incident | $676,517 | Ponemon / DTEX 2025 |
| Insider incidents involving cloud or SaaS | 78% | CrowdStrike / Cybersecurity Insiders |
| Orgs highly concerned about AI misuse by insiders | 60% | Proofpoint |
78% of insider incidents already involve cloud or SaaS platforms, and 60% of organisations report high concern about AI misuse by insiders. If you have an insider risk programme, shadow AI belongs inside it rather than in a separate initiative. Organisations with a formal programme avoided $8.2M in annual breach costs — that figure and the $19.5M annual cost come from the 2026 edition of the Ponemon/DTEX report, while the 75% and 55% splits come from the 2025 edition, which is the most recent to publish that breakdown.
⚠️ Shadow AI Statistics You Should Not Trust
These figures circulate widely. Three of them have no traceable source at all. The other two are real numbers that get restated as something they are not — which is the more common failure, and the harder one to spot. If you see any of them in a vendor deck, ask where it came from.
Check a Shadow AI Statistic
Pick a figure you have seen quoted and find out whether it survives a check against the original source.
| Claim | Why it fails |
|---|---|
| "78% of employees bring their own AI to work" | Misstatement. Microsoft's Work Trend Index figure is 75% of knowledge workers use AI at work — which says nothing about approval. |
| "89% reduction in unauthorised use when a sanctioned alternative is provided" | Untraceable. Attributed to a "Healthcare Brew survey" with no accessible document, no methodology and no sample size. |
| "Fewer than 11% of AI applications are visible to IT" | No primary source located, despite appearing on many vendor pages. |
| "47% access AI via personal accounts" (as a survey finding) | The 47% figure is real but it is LayerX telemetry measuring conversations, not a survey measuring people. We saw it misattributed to Netskope during research. |
| "One enterprise managed 67 unsanctioned AI tools" | An anecdote about a single company, repeated as if it were a population statistic. |
🧪 We Tested the Free Tools. Here Is What They Caught.
Every competing page on this topic cites the same vendor reports. None of them ran a test. We did, in August 2026, and the results are the only figures on this page that are not somebody else's research.
Nathan House's Analysis — tuning the free PII filter
We tested the leading free PII filter (Presidio) against 12 sensitive items in realistic prompts. Out of the box it caught 4. After we wrote custom UK and secrets recognisers it caught 10 — but false positives on legitimate business text rose from 3 to 8. Better recall is not free. Method: StationX original testing, August 2026: 12-item test set at a 0.5 confidence threshold, plus a separate PII-shaped stress set for false positives
Out of the box, Microsoft's Presidio missed API keys, database passwords, AWS keys, UK sort codes and UK postcodes. Writing custom recognisers for UK identifiers and secrets took it from 4 of 12 to 10 of 12 — but false positives on legitimate business text rose from 3 to 8. Better recall is not free, which is why we would run it in warn mode before blocking anything.
Nathan House's Analysis — the free blocklist ceiling
We measured six free AI blocklists against a vendor-sourced ground truth. The best managed 83% recall, two scored 0%, and stacking all five free lists together still caught only 19 of 26 domains. The highest-recall list does not block chatgpt.com — ChatGPT's actual consumer address. Method: StationX original testing, August 2026: recall measured with comm -12 against an 18-domain ground truth excluding any domain derived from the commercial sample
The blocklist result is the more damaging one. We measured six AI blocklists — five free, plus a teaser sample of a commercial feed — against a ground truth built from vendor documentation. The best free list caught 15 of 18 domains, or 83%. Two scored 0%: they target AI-generated websites rather than AI tools, which is not obvious from their descriptions.
Stacking all five free lists together caught 19 of 26 domains on our wider list, which includes eight endpoints only the commercial feed documented. Two denominators, deliberately: 18 is the fair test (scoring a commercial list against domains derived from itself would be circular), and 26 is the fuller picture of what is actually out there. Neither number is close to complete coverage.
The single most useful thing we found
The highest-recall free blocklist blocks chat.openai.com and api.openai.com — but not chatgpt.com, which is where OpenAI moved its consumer product. The best free list does not block ChatGPT's actual address. If you are relying on a free list, check that one domain today.
| Finding | Value | Source |
|---|---|---|
| Presidio out of the box | 4 of 12 | StationX original testing, August 2026 |
| Presidio after custom UK and secrets rules | 10 of 12 | StationX original testing, August 2026 |
| False positives on business text, before and after | 3 → 8 | StationX original testing, August 2026 |
| Best free AI blocklist recall | 83% | StationX original testing, August 2026 |
| All five free blocklists stacked together | 19 of 26 | StationX original testing, August 2026 |
| Lists scoring 0% recall (of the 6 tested) | 2 of 6 | StationX original testing, August 2026 |
| The best free list vs ChatGPT's real address | Misses chatgpt.com | StationX original testing, August 2026 |
⏱️ Shadow AI by the Minute
Annual totals are hard to picture. Broken down, the scale is easier to hold.
Nathan House's Analysis — data volume per minute
Enterprises sent 35 GB of data to AI tools every minute of 2025 — 49 TB a day, every day Method: 18,033 TB × 1,024 / 525,600 minutes per year
Nathan House's Analysis — policy violations per minute
ChatGPT alone triggered 780 data-loss-prevention violations every minute of 2025 — about 13 every second, including attempts to share Social Security numbers, source code and medical records Method: 410,000,000 violations / 525,600 minutes per year
🇬🇧 The UK Picture
71% of UK employees have used unapproved consumer AI tools at work, and 51% do so every week (Microsoft UK / Censuswide, n=2,003, October 2025).
| Finding | Value | Source |
|---|---|---|
| UK employees who used unapproved AI at work | 71% | Microsoft UK / Censuswide, October 2025 |
| UK employees doing so every week | 51% | Microsoft UK / Censuswide, October 2025 |
| UK employees using consumer AI for finance tasks | 22% | Microsoft UK / Censuswide, October 2025 |
| UK employees concerned about the data they enter | 32% | Microsoft UK / Censuswide, October 2025 |
Nathan House's Analysis — usage is high, concern is not
71% of UK employees have used unapproved AI tools at work, while only 32% say they are concerned about the privacy of company or customer data they put into them — 2.2x as many report doing it as report worrying about it. Censuswide did not ask whether these are the same people, and did not test any intervention, so this shows a gap between behaviour and stated concern rather than proving what would close it. Method: 71% / 32% = 2.2x (Microsoft UK/Censuswide, n=2,003, October 2025). Two separate cross-sectional response rates; no overlap data published.
22% used consumer AI assistants for finance-related tasks. Only 32% expressed concern about the privacy of company or customer data they entered, and 29% about the security of their organisation's systems.
Nathan House's Analysis — UK against the global figure
71% of UK employees have used unapproved AI tools at work against 66% in the four-country PagerDuty survey — a 5-point gap. Different samples and question wording, so treat this as two readings of the same problem rather than a like-for-like national comparison. Method: Microsoft UK/Censuswide n=2,003 UK employees vs PagerDuty/Wakefield n=1,250 across US/UK/AU/JP. NOT methodologically matched.
⚠️ Disclosure
Microsoft commissioned this research, and Microsoft sells Copilot — the governed alternative to the behaviour it measures. The methodology is stated and the sample is large, so the figures are usable. But you should know who paid for them.
🏛️ One Shadow AI Incident on the Public Record
Almost every figure above this line comes from a company selling security products. This one comes from a legally required disclosure instead.
On 11 May 2026, CB Financial Services, Inc. (NASDAQ: CBFV) filed a Form 8-K under Item 1.05, Material Cybersecurity Incidents. In the filing's own words, its subsidiary Community Bank had become aware of "an internal incident involving the handling of certain non-public customer information using an unauthorized artificial intelligence-based software application."
Read what the filing does and does not say. It calls the event an internal incident and says it did not involve disruption to operations, customer account access, payment systems or core IT — so no outage, and nothing describing an intrusion. It cites the volume and sensitive nature of the information as the reason the company judged the event material, and names what was disclosed: customer names, Social Security numbers and dates of birth, handled through an AI application the bank had not authorised. The filing does not name the application, say who used it, or explain how the data reached it, and the investigation was still ongoing when the 8-K was filed.
Why one filing outweighs a survey
An Item 1.05 is not research and not marketing. It is a legal disclosure a public company must make when it judges a cybersecurity incident material to investors — here, two days after the bank became aware. CB Financial is a commercial company like any other, but it had no reason to volunteer this and every incentive not to. That makes it a different kind of evidence from a vendor report, and the only disclosure of its kind cited on this page.
We searched EDGAR's full-text index for shadow AI in 8-K filings and found seven. Six are security vendors — Zscaler, CrowdStrike, Tenable, F5, JFrog and AvePoint — using the phrase as marketing language in earnings releases. Searches for “unsanctioned AI” and “unauthorized AI tool” return nothing at all. CB Financial's is the only filing we could find describing an actual incident, though EDGAR full-text search covers 2001 onward and a filing could always describe the same thing in different words.
Law firms covered it in May, correctly, as a disclosure-timing story. The operational lesson is narrower, and it is about where the risk sits rather than which control failed: this happened inside a regulated bank — an organisation with every reason to hold an AI policy, and with more compliance scrutiny than most businesses will ever face. The filing does not tell us which control was missing, and we should not pretend otherwise. What it does tell us is that the 43% figure at the top of this page is not an abstraction. It looks like this when it happens to one named company, and it reaches the SEC.
📐 How Do You Compare?
Four questions, answered in your browser. Nothing is sent anywhere.
How Do You Compare?
Answer four questions and see where you sit against the published rates. Nothing is sent anywhere — this runs entirely in your browser.
❓ Frequently Asked Questions
What percentage of companies have a shadow AI problem?
43% of breached organisations experienced a shadow AI incident in IBM's 2026 Cost of a Data Breach Report, up from 20% the year before. Note the precise wording: that is 43% of organisations that suffered a breach, not 43% of all breaches and not 43% of all companies. The distinction gets lost constantly in secondhand write-ups.
How much does a shadow AI breach cost?
$5.39M on average, against a $4.99M global average for all breaches (IBM 2026). That is roughly 8% higher. IBM compares cohorts rather than running a controlled experiment, so the figure shows an association, not proof that shadow AI caused the extra cost.
Is the $670,000 shadow AI figure real?
The number is real but almost universally misquoted. IBM's 2025 report compared organisations with high levels of shadow AI against those with little or none, and found a $670,000 gap between those two groups. It is not a surcharge added to an individual breach, which is how most articles present it.
How many employees use AI tools without permission?
66% of office professionals told PagerDuty they had used AI at work believing it was against company policy (n=1,250, Wakefield Research, four countries). In the UK specifically, Microsoft and Censuswide put it at 71%, with 51% doing so weekly (n=2,003).
Does blocking AI websites stop shadow AI?
Only partially, and less than most people assume. We measured six AI blocklists — five free, one a sample of a commercial feed — against ground truth built from vendor documentation. The best free list caught 15 of 18 domains (83%) and two scored 0%. Stacking all five free lists caught 19 of 26 on our wider list. The best-performing list does not block chatgpt.com, which is ChatGPT's actual consumer address.
Which AI tool receives the most company data?
Grammarly, not ChatGPT. Zscaler measured 3,615 TB going to Grammarly in 2025 against 2,021 TB to ChatGPT — roughly 1.8 times as much. The tool most policies name by default is not the largest destination.
Are AI governance policies becoming more common?
Going the wrong way. Among organisations that suffered an AI-related incident, 68% had no AI governance policy in IBM's 2026 report, up from 63% the year before — so the share that did have one fell from 37% to 32%, while shadow AI incidents more than doubled. Note the scope: this describes organisations that were already breached via AI, not the wider population.
Has any company officially reported a shadow AI breach?
Yes. CB Financial Services (NASDAQ: CBFV) filed an SEC Form 8-K under Item 1.05 on 11 May 2026, disclosing that its subsidiary Community Bank had handled non-public customer information — names, Social Security numbers and dates of birth — using what the filing calls an unauthorised artificial intelligence-based software application. The filing describes it as an internal incident and reports no disruption to operations or systems. It is the only 8-K we could find describing an actual shadow AI incident rather than using the phrase as vendor marketing.
Where do shadow AI statistics come from?
Almost entirely from companies selling security products. We searched the ICO, DSIT, ENISA, NIST, CISA, the ONS, Eurostat, the US Census Bureau, ISACA, ISC2, the IAPP, peer-reviewed literature and court filings, and found no government, regulator, statistics office or academic body publishing shadow AI prevalence data. Vendor telemetry is the only measurement that exists, which is worth knowing when you read any figure on this page.
🧮 About This Data
This article draws on 55 statistics held in our research database and tagged as shadow AI or directly adjacent to it. Every figure names the organisation that published it and links to that organisation's own page, not to a summary of it.
The primary sources are IBM Cost of a Data Breach 2026 (602 organisations, Ponemon Institute, 17 industries, 16 countries, fieldwork March 2025 to February 2026), PagerDuty's Shadow AI Survey (Wakefield Research, 1,250 office professionals across the US, UK, Australia and Japan, excluding IT roles), LayerX State of AI Usage 2026 (browser telemetry), Zscaler ThreatLabz 2026 (approximately 989 billion AI/ML transactions), Microsoft UK / Censuswide (2,003 UK employees, October 2025), Gartner, the World Economic Forum, Cisco and the Ponemon/DTEX Cost of Insider Risks series.
Statistics marked "Nathan House's Analysis" are derived — computed by cross-referencing entries in our database rather than quoted from a report. Each one shows its arithmetic so you can check it. Where a derived figure is our calculation rather than something the original publisher stated, we say so explicitly.
The Presidio and blocklist results are our own testing, run in August 2026. Method and full results are described in the section above.
Last updated: August 2026.
🚨 The most important caveat on this page
No government, regulator or national statistics office publishes shadow AI prevalence data. We searched the ICO, DSIT's Cyber Security Breaches Survey, ENISA, NIST, CISA, the ONS, Eurostat, the US Census Bureau, ISACA, ISC2 and the IAPP, and found none of them measuring it. Academia has barely started: the most substantial study we found is a 2026 Linnaeus University master's thesis (Kucelin and Vorkapic) which surveyed 176 valid cases and found 44.9% reporting bypass behaviour — useful, but a student thesis with a small sample, not a peer-reviewed national study. So effectively every figure with real scale behind it, including almost every figure on this page, comes from a company selling something adjacent to the problem. The one exception here is CB Financial's SEC filing — a legal disclosure, not research. That is not a reason to discard the rest. Vendor telemetry sees traffic no survey could reach. But it is the single most important thing to know when you read any shadow AI number, including ours.
About the Author
Nathan House, Founder & CEO of StationX
Nathan House has 30 years of hands-on cybersecurity experience and is Cambridge-educated, holding CISSP, CISA, CISM, OSCP, CEH, and SABSA. He founded StationX in 1999 — one of the UK’s first cybersecurity companies — and has secured £71 billion in UK mobile banking transactions and the London 2012 Olympics, advising clients including Microsoft, Cisco, BP, Vodafone, and VISA. He authored the world’s most popular cybersecurity course — a #1 Udemy bestseller taken by over 500,000 students — and was named Cyber Security Educator of the Year 2020, AI Security Educator of the Year, and a UK Top 25 Security Influencer 2025. A DEF CON speaker and featured expert on CNN, Fox News, NBC, and the BBC, Nathan leads StationX’s training of more than half a million students worldwide.