Every AI governance pilot we run starts the same way: a security leader who believes their organisation has "a handful" of AI tools in use, a deployment that takes under thirty minutes through the MDM and identity provider they already have, and then seventy-two hours in which that belief is replaced by a number. The number is never a handful.
This is what those three days look like, hour by hour, with the published 2026 benchmarks beside each finding so you can calibrate your own result. None of the figures below are ours; they come from Check Point, Netskope, Cyberhaven, Harmonic, Zscaler, GitGuardian, 1Password, AvePoint, the Verizon DBIR and IBM. Our contribution is the order in which you will meet them.
Hours 0–24: the inventory
What you'll find: far more AI than you sanctioned, and about half of it on personal accounts.
Discovery runs against a catalog of 1,400+ AI applications and agents across browsers and managed endpoints. By the end of day one you have a first complete inventory with owners, first-seen dates and risk scores. The benchmarks say what to expect. Check Point's July 2026 report found the average organisation adopting around ten AI applications a month, "many without formal approval", and its July data put the average at eight GenAI tools in active use. The long tail is the surprise: Cyberhaven found cautious enterprises using fewer than fifteen GenAI tools and the top 1% of adopters more than 300, and Harmonic observed 665 distinct GenAI and AI-embedded tools across its customer base in 2025. Zscaler saw more than 3,400 applications generating AI traffic on its platform, four times the year before.
Then the account split. Netskope's Cloud and Threat Report 2026 found 47% of GenAI users still on personal, unmanaged accounts — down from 78% a year earlier, but nearly half. The Verizon DBIR 2026 found 67% of AI users on corporate devices using non-corporate accounts; LayerX put personal-identity conversations at 47%. Cyberhaven broke it down by tool: 32% of ChatGPT use, 25% of Gemini, 58% of Claude and 61% of Perplexity through personal accounts. Harmonic's May 2026 usage index adds the detail that stings: 64.5% of activity on personal accounts is business work, and on enterprise-licensed plans 45.6% of activity is personal. Your licences are underused and your business data is on the free tier.
Day-one output: the inventory, the sanctioned-versus-personal ratio, and the list of tools nobody in security had heard of.
Hours 24–48: what is in the prompts
What you'll find: a few percent of prompts carrying the data you least want to lose — concentrated in source code, legal material and financial figures.
With inspection running inline — 40+ classifiers and 60+ secret types at under 38 ms p99, so nobody notices — day two classifies what is actually being sent. State your denominator before you read the benchmarks, because the vendors measure differently. At prompt level, Harmonic found 2.6% of 22.4 million enterprise prompts contained company-sensitive data; LayerX puts it above 6% of conversations; Check Point measured 4% of prompts as high-risk, "one in every 25", rising to one in seventeen in business services, and in its July 2026 data 22% of prompts contained potentially sensitive information. Cyberhaven's figure — 39.7% of data movements into AI tools involve sensitive data — counts pastes and uploads, not prompts, which is why it is higher.
The classes are consistent across sources. Harmonic: legal content 35%, code and technical IP 26.5%, financial 16.6%, M&A 12.6% — "code, legal documents, and financial data comprise 74.5%" of exposures. Check Point's July data, by share of organisations affected: personal data 70%, financial 68%, network and IT infrastructure 68%, legal and regulatory 63%, HR 62%. The DBIR puts source code first "by a significant margin". Count the policy events too: Netskope's average organisation records 223 GenAI data-policy violations a month, the top quartile 2,100.
Day two is also when the redact-versus-block decision gets made with evidence rather than instinct. Most pilots start in observe mode, see the 3–4%, and switch the top data classes to redact by day three — the employee keeps working, the value stays home, the violation count starts falling.
Day-two output: the sensitive-prompt rate by data class, team and tool; the first redactions; the baseline violation count for the board.
Hours 48–72: the developers and the agents
What you'll find: coding assistants in half the engineering org, secrets in their context, and agents already running that nobody counted.
Endpoint discovery reaches the surfaces the browser cannot — desktop apps, IDEs, local models, CLI agents and MCP servers in configuration files. Cyberhaven measured developer coding-assistant adoption at 49.5% by December 2025 and a peak of 60.5% by June 2026, with nearly 90% at leading adopters and 30% of users running two or more tools. GitGuardian's 2026 secrets report found AI-assisted commits leaking secrets at roughly twice the baseline rate — 3.2% against 1.5% — and 24,008 unique secrets in MCP-related configuration files on public GitHub, 2,117 of them still valid; that is a single source, but it matches what shows up in day-three scans. Nobody has published a credible "MCP servers per developer" figure, so we will not either; expect to find them, and expect several to have no owner.
Then the agents. 1Password's July 2026 survey found 46% of developers running agents in production and 33% of IT and security staff unable to say how many agents or non-human identities exist; 41% of organisations had agents reaching unapproved data. AvePoint found 46.9% of employees using agents daily or weekly and 21.1% of organisations unable to account for unsanctioned agent activity. Cyberhaven's mid-year update recorded Claude desktop-agent adoption up 1,233% in six months. The inventory you built on day one grows again on day three, and the new entries have tool access.
Day-three output: the developer-surface inventory, the secrets-in-context count, the agent and MCP server list with owners and scopes — and the first kill-switch test on an agent someone had forgotten.
Hour 72: the report
Three days produce the executive risk report most organisations have never had: the AI estate by surface, the sanctioned ratio, the sensitive-prompt rate and violation baseline, the developer and agent findings, and the governance gap in the same terms IBM used in July — 68% of breached organisations lacked AI governance, 92% of those with AI-related breaches lacked access controls, shadow AI in 43% of incidents at $5.39M each. Against that, a before-and-after: violations falling from the day redaction switched on, and a personal-account share that has started to move. Across pilots the standing result is a 94% reduction in shadow-AI usage within 30 days, because once the sanctioned tool is safe to use people stop needing the other one.
What the pilot does not do is decide anything for you. It replaces estimates with counts. The decision — which tools to sanction, which data classes to redact, which agents to keep — is yours, and by hour 72 it is a decision about evidence rather than a debate about instinct.
What to bring to hour zero
- Your MDM and identity provider admin — deployment is under thirty minutes and needs both.
- Your current AI policy, so the classifiers map to the data classes it names.
- Your list of "known" AI tools, to measure the gap.
- One engineering team willing to be the day-three cohort.
- The board date, because the report is shaped for it.
FAQ
How much of the estate does a 72-hour pilot cover?
Browser and managed endpoints for the groups in scope, usually a few hundred to a few thousand users. The findings scale; the ratios — sanctioned share, sensitive-prompt rate, agents per developer — hold across the organisation in our experience and match the published benchmarks above.
Will employees notice?
Inspection runs under 38 ms at p99 and starts in observe mode. Employees notice when redaction begins, and typically only when it catches something they are glad it caught.
Why do the sensitive-data percentages differ so much between reports?
Denominators. Harmonic (2.6%) and Check Point (4%) count prompts; LayerX (6%+) counts conversations; Cyberhaven (39.7%) counts data movements including pastes and uploads. The pilot reports all three so the board is not confused by the next vendor deck.
What happens after day three?
The pilot evidence is portable: you keep the inventory, the baseline and the report whether or not you continue. Most organisations move the pilot cohort into enforcement and extend discovery to the rest of the estate.
Sources: Check Point Research, AI Security Report 2026 (14 Jul 2026) · Check Point, July 2026 threat update (12 Aug 2026) · Netskope Cloud and Threat Report 2026, via Infosecurity Magazine (7 Jan 2026) · Cyberhaven 2026 AI Adoption and Risk Report (5 Feb 2026) · Cyberhaven on sensitive data in AI tools (11 Feb 2026) · Harmonic Security, 22 million prompts (15 Jan 2026) · Harmonic Security AI Usage Index (20 May 2026) · Zscaler ThreatLabz 2026 AI Security Report (27 Jan 2026) · LayerX State of AI Usage Report 2026 · Verizon DBIR 2026, via National Law Review (28 May 2026) · GitGuardian State of Secrets Sprawl 2026 (17 Mar 2026) · 1Password agent survey (28 Jul 2026) · AvePoint State of AI 2026 (29 Jun 2026) · IBM Cost of a Data Breach 2026, via Cybersecurity Dive (29 Jul 2026).
Related: Endpoint AI Governance · Redact, Don't Block · The AI Governance KPI Pack · Start a 72-hour pilot.