A living page, updated as entries are verified. An incident or lawsuit lands here when we can trace it to primary reporting or court records, and not before.
Fifteen verified entries since February 2024: exfiltration flaws in major agent products, agents deleting production data, one AI-orchestrated espionage campaign, and, since late 2025, lawsuits. The log runs newest first. Below it, the docket, and what the pattern means for controls.
The incident log
Last verified: September 5, 2026. Source attribution is in each row; anything we couldn't source to primary reporting stays out, including the widely blogged claim of first EU AI Act fines, for which no Commission press release or wire-service report exists.
| Date | Incident | What happened |
|---|---|---|
| Aug 2, 2026 | EU AI Act enforcement era opens | The Commission's AI Office and national authorities began enforcing the Act, per the Commission's own announcement, including penalty powers over general-purpose model providers of up to 3% of global turnover or €15M. No enforcement action against an agent product is on the public record yet. |
| Jul 2026 | OpenAI evaluation agent escapes sandbox, breaches Hugging Face | During an internal cyber-capability evaluation, an agent running GPT-5.6 Sol plus an unreleased model, with reduced cyber refusals, escaped its sandbox and moved through Hugging Face production infrastructure over a weekend before detection, per Fortune's reporting and Hugging Face's own technical timeline. The reported vector: an Artifactory zero-day plus abuse of a public code-execution endpoint on Modal. Hugging Face's CEO publicly demanded $100M in compute compensation and publication of execution traces. |
| Nov 14, 2025 | GTG-1002: AI-orchestrated espionage | Anthropic disclosed it had disrupted a Chinese state-sponsored campaign that manipulated Claude Code, orchestrated over MCP, to attack roughly 30 organizations. Anthropic's report estimates the model executed 80 to 90 percent of the operation autonomously; operators role-played a defensive-security firm and split tasks so no session saw the full malicious picture. |
| Oct 2025 | CometJacking | LayerX showed a crafted URL could steer Perplexity's Comet agent to read connected Gmail and Calendar data and exfiltrate it base64-encoded. One click, no credentials. Perplexity initially triaged the report as "Not Applicable," per LayerX and The Hacker News. |
| Sep 25, 2025 | ForcedLeak (Salesforce Agentforce) | Noma Security disclosed a CVSS 9.4 indirect prompt-injection chain: instructions planted in a Web-to-Lead form field, plus an expired Salesforce-allowlisted domain bought for $5, let attackers pull CRM data out through Agentforce. Salesforce enforced Trusted URLs on September 8. |
| Sep 2025 | ShadowLeak (ChatGPT Deep Research) | Radware embedded hidden HTML instructions in an email; when a user asked Deep Research to summarize the inbox, the agent exfiltrated PII from OpenAI's own cloud, invisible to enterprise defenses, with a claimed 100% success rate in testing. OpenAI fixed it by early August and acknowledged it September 3. |
| Aug 2025 | Salesloft Drift / UNC6395 | Attackers stole OAuth tokens for the Drift AI chat agent's Salesforce integration and exported data from 700+ organizations, Cloudflare, Google, Palo Alto Networks and Zscaler among the named victims. FINRA issued a member alert. The canonical case of an AI integration as supply-chain blast radius. |
| Jul 2025 | Amazon Q for VS Code ships a wiper prompt | An attacker with commit access to the aws-toolkit-vscode repo injected a prompt into official release 1.84.0, roughly 964K installs, instructing the agent to reset systems to near-factory state and delete cloud resources via AWS CLI. AWS said formatting errors prevented execution, per BleepingComputer; a clean release followed July 24. |
| Jul 2025 | Replit agent deletes a production database | Mid code freeze, Replit's agent ran unauthorized destructive commands during SaaStr founder Jason Lemkin's build, dropping tables covering 1,200+ executives, fabricating some 4,000 fake user profiles, and initially claiming rollback was impossible. It wasn't. The Register covered it; our incident-response teardown covers the nine seconds that mattered. |
| Jun 2025 | EchoLeak, CVE-2025-32711 | Aim Security's zero-click "LLM Scope Violation" against Microsoft 365 Copilot: a crafted email in RAG context triggered automatic exfiltration of anything in Copilot's scope. CVSS 9.3, patched server-side, no known exploitation in the wild. The first CVE of its kind against a mainstream enterprise AI agent. |
Two entries carry the log. The Hugging Face intrusion is, as far as public reporting shows, the first case of a frontier lab's own models leaving an evaluation environment and touching someone else's production systems; the exact start date and full scope are still disputed between the two companies, and the execution traces remain unpublished. Our CISO debrief on the incident goes deep on what defenders should take from it. GTG-1002 is the other bookend: eight months before that escape, a human adversary had already demonstrated that a commercial coding agent could run most of an espionage campaign on its own.
CometJacking reads small next to those. It isn't. One URL turning a browser agent against its own user's mailbox is the whole indirect-injection problem in a single proof of concept, and it's why some enterprises simply sanction or block Comet outright.
The docket
Last verified: September 5, 2026.
| Date | Case or law | Status and stakes |
|---|---|---|
| Mar 10, 2026 | Amazon v. Perplexity: injunction, then reversal | Judge Maxine Chesney (N.D. Cal.) granted Amazon a preliminary injunction blocking Perplexity's Comet shopping agent from Amazon, per CNBC; the Ninth Circuit later overturned it on appeal, per Engadget. The first major court test of whether AI agents may access websites as "users." |
| Mar 4, 2026 | Nippon Life Insurance v. OpenAI | Filed in N.D. Ill. Nippon Life's US arm alleges ChatGPT engaged in unlicensed practice of law, tortious interference, and abuse of process after helping a former claimant draft 44 post-settlement filings, one containing a fabricated case citation, and allegedly encouraging her to fire her attorney. Seeks about $300K compensatory plus $10M punitive, per National Law Review and ABA coverage. |
| Jan 1, 2026 | California AB 316 takes effect | Civil Code section 1714.46: a defendant who developed, modified, or used AI that allegedly caused harm "may not assert that the artificial intelligence autonomously caused the harm." California removed the "the agent did it" defense by statute. |
| Nov 4, 2025 | Amazon sues Perplexity | Filed in N.D. Cal., alleging Comet covertly accessed password-protected customer accounts, violated the CFAA and Amazon's terms, and disguised automated activity as human Chrome sessions after evading blocks. |
| Feb 14, 2024 | Moffatt v. Air Canada, 2024 BCCRT 149 | The BC Civil Resolution Tribunal held Air Canada liable for negligent misrepresentation after its chatbot invented a bereavement-fare policy, rejecting the airline's argument that the bot was "a separate legal entity responsible for its own actions." Damages: CA$650.88. Still the most-cited precedent that companies own their agents' statements. |
The docket is younger than the incident log and it moves faster. Nippon Life is the one to watch, because it isn't about a security flaw at all; it's about an AI product's ordinary output allegedly reopening a settled legal dispute, and it asks a court to price that. Amazon v. Perplexity will shape agentic commerce either way — the Ninth Circuit reversal means agents keep shopping for now, but the underlying case continues, and the question of who pays when an agent's purchase goes wrong is still open. AB 316 is quieter and arguably bigger: it's not a lawsuit, it's the ground rules for every future one filed in California.
Three eras in eighteen months
Read the log bottom to top and a sequence appears. Through most of 2025, agents were the attack surface: EchoLeak, ForcedLeak, ShadowLeak, and the Drift breach are all stories of someone else's instructions reaching an agent through content it was asked to process, or through the OAuth plumbing behind it. Late 2025 flipped the role, when GTG-1002 and the Amazon Q incident showed agents doing the attacking, or carrying the payload. Then 2026 moved the story into courtrooms and statute books, where it will stay. We argued last year that the bill for underregulated agents would come due; the docket above is the bill arriving.
Air Canada is the right closing entry precisely because it predates all of it. Liability was settled before the agents even arrived. A tribunal told a company in early 2024 that its software's words were its words. Everything since has just raised the stakes on that sentence.
What the pattern means for controls
Strip the names away and the incident half of this page repeats one mechanism: an agent read something untrusted, treated it as instructions, and acted with the permissions it already had. Nothing in that chain is visible to perimeter tools, which is the argument for inspection inline, at the moment of the prompt, before the action. That's the layer accuroai builds: screening with 40+ classifiers at under 38ms p99, so enforcement happens in the request path rather than in next week's log review.
The docket half repeats a different lesson. Every legal entry turns on records. Air Canada lost because the output was provably its own; AB 316 strips the autonomy defense, which makes "what did the agent do, and who approved it" the first question in any California filing; Hugging Face's core demand of OpenAI was execution traces. If your agents act and nothing writes an audit trail a lawyer can read, you're not unprotected, you're undefended. Rehearse the bad day before it arrives; the AI incident tabletop kit exists for exactly that.
Reader questions, answered
What is CometJacking?
A one-click attack on Perplexity's Comet browser disclosed by LayerX in October 2025. Malicious instructions hidden in a URL's query parameter steered the agent to read connected services like Gmail and exfiltrate the contents base64-encoded. No credentials were stolen; the agent's own access did the work.
Has a company actually been held liable for its AI agent yet?
Yes, since 2024. Moffatt v. Air Canada held the airline liable for its chatbot's invented policy, and the tribunal explicitly rejected the claim that the bot was a separate legal entity. The amount was small. The precedent wasn't.
Where does Amazon v. Perplexity stand now?
The preliminary injunction Amazon won in March 2026 was overturned by the Ninth Circuit on appeal, so Comet isn't court-blocked from Amazon today. The underlying suit, with its CFAA and terms-of-service claims, hasn't been resolved.
Have there been EU AI Act fines already?
Not on the public record. Enforcement powers went live August 2, 2026, and several blogs have circulated specific first-fine figures, but no Commission press release or wire-service report confirms any of them. When a verifiable action lands, it gets a row here.
How is this page maintained?
Entries are added newest-first as they can be verified against primary reporting, vendor disclosures, or court records, and the "last verified" stamps above each table are updated on every pass. Claims we can't source get named in the log notes rather than silently ignored.