Blog topic
AI Agent Security: Prompt Injection, Worms and Supply Chain
AI agents became infrastructure: they write code, move tickets, query production. And they brought an attack surface no classic risk matrix covered: where data read becomes executed instruction. This collection gathers our field coverage: prompt injection escalating to RCE, worms spreading across frameworks without a click, skill and MCP supply chains under attack, and the guardrails that are starting to actually work. If your company runs Claude Code, Copilot or any tool-connected agent, start with the first article and follow the trail.
20 articles in this track
- 01
Prompt Injection Is the New SQL Injection, Now It Leads to RCE
73% of AI deployments have prompt injection. Chatbots leak data via markdown rendering. Semantic Kernel enables RCE via Startup folder. Data and defenses.
5/28/2026 · 5 min read
- 02
Prompt Injection Is State Poisoning, Your Agent Is Exposed
CoT Forgery and Trojan Hippo prove prompt injection poisons internal model state. The security boundary is in the wrong place. Here is what changes.
5/27/2026 · 4 min read
- 03
Agentjacking: Hijacking AI Coding Agents via Sentry MCP
Attack discovered by Tenet Security hijacks AI agents via Sentry MCP with 85% success rate. 2,388 orgs with exposed DSNs. Open-source mitigation available.
6/26/2026 · 5 min read
- 04
313 CVEs in the MCP Ecosystem: The Protocol Connecting AI Agents to Tools Has an Architectural Problem
The Model Context Protocol has 313 indexed CVEs in 18 months. The root cause is architectural: there is no separation between data and control. Tool descriptions are processed as instructions by the LLM, with the same authority as system prompts.
8/4/2026 · 4 min read
- 05
SearchLeak: P2P Injection + Bing SSRF Exfiltrates M365 Copilot Data
CVE-2026-42824 discovered by Varonis enables data exfiltration from M365 Copilot Enterprise via prompt injection in the search parameter and SSRF through Bing. One click is enough.
6/17/2026 · 5 min read
- 06
LangGraph: From SQL Injection to RCE Through AI Agent Memory
Check Point Research documented an exploitation chain from SQL injection to remote code execution through the LangGraph checkpointer, the persistence layer that gives the agent memory is the same one that gives the attacker persistence.
6/15/2026 · 4 min read
- 07
n8n-mcp: CVSS 9.9 IDOR Exposes All Tenants' Credentials
CVE-2026-54052 (assigned by Manifold Security; pending NVD publication) in n8n-mcp lets any authenticated tenant read every other tenant's credentials. Fifth multi-tenant security issue in 2026, the persistence layer was never isolated.
6/17/2026 · 4 min read
- 08
PocketOS: The Agent That Deleted Production, the IAM That Failed, and the Confession That Shouldn't Matter
A Claude Opus 4.6 agent in Cursor deleted PocketOS's production volume and backups, causing a ~30 hour outage. The root cause was not the model, it was IAM. Why system prompts are advisory and RBAC is enforcement.
7/9/2026 · 4 min read
- 09
LLM Agent Worms: Zero-Click Propagation Across Frameworks
The first autonomous worm propagating between LLM agents without human interaction. Zero-click, cross-platform, 3 hops. Defense requires a formal theorem.
6/3/2026 · 5 min read
- 10
LLM Self-Replication Worm: From 6% to 81% in One Year
Palisade Research documented the first LLM self-replication worm: 4 hops, 3 continents, zero human intervention. Success rates jumped from 6% to 81% in 12 months.
5/28/2026 · 5 min read
- 11
Miasma Worm: When Trust Infrastructure Becomes the Attack Itself
How the Miasma Worm exploited SLSA, Sigstore, and legitimate credentials to compromise npm, PyPI, and Microsoft repositories in 7 days, without a single CVE.
6/12/2026 · 4 min read
- 12
Miasma and IronWorm: When the Supply Chain Learns to Infect Your AI Agent
Two supply chain worms in June 2026 rewrote AI agent instructions and installed eBPF rootkits. Technical analysis of Miasma and IronWorm, and how to protect your agents.
6/27/2026 · 5 min read
- 13
PoisonedSkills: Skill Docs That Make AI Agents Run Malware
PoisonedSkills uses skill documentation to execute payloads in AI coding agents via DDIPE. 33.5% bypass rate. 4 CVEs. Skill registries are the new supply chain.
6/3/2026 · 4 min read
- 14
MemPoison + MCFA: The Memory Attack Surface in LLM Agents
Memory attacks on LLM agents reach 95% success. MemPoison poisons memory, MCFA hijacks control flow. Current defenses are insufficient.
6/3/2026 · 4 min read
- 15
Malicious LLM API Routers: The Invisible Threat Inside Your AI Agents
428 routers tested, 9 injecting code, 1 draining Ethereum. How malicious LLM API routers compromise AI agents without detection.
6/5/2026 · 3 min read
- 16
Claude Code in Your Pipeline: The Structural Hole and Rule of Two
Claude Code GitHub Action exposes credentials via unsandboxed Read tool. Microsoft steals keys in two steps. RyotaK: 50 bypasses. The Rule of Two you need to adopt.
6/6/2026 · 4 min read
- 17
Claude Code: 6 Disclosures, 1 Real Attack, and the Architectural Disease
A structural vulnerability in the Claude Code GitHub Action produced 6 separate disclosures, 1 source code leak, and 1 real supply chain attack. The patch covers a symptom, the disease is architectural.
6/17/2026 · 5 min read
- 18
Autoguardrails: Karpathy's Autoresearch Transposed to AI Safety
Santander AI Lab turned Karpathy's autoresearch into autoguardrails: an agent that searches over policy.md to minimize Attack Success Rate without destroying benign pass. Same turnstile, different metric.
7/3/2026 · 4 min read
- 19
Mechanical Governance for LLMs: The mech-gov-framework and the EU AI Act
27% of LLM deferrals carry zero decision-relevant information. Santander's mech-gov-framework solves this with 4 mechanical primitives, with direct implications for the EU AI Act.
6/27/2026 · 4 min read
- 20
NVIDIA SkillSpector: The Scanner That Proved AI Skills Security Needs More Than Static Analysis
26% of AI skills contain vulnerabilities and 5% are likely malicious, according to Liu et al. NVIDIA launched SkillSpector, but Trail of Bits proved static scanners can be bypassed in under 1 hour.
6/26/2026 · 4 min read
Frequently asked questions on this topic
In SQL injection you separate data from commands with parameterized queries. In prompt injection, natural language IS the command: there is no "pure data". A document, email or web page read by an agent can carry instructions it executes as if they were yours. Consequences range from context exfiltration to full RCE when the agent has tool access.
Yes: that is the LLM agent worm category: malicious content written by one agent infects the next agent that reads it, in a chain. Research demonstrated zero-click propagation across popular frameworks and self-replication rates growing from 6% to 81% in a year. The vector is shared memory and context, not a network vulnerability.
The three vectors with the most real findings: (1) MCP servers and plugins holding broad credentials: a CVSS 9.9 IDOR and hundreds of published CVEs; (2) CI/CD pipelines where the agent reads external input and executes commands; (3) skill/tool marketplaces where the documentation itself carries the payload. Rule of thumb: least privilege per tool, sandbox per action, human review for irreversible operations.
It does: SearchLeak proved it: injection via search results reached M365 Copilot data with zero custom plugins. Corporate assistants aggregate email, files and web into the same context; one poisoned source contaminates all of them. Governance (who can connect what) and training reduce risk even without customization.