Skip to main content
All editions

Newsletter

One GitHub Issue. And the Agent Can Pay.

7 min read

This month, a single opened issue could hijack every repository running Claude Code’s action. The uncomfortable part isn’t the bug. It’s that 1,578 agent tool-servers can now move money — and the exploit shipped in a tool description.

In January, a security researcher found that one GitHub issue — opened by anyone, on any public repository running Anthropic’s Claude Code action — was enough to hijack it. The agent read the issue, traded its OIDC token for a write-scoped application token, and handed the attacker the keys to the code, the issues and the workflows. Because Anthropic’s own action used the same pattern, a working attack could have pushed malicious code into the action itself and down onto every project pulling it. Anthropic fixed the core bypass in four days and hardened it through the spring; the patched version is v1.0.94. The public disclosure landed this month. Treat it as a fire drill, not a fire.

Because the interesting number isn’t the CVE. In the same window, researchers counted 1,184 malicious skills live on ClawHub, a universal shell-injection design flaw (GuardFall) sitting under more than half a million open-source deployments, and 492 agent tool-servers exposed to the open internet with no authentication at all. None of these attacked a model. Every one attacked the tool chain wrapped around it. And the tool chain, increasingly, can spend.

The Attack Surface Moved While You Watched the Model

For two years, “AI security” meant the model — jailbreaks, prompt injection at the prompt, data leaking out of weights. Enterprises stood up red teams, bought guardrails, wrote model cards. Meanwhile the exploitable surface migrated one layer out, into the connective tissue: the Model Context Protocol servers, the skills, the actions, the tool descriptions the agent reads to decide what to do next. Microsoft’s AI Red Team refreshed its taxonomy this year to name two failure modes outright — supply-chain compromise and excessive agency. Phoenix Security’s mid-year data says the first half of 2026 already produced 2.6 times the campaign volume and 4.5 times the package-compromise volume of all of 2025. IBM’s X-Force puts supply-chain and third-party compromise at roughly four times its 2020 level. The model was never the soft target. The tools were.

Why a Tool Description Is Now an Instruction

The mechanism matters, because it decides where your controls belong. An agent chooses what to do by reading the descriptions of the tools available to it. Poison that description — inside a package, a config file, a marketplace skill, a remote MCP server — and you are not exploiting a bug. You are handing the agent an instruction it believes is legitimate, and it will follow that instruction silently, in every session, for every user, until someone happens to notice. No memory corruption, no payload in the classic sense. The exploit is simply text the agent was built to trust. That is what makes this a supply-chain problem and not a patch-Tuesday problem: the compromise arrives on the very same channel as the capability.

The Blast Radius Now Includes the Money

Now put the two curves on one chart. In January 2025 there were 47 MCP servers with payment-execution capability. By February 2026 there were 1,578 — a thirty-four-fold rise in eighteen months. On Base alone, agents have already settled roughly 169 million transactions. The NSA has flagged MCP security gaps precisely as banks begin wiring agents into payment systems. So the agent reading an untrusted tool description is, more and more often, the same agent holding a payment credential that can move a stablecoin with no human in the loop.

Regular readers know the shape of this. In Edition 51 the compromise was the gateway (AGCR-D). In Edition 56 it was the commit leg, where a payment becomes irreversible (PACT-D). In Edition 59 it was the settlement asset carrying its own legal regime across a two-issuer bridge (SRX-D). Edition 60 is where those threads meet: a compromised agent with payment authority does not leak data, it spends — and on rails engineered for finality, the spend can be unwind-proof. Model risk asks whether the agent can be fooled. This asks what happens to your money when it is.

The Regulator Is Already Standing at the Blast Site

This is not a hypothetical you get to schedule for 2028. Three live regimes already sit on top of it. DORA’s grace period is over: enforcement is active, the ESAs have designated nineteen critical ICT providers — AWS, Azure and Google Cloud among them — under direct oversight, and the first compulsion payments are being issued. If your paying agent runs on one of those clouds, its blast radius is a supervised operational-resilience surface. PSD3 and the PSR — final compromise texts published in April, Official Journal publication imminent — rewrite fraud liability around a human who is “manipulated” or “grossly negligent”; there is no actor category for an agent that authorises the payment itself, so the liability question is open, not settled. And from 2 August, the AI Act’s enforcement powers over general-purpose models switch on, with national authorities able to investigate and sanction. Three regulators, one incident surface — and most architecture diagrams don’t show the agent’s reach at all.

ABR-D — The Agent Blast-Radius Diagnostic

So here is the instrument. Blast-radius risk asks a harder question than model risk: when — not if — one agent is compromised, how far does the damage travel before something stops it? Score every autonomous agent you run on five axes, from 1 (contained) to 5 (unbounded).

Credential & Identity Scope (CIS). What can this agent’s credentials touch if stolen? One repository, or the whole organisation? A single ledger, or the payment rail? Inherited human credentials with full scope are the worst case, and the most common.

Tool-Supply-Chain Integrity (TSI). Where do the agent’s tools, skills and MCP servers come from, and who verifies them? Unpinned marketplace skills and unauthenticated remote servers are open doors — 492 of them stood open this month.

Payment-Execution Authority (PEA). Can this agent move money, and up to what limit, without a human? An agent with unbounded PEA and weak TSI is a wire transfer waiting for the right tool description to arrive.

Human-in-the-Loop Gating (HLG). At what value — or what degree of irreversibility — must a human approve? If the gate is “never,” or “only above a threshold nobody deliberately set,” you do not have a gate.

Settlement Reversibility & Finality (SRF). If the payment is wrong or fraudulent, can it be clawed back — and under whose law? On finality-engineered stablecoin rails the honest answer is often “no,” which is why this axis inherits the PACT-D and SRX-D questions directly.

Read the Number, Then Redraw the Diagram

Add the five. Five to nine: contained — a compromise hurts, but it stops. Ten to seventeen: reachable — one weak axis (usually TSI or HLG) away from a bad day. Eighteen to twenty-five: unbounded — a single poisoned tool description reaches your money and your regulator in the same motion. And any single axis at five is a board-level finding regardless of the total, because blast radius does not average; it takes the widest path it can find.

The reason this is so hard to see from inside is that no single owner holds it. Security owns the model and the perimeter. Platform owns the MCP servers. Treasury owns the payment authority. Nobody owns the line that runs from an untrusted tool description, through an over-scoped credential, to an irreversible settlement on a supervised cloud. Drawing that line — and deciding where to cut it — is not a security task or a payments task or a compliance task. It is an architecture task: the enterprise-architecture question of what any one component is permitted to reach. That cross-boundary call, made by someone with no stake in whose budget the agent sits under, is exactly what a fractional enterprise architect exists to make.

You secured the model. Good. Now find out how far the tool underneath it can reach — before someone opens an issue and finds out for you.

Hawk Nest Newsletter is written by Paulo Falcao. For twenty-five years, helping organisations turn complex technology challenges into measurable business outcomes — payments systems, enterprise architecture, AI, technology. The intersection of strategy and architecture, converted into reliable, revenue-generating reality. ABR-D joins the IP portfolio next to SIRM, AVAEM, SHAD, ACAM, SAVED, GAIA-D, AGCR-D, AASI, SSV, ATOM, PVC, PACT-D, RCS-D, AGP-R, and SRX-D.

  • security
  • AI
  • payments
  • enterprise architecture

Have a similar challenge?

Book a 30-minute call to talk through AI governance, architecture or payments — no pitch, just a senior second opinion.

Book a 30-min call