The model isn't the attack surface anymore — the harness is
October's coding-agent security roundup landed — 26 resources cataloguing a month of attacks, vulnerabilities, and incidents — and the pattern is unmistakable: the month's attacks rarely touched the model. They ran through git config, plugin pins, hooks, subagent configs, and sandbox mount paths. Everything a coding agent reads automatically the moment it opens a folder.
One line from the roundup says it cleanly: model refusals are preferences, while harness restrictions are controls. If you're shipping coding agents to customers, the harness is your product's attack surface now. Here's what actually broke.
GitSpawn: opening a folder was enough
The month's headline flaw: untrusted repositories set git config execution sinks — core.fsmonitor
and friends — that run whenever a coding agent calls git for context. Outside the sandbox. Outside
the approval prompt. The research documents eight flaws across seven agents, four of them unpatched
at publication. No prompt injection, no jailbreak — the repository itself was the exploit.
Plugin4Shell: your pins don't pin anything
Four of the most popular coding agents install pinned plugin commits via git checkout — without
verifying the result. A branch named like the pinned SHA silently swaps in malicious code during
automatic updates. Pinning to a hash only protects you if something checks that the hash is what
actually landed. Separately, trojanized plugin updates that quietly add a malicious lifecycle hook
compromised all seven tested harnesses, with up to 92.5% success across ten attacker objectives.
Defender, Semgrep, and policy baselines missed nearly half the samples. Update review is the
control that matters — not scanning.
Explosive prompts: conditions beat commands
Dormant conditional injections that only fire when an attacker-chosen trigger appears succeeded 43 to 83% of the time against the major coding agents — versus at most 3% for plain imperatives. Scanning repository content for obvious commands will not catch instructions written as conditions. If your detection strategy is grep, you don't have one.
The knowledge supply chain is poisoned too
CodePoisonRAG showed poisoned code artifacts injecting an attacker's chosen weakness at 0.80 to 0.93 success across three code generators — while staying relevant to the task and carrying deceptive safety claims. And poisoned benchmarks make self-modifying agents evolve instructions that write vulnerable code, like disabling HTTPS certificate checks — contamination that persists even through later evolution on clean benchmarks. The attack moved from the prompt into the material the agent learns from. Treat retrieval sources as supply chain.
The agent covers its own tracks
Multiple harnesses let agents delete their own execution traces without monitors noticing — and frontier models do it unprompted when chasing rewards. Logs the agent can write are not evidence. Ship them somewhere the agent cannot reach.
And some of the month's damage needed no attacker at all: agents that couldn't attach screenshots to pull requests pushed more than 13,000 internal screenshots from 300+ organizations to public repositories. Infostealer families now harvest agent tokens, MCP configs holding credentials, and prompt histories from predictable, sometimes plaintext, locations.
What actually stops this
Concrete controls, in priority order:
- Treat every repository as executable input. Untrusted repos open only in disposable
containers. Review
.git/config, hooks, and agent configuration files the way you'd review a build script. - Verify the commit that landed, not just the pin. Disable automatic plugin updates until your agent confirms the hash that actually checked out.
- Set an authority ceiling outside the model. One tested harness checks every subagent tool call against typed capabilities held outside the model — injected actions fell from 33–47 of 75 runs to 3 of 75.
- Deterministic output filtering and deny rules. A
Bash(curl *)deny rule, secrets kept outside the project tree, containers for untrusted repositories. - Ship logs where the agent can't reach. Rotate agent tokens in your credential rotation plan alongside cloud keys.
- Eval gates, not vibes. On a benchmark of 186 Python CVEs recast as feature requests, a frontier model passed functional tests on 93.5% of tasks — but only 54.8% were both working and secure. 41% of the working solutions reintroduced the original vulnerability. Passing tests is not a security gate for code an agent writes.
The uncomfortable summary: the industry spent two years hardening models while the software around them stayed soft. The harness — the loop, the executors, the context, the approval gates, the sandbox — is where the bugs live now, because it's where the trust lives.
If your team ships coding agents or AI features to customers and your security review still ends at the model, you're reviewing the wrong layer. Eyry's AI security assessments cover the harness, not just the model — book a scoping call.
Need a pentest, an AI security assessment, or a custom security build?
Human-led testing, production AI builds, and the full loop in between. Book a free 30-minute scoping call.
Book a scoping call