Skip to content
Deployment & Production

45% Of AI-Generated Code Fails A Security Scan: Provenance, Risk Tiers And The In-VPC Deployment Pattern For Shipping Agent-Written Code In A Regulated Pipeline

Veracode's testing of more than 100 language models across 80 real tasks found 45% of generated code introducing OWASP Top 10 vulnerabilities - a pass rate that its Spring 2026 update shows has barely moved from the first report, with Java failing 72% of the time. In the same year 89% of platform engineering teams use AI daily, only 69.7% have a governance policy, 59% cannot quickly shut an AI system down and 68% cannot distinguish an agent's actions from a human's. A trading firm's security team found an auto-updated policy had let non-compliant operation run for two weeks; a Fortune 100 institution's agent deleted a non-compliant service, audit logs and all. With the EU AI Act's high-risk obligations live since August and DORA and SR 11-7 already in force, the question is not whether to use coding agents in finance. It is how to make their output pass the same gates as a human's. Here is the pipeline.

AlchmAI Engineering14 min read

45%

Of AI-generated code samples failed security testing and introduced OWASP Top 10 vulnerabilities, across 100+ models and 80 tasks (Veracode)

72%

Security failure rate for Java, the riskiest language tested - and the language of a great deal of bank middleware

59% / 68%

Platform teams that cannot quickly shut an AI system down, and that cannot tell an agent's actions from a human's

2 weeks

How long a trading firm ran non-compliant after an auto-updated agent policy enabled it - discovered by the security team, not the pipeline

The most useful number in AI-assisted software engineering has been stable for a year, and its stability is the point. Veracode tested over 100 large language models across Java, JavaScript, Python and C# on 80 real-world coding tasks, ran the output through production-grade static analysis, and found 45% of samples failing security tests with OWASP Top 10 vulnerabilities. Java failed 72% of the time. The Spring 2026 update found the average pass rate at 56%, barely changed from 55% in the first report - so as models got dramatically better at producing code that works, they got no better at producing code that is safe. Teams report generating code four times faster and ten times more security findings.

That would be a manageable engineering problem if the code were arriving through the same gates as human code. Increasingly it is not. A survey of platform engineering teams found 89% using AI daily but only 69.7% with a governance policy; 84% call AI governance a serious concern; 59% cannot quickly shut an AI system down in an emergency; and 68% cannot distinguish an AI agent's actions from a human's in their own systems. The examples from regulated firms are the instructive part: a Fortune 100 financial institution's agent deleted a non-compliant service, including its audit logs, and triggered a compliance investigation; a trading firm's security team discovered that auto-updated agent policies had inadvertently enabled non-compliant operation for two weeks.

Principle One: Provenance Is A Build Artefact

The 68% who cannot tell an agent's action from a human's have a data problem before they have a governance problem. Every change that reaches a regulated pipeline must carry a provenance record - which agent, which model snapshot, which harness configuration, which prompt or task, which human directed it and which human approved it - attached to the commit and the build, not stored in a chat log. We generate it as a signed attestation in the same family as supply-chain attestations, because that is the infrastructure regulated pipelines already verify.

typescriptprovenance/attest.ts
import { createHash, sign } from "node:crypto";

/**
 * One attestation per agent-authored change. Attached to the commit as a note
 * and to the build as an artefact; verified by the pipeline before any gate
 * that depends on risk tier. The point is that "who wrote this" becomes a
 * query against signed data, not an archaeology exercise in a chat history.
 */
export interface AgentChangeAttestation {
  schema: "alchmai.dev/agent-change/v1";
  commit: string;
  author: {
    kind: "agent";
    agent: string;                 // e.g. "claude-code"
    modelSnapshot: string;         // dated snapshot, never a floating alias
    harnessConfigHash: string;     // policy.yaml + AGENTS.md + tool registry
    taskId: string;                // links to the ticket and the full trace
    promptHash: string;            // hash, not content - the content stays in the trace store
  };
  humans: {
    directedBy: string;            // who set the task
    reviewedBy: string[];          // who approved, must be non-empty for tier >= 2
  };
  riskTier: 1 | 2 | 3;             // see classify()
  toolCalls: { count: number; refused: number; monitorVerdictSummary: string };
  scans: { sast: string; sca: string; secrets: string };   // result refs, not booleans
  createdAt: string;
}

export function attest(a: Omit<AgentChangeAttestation, "schema" | "createdAt">, signingKey: string) {
  const body: AgentChangeAttestation = { schema: "alchmai.dev/agent-change/v1", createdAt: new Date().toISOString(), ...a };
  const canonical = JSON.stringify(body, Object.keys(body).sort());
  const digest = createHash("sha256").update(canonical).digest();
  return { body, signature: sign(null, digest, signingKey).toString("base64") };
}

Principle Two: Risk Tiers Decide The Gate, Not The Author

The instinct in regulated firms is to review agent-written code harder than human-written code. That is the wrong axis. The 45% finding is about the code, and the gate should be about what the code touches. A change to a logging format and a change to the pre-trade risk gate deserve different scrutiny regardless of who - or what - wrote them. We classify every change into three tiers from the paths it touches and the resilience service map, and the tier decides the gates.

pythonpipeline/classify.py
from dataclasses import dataclass
from fnmatch import fnmatch

TIER_3 = ["risk/**", "oms/**", "surveillance/**", "compliance/**", "payments/**", "auth/**"]
TIER_2 = ["marketdata/**", "reporting/**", "onboarding/**", "infra/**", "*.tf", "policy/**"]

@dataclass(frozen=True)
class Gates:
    sast_block_on: str           # "high" | "medium" | "any"
    human_reviewers: int
    require_change_ticket: bool
    require_parity_harness: bool
    require_eval_gate: bool
    require_smf_signoff: bool     # accountable Senior Manager, UK SM&CR

def classify(changed_paths: list[str]) -> tuple[int, Gates]:
    """Tier is a property of what the change TOUCHES, independent of author.
    An agent and a human editing risk/gate.py face the same gates."""
    if any(fnmatch(p, g) for p in changed_paths for g in TIER_3):
        return 3, Gates("any", 2, True, True, True, True)
    if any(fnmatch(p, g) for p in changed_paths for g in TIER_2):
        return 2, Gates("medium", 1, True, False, True, False)
    return 1, Gates("high", 1, False, False, False, False)

def enforce(attestation: dict, gates: Gates, sast_findings: list[dict]) -> list[str]:
    """Returns the reasons a change may not merge. Empty means it may."""
    reasons = []
    worst = max((f["severity"] for f in sast_findings), default="none")
    order = ["none", "low", "medium", "high", "critical"]
    if gates.sast_block_on == "any" and sast_findings:
        reasons.append("SAST_FINDINGS_PRESENT")
    elif order.index(worst) >= order.index(gates.sast_block_on) and worst != "none":
        reasons.append("SAST_SEVERITY:" + worst)
    if len(attestation["humans"]["reviewedBy"]) < gates.human_reviewers:
        reasons.append("INSUFFICIENT_HUMAN_REVIEW")
    if gates.require_change_ticket and not attestation["author"]["taskId"]:
        reasons.append("NO_CHANGE_TICKET")
    if attestation["author"]["modelSnapshot"].endswith("-latest"):
        reasons.append("FLOATING_MODEL_ALIAS")     # unreproducible; refuse at any tier
    return reasons

Principle Three: The Agent Runs Inside Your Perimeter

The architectural shift of 2026 is that the most defensible AI coding tooling for regulated industries is not a SaaS platform. It is infrastructure-as-code that deploys into the buyer's already-compliant cloud account, processes requests inside the buyer's VPC, and writes audit logs to the buyer's own database under the buyer's own keys - in one representative implementation, Terraform into the customer's AWS account, with every interaction logged to customer-owned DynamoDB tables encrypted with customer-managed KMS keys: who asked, what they asked, what the model returned, token count and cost. The platform-engineering pattern for financial services and government reduces to four mechanisms: provisioning through infrastructure-as-code templates, policy as version-controlled role-based access, comprehensive audit of prompts, tool calls and resource access, and a central proxy for model usage authenticated against the identity provider.

  • Proxy every model call through a gateway you run, authenticated by your identity provider. This is where you enforce which models, which snapshots, which data classes may leave, and where the audit record is written. It is also the kill switch the 59% do not have.
  • Keep the trace store inside the perimeter and under your keys. The prompt content, the tool calls and the model output are evidence; under DORA and SR 11-7 they need the same retention and access controls as any other record of a change to a regulated system.
  • Workspace-level isolation, ephemeral per task, with automated lifecycle. A durable agent environment is a durable attack surface; the Fortune 100 deletion happened in an environment that had accumulated the permissions to do it.
  • Deterministic guardrails at the workspace boundary - what paths, what commands, what network - enforced before execution. The auto-updated policy incident is the argument for policy as code with review: a policy change is a change, and it goes through the tiers.

The Pipeline, End To End

  1. 01The agent works in an ephemeral, in-VPC workspace, every model call through the proxy, every tool call graded and logged to the trace store.
  2. 02On proposing a change, the harness emits a signed attestation: agent, model snapshot, harness hash, task, humans, tool-call summary, scan references.
  3. 03The pipeline verifies the signature, classifies the change by touched paths into a tier, and selects the gates.
  4. 04SAST, SCA and secrets scanning run as for any change; the tier decides what blocks. Tier 3 blocks on anything. A floating model alias blocks at any tier because the change is not reproducible.
  5. 05Human review to the tier's count, with the attestation and the trace link in the PR. Reviewers see the plan, the tool calls and the scan results, not just the diff.
  6. 06For tier 3, the parity harness and the eval gate run, and an accountable Senior Manager signs off - the SM&CR requirement that does not change because the author was a model.
  7. 07The attestation ships with the build and lands in the change record, so 'who wrote this and on what basis' is a query for the life of the system.

What To Measure

  • SAST findings per agent-authored change versus human-authored, by tier and language. The Veracode 45% is a baseline; your number should fall as the harness is tuned to your scanner.
  • Tier 3 block rate and time-to-disposition. High at first is correct; rising later is a harness regression.
  • Attestation coverage. What share of merged changes carry a valid signed attestation. Anything under 100% is the 68% problem in your own pipeline.
  • Kill-switch drill time. From decision to no model calls proxied, measured quarterly. The 59% have never measured it.
  • Policy change volume through the tiers. If agent policy edits are bypassing the pipeline, you have the two-week problem and do not know it yet.

The Bottom Line

AI-generated code fails security scans 45% of the time and has for a year, while its volume inside financial firms has grown faster than the governance around it - 89% daily use against 69.7% with a policy, and majorities that cannot shut the system down or tell its actions from a person's. Neither banning agents nor trusting them is a pipeline. The pipeline is: a signed provenance attestation on every agent-authored change, risk tiers decided by what the change touches rather than by who wrote it, gates that block tier-3 paths on any finding and refuse floating model aliases everywhere, an agent that runs inside your VPC through a proxy you control with traces under your own keys, ephemeral workspaces with policy-as-code guardrails, and a Senior Manager's signature where the regime already requires one. Built that way, agent-written code passes the same gates as human code and leaves a better record. That is the deployment engineering we do for financial teams in London, and with the EU AI Act live, DORA in force and the FCA asking for evidence, it is the difference between using coding agents and being able to prove you used them safely.

References & Further Reading

AI-generated code securityWorkflow automation code samplesAI Automation London codecode provenanceAI Agency Developer LondonWorkflow Automation Agency architectureregulated CI/CD
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information