Skip to content
AI Strategy & ROI

No DSIT, No AI Safety Law, And A Frontier Model The AI Security Institute Never Saw: Britain Now Runs On Post-Release Assurance - So UK Fintechs Must Build It

In two months Britain's AI governance changed shape. Andy Burnham's government scrapped the Department for Science, Innovation and Technology, dropped the planned law that would have compelled frontier developers to submit models for testing, and made Kanishka Narayan AI minister without a dedicated department. Anthropic's Mythos 5.1 became the first of its models the AI Security Institute was not allowed to test before launch, while OpenAI gave access for GPT-6 Astra. In the Lords, the government rejected amendments to bring AI vendors into the Cyber Security and Resilience Bill, pointing instead to the voluntary AI Cyber Security Code of Practice and ETSI EN 304 223. We think the UK's flexible approach can be an advantage - but only if firms do the assurance that the state is no longer doing up front. Here is the political context from a British point of view, and the engineering: a model supplier register, fingerprint-triggered regression evals and incident clauses, in code.

AlchmAI Engineering15 min read

0

Pre-release AISI tests of Anthropic's Mythos 5.1 - the first Anthropic model the institute was not allowed to test before launch

1

Frontier model AISI did test pre-release this month: OpenAI's GPT-6 Astra, launched 3 September

No

Dedicated department for AI after DSIT was scrapped; the AI minister now spans the Cabinet Office and the business department

Voluntary

The AI Cyber Security Code of Practice and ETSI EN 304 223 - the government's answer to calls to regulate AI vendors in the cyber bill

Britain entered the summer with a plan: a department leading AI policy, an AI Security Institute with privileged pre-release access to frontier models, and groundwork laid in Keir Starmer's final months for a law that would compel developers to have their products tested before release. It leaves September with a different settlement. Andy Burnham's government scrapped the Department for Science, Innovation and Technology. The planned safety law was dropped. Kanishka Narayan became AI minister, working across the Cabinet Office and the business department without a department of his own, and wrote in the Times that nothing is off the table. And the institute's access now depends on each lab: OpenAI let it test GPT-6 Astra before its 3 September launch, while Anthropic's Mythos 5.1 was the first Anthropic model it could not test beforehand.

The same approach runs through the Cyber Security and Resilience Bill. In the Lords, cybersecurity minister Baroness Lloyd argued that regulating AI vendors would not address the harms some AI products can pose, and the government rejected amendments that would have required vendors to demonstrate safety guardrails or given emergency shutdown powers over AI systems. Its answer is the voluntary AI Cyber Security Code of Practice, which informed the first global AI cybersecurity standard, ETSI EN 304 223, plus the AISI's cooperative testing. Baroness Kidron put the critics' case sharply: the NHS must protect itself, but the AI attacking it has no requirement to check itself.

What Post-Release Assurance Means For An Engineering Team

If nobody outside the lab has necessarily tested the model you call, your own tests are the assurance. For regulated UK firms that is not new in principle - the FCA expects firms to understand and monitor the systems they use, and the resilience regime expects supplier concentration and incident response to be managed - but it is new in practice, because the models change underneath you and the upstream safety net is thinner. Three pieces of engineering cover most of it.

1. A Model Supplier Register That Is Code, Not A Spreadsheet

yamlgovernance/model-register.yaml
- id: claude-opus-5-5
  provider: anthropic
  jurisdiction: US
  hosting: [anthropic-api, aws-bedrock-eu-west-2]
  aisi_pre_release_tested: unknown     # record what you can evidence
  retire_not_before: 2027-09-22
  workloads: [research-summary, code-review]
  data_classes_allowed: [public, internal]
  exit_plan: gpt-6-sol                 # evaluated alternative, same golden set
  incident_contact: vendor-security@...
  last_regression_eval: 2026-09-24

- id: gpt-6-luna
  provider: openai
  jurisdiction: US
  hosting: [openai-api]
  aisi_pre_release_tested: unknown
  workloads: [kyc-extract]
  data_classes_allowed: [public, internal, confidential-redacted]
  exit_plan: claude-sonnet-5
  last_regression_eval: 2026-09-23

Keep it in the repository, validate it in CI, and fail the build if a workload calls a model that is not registered or sends a data class the register does not allow. That single check turns 'which models do we depend on?' - the first question on any supervisory call - into a query.

2. Regression Evals Triggered By Model Change, Not The Calendar

Hosted models change behind stable names: safety patches, routing changes, new snapshots. Record a fingerprint of every response's model metadata and run the workload's golden set whenever it changes, before the change can reach customers at scale.

pythonassurance/fingerprint.py
import hashlib, json

def fingerprint(resp_meta: dict) -> str:
    # Whatever the provider returns that identifies the serving model:
    # model id, snapshot/version, system fingerprint, region.
    keys = ("model", "model_version", "system_fingerprint", "region")
    blob = json.dumps({k: resp_meta.get(k) for k in keys}, sort_keys=True)
    return hashlib.sha256(blob.encode()).hexdigest()[:16]

class ModelWatch:
    def __init__(self, store, evals, alert):
        self.store, self.evals, self.alert = store, evals, alert

    def observe(self, workload: str, resp_meta: dict) -> None:
        fp = fingerprint(resp_meta)
        known = self.store.get(workload)
        if known == fp:
            return
        self.store.set(workload, fp)
        # New serving model: throttle to canary share and run the golden set.
        self.store.set_canary(workload, percent=5)
        result = self.evals.run(workload)
        if result.pass_rate < result.floor or result.safety_regressions:
            self.store.set_canary(workload, percent=0)   # route to exit_plan model
            self.alert(workload, fp, result)
        else:
            self.store.set_canary(workload, percent=100)

3. Incidents That Include The Vendor's AI

  • Add AI-specific incident notification to supplier contracts at renewal: model behaviour incidents, agent misbehaviour and security disclosures, not only outages.
  • Map AI incidents to the 24-hour notification and 72-hour report clock the UK financial regulators already run, so an AI failure is classified like any other.
  • Track vendor disclosures actively. This month alone brought disclosures of agent misbehaviour from OpenAI and Google; your incident process should ask 'are we exposed?' the same day.
  • Align with ETSI EN 304 223 and the AI Cyber Security Code of Practice now. They are voluntary, which means they are the evidence you can show that you did the reasonable thing.

“Britain chose not to test every model at the border. That makes every UK firm that deploys one part of the testing regime, whether it has noticed or not.”


What Government Should Do Next

From a UK industry perspective, three moves would strengthen the settlement without abandoning its flexibility. First, make AISI pre-release access an expectation for any frontier model marketed to UK critical sectors, including financial services - access that OpenAI already gives and that Anthropic gave until this month. Second, publish the AISI's findings in a form firms can use in their own supplier assessments. Third, give the AI minister the policy capacity a department used to provide; flexibility without capacity is just absence. The UK's approach can be a competitive advantage. It needs to be run like one.

The Bottom Line

With DSIT scrapped, the AI safety law dropped, AI vendors kept out of the cyber bill, and Mythos 5.1 launched without AISI pre-release testing, Britain now relies on voluntary codes, cooperative lab access and post-release monitoring. For UK fintechs and financial firms, that makes their own assurance the safeguard: a model supplier register enforced in CI, regression evaluations triggered whenever the serving model changes with automatic fallback to an evaluated exit model, and AI-specific incident clauses mapped onto the existing 24-hour and 72-hour clock. We back the UK's flexible approach and we think it can win - and as an AI agency in London, this is the engineering we build so that our clients are the proof that it works.

References & Further Reading

AI Agency LondonUK AI policyAI Security InstituteWorkflow Automation LondonAI Automation London codemodel assuranceCyber Security and Resilience Bill
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information