147 Pages, Seven Recommendations, No New Rulebook: What The FCA's Mills Review Actually Asks UK Engineering Teams To Build
On 6 July 2026 the FCA published the Mills Review - a 147-page landmark study of AI in retail financial services with seven recommendations, commissioned in January and delivered in six months. It identifies four shifts: how firms operate, how consumer journeys evolve, how competition and market power reshape, and how fraud and cyber risk amplify. What it conspicuously does not do is create a new AI rulebook, and a good and poor practice publication is still to come later this year. For UK engineering teams that combination is the whole story - and it is considerably more demanding than a rulebook would have been, because you cannot satisfy it with a compliance matrix. You satisfy it with instrumentation.
AlchmAI Engineering13 min read
147 pages
The Mills Review, published 6 July 2026, six months after the FCA commissioned it in January
7
Recommendations on how the UK should harness and regulate AI in retail financial services
4 shifts
Firm operations, consumer journeys, competition and market power, and the amplification of fraud and cyber risk
0
New AI rulebooks created - with a good and poor practice publication still to come later in 2026
British financial regulation has a particular habit that foreign observers routinely misread as indecision. When something important arrives, the FCA commissions a review, publishes it, and then tells firms to apply the rules they already have. In January 2026 it commissioned the Mills Review to consider artificial intelligence in the context of retail financial services. The resulting 147-page report landed on 6 July with seven recommendations, and it has been widely described as a landmark - accurately, because of what it sets in motion rather than because it changes the perimeter.
We build AI systems for UK financial firms, and we have now read this document the way an engineering team has to read it: not asking what is prohibited, but asking what evidence a supervisor will expect to see when they arrive. On that reading the review is more demanding than a prescriptive rulebook would have been, and considerably more useful. It identifies four AI-driven shifts likely to reshape retail financial services - the transformation of firm operations, the evolution of consumer journeys, the reshaping of competition and market power, and the amplification of fraud and cyber risks - and it declines to convert any of them into a checklist.
The Four Shifts, Read As Engineering Requirements
Taking each of the review's four shifts and asking what it implies for a system you have to build and operate:
- 01Transformation of firm operations. This is the shift most firms are already living, and the engineering requirement is attribution: when a process moves from human to automated, you must still be able to say who is accountable for a given outcome and on what basis it was reached. Under the Senior Managers regime, accountability does not become diffuse because the work became automated. Practically: a decision record per consequential output, naming the model version, the inputs, the checks applied and the human who could have intervened.
- 02Evolution of consumer journeys. Consumer Duty requires firms to monitor outcomes, not just processes, and the FCA has separately published good and poor practice examples on exactly that monitoring. An AI-driven journey makes this harder because the journey differs per consumer. The requirement is outcome telemetry segmented by cohort - if your system performs well on average and poorly for one group, an average is not evidence of a good outcome, and increasingly it is evidence of the opposite.
- 03Reshaping of competition and market power. Largely a policy concern, with one sharp engineering edge: concentration on a small number of model providers. If a single provider going down or changing terms would halt a customer-facing journey, that is an operational resilience exposure you are expected to have identified, and the Bank of England has already flagged the systemic version of it.
- 04Amplification of fraud and cyber risk. The most immediate. Deepfake-enabled fraud has moved from a novelty to a volume crime, and the same generative capability that improves your onboarding also improves an attacker's. The requirement is that controls designed when synthetic media was expensive get re-examined now that it is not - voice authentication being the obvious casualty.
The Good And Poor Practice Publication Is The Thing To Prepare For
The most actionable sentence in the surrounding coverage is that the FCA will launch an AI good and poor practice publication later this year, informed by direct engagement with firms about what is working well, where firms face challenges, and where further clarity would help. Engineering teams should treat that as the real deliverable. A review sets direction; a good and poor practice document is what supervisors will hold up in a meeting, and it is being assembled right now from what firms are doing.
This is consistent with everything else the UK regulator has done in this space. AI Live Testing put eight firms including Barclays, Lloyds through Scottish Widows, UBS and Experian into supervised production trials with a technical assurance partner examining the evidence. The revealed preference could not be clearer: the FCA wants to see systems running, with controls that produce artefacts, rather than policies describing controls that might.
“An outcomes-based regulator supervising by live testing and publishing good practice from real firms is, in effect, asking one question: show me your telemetry. Firms that have it will find this environment comfortable. Firms with a policy library will not.”
The Six Artefacts, Made Concrete
We have written before about the artefacts a UK financial AI system should emit. The Mills Review sharpens two of them enough to be worth restating with implementation detail - the outcome telemetry that Consumer Duty now effectively requires of an AI-driven journey, and the accountability record that survives a Senior Managers conversation.
/**
* One record per consumer-affecting AI decision. This is simultaneously the
* Consumer Duty outcome evidence and the SM&CR accountability trail - and
* writing it once, at the decision point, is the only way it stays accurate.
*/
export interface OutcomeRecord {
decisionId: string;
occurredAt: string; // ISO 8601, UTC
// WHAT decided, pinned so the behaviour is reconstructible years later.
modelSnapshot: string; // dated snapshot, never a floating alias
promptVersion: string;
policyBundleVersion: string;
retrievalIndexVersion: string | null;
// WHAT it saw. Identifiers, not payloads - the payloads live in the store
// they came from, with their own retention and access controls.
inputRefs: string[];
asOf: string; // point-in-time basis for any market or
// reference data used
// WHAT it decided, and what the deterministic layer said about it.
outcome: { kind: string; value: unknown };
deterministicChecks: { code: string; passed: boolean }[];
// WHO was accountable, and whether oversight was real.
reviewerId: string | null;
reviewerSawEvidence: boolean; // did the UI actually show them the basis?
reviewerOverrode: boolean;
smfOwner: string; // the accountable Senior Manager function
// Consumer Duty cohort dimensions. Segmented outcome monitoring is not
// possible retrospectively if these are not captured at decision time.
cohort: {
vulnerabilityFlags: string[];
productCategory: string;
channel: "app" | "web" | "phone" | "branch" | "adviser";
tenureBand: string;
};
}Segmented Outcome Monitoring, Which Averages Will Not Give You
The Consumer Duty implication of an AI-driven journey is the one most likely to catch firms out, because the failure is invisible in the headline metric. A model that performs well overall can perform materially worse for a cohort, and Consumer Duty is explicit that outcomes must be monitored rather than processes. The engineering answer is unglamorous and cheap if built in, expensive if retrofitted.
-- Run on a schedule; alert on the DELTA, not the level. The question a
-- supervisor asks is not "is this good?" but "is this worse for anyone?"
WITH windowed AS (
SELECT
cohort_product_category AS product,
cohort_channel AS channel,
has_vulnerability_flag AS vulnerable,
date_trunc('week', occurred_at) AS week,
count(*) AS decisions,
avg(CASE WHEN outcome_kind = 'accepted' THEN 1 ELSE 0 END) AS accept_rate,
avg(CASE WHEN reviewer_overrode THEN 1 ELSE 0 END) AS override_rate,
avg(CASE WHEN complaint_within_90d THEN 1 ELSE 0 END) AS complaint_rate
FROM outcome_records
WHERE occurred_at >= now() - interval '26 weeks'
GROUP BY 1, 2, 3, 4
)
SELECT
product, channel, week,
accept_rate FILTER (WHERE vulnerable) AS vuln_accept,
accept_rate FILTER (WHERE NOT vulnerable) AS base_accept,
accept_rate FILTER (WHERE vulnerable)
- accept_rate FILTER (WHERE NOT vulnerable) AS accept_gap,
complaint_rate FILTER (WHERE vulnerable)
- complaint_rate FILTER (WHERE NOT vulnerable) AS complaint_gap
FROM windowed
GROUP BY product, channel, week
HAVING abs(accept_rate FILTER (WHERE vulnerable)
- accept_rate FILTER (WHERE NOT vulnerable)) > 0.05
ORDER BY week DESC, abs(accept_gap) DESC;The Case For Britain, With The Counterweight
We will state our bias plainly, as we have before: for building financial AI, we think the UK is currently the best environment in the world, and the Mills Review reinforces rather than weakens that. A regulator that commissions a serious 147-page study, delivers it in six months, runs supervised live testing of real systems with an assurance partner, and then publishes good and poor practice drawn from actual firms is doing something no other major jurisdiction is doing at that maturity. It gives builders something far more useful than permission: it gives them a moving picture of what good looks like, assembled from peers.
The honest counterweight is real and should be said in the same breath. Outcomes-based regulation transfers judgement to the firm, and judgement is expensive. A large bank has a model risk function that can absorb that. A forty-person wealth manager does not, and for them the absence of a checklist is a genuine cost rather than a flexibility. There is a live risk that the UK's approach works beautifully for institutions with in-house capability and leaves the mid-market unable to deploy anything with confidence - which would be a poor outcome for competition, and notably one of the four shifts the review itself flags. The good and poor practice publication is the obvious remedy, and the more concrete it is, the more useful it will be to exactly the firms that need it most.
The other counterweight worth noting is dual-jurisdiction cost. Firms serving EU customers are simultaneously subject to the AI Act's high-risk obligations, in force since 2 August 2026, which are prescriptive in the way the UK's approach is not. Our recommendation there is unchanged: build to the stricter documentary standard once and produce two reports from one evidence base. Maintaining separate UK and EU paths is a false economy, because the second path always rots and it rots quietly.
What To Do In The Next Quarter
- 01Capture cohort dimensions at decision time, now. Segmented outcome monitoring cannot be reconstructed from records that did not record the segment. This is the single change with the shortest window - every day you do not capture it is a day you cannot later analyse.
- 02Instrument the oversight interface, not just the decision. Record whether the reviewer was shown the evidence, and track override rates as a health metric. A rate near zero is a finding about your process, not a compliment to your model.
- 03Name the accountable Senior Manager per automated decision path, in the system rather than in a document. If the SM&CR mapping lives only in a governance pack, it is out of date.
- 04Re-examine every control designed when synthetic media was expensive. Voice authentication first, then any process where one person can authorise something consequential after one conversation.
- 05Write down your model-provider failure mode and test it. Outage, deprecation, price change. This is ordinary operational resilience work that the AI framing has let too many firms skip.
- 06Watch for the good and poor practice publication and read it as an engineering specification. It is the closest thing to a checklist the UK regime will produce, and it is being written from what firms are doing right now.
The Bottom Line
The Mills Review is a landmark that creates no new rules, and engineering teams should understand why that is the demanding outcome rather than the lenient one. Seven recommendations, four identified shifts, and an explicit intention to publish good and poor practice later in the year add up to a regulator that will supervise AI through the rules it already has - Consumer Duty, SM&CR, operational resilience - and will judge firms on demonstrated outcomes rather than documented intentions. That is an environment in which instrumentation is the compliance strategy: a decision record with pinned versions and point-in-time inputs, an oversight interface that proves the reviewer saw the basis, cohort dimensions captured at decision time so segmented monitoring is possible at all, and a tested answer to what happens when your model provider has a bad week. Firms that build those artefacts will find the UK the easiest major jurisdiction in the world to ship financial AI in. Firms that build a policy library will find it the hardest, and will not discover which until someone asks to see the telemetry.
References & Further Reading
- FCA - FCA publishes landmark review into impact of AI on retail financial services (Mills Review, 6 July 2026). fca.org.uk/news/press-releases/fca-publishes-landmark-review-impact-ai-retail-financial-services
- A&O Shearman - Report issued by the Mills Review: the future of AI in retail financial services. aoshearman.com/en/insights/report-issued-by-the-mills-review-the-future-of-ai-in-retail-financial-services
- Skadden - The FCA's Mills Review charts a course for AI-enabled retail financial services. skadden.com/insights/publications/2026/07/fcas-mills-review-charts-a-course-for-ai-enabled
- Global Regulation Tomorrow - FCA good and poor practice examples on monitoring consumer outcomes under the Consumer Duty. regulationtomorrow.com/2026/07/fca-good-and-poor-practice-examples-in-relation-to-monitoring-consumer-outcomes-under-the-consumer-duty
- FCA - AI Lab and AI Live Testing. fca.org.uk/firms/innovation/ai-lab
- Covington - UK financial services regulators' approach to artificial intelligence in 2026. globalpolicywatch.com/2026/04/uk-financial-services-regulators-approach-to-artificial-intelligence-in-2026
- JURIST - UK financial regulator publishes landmark AI review. jurist.org/news/2026/07/uk-financial-regulator-publishes-landmark-ai-review
- BCLP - AI regulation in financial services: turning principles into practice. bclplaw.com/en-US/events-insights-news/ai-regulation-in-financial-services-turning-principles-into-practice.html
- European Commission - The EU Artificial Intelligence Act. artificialintelligenceact.eu
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information