120,000 Client Emails A Day: The Classify-Enrich-Route Pipeline Behind Barclays' Global Markets Rollout, And Why Bounded Decisions Beat Generation
On 1 October Barclays said Claude now classifies, enriches and chooses a processing path for roughly 120,000 client emails a day in its Global Markets business, alongside a knowledge assistant used by 16,000 colleagues and a plan to put Claude Code in half its developers' hands by year-end. The same week OpenAI previewed a Decisions API: a constrained Luna model that picks one answer from a developer-defined set in well under a second, built for exactly this kind of classification and routing. The two announcements describe the same engineering truth: the highest-volume AI workload in a bank is not generating text, it is deciding - which desk, which priority, which workflow, which action. This is the pipeline: schema-first extraction, bounded decisions with enumerated answers, confidence calibration against a labelled set, a human-review band, and cost per routed email at 120,000 a day. With code.
AlchmAI Engineering15 min read
120,000
Client emails a day that Barclays' Global Markets platform classifies, enriches and routes with Claude (announced 1 October)
16,000+
Barclays colleagues using the Claude-based knowledge assistant, with more than a million searches handled since 2025
<1 sec
Latency class of OpenAI's previewed Decisions API, which returns one answer from a developer-defined set for classification and routing
50%
Of Barclays developers expected to be using Claude Code by the end of 2026, a majority of engineers during 2027
Barclays' announcement contained the single most useful production number of the week: about 120,000 client emails a day in Global Markets, classified, enriched and routed by Claude to the right processing path, reducing manual handling for operations staff. It sits alongside a knowledge assistant with 16,000 users and a Claude Code rollout targeting half the bank's developers by year-end, all 'within a secure and well-governed environment', in the words of co-COO Anne Marie Darling. Two days earlier OpenAI had previewed a Decisions API at DevDay - a smaller, constrained Luna model that takes context, a question and a finite list of allowed answers and returns one of them in well under a second - for classification, routing and choosing an agent's next step.
Put the two together and you have a design principle for the most common AI workload in finance. Most of what a bank needs from a model at scale is not prose. It is a decision from a known set: which desk owns this email, what priority, which workflow, does it need a human, what is the client asking for. Treating those as bounded decisions rather than open generation makes them faster, cheaper, measurable and auditable - and it is how a pipeline survives at 120,000 a day.
The Pipeline
- 01Ingest and normalise: strip signatures and disclaimers, thread replies, extract attachments' text, redact what the model does not need to see.
- 02Extract: a schema-first pass that pulls the structured facts - client, product, instrument identifiers, dates, amounts, request type - into typed fields.
- 03Decide: bounded decisions for desk, priority, workflow and human-review need, each from an enumerated set, each with a confidence.
- 04Enrich: look up the client, the account team, open tickets and SLAs from systems of record using the extracted identifiers - deterministic, not generated.
- 05Route: apply a versioned policy that uses decisions, confidences and enrichment to pick the queue - or the human-review band.
// Every decision is an enum. The model cannot invent a desk or a priority.
export const DESKS = ["fx_sales", "rates_sales", "equities_sales", "credit_sales", "prime_services", "operations_settlements", "operations_confirmations", "onboarding_kyc", "unknown"] as const;
export const PRIORITY = ["p1_market_moving", "p2_same_day", "p3_two_days", "p4_informational"] as const;
export const REQUEST = ["price_request", "trade_confirmation_query", "settlement_break", "document_request", "complaint", "onboarding", "research_question", "other"] as const;
export const TRIAGE_TOOL = {
name: "triage_email",
description: "Classify a client email for routing. Choose ONLY from the allowed values.",
input_schema: {
type: "object",
properties: {
desk: { type: "string", enum: [...DESKS] },
priority: { type: "string", enum: [...PRIORITY] },
request_type: { type: "string", enum: [...REQUEST] },
needs_human: { type: "boolean", description: "True if the email asks for a price, a trade or a change to standing instructions" },
client_name: { type: ["string", "null"] },
identifiers: { type: "array", items: { type: "string" }, description: "ISINs, trade ids, account numbers found verbatim" },
confidence: { type: "number", minimum: 0, maximum: 1 },
},
required: ["desk", "priority", "request_type", "needs_human", "confidence"],
},
strict: true,
} as const;import { TRIAGE_TOOL } from "./schema";
export interface Triage {
desk: string; priority: string; request_type: string; needs_human: boolean;
client_name: string | null; identifiers: string[]; confidence: number;
}
// Works with any provider that supports strict tool schemas; swap for a Decisions API
// call when it exits preview - the enum set and the downstream policy do not change.
export async function decide(email: NormalisedEmail, llm: LLM): Promise<Triage> {
const res = await llm.call({
system: "You route client emails for a bank's markets business. Use only the tool. Never guess identifiers.",
messages: [{ role: "user", content: email.redactedText }],
tools: [TRIAGE_TOOL],
maxTokens: 200,
});
const t = res.toolInput as Triage;
// Belt and braces: validate enums even with strict schemas, and never trust confidence blindly.
if (!TRIAGE_TOOL.input_schema.properties.desk.enum.includes(t.desk as any)) t.desk = "unknown";
return t;
}Calibration: Make The Confidence Mean Something
A model's self-reported confidence is a hint, not a probability. Before you route on it, calibrate it against a labelled set of a few thousand historical emails: bucket predictions by stated confidence and measure actual accuracy per bucket. Then set the human-review threshold where accuracy drops below what the business accepts - and recheck weekly, because the mix of emails changes.
from collections import defaultdict
def calibration_table(preds, buckets=10):
"""preds: list of (stated_confidence, correct: bool). Returns per-bucket accuracy and count."""
table = defaultdict(lambda: [0, 0])
for conf, ok in preds:
b = min(int(conf * buckets), buckets - 1)
table[b][0] += 1
table[b][1] += int(ok)
return {f"{b/buckets:.1f}-{(b+1)/buckets:.1f}": {"n": n, "accuracy": (c / n if n else None)}
for b, (n, c) in sorted(table.items())}
def review_threshold(table, min_accuracy=0.97):
"""Lowest confidence bucket whose accuracy meets the floor; route below it to humans."""
for rng, row in table.items():
if row["n"] and row["accuracy"] is not None and row["accuracy"] >= min_accuracy:
return float(rng.split("-")[0])
return 1.01 # nothing qualifies: everything to humans until the model improvesRouting Policy And The Human Band
export function route(t: Triage, enrich: Enrichment, threshold: number): Route {
// Hard rules first: anything that could move money or a price sees a person.
if (t.needs_human || t.request_type === "price_request" || t.request_type === "complaint") {
return { queue: "human_review:" + t.desk, reason: "consequential request type" };
}
if (t.confidence < threshold) {
return { queue: "human_review:triage", reason: "below calibrated confidence " + threshold };
}
if (t.desk === "unknown") return { queue: "human_review:triage", reason: "unknown desk" };
// Enrichment is deterministic: SLA and owner come from systems of record, not the model.
const sla = enrich.clientTier === "platinum" ? "2h" : t.priority === "p1_market_moving" ? "15m" : "same_day";
return { queue: t.desk, owner: enrich.accountTeam, sla, reason: "auto-routed" };
}Cost At 120,000 A Day
Bounded decisions are cheap because the output is tiny - a tool call of a few dozen tokens - and because the prompt prefix is identical on every call and therefore cacheable. At mid-tier model prices with cached system prompts, a 1,500-token redacted email costs a fraction of a cent to triage; 120,000 a day is hundreds of dollars, not tens of thousands. The expensive part is the human band, which is why calibration that keeps it near the genuinely ambiguous few per cent is the real cost lever. Measure cost per correctly routed email, not per token.
“The most valuable thing a model does at 120,000 a day is not write. It is choose - from a list you wrote, with a confidence you measured, under a policy you can show the regulator.”
Operating It
- Log every decision with the model version, prompt version, the enum chosen, the confidence and the route - the audit trail for a routing error is the same as for any operational control.
- Sample the auto-routed band daily for silent errors; the human band tells you about ambiguity, not about confident mistakes.
- Treat taxonomy changes as releases: a new desk means a new enum value, a recalibration and a canary.
- Keep the extraction and routing provider-agnostic. The enum set, the policy and the calibration are yours; the model behind the tool call is a configuration.
The Bottom Line
Barclays routing 120,000 client emails a day with Claude and OpenAI previewing a Decisions API in the same week point at the same lesson: the highest-volume AI workload in a bank is deciding, not generating. The pipeline that scales is schema-first extraction, bounded decisions from enumerated sets enforced by strict tool schemas, confidence calibrated against labelled history with a human-review band set where accuracy demands it, deterministic enrichment from systems of record, and a versioned routing policy that sends anything consequential to a person. It is cheap at scale, measurable, auditable and provider-agnostic. That is the LLM engineering we build for markets and operations teams in London, and Barclays has just put a production number on it.
References & Further Reading
- Anthropic - Barclays scales Claude to upgrade operations and improve client experience (1 October 2026). anthropic.com/news/barclays-scales-claude
- Resultsense - Barclays widens Claude use, targets half of developers. resultsense.com/news/2026-10-02-barclays-claude-code-rollout
- Hugging Face blog - What is OpenAI Decisions API? A practical guide. huggingface.co/blog/sora-2/what-is-openai-decisions-api-a-practical-guide
- InfoQ - OpenAI DevDay 2026 recap for developers. infoq.com/news/2026/10/openai-devday-2026
- Anthropic Docs - Tool use with Claude (strict schemas). docs.anthropic.com/en/docs/build-with-claude/tool-use
- Guo et al. - On Calibration of Modern Neural Networks. arxiv.org/abs/1706.04599
- Anthropic Docs - Prompt caching. docs.anthropic.com/en/docs/build-with-claude/prompt-caching
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information