BNY's Agents Read Payment Fields For Meaning, BNP Paribas Cut A Workflow From Ten Steps To Six: Building A Payment-Repair Agent That Learns Exceptions From Humans, In Code
At Sibos on 29 September the production examples finally arrived. BNY described agents that read incoming payment fields for meaning - not just format - check them against ISO 20022 and flag what a human needs to resolve. BNP Paribas said agents took a securities-services workflow from ten steps to six and were doing 80 to 85 per cent of the work within weeks, with staff teaching the system how to handle exceptions rather than processing them. HSBC's Smart Checking verifies trade documents against rules; Deutsche Bank's Ada framework reuses agents across divisions. Payment repair is the ideal first agent for any bank: enormous volume, rule-shaped, exception-driven, and a repair team that already reviews every case. This is the architecture we build for it: deterministic ISO 20022 validation first, a semantic repair proposal second, a confidence-gated apply-or-queue policy, and an exception-teaching loop that turns each human correction into a reusable rule. With code.
AlchmAI Engineering16 min read
10 → 6
Steps in a BNP Paribas securities-services workflow after agents took over; 80-85% of the work automated within weeks
Meaning
What BNY's agents read in payment fields - not only format - before checking against ISO 20022 and flagging for a person
1 month → 1 day
Deutsche Bank lending decision time with its Ada agent framework and knowledge graphs
<50%
Of banking customers who would let AI execute a transaction (Temenos) - the reason repair agents propose and people apply
Payment repair is where straight-through processing goes to die. A payment arrives with a beneficiary name that does not match the account, a BIC that has been replaced, an address in the wrong field, a purpose code missing, a reference that fails a sanctions screen on a false positive. Every one drops into a queue where an operations analyst reads it, works out what was meant, fixes it and releases it. At scale that is thousands of items a day, and it is exactly the work BNY described automating at Sibos: agents that read the fields for meaning, check syntax against ISO 20022 and flag the issues a human needs to see. BNP Paribas' account of its securities-services workflow - ten steps to six, 80-85% automated within weeks, staff teaching exceptions rather than processing them - is the same pattern applied one desk over.
The architecture that makes this safe in a bank has four layers, and the order matters. Deterministic validation runs first and alone decides what is structurally wrong. The model proposes a repair, with evidence, for what validation cannot fix. A policy decides whether the proposal is applied automatically or queued for a person, based on confidence, value and type. And every human correction is captured as a candidate rule, so the system learns the bank's exceptions from the people who know them.
Layer 1: Deterministic Validation
from dataclasses import dataclass, field
import re
IBAN_RE = re.compile(r"^[A-Z]{2}[0-9]{2}[A-Z0-9]{11,30}$")
BIC_RE = re.compile(r"^[A-Z]{6}[A-Z0-9]{2}([A-Z0-9]{3})?$")
PURPOSE_CODES = {"SALA", "SUPP", "TRAD", "INTC", "DIVI", "LOAN", "TAXS"} # from the external code list
@dataclass
class Finding:
path: str; code: str; severity: str; detail: str = ""
@dataclass
class ValidationResult:
findings: list = field(default_factory=list)
@property
def blocking(self): return [f for f in self.findings if f.severity == "error"]
def iban_checksum_ok(iban: str) -> bool:
s = iban[4:] + iban[:4]
digits = "".join(str(int(c, 36)) for c in s)
return int(digits) % 97 == 1
def validate_pacs008(tx: dict) -> ValidationResult:
r = ValidationResult()
cdtr_acct = tx.get("CdtrAcct", {}).get("Id", {}).get("IBAN")
if not cdtr_acct:
r.findings.append(Finding("CdtrAcct.Id.IBAN", "MISSING", "error"))
elif not IBAN_RE.match(cdtr_acct) or not iban_checksum_ok(cdtr_acct):
r.findings.append(Finding("CdtrAcct.Id.IBAN", "INVALID_IBAN", "error", cdtr_acct))
bic = tx.get("CdtrAgt", {}).get("FinInstnId", {}).get("BICFI")
if bic and not BIC_RE.match(bic):
r.findings.append(Finding("CdtrAgt.FinInstnId.BICFI", "INVALID_BIC", "error", bic))
purpose = tx.get("Purp", {}).get("Cd")
if purpose and purpose not in PURPOSE_CODES:
r.findings.append(Finding("Purp.Cd", "UNKNOWN_PURPOSE", "warning", purpose))
if not tx.get("Cdtr", {}).get("Nm"):
r.findings.append(Finding("Cdtr.Nm", "MISSING_NAME", "error"))
addr = tx.get("Cdtr", {}).get("PstlAdr", {})
if addr and not addr.get("Ctry"):
r.findings.append(Finding("Cdtr.PstlAdr.Ctry", "MISSING_COUNTRY", "warning"))
return rLayer 2: Semantic Repair Proposals
For each finding validation cannot resolve, the model is given the payment, the finding, the relevant reference data - the beneficiary's known accounts, the BIC directory entry, the sender's history with this beneficiary - and the bank's accumulated exception rules. It returns a proposal: the field, the proposed value, the evidence it relied on and a confidence. It does not touch the payment.
import json
from dataclasses import dataclass
@dataclass(frozen=True)
class RepairProposal:
path: str; current: str; proposed: str; evidence: list; confidence: float; rationale: str
PROMPT = (
"You are a payments operations analyst. A payment failed a check. Using ONLY the reference data "
"and exception rules provided, propose the single most likely intended value for the failing field. "
"Cite the reference-data ids you relied on. If the evidence is insufficient, return proposed=null. "
"Respond as JSON with fields: proposed, evidence (list of ids), confidence (0-1), rationale."
)
def propose_repair(tx: dict, finding, refdata: dict, rules: list, llm) -> RepairProposal:
ctx = {"payment": tx, "finding": finding.__dict__, "reference_data": refdata, "exception_rules": rules}
out = json.loads(llm(PROMPT, json.dumps(ctx)))
return RepairProposal(
path=finding.path, current=finding.detail, proposed=out.get("proposed"),
evidence=out.get("evidence", []), confidence=float(out.get("confidence", 0)),
rationale=out.get("rationale", ""),
)
# Example of the reference data the model sees for an INVALID_BIC finding:
# {"bic_directory": [{"id": "bicdir:MIDLGB22", "bic": "MIDLGB22", "status": "active", "replaces": "MIDLGB2L"}],
# "beneficiary_history": [{"id": "hist:8812", "iban": "GB29...", "last_bic": "MIDLGB22", "count": 14}]}Layer 3: Apply Or Queue - A Policy, Not A Model Decision
from decimal import Decimal
AUTO_APPLY = {
# path -> (min confidence, max amount, required evidence kinds)
"CdtrAgt.FinInstnId.BICFI": (0.95, Decimal("250000"), {"bicdir"}),
"Cdtr.PstlAdr.Ctry": (0.90, Decimal("1000000"), {"hist", "bicdir"}),
"Purp.Cd": (0.90, Decimal("1000000"), {"hist"}),
}
NEVER_AUTO = {"CdtrAcct.Id.IBAN", "Cdtr.Nm", "IntrBkSttlmAmt"} # account, name, amount: always a person
def decide(proposal, amount: Decimal, sanctions_clear: bool) -> str:
if proposal.proposed is None or not sanctions_clear or proposal.path in NEVER_AUTO:
return "queue"
rule = AUTO_APPLY.get(proposal.path)
if not rule:
return "queue"
min_conf, max_amt, kinds = rule
evidence_kinds = {e.split(":")[0] for e in proposal.evidence}
if proposal.confidence >= min_conf and amount <= max_amt and evidence_kinds & kinds:
return "apply"
return "queue"Account identifiers, beneficiary names and amounts are never auto-applied, whatever the confidence: those are the fields fraud targets, and they are the ones Temenos' research says customers will not let an AI touch. Everything else is applied only when confidence, value and evidence kind all clear thresholds that the bank sets and versions.
Layer 4: Teaching Exceptions
This is the part BNP Paribas described and most pilots skip. When an analyst corrects a queued item - accepts the proposal, edits it or rejects it - the system records the correction as a candidate rule: the finding pattern, the context that distinguished it, and the resolution. Candidates that recur are promoted, after review, into the exception rules the model is given next time. The analysts stop being the processing capacity and become the teachers.
from collections import defaultdict
from dataclasses import dataclass
@dataclass(frozen=True)
class Correction:
finding_code: str; path: str; context_key: str; resolution: str; analyst: str; proposal_accepted: bool
class ExceptionTeacher:
def __init__(self, promote_after=3):
self.candidates = defaultdict(list)
self.rules = [] # promoted, reviewed rules fed to propose_repair()
self.promote_after = promote_after
def record(self, c: Correction) -> None:
key = (c.finding_code, c.path, c.context_key)
self.candidates[key].append(c)
same = [x for x in self.candidates[key] if x.resolution == c.resolution]
if len(same) >= self.promote_after and len({x.analyst for x in same}) >= 2:
self.rules.append({
"when": {"finding": c.finding_code, "path": c.path, "context": c.context_key},
"then": c.resolution,
"evidence": f"{len(same)} corrections by {len({x.analyst for x in same})} analysts",
"status": "pending_review", # a supervisor approves before it is used
})
# context_key examples: "sender:ACME-LTD|cdtr:GB29..." or "bic_replaced:MIDLGB2L->MIDLGB22"
# Each promoted rule is a sentence the model reads next time, with the human evidence behind it.“The analysts were never the bottleneck because they were slow. They were the bottleneck because the system could not learn what they knew. Teach it, and ten steps become six.”
Measuring It
- Auto-apply rate and its reversal rate: how often an applied repair was later corrected. The second number is the one that matters.
- Queue time per item, and the share of items where the analyst accepted the proposal unchanged - BNP Paribas' 80-85% is this metric.
- Rules promoted per week and the drop in recurrence of the findings they cover.
- Every applied repair logged with the proposal, evidence ids, policy version and the validation findings - the audit trail that lets the bank show exactly why a payment was changed.
The Bottom Line
BNY reading payment fields for meaning against ISO 20022, BNP Paribas cutting a workflow from ten steps to six with staff teaching exceptions, HSBC checking documents and Deutsche reusing agents across divisions are the same pattern in production at four of the world's largest banks. The architecture that reproduces it is deterministic validation first, semantic repair proposals with cited evidence second, a versioned apply-or-queue policy that never auto-applies accounts, names or amounts, and an exception-teaching loop that turns analyst corrections into reviewed rules. Build it on one payment queue and the result is measurable in weeks. That is the workflow automation we deliver for banks and payment firms in London, and Sibos has just shown the ceiling.
References & Further Reading
- The Fintech Times - Sibos 2026 day two: banks put AI agents to work, but keep the judgement for people. thefintechtimes.com/sibos-2026-day-two-banks-put-ai-agents-to-work-but-keep-the-judgement-for-people
- ISO 20022 - message definitions (pacs.008 FI to FI customer credit transfer). iso20022.org/iso-20022-message-definitions
- ISO 20022 - external code sets (purpose codes). iso20022.org/catalogue-messages/additional-content-messages/external-code-sets
- Swift - ISO 20022 for payments. swift.com/standards/iso-20022
- Pay.UK - ISO 20022 for UK payments. wearepay.uk/what-we-do/standards/iso-20022
- The Fintech Times - Sibos 2026 day one: Fraser tells Miami to move fast without breaking trust. thefintechtimes.com/sibos-2026-day-one-fraser-tells-miami-to-move-fast-without-breaking-trust
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information