Skip to content
Workflow Automation

Goldman's Juniors Now 'Manage A Virtual Army' Of Agents. Nobody Has Built Them The Tools: The Agent Supervision Console - Run Records, Risk-Ranked Review Queues, Sampling And Span Of Control, In Code

On 8 October Goldman Sachs' Kevin Sneader told the Milken Institute Asia Summit that the bank's new recruits now start by managing AI agents - 'this virtual army that they've now got'. Bank job postings mentioning AI are up 49% this year, and mentions of 'agent orchestration' rose from 108 to 1,967. The management problem is real and the tooling is not: most firms give a junior a chat window and a Slack channel and call it supervision. Managing agents at volume needs what managing people at volume needs - a record of what each agent was asked and did, a queue that puts the riskiest work in front of a human first, a sampling policy that checks confident work too, a limit on how many agents one person can meaningfully oversee, and metrics that show whether supervision is working. This is the architecture we build for it, with code.

AlchmAI Engineering15 min read

'Virtual army'

Goldman's description of the agents its new recruits manage from day one (Kevin Sneader, Milken Asia Summit, 8 October)

+49%

Bank job postings mentioning AI in 2026 - almost 140,000 - per hiring analytics firm Draup

108 → 1,967

Postings mentioning 'agent orchestration', 2025 to 2026

5

Components a supervision console needs: run records, a risk-ranked queue, sampling, span-of-control limits and supervision metrics

Kevin Sneader's line from Singapore - 'when our young folks now start work, they're managing agents' - has been quoted all week as a story about careers. It is also a statement about software. Managing work at volume has always required tools: case management systems for operations teams, order management systems for traders, ticketing systems for engineers. Managing agents is no different, and the hiring data says demand is arriving fast - Draup counted almost 140,000 AI-related bank postings this year, up 49%, with 'agent orchestration' mentioned seventeen times more often than last year and responsible AI and governance close behind.

What most junior staff actually get is a chat interface, an output and an instruction to check it. That does not scale past a handful of agents, it gives no view of risk, and it leaves no record of who checked what. Sneader was candid that the role of middle managers is unresolved; part of the answer is that the supervision middle managers used to provide has to be built into the tools juniors use.

1. The Run Record

typescriptsupervision/model.ts
export interface AgentRun {
  id: string;
  agentId: string;                    // which agent, which version
  model: string;                      // provider model id
  task: { type: string; brief: string; requestedBy: string };
  inputs: { kind: "document" | "system" | "prior_run"; ref: string; hash: string }[];
  actions: { tool: string; args: unknown; at: string; result: "ok" | "error" }[];
  output: { kind: string; body: unknown; citations: string[] };
  selfReported: { confidence: number; uncertainties: string[] };
  risk: { score: number; drivers: string[] };          // computed by the platform, not the agent
  status: "queued" | "in_review" | "accepted" | "edited" | "rejected" | "escalated" | "auto_released";
  review?: { by: string; at: string; decision: string; diff?: string; reason?: string; minutes: number };
}

Everything a reviewer or an auditor needs is on the record: what the agent was asked and by whom, what it read, which tools it called, what it produced with what citations, what it said it was unsure of, and the risk the platform assigned. The agent's self-reported confidence is kept, but it never decides anything by itself.

2. A Queue Ranked By Risk, Not Arrival

pythonsupervision/risk.py
RISK_WEIGHTS = {
    "touches_money": 0.35,           # proposes a payment, trade, limit or price
    "client_facing": 0.25,           # output will reach a client
    "low_citation_coverage": 0.15,   # claims without sources
    "novel_case": 0.10,              # unlike anything in the last 90 days of runs
    "agent_uncertain": 0.10,         # the agent flagged uncertainty
    "tool_errors": 0.05,
}

def risk_score(run) -> tuple:
    drivers = []
    if run.touches_money: drivers.append("touches_money")
    if run.client_facing: drivers.append("client_facing")
    if run.citation_coverage < 0.9: drivers.append("low_citation_coverage")
    if run.novelty > 0.8: drivers.append("novel_case")
    if run.self_confidence < 0.7 or run.uncertainties: drivers.append("agent_uncertain")
    if any(a.result == "error" for a in run.actions): drivers.append("tool_errors")
    return round(sum(RISK_WEIGHTS[d] for d in drivers), 2), drivers

def queue_priority(run, now) -> float:
    age_hours = (now - run.created_at).total_seconds() / 3600
    sla_pressure = min(age_hours / run.sla_hours, 1.5)
    return run.risk_score * 0.7 + sla_pressure * 0.3      # risky first, but nothing waits forever

3. Sample The Confident Work Too

A queue that only shows risky work trains humans to trust everything else - and agents are most dangerous when they are confidently wrong. A sampling policy routes a fraction of low-risk, high-confidence output to review regardless, with the rate rising automatically when reviewers start finding problems in the sample.

pythonsupervision/sampling.py
import random

class AdaptiveSampler:
    """Review a share of low-risk runs; raise the share when the sample finds errors."""
    def __init__(self, base_rate=0.05, max_rate=0.5):
        self.rate, self.base, self.max = base_rate, base_rate, max_rate

    def should_review(self, run) -> bool:
        if run.risk_score >= 0.35:            # always reviewed via the queue
            return True
        return random.random() < self.rate

    def observe(self, sampled_review_found_error: bool):
        if sampled_review_found_error:
            self.rate = min(self.max, self.rate * 2)        # errors in 'safe' work: look harder
        else:
            self.rate = max(self.base, self.rate * 0.95)    # decay slowly back to base

4. Span Of Control

A 'virtual army' still needs officers in proportion. If one junior is nominally supervising forty agents producing four hundred outputs a day, nobody is supervising them. Set a span of control in review-minutes, not agent count, and stop assigning work to a supervisor whose queue cannot be cleared in their working day.

pythonsupervision/assignment.py
def assign(run, supervisors, est_minutes):
    """Assign to the qualified supervisor with capacity; escalate if nobody has any."""
    qualified = [s for s in supervisors if run.task_type in s.qualified_tasks and run.risk_score <= s.max_risk]
    with_capacity = [s for s in qualified if s.queued_minutes + est_minutes(run) <= s.daily_review_minutes]
    if not with_capacity:
        return escalate(run, reason="no supervisor capacity - throttle agent intake")   # back-pressure on the agents
    s = min(with_capacity, key=lambda s: s.queued_minutes)
    s.queued_minutes += est_minutes(run)
    return s.id

The escalation branch is the important one: when supervision capacity is exhausted, the platform slows the agents rather than letting unreviewed work pile up or auto-release. Throughput is limited by what humans can genuinely check.

5. Is Supervision Working?

sqlsupervision/metrics.sql
SELECT
  r.review_by                                                     AS supervisor,
  COUNT(*)                                                        AS reviewed,
  AVG(CASE WHEN r.decision = 'accepted' THEN 1 ELSE 0 END)        AS accept_rate,
  AVG(CASE WHEN r.decision IN ('edited','rejected') THEN 1 ELSE 0 END) AS intervention_rate,
  AVG(r.minutes)                                                  AS avg_minutes,
  -- Did errors escape review? Defects found later in outputs this supervisor accepted:
  SUM(CASE WHEN d.run_id IS NOT NULL THEN 1 ELSE 0 END)           AS escaped_defects
FROM agent_reviews r
LEFT JOIN downstream_defects d ON d.run_id = r.run_id AND r.decision = 'accepted'
WHERE r.at >= now() - interval '30 days'
GROUP BY 1
ORDER BY escaped_defects DESC, accept_rate DESC;
-- A 99% accept rate with escaped defects is rubber-stamping. A 0% accept rate is a broken agent.

“Giving a junior an army is easy. Giving them the console that shows which soldier is about to do something expensive is the actual job.”


The Training Dividend

Built this way, supervision is also how juniors learn. The queue shows them the hardest cases first, each with the agent's reasoning and sources; the sampling policy shows them what plausible-but-wrong looks like; the metrics tell them, and their managers, whether their judgement is improving. That is a better apprenticeship than formatting decks, and it answers part of Sneader's question about where middle management goes: into designing the risk weights, sampling policies and spans of control that the console enforces.

The Bottom Line

Goldman's new recruits managing a 'virtual army' of agents, and the 49% rise in bank AI postings with agent orchestration up seventeen-fold, make supervision tooling a production requirement. The console that works treats supervision as a queue: complete run records, a queue ranked by platform-computed risk with SLA pressure, adaptive sampling of confident work, span of control measured in review-minutes with back-pressure on the agents, and metrics that expose rubber-stamping and escaped defects. That is the workflow automation architecture we build for financial firms in London, and it is what turns a virtual army into a supervised workforce.

References & Further Reading

Workflow Automation Agency architectureWorkflow automation code samplesagent supervisionAgentic AI finance codeWorkflow Automation Londonhuman in the loopAI Agency fintech
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information