Microsoft Cut Per-Employee AI Budgets From $100,000 To $10,000 A Month; Haiku 5.5 Costs Five Times More Above 100K Tokens. Building The AI Spend Gateway A Bank Needs Before Its First Budget Review
Two numbers from this week define the AI cost problem. Microsoft reportedly cut monthly AI spending limits in its cloud and AI division from $100,000 per employee to around $10,000 and trimmed its projected internal spend on Anthropic by more than a third, while Meta halved its Claude Code users to 30,000. And Anthropic's Claude Haiku 5.5, released on 7 October, costs $0.10 per million input tokens for prompts up to 100,000 tokens - and $0.50 above it, with output rising from $0.50 to $2.50. Price tiers, per-person caps and runaway agent runs are now the three things that decide an AI bill, and research cited by the Wall Street Journal says only 11% of businesses can predict theirs. This is the gateway we put in front of every model call in a financial firm: a pricing engine that understands tiers, budgets enforced per person, team and run, routing to the cheapest model that passes, and the monthly report that keeps the programme funded - in code.
AlchmAI Engineering15 min read
$100k → $10k
Reported monthly AI spending limit per employee in Microsoft's cloud and AI division, before and after, in most cases
$0.10 → $0.50
Claude Haiku 5.5 input price per million tokens for prompts up to and above 100,000 tokens; output rises from $0.50 to $2.50
30,000
Meta employees still using Claude Code after a reported cut from about 60,000
11%
Of 400 businesses that accurately predict their AI costs, per research cited by the Wall Street Journal
The Information's report on 5 October put numbers on what every large AI user is discovering. Microsoft had expected internal spending on Anthropic technology of at least $1bn a year; after asking staff to use GitHub Copilot and OpenAI models instead, that projection has reportedly fallen by more than a third, and monthly AI spending limits in its cloud and AI division have dropped from $100,000 per employee to around $10,000 in most cases. Meta has reportedly halved its Claude Code users to about 30,000. Both remain large Anthropic customers; both are also its competitors. The point for everyone else is that the companies best placed to run AI cheaply have decided they need hard limits.
Two days later Anthropic released Claude Haiku 5.5, with a pricing structure that rewards engineering. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output - 90% and 50% below Haiku 4.5. Above 100,000 tokens, the whole request is priced at $0.50 input and $2.50 output. It is the first Haiku with adjustable effort, scores 43 on Artificial Analysis' intelligence index at maximum effort against 17 for Haiku 4.5, and offers a one-million-token context. A request at 99,000 tokens and one at 101,000 tokens can differ in cost by a factor of five. Gateways that do not know that will spend it.
1. A Pricing Engine That Knows About Tiers And Caching
from dataclasses import dataclass
@dataclass(frozen=True)
class PriceBand:
max_prompt_tokens: int | None # None = no upper bound
input_per_m: float
output_per_m: float
cache_read_per_m: float
# Verify against the provider's current price list before relying on these.
PRICES = {
"claude-haiku-5-5": [PriceBand(100_000, 0.10, 0.50, 0.01), PriceBand(None, 0.50, 2.50, 0.05)],
"claude-sonnet-5-5": [PriceBand(None, 2.00, 10.00, 0.20)],
"claude-opus-5-5": [PriceBand(None, 4.00, 20.00, 0.20)],
}
def band_for(model: str, prompt_tokens: int) -> PriceBand:
for band in PRICES[model]:
if band.max_prompt_tokens is None or prompt_tokens <= band.max_prompt_tokens:
return band
raise ValueError("no price band")
def estimate_cost(model: str, prompt_tokens: int, cached_tokens: int, output_tokens: int) -> float:
b = band_for(model, prompt_tokens) # the WHOLE request is priced by its prompt size
fresh = prompt_tokens - cached_tokens
return (fresh * b.input_per_m + cached_tokens * b.cache_read_per_m + output_tokens * b.output_per_m) / 1_000_000
def tier_cliff_warning(model: str, prompt_tokens: int, margin: float = 0.05):
"""Flag prompts just over a tier boundary - usually a few retrieved chunks too many."""
for band in PRICES[model]:
lim = band.max_prompt_tokens
if lim and lim < prompt_tokens <= lim * (1 + margin):
return f"{prompt_tokens} tokens is {prompt_tokens - lim} over the {lim} tier: trim retrieval to save ~5x"
return None2. Budgets Per Person, Per Team And Per Run
Per-employee caps stop slow overspend. Per-run caps stop the agent that loops at 3am. Team budgets make cost visible where decisions are made. The gateway checks all three before a call and records the actual cost after it, so the ledger is exact rather than estimated.
interface Budgets { userMonthlyUsd: number; teamMonthlyUsd: number; runUsd: number }
export async function guardedCall(req: ModelRequest, ctx: { userId: string; teamId: string; runId: string }, ledger: Ledger, budgets: Budgets) {
const estimate = estimateCost(req.model, req.promptTokens, req.cachedTokens, req.maxOutputTokens); // worst case
const [user, team, run] = await Promise.all([
ledger.monthToDate("user", ctx.userId), ledger.monthToDate("team", ctx.teamId), ledger.total("run", ctx.runId),
]);
if (run + estimate > budgets.runUsd) throw new BudgetError("run budget exceeded - agent paused for review");
if (user + estimate > budgets.userMonthlyUsd) throw new BudgetError("monthly personal AI budget reached");
if (team + estimate > budgets.teamMonthlyUsd) throw new BudgetError("team budget reached - ask your budget owner");
const warning = tierCliffWarning(req.model, req.promptTokens);
if (warning) await telemetry.warn("tier_cliff", { ...ctx, model: req.model, warning });
const res = await provider.call(req);
const actual = estimateCost(req.model, res.usage.inputTokens, res.usage.cachedTokens, res.usage.outputTokens);
await ledger.record({ ...ctx, model: req.model, usd: actual, task: req.taskType, outcomeId: req.outcomeId });
return res;
}3. Route To The Cheapest Model That Passes
Haiku 5.5's jump in capability at a tenth of the price changes the default. For classification, extraction, routing and summarisation, the right model is the cheapest one that clears the task's quality floor on your own evaluation set - and for many tasks that is now a small model. The router records the choice so the monthly report can show what routing saved.
# Pass rates come from your own golden sets, per task type, refreshed on every model release.
EVALS = {
"classify_email": {"claude-haiku-5-5": 0.985, "claude-sonnet-5-5": 0.990},
"extract_kyc": {"claude-haiku-5-5": 0.962, "claude-sonnet-5-5": 0.981},
"draft_credit_memo": {"claude-sonnet-5-5": 0.94, "claude-opus-5-5": 0.96},
}
FLOORS = {"classify_email": 0.98, "extract_kyc": 0.975, "draft_credit_memo": 0.93}
def route(task: str, prompt_tokens: int, expected_output: int) -> str:
eligible = [m for m, p in EVALS[task].items() if p >= FLOORS[task]]
if not eligible:
raise RuntimeError("no model meets the quality floor for " + task)
# Cost is per request at this prompt size, so tiers are respected automatically.
return min(eligible, key=lambda m: estimate_cost(m, prompt_tokens, 0, expected_output))
# classify_email -> haiku; extract_kyc -> sonnet (haiku below floor); draft_credit_memo -> sonnet (cheaper than opus, passes)4. The Report That Keeps The Programme Funded
-- Cost per outcome by team and task, with what routing saved versus sending everything to the frontier model.
SELECT
l.team_id,
l.task_type,
COUNT(DISTINCT l.outcome_id) AS outcomes,
SUM(l.usd) AS spend_usd,
SUM(l.usd) / NULLIF(COUNT(DISTINCT l.outcome_id), 0) AS usd_per_outcome,
SUM(l.frontier_equivalent_usd) - SUM(l.usd) AS routing_saving_usd,
SUM(CASE WHEN l.tier_cliff THEN 1 ELSE 0 END) AS tier_cliff_calls,
SUM(CASE WHEN l.run_budget_hit THEN 1 ELSE 0 END) AS runs_stopped
FROM ai_ledger l
WHERE l.at >= date_trunc('month', now())
GROUP BY 1, 2
ORDER BY spend_usd DESC;“A per-employee cap tells you who spent the money. A gateway tells you what it bought - and that is the only number that survives a budget review.”
The Bottom Line
Microsoft's reported fall from $100,000 to around $10,000 a month per employee, Meta's halving of Claude Code users and Claude Haiku 5.5's fivefold price step above 100,000 prompt tokens show where AI cost control is heading: hard limits, tier-aware pricing and routing to the cheapest model that passes. For banks and fintechs scaling coding agents and AI tools, the gateway is the control - a pricing engine that understands tiers and caching, budgets enforced per person, team and run before every call, a router driven by your own evaluation floors, and a monthly report of cost per outcome and routing savings. That is the AI automation and production engineering we deliver in London, and it is what turns an AI programme from a cost line into a business case.
References & Further Reading
- PYMNTS - Microsoft and Meta steer staff from Anthropic Claude to in-house AI (citing The Information, 5 October 2026). pymnts.com/news/artificial-intelligence/2026/microsoft-meta-steer-staff-from-anthropic-claude-in-house-ai
- Cyber Security News - Meta and Microsoft are actively cutting employee use of Claude AI. cybersecuritynews.com/meta-microsoft-claude-ai
- Developers Digest - Claude Haiku 5.5: pricing, migration changes and when to use it. developersdigest.tech/blog/claude-haiku-5-5-release-guide-2026
- Technology Org - Claude Haiku 5.5 cuts token prices by up to 90% (8 October 2026). technology.org/2026/10/08/claude-haiku-5-5-anthropic-price-benchmarks
- Anthropic - Claude Haiku 5.5. anthropic.com/claude-haiku-5-5
- Ramp - AI Index. ramp.com/data/ai-index
- FinOps Foundation - FinOps for AI. finops.org/wg/finops-for-ai-overview
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information