Opus 5.5 And GPT-6 Sol And Luna Launched The Same Day: A CTO's Guide To Breaking Changes, Cost Per Task And Eval-Gated Model Routing
On 22 September Anthropic released Claude Opus 5.5 at $4/$20 per million tokens with cache reads cut to $0.20, a 1M-token context and 128K output - and a list of breaking API changes. The same day OpenAI released GPT-6 Sol at $2/$10 and Luna at $0.10/$0.50, permanent cuts of 50% or more. Hacker News gave the two launches roughly 3,500 points between them and every engineering Slack filled with benchmark screenshots. For a CTO in financial services the right response is neither to switch nor to ignore it. It is to have a pipeline that turns a model launch into a measured decision within a week. This is the education piece: what actually changed, why benchmark margins no longer decide anything, how to compute cost per successful task, and the eval-gated router - with code - that makes every future launch routine.
AlchmAI Engineering15 min read
$4 / $20
Opus 5.5 per million input/output tokens, down 20%; cache reads $0.20, down 60% (22 September)
$2 / $10
GPT-6 Sol per million tokens - matching Sonnet 5 and half Opus 5.5; OpenAI says the prices are permanent
$0.10 / $0.50
GPT-6 Luna, positioned for summarisation, extraction and straightforward high-volume questions
5
Breaking or behavioural API changes in Opus 5.5, from thinking configuration to forced tool choice
Model launches used to be quarterly events. They are now weekly, and 22 September was a double: Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol and Luna within hours of each other. Opus 5.5 leads on agentic coding and knowledge-work benchmarks - 66.4% on Terminal-Bench 4.0, around 1846 Elo on GDPval-AA - at a 20% lower list price and a 60% lower cache-read price than Opus 5, with a 1M-token context and a retirement commitment no earlier than 22 September 2027. GPT-6 Sol and Luna cut prices by half or more, with OpenAI attributing the reduction to inference and caching improvements and positioning Sol for coding and agents and Luna for high-volume extraction and summarisation.
The commentary focused on which is 'better'. That is the wrong question for a firm running AI in production, and the most honest line in the launch coverage says why: at this capability level, benchmark margins have become a less reliable guide to real-world differences. Opus 5.5 leads some benchmarks and trails GPT-6 Astra on others, often within measurement noise. The only numbers that decide anything are yours.
Step 1: Know What Breaks
Opus 5.5 is not a drop-in replacement. The launch notes list changes that will return errors or change behaviour in code written for earlier models:
- Disabled thinking is no longer accepted; thinking is adaptive, with an effort control. Code that turned thinking off for latency needs to set a low effort instead.
- Forcing tool use with tool_choice 'any' or a named tool returns a 400; strict tool definitions are the replacement. Structured-extraction pipelines that forced a single tool are the most likely casualty.
- Thinking blocks cannot be replayed after the system prompt or tools change - relevant to agents that edit their own toolset mid-run.
- The older computer-use tool version is rejected in favour of the new toolset.
- Changing effort mid-session invalidates the prompt cache, which matters for cost as much as for correctness.
None of this is exotic, but it is exactly the kind of change that passes a smoke test and fails at 3am on a batch job. The defence is a thin internal wrapper so that model-specific parameters live in one place, and a contract test per workload that runs against any candidate model before it is allowed near production traffic.
// One place for model-specific quirks. Application code asks for a
// capability ("extract", "reason", "summarise"), never for raw parameters.
export interface ModelProfile {
id: string;
provider: "anthropic" | "openai";
price: { in: number; out: number; cacheRead: number }; // USD per 1M tokens
supportsForcedTool: boolean; // false for Opus 5.5 - use strict tools
thinking: "adaptive" | "configurable" | "none";
retireNotBefore?: string;
}
export const PROFILES: Record<string, ModelProfile> = {
"opus-5.5": { id: "claude-opus-5-5", provider: "anthropic",
price: { in: 4, out: 20, cacheRead: 0.2 },
supportsForcedTool: false, thinking: "adaptive",
retireNotBefore: "2027-09-22" },
"gpt-6-sol": { id: "gpt-6-sol", provider: "openai",
price: { in: 2, out: 10, cacheRead: 0 /* set from your contract */ },
supportsForcedTool: true, thinking: "configurable" },
"gpt-6-luna":{ id: "gpt-6-luna", provider: "openai",
price: { in: 0.1, out: 0.5, cacheRead: 0 },
supportsForcedTool: true, thinking: "configurable" },
};
export function extractionRequest(p: ModelProfile, schemaTool: object) {
// Opus 5.5 rejects forced tool choice: mark the tool strict and let the
// model call it; validate the output either way.
return p.supportsForcedTool
? { tools: [schemaTool], tool_choice: { type: "tool", name: "emit" } }
: { tools: [{ ...schemaTool, strict: true }] };
}Step 2: Price Per Successful Task, Not Per Token
A token price is an input to a cost, not a cost. What a workload costs depends on how many tokens the model uses to finish, how much of the prompt is cached, and - most importantly - how often it succeeds. A model at half the token price that needs a retry one time in five, or a human correction one time in ten, can be the more expensive choice. Launch coverage made the same point from the other direction: on some coding benchmarks, a medium effort setting finished for under a dollar while scoring higher than maximum effort.
from dataclasses import dataclass
@dataclass
class RunResult:
input_tokens: int
cached_tokens: int
output_tokens: int
passed: bool # judged against YOUR golden set, not a benchmark
human_minutes: float # review/correction time when it failed
def cost_per_success(results, price, analyst_rate_per_min=1.5):
spend = 0.0
for r in results:
fresh = r.input_tokens - r.cached_tokens
spend += (fresh * price["in"] + r.cached_tokens * price["cacheRead"]
+ r.output_tokens * price["out"]) / 1_000_000
spend += r.human_minutes * analyst_rate_per_min
wins = sum(1 for r in results if r.passed)
return float("inf") if wins == 0 else spend / wins
# Run the same 200-case golden set per workload against each candidate at
# two effort levels. Rank by cost_per_success, subject to a quality floor.Step 3: Let Evaluations, Not Launch Posts, Move Traffic
The router is where the discipline lives. Each workload has a current model and a set of candidates. A candidate is promoted only if it clears the workload's quality floor on the golden set, does not regress on the safety cases (refusals, hallucinated figures, leaked data), and lowers cost per successful task. Promotion goes through shadow traffic first, then a small percentage, then full traffic, with automatic rollback on a quality alarm.
interface WorkloadPolicy {
name: string; // e.g. "kyc-extract", "research-summary"
current: string; // profile key
candidate?: string;
canaryPercent: number; // 0 = shadow only
qualityFloor: number; // min pass rate on golden set
}
interface EvalScore { passRate: number; safetyRegressions: number; costPerSuccess: number }
export function promote(p: WorkloadPolicy, cur: EvalScore, cand: EvalScore): WorkloadPolicy {
const eligible =
cand.passRate >= p.qualityFloor &&
cand.passRate >= cur.passRate - 0.01 && // no meaningful quality loss
cand.safetyRegressions === 0 &&
cand.costPerSuccess < cur.costPerSuccess;
if (!eligible) return { ...p, candidate: undefined, canaryPercent: 0 };
const next = p.canaryPercent === 0 ? 5 : Math.min(100, p.canaryPercent * 4);
return next === 100
? { ...p, current: p.candidate!, candidate: undefined, canaryPercent: 0 }
: { ...p, canaryPercent: next };
}
export function pick(p: WorkloadPolicy, requestHash: number): string {
if (p.candidate && requestHash % 100 < p.canaryPercent) return p.candidate;
return p.current;
}“A launch day should not produce a migration. It should produce an evaluation run, and the evaluation run should produce the migration - or not.”
What We Would Expect To See In Financial Workloads
- High-volume extraction and classification - KYC documents, trade confirmations, email triage - is where Luna-class pricing is transformative, provided it clears the quality floor. Test it first there.
- Complex agentic work - multi-step research, code changes, reconciliations - is where Opus 5.5's agentic gains and lower cache price matter, particularly with long, stable, cacheable context.
- Customer-facing answers remain grounded in deterministic calculation whatever the model; a new model does not change that rule.
- Keep at least two providers evaluated for every critical workload. Same-day launches are a reminder that the market moves fast; provider incidents are a reminder that it also fails.
The Bottom Line
Claude Opus 5.5 at $4/$20 with cheaper caching, and GPT-6 Sol at $2/$10 and Luna at $0.10/$0.50, arrived on the same day with benchmark margins too narrow to decide anything and API changes that will break code written for earlier models. The mature response is a pipeline: model-specific quirks isolated in one profile layer with contract tests, cost measured per successful task on your own golden sets including human correction time, and a router that promotes candidates only through shadow and canary traffic when quality holds and cost falls. Build that once and every future launch becomes a routine evaluation run. That is the model-operations work we do as an AI agency for fintech teams in London, and this week showed exactly why it pays for itself.
References & Further Reading
- Digital Applied - Claude Opus 5.5: pricing, benchmarks and breaking changes. digitalapplied.com/blog/claude-opus-5-5-launch-pricing-benchmarks-2026
- MindStudio - Claude Opus 5.5: benchmarks, pricing and real-world performance. mindstudio.ai/blog/claude-opus-5-5-release
- AI Pricing Guru - Claude Opus 5.5 pricing: 40% lower typical cost. aipricing.guru/news/claude-opus-5-5-api-pricing-september-2026
- VentureBeat - OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more. venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more
- The New Stack - OpenAI releases GPT-6 Sol and Luna, and cuts token prices in half. thenewstack.io/openai-gpt-6-sol-luna-release
- Requesty - GPT-6 Sol and Luna: pricing, release date and API access. requesty.ai/blog/gpt-6-sol-luna-pricing-release-api
- OpenTelemetry - Semantic conventions for generative AI. opentelemetry.io/docs/specs/semconv/gen-ai
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information