AI Coding Made CI The Bottleneck: Linear's Rework, CircleCI's 70.8% Main-Branch Success Rate, And A Tiered Validation Pipeline For A Regulated Repository
Linear's engineering post on 21 September - 136 points and 135 comments on Hacker News - says the quiet part: agents made shipping code exponentially faster, and validating it did not keep up. CircleCI's 2026 report, from 28 million workflows, has the industry-wide version: throughput up 59% on average, but for the median team main-branch throughput fell 7%, main-branch success dropped to 70.8% - the lowest in over five years against a 90% benchmark - and recovery time rose 13% to 72 minutes. Fewer than one team in twenty grew both curves. Linear's fixes are concrete and transferable: faster runners, tsgo cutting typecheck 73%, Oxlint cutting lint 68%, sparse checkouts, 8 shards, 87,000 runner-minutes saved by batching. For a regulated financial repository the same problem needs one more thing - validation tiered by what a change touches, so the pipeline spends its minutes where the risk is.
AlchmAI Engineering14 min read
70.8%
Main-branch build success rate in CircleCI's 2026 report - the lowest in over five years, against the platform's 90% benchmark
+59% / -7%
Average throughput growth year on year - but for the median team, feature branches +15% and main branch -7%
72 min
Mean recovery time from a failed main build, up 13% - the cost of every AI-generated change that breaks the trunk
87,000
Runner-minutes a month Linear saved by batching short checks - 11.8% of its CI - while quadrupling its test suite
There is a sentence in Linear's 21 September engineering post that every team running coding agents will recognise, from Mufeez Amjad: agents have made it exponentially faster to ship code, but validating those changes has not kept up at the same rate. The post went to the front of Hacker News with 136 points and 135 comments because it describes a bottleneck that has moved. For a decade the constraint on shipping was writing the code. Now the code arrives in minutes and waits on a pipeline that was sized for humans.
CircleCI's 2026 State of Software Delivery report gives the industry-wide picture from more than 28 million workflows, and it is worse than the average suggests. Throughput is up 59% year on year on average - but that figure is pulled by the top 5% of teams. For the median team, feature-branch throughput rose 15% while main-branch throughput, the thing that actually ships, fell 7%. Main-branch success dropped to 70.8%, the lowest in over five years against CircleCI's own 90% benchmark, and mean recovery time from a failed build rose 13% to 72 minutes. Fewer than one team in twenty grew both curves at once; that slice hit +26% on main and +85% on feature branches. The rest generated more code and delivered less of it.
What Linear Actually Did, And What Transfers
Linear's numbers are worth listing because almost all of them transfer to any TypeScript-heavy financial platform, and because they show that the gains come from a dozen unglamorous changes rather than one clever one.
- Runners: moved from GitHub Actions hosted runners to third-party runners with faster CPUs and better storage - jobs 34% faster.
- Typecheck: switched to tsgo, the native TypeScript compiler - 73% faster.
- Lint: switched to Oxlint - API lint 68% faster.
- Checkout: sparse, blobless checkouts with limited history, and a composite action with retry logic and git caching in place of actions/checkout. Cache marker writes removed from the merge path.
- Setup: Postgres clients pre-installed in base images; dependency installs restricted to required packages, taking pnpm install from 44-73 seconds to 16-18; node_modules caching abandoned in favour of faster clean installs.
- Tests: shards from 4 to 8 under Vitest; oversized files split to balance shards; opt-in module state sharing with strict isolation rules, which Linear calls its largest single improvement at roughly 17% of monthly cost.
- Batching: short checks batched into fewer jobs - about 87,000 runner-minutes a month, 11.8% of total CI.
- Result: PR wait down from over six minutes to just over five, runner time per test roughly halved, test suite quadrupled - and a suite that would otherwise take about eleven minutes today.
“The lesson from Linear is not any single optimisation. It is that a team quadrupled its test coverage while making CI faster, because it treated the pipeline as a product with a latency budget rather than a chore with a queue.”
The Missing Piece For Regulated Repositories: Tier The Validation
Linear's changes make every run cheaper. A financial repository also needs to make the expensive runs rarer, by spending validation minutes in proportion to risk. The code-governance playbook classified changes into three tiers by the paths they touch; the same classification decides how much of the pipeline a change must pay for. A logging tweak in tier 1 runs typecheck, lint and impacted tests in the merge queue. A change under risk/ or oms/ runs the whole suite, the parity harness and the eval gate, and cannot be batched with anything.
name: validate
on:
pull_request:
merge_group: # merge queue: batches compatible PRs, revalidates the batch
jobs:
classify:
runs-on: ubuntu-latest
outputs:
tier: steps.tier.outputs.tier
impacted: steps.tier.outputs.impacted
steps:
- uses: ./.github/actions/sparse-checkout # blobless, limited history, retry + cache
- id: tier
run: |
# Tier from touched paths (same rules as pipeline/classify.py). Emits the
# impacted test set from the dependency graph so tier 1 runs only what changed.
uv run python -m pipeline.classify --emit-github-outputs
fast: # every tier, every PR - the floor
needs: classify
runs-on: [self-hosted, fast] # faster CPUs and storage: -34% in Linear's case
steps:
- uses: ./.github/actions/sparse-checkout
- run: pnpm install --filter ...[HEAD^] --frozen-lockfile # only affected packages
- run: pnpm tsgo --noEmit # native typecheck
- run: pnpm oxlint # native lint
- run: pnpm vitest run --shard=1/8 --changed # impacted tests only, tier 1
if: needs.classify.outputs.tier == '1'
full: # tier 2 and 3: the whole suite, sharded
needs: [classify, fast]
if: needs.classify.outputs.tier != '1'
runs-on: [self-hosted, fast]
strategy: { matrix: { shard: [1,2,3,4,5,6,7,8] } }
steps:
- uses: ./.github/actions/sparse-checkout
- run: pnpm install --frozen-lockfile
- run: pnpm vitest run --shard=matrix.shard/8
regulated: # tier 3 only: cannot be batched, cannot be skipped
needs: [classify, full]
if: needs.classify.outputs.tier == '3' && github.event_name != 'merge_group'
runs-on: [self-hosted, fast]
steps:
- uses: ./.github/actions/sparse-checkout
- run: uv run python -m evals.run --suite golden --out results.json
- run: uv run python -m evals.gate results.json --hard-fail confirmed_ranked_low=0 --hard-fail invalid_citations=0
- run: uv run pytest tests/test_parity.py -x # backtest vs live decisions
- run: uv run python -m pipeline.sast_gate --block-on any
- run: uv run python -m pipeline.attestation --verify # signed agent provenance present and validTest Impact Analysis Without Fooling Yourself
Running only impacted tests is the biggest tier-1 saving and the easiest to get wrong. Two rules keep it honest in a regulated repository.
- 01Impact is computed from the module dependency graph, not from file paths alone. A change to a shared Decimal helper impacts every test that imports it transitively; a path-based filter would miss that and let a rounding change into the OMS untested.
- 02Impacted-only runs are for pull requests, never for the merge queue or the trunk. The batch that lands runs the full suite. Impact analysis buys speed on the way in; it must not buy silence on the way through.
- 03Flaky tests are quarantined, not retried into passing. CircleCI's 72-minute recovery time is mostly humans re-running things that were red for reasons unrelated to the change. A quarantine list with an owner and an expiry turns that into a queue with a metric.
- 04Coverage of tier-3 paths is a gate. If a change under risk/ reduces line or branch coverage on risk/, it does not merge, whatever the test results say. Agents are good at making tests pass and indifferent to whether the tests mean anything.
The Cost Dimension The Report Does Not Mention
More runs at a lower success rate is also a bill. Linear's 87,000 saved runner-minutes a month is 11.8% of a large CI spend, and the post's projected eleven-minute baseline is a reminder that unmanaged growth is exponential in agent-era repositories. Three practices from our own engagements keep it linear.
- Cost per merged change, split by tier, on a dashboard with an owner. This is the CI equivalent of cost per completed task for agents, and it is the number that reveals when a harness change started generating churn.
- A latency budget per tier, enforced as a test. Tier 1 under four minutes, tier 2 under ten, tier 3 whatever the parity harness needs. A PR that pushes a tier over budget gets a failing check that says so, which is how the pipeline stays a product.
- Agent-specific rate limits at the queue. An agent that opens forty PRs in an hour is not forty times more productive; it is generating forty validation runs. A per-author concurrency limit on the merge queue is a control on the harness expressed in CI.
The Bottom Line
The constraint on shipping has moved from writing code to validating it, and the numbers are now industry-wide: 59% more throughput on average, 7% less on the median team's main branch, a 70.8% main-branch success rate that is the worst in over five years, and 72 minutes to recover from each failure. Linear's rework shows the pipeline can be made fast enough to keep up - faster runners, native typecheck and lint, sparse checkouts, minimal installs, eight shards, batched short checks, a quadrupled suite at lower cost - and a regulated financial repository needs one thing more: validation tiered by what the change touches, with a merge queue for everything except the tier-3 paths that must be validated alone, impact analysis that is honest about the dependency graph and never applied to the trunk, quarantine instead of retry, and cost and latency budgets per tier with an owner. Do that and the gates that make a bank's pipeline defensible stop being the thing people ask to remove. That is the workflow engineering we build for financial teams in London, and this week's front page is why.
References & Further Reading
- Linear - AI coding has made CI a bottleneck, so we reworked ours to keep up (21 September 2026). linear.app/now/ci-bottleneck-reworked
- CircleCI - Five takeaways from the 2026 State of Software Delivery report. circleci.com/blog/five-takeaways-2026-software-delivery-report
- Thoughtworks - A Thoughtworks perspective on CircleCI's 2026 State of Software Delivery report. thoughtworks.com/en-us/insights/blog/generative-ai/a-thoughtworks-perspective-on-circleci-s-2026-state-of-software-
- Rob Bowley - More code, less delivery: does the CircleCI 2026 report really show 1 in 20 teams are benefiting?. blog.robbowley.net/2026/04/02/more-code-less-delivery-but-does-the-circleci-2026-report-really-show-1-in-20-teams-are-benefiting
- Josh Tuddenham - 59% more code, 30% more failures: the toolchain bottleneck is here. joshtuddenham.dev/blog/toolchain-bottleneck
- DevOps.com - The validation gap is costing you more than you think. devops.com/the-validation-gap-is-costing-you-more-than-you-think
- DEV Community - Breaking the CI bottleneck: architecting pipelines for the AI code era. dev.to/tamizuddin/breaking-the-ci-bottleneck-architecting-pipelines-for-the-ai-code-era-hm
- Hacker News AI Digest, 2026-09-22 (Linear post, 136 points). github.com/yaojiejia/agents-radar/issues/177
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information