Skip to content
Agentic AI

One Sign Error Took Down Three AI-Written Papers In A Day. What OpenAI's Maths Release Teaches CTOs About Verifying Agent Output - Dependency Chains, Verification Tiers And Property Tests

On 6 October OpenAI published 722 AI-generated manuscripts on 372 open mathematical problems - the most-discussed AI story on Hacker News that week. The next day it withdrew three. A sign error in one paper invalidated a cancellation argument and the construction that two other papers depended on, so all three fell together; fourteen more were repaired. About 42% of the results had machine-checked proofs in Lean. A Cornell mathematician said OpenAI should have announced the verified manuscripts first; an MIT mathematician said results nobody can reproduce should be treated as unverified. This is the education piece for senior engineers and CTOs, because the lesson is not about mathematics. It is about what happens when an agent produces many interdependent outputs quickly, some checked and some not, and one is wrong. In finance that is a reconciliation, a model, a report. Here are the three disciplines that would have contained it - dependency tracking, verification tiers and property tests - with code.

AlchmAI Engineering14 min read

722

AI-generated manuscripts OpenAI released on 6 October, covering 372 open mathematical problems

3

Withdrawn the next day: one sign error invalidated an argument and the construction two dependent papers relied on

42%

Of the remaining 719 results with machine-checked Lean proofs (about 300); 14 more manuscripts needed repairs

1,235

Hacker News points for OpenAI's announcement - the week's most-discussed AI story among engineers

The scale was the point. On Tuesday 6 October OpenAI posted more than 700 preprints describing progress on 372 open problems across geometry, computer science and algebra, produced by an internal model it has not released, at an average of about three hours of compute per result. On Wednesday it withdrew three. According to the repository's change history, a sign error in 'Algebraicity of Weil classes on split abelian eightfolds' invalidated a stabilization-trace cancellation argument and the construction used by two dependent papers - so the error in one manuscript took down two more that built on it. Fourteen others received proof repairs, corrected statements or clarified dependencies. About 300 of the remaining 719 results, roughly 42%, had machine-checked proofs in the Lean proof assistant.

The reaction from mathematicians was measured but pointed. Cornell's Alex Townsend said OpenAI should have announced the Lean-verified manuscripts first and suspects more errors will surface. MIT's Andrew Sutherland credited OpenAI for withdrawing quickly, but said that until the model is released and results can be replicated, claims of one-shot solutions should be treated as unverified. OpenAI did not publish the specific prompts or per-result compute costs that an Institute for Advanced Study advisory group had recommended.

1. Track Dependencies, So One Failure Invalidates Exactly What It Should

OpenAI could withdraw the two dependent papers within a day because the dependency was recorded. Most agent pipelines do not record it: a derived figure is copied into a report and the link to its inputs is lost. Treat every agent output as an artefact with explicit inputs, and when any input fails verification, invalidate everything downstream automatically.

pythonlineage/artefacts.py
from dataclasses import dataclass, field
from collections import defaultdict, deque

@dataclass
class Artefact:
    id: str
    kind: str                       # "calc", "summary", "report", "filing"
    inputs: list = field(default_factory=list)
    status: str = "unverified"      # unverified | verified | failed | invalidated
    tier: str = "unverified"        # machine_checked | reviewed | self_checked | unverified
    verified_by: str | None = None  # "machine", "reviewer:<id>", None

class Lineage:
    def __init__(self):
        self.items = {}
        self.children = defaultdict(set)

    def add(self, a: Artefact):
        self.items[a.id] = a
        for i in a.inputs:
            self.children[i].add(a.id)

    def fail(self, artefact_id: str, reason: str) -> list:
        """Mark one artefact failed and invalidate every descendant. Returns what was pulled."""
        self.items[artefact_id].status = "failed"
        pulled, queue = [], deque(self.children[artefact_id])
        while queue:
            cid = queue.popleft()
            child = self.items[cid]
            if child.status != "invalidated":
                child.status = "invalidated"
                pulled.append(cid)
                queue.extend(self.children[cid])
        audit.write({"event": "invalidation", "root": artefact_id, "reason": reason, "pulled": pulled})
        return pulled

# Usage: a fee calculation with a sign error is failed -> every client report and the monthly
# return that consumed it are invalidated in one call, and the audit shows exactly why.

2. Verification Tiers, And Release Only What You Have Verified

Townsend's suggestion - announce the Lean-verified results first - is a release policy any engineering organisation can adopt. Give every artefact a verification tier, and let the tier decide where it may go. Machine-checked results can flow automatically; human-reviewed ones can go to internal users; unverified ones stay in a draft lane. The tier must be a property the release system enforces, not a label someone remembers to read.

pythonlineage/release_gate.py
TIERS = {
    "machine_checked": 3,   # reproduced by a deterministic checker: recomputation, reconciliation to source, formal proof
    "reviewed": 2,          # a named, qualified human checked it against sources
    "self_checked": 1,      # the agent's own consistency checks passed
    "unverified": 0,
}

DESTINATIONS = {
    "client_report": 3,
    "regulatory_return": 3,
    "internal_dashboard": 2,
    "draft_workspace": 0,
}

def ancestors(artefact_id: str, lineage) -> set:
    seen, stack = set(), list(lineage.items[artefact_id].inputs)
    while stack:
        i = stack.pop()
        if i not in seen:
            seen.add(i); stack.extend(lineage.items[i].inputs)
    return seen

def may_release(artefact, destination: str, lineage) -> tuple:
    need = DESTINATIONS[destination]
    # An artefact is only as verified as its least-verified input.
    chain = [artefact] + [lineage.items[i] for i in ancestors(artefact.id, lineage)]
    weakest = min(TIERS[a.tier] for a in chain)
    if any(a.status in ("failed", "invalidated") for a in chain):
        return False, "depends on a failed or invalidated artefact"
    if weakest < need:
        return False, f"weakest link is tier {weakest}, destination needs {need}"
    return True, "ok"

The 'weakest link' rule is the one most teams miss. A report reviewed by a human is not tier 2 if it contains a number no one recomputed. That is precisely how a sign error in one paper reached two others that were otherwise sound.

3. Property Tests For The Errors Agents Actually Make

A sign error is the canonical agent mistake in finance too: a fee subtracted instead of added, a debit treated as a credit, a short position's P&L inverted. Example-based tests rarely catch them because the author's examples share the author's assumption. Property-based tests state invariants that must hold for any input and let the framework search for counter-examples - the closest thing most finance code has to a machine-checked proof.

pythontests/test_pnl_properties.py
from decimal import Decimal
from hypothesis import given, strategies as st

from pnl import position_pnl, net_proceeds   # functions an agent wrote or modified

qty = st.decimals(min_value=Decimal("-1000000"), max_value=Decimal("1000000"), places=0).filter(lambda q: q != 0)
price = st.decimals(min_value=Decimal("0.01"), max_value=Decimal("10000"), places=4)
fee = st.decimals(min_value=Decimal("0"), max_value=Decimal("100"), places=2)

@given(qty, price, price)
def test_pnl_sign_follows_direction(q, entry, exit_):
    pnl = position_pnl(q, entry, exit_)
    if exit_ > entry:
        assert (pnl > 0) == (q > 0)          # long gains when price rises, short loses
    if exit_ == entry:
        assert pnl == 0

@given(qty, price, price)
def test_pnl_antisymmetric(q, entry, exit_):
    assert position_pnl(q, entry, exit_) == -position_pnl(-q, entry, exit_)

@given(qty, price, fee)
def test_fees_always_reduce_proceeds(q, p, f):
    assert net_proceeds(q, p, f) <= net_proceeds(q, p, Decimal("0"))   # a fee can never help the client
  • Write properties, not examples: direction, antisymmetry, monotonicity in fees and rates, conservation (debits equal credits), bounds.
  • Run them on every change an agent makes to calculation code, and make a failing property a blocking CI result.
  • Record the falsifying example the framework finds; it is usually the clearest bug report you will ever get.

“OpenAI's model produced hundreds of results in a few days. The scarce resource was never generation. It was verification - and knowing what depended on what.”


The Reproducibility Point

Sutherland's objection - results from a model nobody else can run should be treated as unverified - applies inside firms as well. An agent-produced figure is reproducible only if you kept the model version, the prompt, the inputs and the tool calls. Store them with the artefact. A finding you cannot reproduce is a finding you cannot defend to an auditor, a client or a regulator, however plausible it looks.

The Bottom Line

OpenAI's release of 722 AI-written manuscripts and its withdrawal of three the next day - one sign error breaking an argument that two other papers depended on, with 14 more repaired and about 42% machine-checked - is the clearest public lesson yet in what large-scale agent output needs. For engineering leaders the transferable disciplines are concrete: record dependencies so a failure invalidates exactly its descendants, assign verification tiers and release by the weakest link, and use property-based tests to catch the sign and direction errors agents make. Keep everything needed to reproduce a result. That is how we build agentic systems for financial firms as an AI agency in London, and it is the difference between withdrawing three papers and withdrawing a filing.

References & Further Reading

Agentic AI finance codeverificationproperty-based testingAI Agency Developer LondonAgentic AI London codeengineering educationdependency tracking
Share Email
AI

AlchmAI Engineering

Engineering, London

Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.

Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information