OpenAI’s Agents Are Training In Software. Ironclad’s Terms Deserve A Read
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Agents Are Training In Software. Ironclad’s Terms Deserve A Read on ThorstenMeyerAI.com

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training and evaluating a frontier model in hosted copies of Ironclad’s contract-management software. The model met an average 55% of rubric criteria across 11 tasks; OpenAI says its time estimates were simulated, and the results do not establish customer productivity gains or readiness for unsupervised use.

OpenAI published details on October 6 of a collaboration with Ironclad, a contract-management software company, to train and evaluate a frontier model on tasks inside hosted copies of Ironclad’s product. The work tested 11 legal, commercial and procurement workflows, but the reported model met an average 55% of evaluation criteria—a research result, not evidence that agents can safely run contract processes without human review.

OpenAI said Ironclad staff and OpenAI employees who use the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and adapting reusable clauses to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. The tasks were graded against rubrics with 8 to 50 criteria, depending on complexity.

For training, Ironclad supplied hosted copies of its software in which models could practise. OpenAI said it generated synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. The company said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at the stated high setting. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used in Astra’s development scored 63.7%. On one showcase task, Astra met about 94% of the criteria. These figures describe performance on this set of research tasks, not a general measure of contract work.

At a glance
reportWhen: Published October 6; the source does no…
The developmentOpenAI published details of work with Ironclad to train and evaluate an AI model on legal, commercial and procurement workflows inside contract-management software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The reported average is a share of rubric criteria met, not a percentage of tasks completed successfully. That difference matters in workflows where a missed requirement can invalidate the result. A procurement process might need Finance approval above a spending threshold, Security review for specified requests and Legal review for nonstandard terms. Getting two of those three rules right does not make the workflow two-thirds safe: the missing approval could let a purchase proceed without a required check.

OpenAI’s account itself identifies the risk of an agent losing track of a business rule during a multi-step task and says human oversight remains necessary. The comparison with an experienced user also cannot yet establish a productivity gain. OpenAI’s time figures are simulated estimates based on assumed processing and generation speeds, not observed customer time savings. The results suggest progress on a difficult software task, while leaving accuracy and verification central to any deployment decision.

The collaboration also points to a potential change in how AI systems are developed: software companies may provide controlled environments and domain expertise so models can be tested against real product workflows. That may make agents more capable inside specialised applications. It also puts greater weight on the software’s underlying business rules, records and controls, especially if customers increasingly interact through agents rather than product screens.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Ironclad Became the Test Environment

The OpenAI post was one of two publications highlighted in the supplied source for October 6. The source says another OpenAI release—722 mathematics manuscripts—received more attention, while the Ironclad post drew less. Some AI news trackers reportedly interpreted “Ironclad” as the name of an agent framework; in this account, it refers to the contract-management vendor.

OpenAI framed the project as training models to understand business rules, carry out multi-step work in specialised software and check completed work against original requirements. The setup combined hosted product copies, task rubrics and synthetic exercises based on public filings. The evaluation therefore offers a bounded test of selected workflows; it does not establish performance across Ironclad’s full product or other companies’ software.

OpenAI also said it is inviting a small number of software companies to work on tasks that current agents cannot reliably complete. It asked prospective partners to bring a concrete example of a failing task, people with detailed knowledge of the work, a secure test environment and data suitable for research. The source identifies this as an invitation, not a broad, completed program or a commitment from named additional partners.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Evaluation

The supplied account does not give enough detail to determine how the criteria were weighted, which specific requirements the models missed, or how scores varied across all 11 tasks. An average can conceal failures that would matter more than others in a legal or procurement setting. The 94% result applies to one showcase task and cannot establish typical performance.

It is also unclear whether the results have been independently evaluated, how the models would perform on live customer workflows, or what levels of human checking would be required in practice. OpenAI’s reported time estimates are simulated, and the source provides no measured customer outcomes. The account does not specify the year of the October 6 publication or provide a date for further results.

Amazon

AI document review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Partner Testing Could Establish

OpenAI says it plans to work with a small number of software companies on tasks that current agents struggle to complete. The next useful evidence would include task-by-task results, a clear account of which criteria were missed, and tests of whether agents can preserve approvals and other controls across longer workflows.

For companies considering agents in contract, finance or customer-record systems, the immediate implication is to ask vendors for the evaluation details behind summary scores: which rules failed, how errors are caught, what data was used, and where people must review work. The source does not announce a deployment timetable or show that Ironclad customers have received these capabilities. Until such details and real-world results are available, the collaboration remains a research and product-testing development.

Amazon

contract analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this report?

Ironclad is a contract-management software company, not the name of a new AI agent framework. OpenAI described using hosted copies of its product for training and evaluation.

What does the 55% result measure?

It is GPT-6 Astra’s average share of evaluation rubric criteria met across the reported tasks. It is not the share of tasks completed successfully, and the source does not provide a full breakdown of missed criteria.

Did the agent save customers time?

The report does not establish customer time savings. OpenAI’s figures of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol are simulated estimates, not measured results from customer use.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can these agents run contract workflows without human review?

The reported findings do not establish that. OpenAI’s account says oversight remains necessary when an agent may lose track of a business rule, and the average rubric score was 55%. The source does not describe a customer deployment or a validated standard for unsupervised use.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Customer service + BPO. The operational-scale displacement.

Empirical evidence shows customer service and BPO sectors are experiencing widespread AI-driven workforce displacement, with hybrid models emerging as the operational norm.

Google’s Quantum Computer Breaks 1000-Qubit Barrier in Computing Milestone

Google’s quantum breakthrough surpassing 1,000 qubits signals a new era in computing, revealing the potential and challenges ahead.

Best AI Mini PC Picks For 2026: Top 10 List

Explore the top 10 AI mini PCs for 2026, featuring powerful processors, expandability, and connectivity options tailored for AI workloads and future-proofing.

How Mixed Reality Differs From Virtual Reality in Daily Use

Guided by the differences between mixed reality and virtual reality, discover how each technology can transform your daily experiences—continue reading to find out more.