🔍 Read the full analysis: OpenAI’s Agents Are Training In Software. Ironclad’s Terms Deserve A Read on ThorstenMeyerAI.com
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
TL;DR
OpenAI described training and evaluating a frontier model in hosted copies of Ironclad’s contract-management software. The model met an average 55% of rubric criteria across 11 tasks; OpenAI says its time estimates were simulated, and the results do not establish customer productivity gains or readiness for unsupervised use.
OpenAI published details on October 6 of a collaboration with Ironclad, a contract-management software company, to train and evaluate a frontier model on tasks inside hosted copies of Ironclad’s product. The work tested 11 legal, commercial and procurement workflows, but the reported model met an average 55% of evaluation criteria—a research result, not evidence that agents can safely run contract processes without human review.
OpenAI said Ironclad staff and OpenAI employees who use the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and adapting reusable clauses to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. The tasks were graded against rubrics with 8 to 50 criteria, depending on complexity.
For training, Ironclad supplied hosted copies of its software in which models could practise. OpenAI said it generated synthetic tasks from publicly filed contracts in the SEC’s EDGAR database and filtered them to remove personal information. The company said the work did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol at the stated high setting. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal OpenAI model used in Astra’s development scored 63.7%. On one showcase task, Astra met about 94% of the criteria. These figures describe performance on this set of research tasks, not a general measure of contract work.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The reported average is a share of rubric criteria met, not a percentage of tasks completed successfully. That difference matters in workflows where a missed requirement can invalidate the result. A procurement process might need Finance approval above a spending threshold, Security review for specified requests and Legal review for nonstandard terms. Getting two of those three rules right does not make the workflow two-thirds safe: the missing approval could let a purchase proceed without a required check.
OpenAI’s account itself identifies the risk of an agent losing track of a business rule during a multi-step task and says human oversight remains necessary. The comparison with an experienced user also cannot yet establish a productivity gain. OpenAI’s time figures are simulated estimates based on assumed processing and generation speeds, not observed customer time savings. The results suggest progress on a difficult software task, while leaving accuracy and verification central to any deployment decision.
The collaboration also points to a potential change in how AI systems are developed: software companies may provide controlled environments and domain expertise so models can be tested against real product workflows. That may make agents more capable inside specialised applications. It also puts greater weight on the software’s underlying business rules, records and controls, especially if customers increasingly interact through agents rather than product screens.
As an affiliate, we earn on qualifying purchases.
How Ironclad Became the Test Environment
The OpenAI post was one of two publications highlighted in the supplied source for October 6. The source says another OpenAI release—722 mathematics manuscripts—received more attention, while the Ironclad post drew less. Some AI news trackers reportedly interpreted “Ironclad” as the name of an agent framework; in this account, it refers to the contract-management vendor.
OpenAI framed the project as training models to understand business rules, carry out multi-step work in specialised software and check completed work against original requirements. The setup combined hosted product copies, task rubrics and synthetic exercises based on public filings. The evaluation therefore offers a bounded test of selected workflows; it does not establish performance across Ironclad’s full product or other companies’ software.
OpenAI also said it is inviting a small number of software companies to work on tasks that current agents cannot reliably complete. It asked prospective partners to bring a concrete example of a failing task, people with detailed knowledge of the work, a secure test environment and data suitable for research. The source identifies this as an invitation, not a broad, completed program or a commitment from named additional partners.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Evaluation
The supplied account does not give enough detail to determine how the criteria were weighted, which specific requirements the models missed, or how scores varied across all 11 tasks. An average can conceal failures that would matter more than others in a legal or procurement setting. The 94% result applies to one showcase task and cannot establish typical performance.
It is also unclear whether the results have been independently evaluated, how the models would perform on live customer workflows, or what levels of human checking would be required in practice. OpenAI’s reported time estimates are simulated, and the source provides no measured customer outcomes. The account does not specify the year of the October 6 publication or provide a date for further results.
As an affiliate, we earn on qualifying purchases.
What Partner Testing Could Establish
OpenAI says it plans to work with a small number of software companies on tasks that current agents struggle to complete. The next useful evidence would include task-by-task results, a clear account of which criteria were missed, and tests of whether agents can preserve approvals and other controls across longer workflows.
For companies considering agents in contract, finance or customer-record systems, the immediate implication is to ask vendors for the evaluation details behind summary scores: which rules failed, how errors are caught, what data was used, and where people must review work. The source does not announce a deployment timetable or show that Ironclad customers have received these capabilities. Until such details and real-world results are available, the collaboration remains a research and product-testing development.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in this report?
Ironclad is a contract-management software company, not the name of a new AI agent framework. OpenAI described using hosted copies of its product for training and evaluation.
What does the 55% result measure?
It is GPT-6 Astra’s average share of evaluation rubric criteria met across the reported tasks. It is not the share of tasks completed successfully, and the source does not provide a full breakdown of missed criteria.
Did the agent save customers time?
The report does not establish customer time savings. OpenAI’s figures of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol are simulated estimates, not measured results from customer use.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It said the work used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can these agents run contract workflows without human review?
The reported findings do not establish that. OpenAI’s account says oversight remains necessary when an agent may lose track of a business rule, and the average rubric score was 55%. The source does not describe a customer deployment or a validated standard for unsupervised use.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
