🔍 Read the full analysis: What To Look For In Ironclad’s Fine Print On OpenAI Agent Training on ThorstenMeyerAI.com
Get smart everyday buys delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI’s October 6 post describes training GPT-6 Astra on legal, commercial and procurement workflows in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while estimated completion times were simulated rather than measured customer savings. The work also marks an invitation for other software companies to help train agents on specialized workflows.
OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. Astra met an average 55% of the criteria across 11 tasks, according to OpenAI—a research result that shows progress on specialized software workflows, but does not establish that agents can safely complete contract work without human review.
OpenAI’s post, titled “Advancing computer use with Ironclad,” describes a collaboration with the contract-management software company, not the launch of a new agent framework. Ironclad staff and OpenAI employees who use the product selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.
OpenAI graded each task against a rubric of 8 to 50 criteria, with the number depending on complexity. Its reported average share of criteria met was 41.6% for GPT-5.6 Sol at high effort and 55.0% for GPT-6 Astra at maximum effort. An internal model used during Astra’s development reached 63.7%. OpenAI also said Astra met about 94% of the criteria on one showcase task. Those figures describe rubric criteria, not the percentage of tasks fully completed or approved.
For training, Ironclad provided hosted copies of its product in which models could practise. OpenAI said it created synthetic tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. OpenAI’s reported time estimate fell from 37 minutes for GPT-5.6 Sol to 19.2 minutes for Astra, but the post says these are simulated estimates based on assumed processing and generation speeds—not measured time savings for customers.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Contract Work Matters
The results matter because contract workflows are judged not only by whether an agent completes several steps, but by whether it applies every required control. A procurement process might need Finance approval above a spending threshold, a Security review for certain requests and Legal review for nonstandard terms. If an agent misses one rule, the workflow could route a purchase without a required approval. An average score of 55% of criteria met cannot show whether the missed requirements were minor or consequential.
OpenAI’s post acknowledges that an agent can lose track of a business rule during a multi-step task and says human oversight remains necessary. For companies considering agents in contract, finance or customer-record systems, the practical question is not just whether the model can use the software. Buyers need to know which requirements it misses, how errors are detected and whether the system blocks unsafe or incomplete work before it is acted on.
The partnership model also has implications for software companies. Working with an AI developer could help improve agent performance on difficult workflows within a vendor’s product. But as agents become more capable of operating an interface, the vendor’s lasting value may depend less on its screens and more on the business rules, data, audit records and controls behind them. That is an industry implication, not a result established by the 11-task study.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Built
OpenAI framed the research around teaching models to understand company rules, carry out multi-step work in specialized software and check completed work against original requirements. The Ironclad test is notable because it places training and evaluation inside a vendor’s product environment, using tasks selected by people familiar with the work, rather than treating computer use as a general ability measured only through abstract tests.
OpenAI said it is inviting a small number of software companies to partner on workflows that current agents cannot reliably complete. It asked prospective partners to bring a concrete example of a failing task, people with deep knowledge of the work, a secure test environment and data suitable for research. The post presents the Ironclad project as an example of that approach; it does not announce a broad program with named partners beyond Ironclad.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
OpenAI’s published average does not identify which criteria Astra missed on each task or how severe those misses were. The source material also does not establish whether the results were independently evaluated, how performance changes across different organizations’ rules, or how often the model would make errors in routine customer use. The 11 selected research tasks are not evidence of performance across all Ironclad workflows.
The time figures also have limits: OpenAI says they are simulated, based on assumed model processing and generation speeds, and are not observed customer savings. The post does not report a customer deployment, a measured comparison of completed work quality in live use, or a basis for concluding that Astra is faster than an experienced person producing a correct result. Data-use statements in the post are OpenAI’s account of its research process.
As an affiliate, we earn on qualifying purchases.
The Next Test Is Reliable Execution
OpenAI says it plans to work with a small number of software companies on difficult agent tasks, but the post does not give a schedule, name additional partners or describe a public release plan for the Ironclad-trained capability. Any later announcement would need to clarify how partner data is handled, how tasks are selected and evaluated, and whether results extend beyond research environments.
For software vendors and prospective buyers, the next useful evidence would be task-level results showing which rules agents miss, how often they make those errors and what safeguards prevent an incomplete workflow from being treated as complete. Until those details and real-world performance measures are available, the Ironclad figures are best read as early research results, not proof that contract tasks can be delegated without review.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Ironclad in OpenAI’s October 6 post?
Ironclad is a contract-management software company. OpenAI’s post describes training and testing a model on workflows inside hosted copies of its product.
What does Astra’s 55% score mean?
It is the average share of rubric criteria met across 11 tasks, according to OpenAI. It does not mean Astra fully completed 55% of the tasks.
Did the study show customers save time?
No measured customer time savings were reported. OpenAI said the 19.2-minute Astra estimate was simulated, based on assumed processing and generation speeds.
Did OpenAI say it used Ironclad customer contracts?
OpenAI said it did not use non-public Ironclad customer data. It said synthetic tasks were built from publicly filed contracts in the SEC’s EDGAR database after filtering to remove personal information.
Can companies use agents for contract work without review?
The reported results do not establish that. OpenAI’s post says human oversight remains necessary, and the average score does not reveal which specific requirements were missed.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
