🔍 Read the full analysis: What Users Should Know About OpenAI Agents Training Inside Software on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training and testing GPT-6 Astra inside hosted copies of contract-management company Ironclad’s software, using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while its estimated completion times were simulated rather than measured customer savings. OpenAI is inviting a small number of software companies to explore similar partnerships.
OpenAI said on Oct. 6 that it trained and evaluated GPT-6 Astra on selected legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The report describes a model-development approach that places agents in specialized business software, but its results show substantial room for improvement: Astra met an average 55% of task criteria, and its completion-time estimates were simulated, not observed customer savings.
OpenAI said employees from both companies selected 11 tasks covering work such as setting up nondisclosure agreements, creating procurement approval processes and adjusting reusable contract clauses to match a requester’s jurisdiction. OpenAI estimated that an experienced user would need 30 to 40 minutes for each task. Tasks were scored against rubrics containing between eight and 50 criteria, depending on complexity.
Ironclad provided hosted copies of its product for models to practise in. OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average 55.0% of rubric criteria, compared with 41.6% for GPT-5.6 Sol at a high setting. Astra’s reported average time per attempt was 19.2 minutes, compared with 37 minutes for Sol. An internal model used in Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria. These are rubric results across the research tasks, not a measure of the share of tasks completed correctly.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Partial Workflow Success Matters
The work points to a possible change in how software agents are developed: instead of learning only general computer actions, models can be trained and evaluated against specific business workflows and their rules. That could matter to organizations hoping to automate work across legal, procurement and other specialized systems. It also gives software vendors a way to test where agents fail in the products their customers use.
But a rubric average does not show whether an agent can safely handle a high-consequence workflow. A procurement process that misses a required Finance approval, Security review or Legal check may be unusable even if it satisfies other criteria. OpenAI’s reported 55% average should not be read as 55% of tasks completed, or as proof that the system is ready to operate unsupervised.
OpenAI’s post acknowledges that an agent can lose track of a business rule during a task and says human oversight remains important. For buyers, the practical issue is not just whether an agent can operate an interface, but whether it can consistently preserve approvals, controls and auditability. For software companies, partnering may help improve agents in their products, while increasing the importance of the rules and records behind those interfaces.
As an affiliate, we earn on qualifying purchases.
Inside Ironclad’s Training Setup
The post was titled “Advancing computer use with Ironclad.” Ironclad is a contract-management software company, not the name of a new agent framework. OpenAI presented the work as a way to train models to understand business rules, carry out multi-step tasks in specialized software and check completed work against original requirements.
OpenAI’s time figures need a separate qualification. The company said the estimated 19.2 minutes for Astra and 37 minutes for Sol were calculated using assumed processing and generation speeds. They are simulated estimates, not time measured in customer deployments, and apply to the 11 research tasks rather than Ironclad workflows in general. The comparison with experienced users’ estimated 30-to-40-minute task times therefore does not establish a real-world productivity gain.
OpenAI also said it is inviting a small number of software companies to partner on tasks current agents cannot reliably complete. The company said prospective partners should bring concrete examples of failures, people with deep knowledge of the work, a secure test environment and data suitable for research.
Deployment Readiness Still Unproven
The report does not establish how GPT-6 Astra would perform across Ironclad’s full product, varied customer data or everyday production use. The 11 selected tasks are a limited research set, and an average rubric score does not disclose which individual requirements were missed across all attempts. OpenAI’s example of about 94% on one task does not resolve that broader question.
It is also unclear how the results would change with repeated trials, different task wording or workflows involving unusual contract terms. The source material provides no measured customer time savings, deployment results or evidence that Astra can reliably complete these processes without human review. OpenAI’s data-use description is its own account; the report does not independently verify those practices.
Software Partnerships and Further Testing
OpenAI says it plans to work with a small group of software companies on tasks that agents still cannot reliably complete. The next useful evidence would include results across more workflows, clearer reporting of which rubric criteria fail, and evaluations showing how models handle business rules over an entire process. Any claims about real productivity should be based on observed use, rather than simulated task times.
Organizations considering agents in contract or procurement systems can use the report as a prompt to ask vendors which requirements are tested, how missed approvals are detected, what human review is required and what records are retained. OpenAI has not provided a public timeline for the prospective partnerships or announced customer deployments arising from them.
Key Questions
What did OpenAI report doing with Ironclad?
OpenAI said it trained and evaluated GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software.
Does Astra’s 55% score mean it completed 55% of tasks?
No. OpenAI reported that Astra met an average 55% of the rubric criteria across the tasks. That is not the percentage of tasks completed successfully.
Were the reported time savings measured in companies?
No. OpenAI described the model times as simulated estimates based on assumed processing and generation speeds. The report did not measure customer time savings.
Did OpenAI say it used private customer contracts?
OpenAI said it used synthetic tasks based on publicly filed contracts from the SEC’s EDGAR database and did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
Can companies use agents for these workflows without review?
The reported results do not show that. OpenAI said human oversight remains important, and the average criteria score leaves unanswered whether agents can reliably preserve every required approval and control.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
