AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Agents Are Training In Your Software. Is Ironclad’s Fine Print Clear? on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get audio and creator gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average of 55% of task criteria, while its reported time estimate was simulated rather than measured customer productivity. The project also signals OpenAI’s interest in working with software companies to train agents on specialized workflows.

OpenAI said it trained its frontier model GPT-6 Astra on 11 legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software, reporting that the model met an average of 55% of the criteria used to grade those tasks. The results describe a research test, not a customer deployment: OpenAI says its time estimates were simulated, and its own account says human oversight remains necessary.

Ironclad staff and OpenAI employees who use the product selected tasks such as creating nondisclosure agreements, setting up procurement approval processes and updating reusable contract language based on a requester’s chosen jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task. Depending on complexity, the researchers graded results against 8 to 50 criteria.

OpenAI reported that GPT-6 Astra met an average of 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol, a model listed as a high-effort comparison. An internal model used in Astra’s development reached 63.7%. On one example task, Astra met about 94% of the criteria. These are rubric scores across the test tasks; they do not mean Astra completed 55% of tasks or that its output was ready for use.

OpenAI said it constructed synthetic training tasks using contracts filed publicly in the U.S. Securities and Exchange Commission’s EDGAR database, with personal information filtered out. The company said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. It also reported estimated completion times of 19.2 minutes for Astra and 37 minutes for the comparison model, while specifying that these were simulated estimates based on assumed processing and generation speeds, not measured time savings for customers.

At a glance
reportWhen: Described in an OpenAI post published O…
The developmentOpenAI described training GPT-6 Astra on selected workflows inside Ironclad’s contract-management product and invited other software companies to partner on similar research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Rules Matter to Agents

The test addresses a practical question for businesses: can an AI agent follow a series of rules inside the specialized software where contracts and approvals are managed? OpenAI’s results suggest that agents can improve on selected tasks, but an average score of 55% of criteria met leaves substantial room for missed requirements. In contract and procurement work, partial completion can still create a consequential error.

For example, a purchasing workflow may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. An agent that handles two rules but misses the third has not produced a safe, nearly complete workflow. The missing control could allow a purchase to proceed without a required review. OpenAI’s account acknowledges that losing track of a business rule limits the work a company can confidently delegate, and says human review is still needed.

The project also matters to software vendors and their customers because the product itself becomes a place to train agents. Vendors may gain models that handle difficult tasks within their systems, while customers will need to know what the agent can do, what it can miss and how its actions are checked. If agents increasingly operate software on users’ behalf, a vendor’s value may depend less on its screens and more on the rules, records, audit trails and controls behind them. That is an implication of the project, not a result established by the test.

Amazon

contract management software with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Built

OpenAI’s October 6 post, titled “Advancing computer use with Ironclad,” describes a collaboration to train models on workflows in a real software product. Ironclad is a contract-management software company, not an agent framework. OpenAI framed the goal as teaching models to understand business rules, carry out multi-step work in specialized software and check the final result against the original requirements.

Ironclad provided hosted versions of its product for model practice. The selected tasks were designed by people familiar with the software and the work. OpenAI says it used publicly filed contracts to create synthetic tasks and that no non-public Ironclad customer data was used. The post presents Astra as the first frontier model trained through this approach, and invites a small number of other software companies to propose difficult tasks, provide domain experts and supply secure test environments and research-usable data.

The collaboration is therefore both a performance test and an outline of a possible research model: software companies contribute realistic workflows and environments, while OpenAI trains and evaluates models against them. The published figures cover the 11 selected tasks; they do not establish how agents perform across Ironclad’s product generally or across other companies’ systems.

What the Scores Do Not Show

The post does not establish how often Astra would complete an entire workflow correctly in routine customer use. A mean share of rubric criteria met can conceal which specific requirements failed, and the reported 94% result applies to one showcase task rather than the whole test. The source material does not give a task-by-task breakdown of missed controls or show how performance changes across a broader range of contracts and organizations.

The estimated times are also not observed working times or measured productivity gains. OpenAI says they are simulations for the 11 research tasks, not estimates covering Ironclad workflows generally. The account says personal information was filtered from public EDGAR filings and that specified customer and internal data were not used, but it does not describe every detail of data handling, model access or deployment safeguards. No customer rollout, independent evaluation or commercial performance result is established in the material.

The Next Test Is Safe Deployment

OpenAI says it is inviting a small number of software companies to partner on tasks that agents cannot reliably complete today. It asks prospective partners to bring a concrete example of failure, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research. The post does not name other partners or provide a schedule for additional results.

For buyers, the next useful evidence would include task-level results showing exactly which criteria agents miss, tests of complete workflows and details of how approvals, permissions and audit records are enforced. Until such evidence is available, the published Ironclad results support continued evaluation and supervised testing—not treating the time estimates as realized savings or letting an agent’s partial score stand in for a correct contract process.

Key Questions

What did OpenAI test with Ironclad?

OpenAI said it trained GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management product, then scored results against task-specific criteria.

Does the 55% score mean Astra completed 55% of the tasks?

No. OpenAI reported an average of 55% of rubric criteria met. That is not the share of tasks completed, and it does not show that the work was ready to use without review.

Did Astra save customers time?

The report does not show measured customer savings. OpenAI’s 19.2-minute estimate for Astra is simulated, based on assumed processing and generation speeds, and covers the 11 research tasks.

Did OpenAI use private Ironclad customer contracts?

OpenAI said it used synthetic tasks built from publicly filed SEC EDGAR contracts, with personal information filtered out. It said it did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.

Can companies let agents manage contract workflows now?

The reported results do not establish that agents can safely run these workflows without supervision. OpenAI’s account says human oversight remains necessary, and the average score leaves questions about missed requirements and controls.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Synthetic Data Pipelines: Generation, Labeling, and Governance

Ineffective data management hampers AI progress—discover how synthetic data pipelines for generation, labeling, and governance can transform your approach.

Technology Operations Signal Monitor: The Future Of Flipper Zero Development

A new role-filtered monitoring tool is being tested to track updates on Flipper Zero, aiding small software teams in early decision-making.

IdeaNavigator AI: One Evidence-Mined Idea a Day

IdeaNavigator AI now publicly ships one evidence-mined software idea daily, transforming problem complaints into validated development opportunities.

The Twelve Real Complaints About AI Tools in 2026 — A Reddit, Twitter, and GitHub Synthesis

In 2026, users on Reddit, Twitter, and GitHub report widespread issues with AI tools, highlighting discrepancies between marketed and actual performance.