OpenAI’s October 6 research report describes training and evaluating computer-use agents on contracting tasks developed with Ironclad. The tasks include configuring agreements, procurement approvals and reusable legal terms. The important unit of evaluation is not a click or a completed screen: it is whether the configured process meets the requirements it was given.
That is a useful way to approach a business workflow. It also makes the reported average score something to inspect, not a reason to enable the finished process. A workflow can satisfy many criteria while missing the approval rule that matters most.
What the research measured
OpenAI says Ironclad employees and Ironclad users at OpenAI helped identify 11 tasks across legal, commercial and procurement work. Each task was evaluated against 8 to 50 criteria, depending on complexity. Ironclad supplied hosted software environments, and OpenAI developed synthetic training tasks around representative workflows.
The report compares GPT-6 Astra at Max reasoning with GPT-5.6 Sol at High reasoning, the settings where each scored highest. Astra’s mean rubric score was 55.0%, compared with 41.6% for Sol. Estimated average time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol.
Those are results on the 11 research tasks, not all Ironclad workflows. The time figures are simulated estimates based on assumed model processing and generation speeds, not measured customer time savings. The report also describes a stronger internal development model; that result should not be substituted for the released model’s score.
A mean rubric score is not a complete-task success rate. The report does not establish that 55.0% of production workflows can be safely run without review. It measures satisfaction of criteria across this research evaluation. It does not publish a universal deployment threshold or certify a particular customer’s approval configuration.
Turn the business rule into cases that can fail
OpenAI uses a procurement example: Finance may need to approve purchases above a spending threshold, Security may need to review certain requests, and Legal may need to review nonstandard terms. An agent must turn the requirements into intake forms, templates, approval rules and a final agreement record.
For the Finance rule, the report explicitly calls for checking requests above and below the threshold. This is more informative than checking that an approval step exists somewhere in the interface. The step must occur for the requests that require it and follow the intended path for those that do not.
For your own pilot, ask the process owner to specify the exact threshold behavior, including what happens at the boundary. Then test those cases in an approved nonproduction environment. This is BaristaLabs guidance, not an additional Ironclad feature or a disclosed OpenAI test protocol. Do not invent a boundary rule because the prose says only “above.”
Use the same approach for an exception. Identify the condition that should route work to a reviewer, what that reviewer must receive and what must not happen before approval. Give the test an observable result: the configured route and resulting record, rather than an agent’s statement that it completed the task.
Keep essential rules separate from the mean
A summary score helps compare attempts. It does not tell the process owner which unmet criteria are harmless presentation issues and which prevent safe use. OpenAI’s report does not disclose a release policy that assigns those consequences for you.
Before the pilot, designate the rules that must hold for the proposed scope. Keep each required condition visible beside its actual test result and supporting evidence. If an essential approval path fails, do not let other successful criteria cancel that failure in an average. Whether the process can proceed is a business control decision, not just a scoring decision.

This does not mean every criterion has the same consequence or that there is a single suitable threshold for every organization. It means the owner must decide what is required and inspect those conditions directly. The research report’s emphasis on full workflow behavior gives a reason to do that work; it does not replace it.
Our Chatham reviewer-comparison guide addresses whether faster evidence preparation agrees with experienced reviewers on the same cases. This Ironclad research poses a different question: did an agent configure the business rules so the resulting software behaves correctly across the required cases?
Separate research progress from a local deployment decision
OpenAI says its synthetic tasks used contracts publicly available in the SEC’s EDGAR database after filters designed to remove personal information. It says it did not use OpenAI customer data, internal contracts or nonpublic Ironclad customer contracts for training or evaluation. That describes this research dataset; it is not permission to send your contracts to an unapproved tool.
Keep access, retention and testing arrangements appropriate to your organization. Let the agent work in a bounded environment, have an authorized owner review the configuration, and verify the resulting behavior before it affects real requests. If a required rule fails, preserve the case and evidence so a revision can be tested against it.
The research shows progress at meeting detailed task requirements with less estimated time per attempt. For a team considering computer-use automation, the practical next step is smaller: choose one workflow, write down the essential rules and show that the configured process follows them. A higher average is useful evidence about model progress. It is not a release decision for your workflow.
Sources
- OpenAI: Advancing computer use with Ironclad, October 6, 2026; checked October 7, 2026. Includes research-evaluation and dataset footnotes.
The task descriptions, rubric results and time limitations above come from OpenAI’s research report. The required-rule review and local pilot design are BaristaLabs recommendations. We have not run the research evaluation or tested an Ironclad customer configuration. This is workflow evaluation guidance, not legal advice.
Computer-use workflow evaluation
Test one workflow’s essential rules
BaristaLabs can help turn a bounded workflow into test cases, required-rule checks and a reviewable deployment decision.
Bring a sanitized process outline, not contracts or credentials.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
