Chatham Financial used Codex to build a trade validation application that gathers transaction evidence, compares key terms, and flags discrepancies for review. In an October 2 customer story published by OpenAI, Chatham reports that early measurement reduced review time from approximately 30 minutes to under 4 minutes.
That is a reported timing result, not a published accuracy rate. The same story says Chatham is comparing the application’s results with those of experienced reviewers before expanding automation. For a team considering a similar workflow, that comparison is the more useful next step: determine what the faster process finds, what it misses, and how much work remains for the reviewer.
What the application is reported to do
Chatham’s Controls and Data Integrity team validates that transaction records reflect what the client authorized and what was executed. OpenAI describes the application as collecting supporting evidence, comparing key terms, and flagging discrepancies. The story places it within Chatham’s Process Zero work, which starts by identifying a workflow’s minimum inputs and evidence and deciding where human judgment is essential.
The distinction matters for a pilot. An application that prepares evidence and identifies possible discrepancies has a different responsibility from a system that approves a transaction or changes a record. The story does not establish that this application can independently make those decisions.
OpenAI also describes Chatham using GPT-5.6 for AI features across its business and Codex to build tools. It does not identify the inference model specifically used in the trade validation application. Building an application with Codex is not evidence that every comparison inside it is performed by a particular model.
Read the timing result within its stated limits
The reported reduction is attributed to Alex Nordlinger, co-head of Chatham’s AI Advisory practice. He describes it as early measurement and says the company is validating performance against real transactions and experienced reviewers before expanding automation.
The inspected story does not disclose the sample size, an accuracy rate, a missed-discrepancy rate, or an independent audit. It also does not define the exact start and end of the measured review. Those gaps do not invalidate the timing observation, but they prevent it from answering every question a prospective user should ask.
For your own trial, define the timing boundary before collecting results. If the existing process includes gathering documents, checking terms, investigating exceptions, and recording a decision, measure the proposed process over the same boundary. An application’s processing time alone would not be comparable to the full human workflow.
Keep setup, integration maintenance, and reviewer correction work visible as separate costs. This is BaristaLabs evaluation guidance, not additional measurement from Chatham. A shorter preparation step may be valuable even when review still takes substantial effort; the important point is to describe the actual work saved rather than claim a complete workflow saving from one timer.
Compare the app and reviewers on the same cases
A useful local evaluation would give the application and experienced reviewers the same approved evidence for the same cases. Start with a defined case set from one trade type or one similarly bounded business process. Include ordinary cases and known exceptions, without assuming a few convenient examples represent the wider workload.
Record which discrepancies each path identifies, which source evidence supports them, and which cases cannot be resolved from the available material. Have a qualified owner examine disagreements. The reviewer’s initial result is a comparison point, not automatically an infallible answer; a disagreement may reveal missing evidence or a mistake on either side.

Pay particular attention to a missed discrepancy that the normal process would have caught. Also inspect false alarms, because a fast application that produces many unnecessary investigations can move work into the review queue. Track unresolved cases explicitly rather than treating silence as agreement.
This same-case comparison is a proposed test design. OpenAI reports that Chatham is comparing results with experienced reviewers, but it does not publish the protocol described here. We have not run the application or validated the company’s measurement.
Decide what must remain with the reviewer
For a first local pilot, keep the existing decision process in control while the application prepares its findings. Give the reviewer the source evidence and the proposed discrepancy, not only a summary that says the case passed. Define what happens when a document is missing, two records conflict, or a term needs professional interpretation.
Our finance-agent workflow guide covers the broader boundary between preparing work and authorizing action. This case adds a narrower evaluation question: does a faster comparison step leave the reviewer with enough evidence to make the same decision, and where do the two paths disagree?
Do not send customer transactions to an unapproved tool merely to reproduce a vendor story. Use the access, retention, and review arrangements appropriate to your organization. This article is workflow evaluation guidance, not financial, investment, or legal advice.
Expand only after the comparison is understood
Chatham says it plans to extend the application to additional trade types while maintaining controls and professional oversight. That is a plan, not a reported result across every trade type.
For another organization, a new document format, product type, or exception can change what the application needs to recognize. Review the evidence coverage and disagreements again when the scope changes. Preserve corrections with the case reference and reviewer rationale so the team can test whether a revision actually fixes the problem. Our Codex feedback-loop guide develops that correction-to-evaluation pattern for document-heavy work.
Before a pilot starts, agree on who resolves disagreements, which kinds of misses stop expansion, and what review effort is acceptable. Choose those criteria for the consequences of the work; there is no universal accuracy threshold in the Chatham story to copy.
The early timing result gives teams a reason to examine this kind of evidence preparation. The ongoing reviewer comparison gives them a way to examine it responsibly. Start with one bounded workflow and a record of what both paths found before deciding that faster review is ready for wider use.
Sources
- OpenAI: Chatham scales its capital markets expertise with OpenAI, October 2, 2026; checked October 4, 2026. This is a vendor-published customer story containing attributed early company measurement, not an independent evaluation.
Product use and the early timing observation above come from that story. The same-case comparison design, timing-boundary recommendations, and pilot acceptance criteria are BaristaLabs guidance.
Evidence-led workflow pilot
Evaluate one review workflow before expanding it
BaristaLabs can help map the inputs, reviewer comparison, exception handling, and timing boundary for a bounded pilot.
Bring a sanitized process outline, not customer transactions.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
