OpenAI introduced Astra for Law on September 17, combining GPT-6 Astra with a legal search index and instructions for legal analysis and writing. For law firms and legal software teams, the useful development is the configuration: a general model gains sources and guidance designed for a specific kind of work.
OpenAI reports better legal research results than the same model using web search alone. That makes the product worth evaluating, but the launch's largest numbers describe different things. A large source collection tells a firm what may be searchable. It does not tell a lawyer whether a particular answer is correct.
What the legal configuration adds
According to OpenAI's announcement, Astra for Law combines the model with tools and instructions intended to find relevant authorities, locate supporting passages, and apply them to legal questions. Its legal search index covers U.S. case law, statutes, regulations, court rules, and administrative decisions across more than 230 million URLs.
The instructions also address how the system uses the research. OpenAI describes tasks such as distinguishing a court's holding from other observations, addressing cases that weaken an argument, and explaining how a contract exception shifts risk. Those are descriptions of the intended behavior, not evidence that every generated analysis performs them correctly.
OpenAI positions the index as complementary to licensed content and specialist products, including those from Thomson Reuters. A firm should therefore evaluate how the new research tool fits with its existing sources, rather than assume that access to a larger index removes the need for them.
Read the benchmark as a comparison of complete setups
OpenAI tested Astra for Law on 200 U.S. legal research questions from the private validation set of Vals AI's Legal Research Bench. At the highest reasoning effort for both systems, it reports 54.0% overall correctness for Astra for Law, compared with 38.7% for GPT-6 Astra using web search alone.
That comparison supports a narrower conclusion than “the model is ready for legal work.” On this evaluation, OpenAI's legal configuration performed better than its web-search configuration. The result does not isolate the effect of the index from the instructions or establish how either setup would perform on a firm's own matters.
It is also a vendor-reported evaluation, not a BaristaLabs hands-on test or an independent reproduction. The relevant question for a buyer is whether the improvement survives on the jurisdictions, question types, and source requirements that its lawyers actually use.
For a local evaluation, keep the research questions and review criteria the same across the tools being compared. Ask a qualified reviewer to examine the cited authority, the quoted passage, and the conclusion drawn from it. Record a useful source that was missed separately from a source that was found but misread. Those failures need different remedies: adding access to documents cannot, by itself, fix faulty interpretation.
Source coverage and answer quality have different denominators
OpenAI says its work with Free Law Project brings CourtListener's case-law collection into the research experience. Free Law Project's coverage page describes its database as covering more than 99.9% of precedential case law published in the United States. The page frames its completeness assessment as of 2023 and describes ongoing collection work.
That is a statement about a collection and a defined class of legal material. It is not a claim that Astra for Law answers 99.9% of legal questions correctly, nor that every relevant legal document is included. The 230 million figure, meanwhile, counts URLs across the broader search corpus; it is not a count of unique court decisions.

Keep those distinctions in a procurement discussion. Coverage can help identify whether the needed material is available. A lawyer still needs to assess whether the answer found the right material, used the right passage, and treated the authority appropriately. A cited opinion can be present in a collection without supporting the proposition attached to it.
Start with a research evaluation, not a replacement plan
Availability also limits the immediate decision. OpenAI says Astra for Law is initially offered to selected law firms through Trusted Access in ChatGPT and Codex, with API access coming soon. The announcement does not establish general API availability or a date when every firm can use it.
For a firm with access, BaristaLabs recommends starting with a bounded set of research questions whose answers experienced lawyers can check. Use material approved for that evaluation. Our guide to a permission-bound legal AI pilot discusses that separate access decision in the context of Google's legal tools. Preserve the questions, retrieved passages, answers, and reviewer findings so the team can distinguish a helpful research result from one that merely looks complete.
For a legal software team waiting for the API, the same work can define acceptance criteria before integration begins. Identify which sources an answer must use, what a reviewer must be able to trace, and which errors would keep the feature out of client-facing work. Confirm access and product terms before committing to a delivery date.
The launch offers a concrete reason to test a specialized research configuration. The decision to adopt it should follow the quality of its answers on the firm's work, not the size of the collection it can search.
Sources checked September 21, 2026: OpenAI's September 17 launch announcement and Free Law Project's CourtListener coverage documentation. Product capabilities, availability, and benchmark results above are source claims; evaluation practices are BaristaLabs recommendations.
AI evaluation
Define what a useful research answer must contain
BaristaLabs can help structure the evaluation workflow, source traceability, and review records. Your legal team supplies the legal judgment and acceptance criteria.
Bring the workflow and evaluation goals, not confidential client material.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
