An agent trace can contain more than the answer you saw on screen. That matters when a team uploads a failed run to a public repository, pastes a debugging transcript into a forum, or sends a full log to a supplier.
On September 30, OpenAI reported disrupting a coordinated campaign that extracted protected model reasoning through manipulated interactions. The immediate decision for a business using agents is not whether to abandon a model. It is whether its team treats unreadable reasoning blocks as safe to share. The report and the research it cites give a concrete reason not to make that assumption.
What OpenAI reported, and what it did not
OpenAI says the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated interactions so that reasoning normally withheld from the final answer became visible to the requester.
One observed pattern involved copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe it. The distinction matters: a protected artifact can become dangerous through how a system accepts and processes it, even when the underlying encryption has not been broken.
OpenAI dates the initial activity to July 1. It reports high-volume spikes on July 24 and 25 involving 16,000 requests from more than 4,000 users, and related prompt-pattern activity across a cluster of more than 15,000 users. It says that cluster was fully disrupted by July 28. Those request counts describe attempted, not necessarily successful, extractions.
The company attributes a core cluster to individuals associated with Moonshot AI, which develops Kimi, while saying it is unclear whether all observed operators came from a single actor. That is OpenAI’s attribution, not an independently established finding in this article. The operational lesson does not depend on accepting a broader claim about every user or every competing model.
Why unreadable blocks still belong in private records
OpenAI links to the August research paper “Stealing Reasoning Traces from Proprietary LLM APIs”. The researchers describe encrypted reasoning blocks returned to clients and passed back in subsequent requests. Their reported attack took advantage of compatibility across sessions, users, and models within a provider’s ecosystem.
The researchers also report extracting private information from reasoning blocks found in public repositories. These are research findings from the tested systems, not a claim that every current API remains vulnerable. OpenAI says it confirmed the disclosed attack paths and used the findings to accelerate mitigations.
For a team sharing logs, the practical point is simpler than the attack mechanics. If you cannot read a block, you cannot verify from its appearance that it contains no customer information, credentials, or sensitive intermediate work. Publishing it also hands another party an artifact that your application may normally send back to a model.
An unreadable reasoning block is not a reviewed public record. Treat it as sensitive workflow state unless your supplier provides a supported basis for doing otherwise. The same caution applies when a log is labeled a debugging trace rather than a customer conversation.
Preserve the original; share a reviewed copy
This is not an argument to delete useful incident evidence. Our earlier SAFE incident-evidence analysis explains why reconstruction needs connected records. The separate question here is who receives those records and in what form.
BaristaLabs recommends starting with one agent workflow that produces exportable traces:
- Locate the records. Check the application logs, trace store, exported transcripts, repository fixtures, and supplier support attachments. Establish which ones include reasoning blocks or complete request and response payloads.
- Keep the original private. Preserve necessary evidence in an access-controlled location under the organization’s retention policy. Record who can retrieve it and why. Do not silently overwrite an incident record with a sanitized version.
- Create a separate sharing copy. Remove reasoning blocks and secrets, minimize customer data, and review the remaining visible text. A readable final answer can still contain sensitive information. Share only what the recipient needs through an approved channel.

Test that process on a synthetic, non-sensitive failed run before a real support incident. Ask someone other than the person exporting the trace to inspect the sharing copy. The useful result is evidence that the team can produce a small, understandable reproduction without exposing the full run. Do not test by replaying another person’s reasoning or uploading a real customer trace publicly.
If sensitive traces are already public, restrict further sharing and involve the responsible security or privacy owner. Assess what was exposed, preserve necessary evidence, and rotate any credentials found to be disclosed. Removing a public file is not proof that copies no longer exist.
Ask about the endpoint you actually use
OpenAI says it strengthened protections across users, workspaces, organizations, and model families. It reports closing a pathway for replaying another user’s encrypted reasoning and recovering its contents, and adding checks that can hold streamed output that might expose reasoning. It also describes account enforcement, infrastructure controls, expanded monitoring, and coordination with third-party services.
The company says the work is continuing. Its report specifically notes that partner-hosted deployments need the same protections as first-party services and that relevant controls are being propagated across cloud partners. That is not a blanket guarantee about every endpoint on the day you read this.
For a hosted deployment or an agent platform between your application and the model, ask the supplier which protections apply to your endpoint, how reasoning artifacts are scoped, what its logs retain, and how it handles support exports. Keep the response attached to the actual service and configuration. A general announcement is not evidence that your particular integration received every relevant change.
A provider’s mitigation claim does not replace your log-sharing controls. Those controls also reduce ordinary accidental disclosure, independent of whether the reported attack still works.
The useful next step is a small records review
Choose one workflow, one trace store, and one person responsible for sharing records. Confirm what is retained, who can access it, and how a reviewed copy leaves the organization. Then document the supplier questions you cannot answer locally.
That is a narrower decision than the stop-and-restart controls for an agent containment incident. Here, a workflow can run within its intended permissions and still produce records that should not be public. Keeping those records private is part of operating the workflow, not merely a response after something goes wrong.
Sources
- OpenAI, “Disrupting a coordinated model-distillation campaign,” September 30, 2026. Source for the campaign account, attribution, attempted-extraction counts, mitigations, and continuing work.
- Panfilov and coauthors, “Stealing Reasoning Traces from Proprietary LLM APIs,” submitted August 10, 2026. Source for the researchers’ reported reasoning-block vulnerability and public-log exposure findings. Research findings are not a verification of every current deployment.
The log-handling recommendations are BaristaLabs analysis, not new provider requirements or a certification of safety.
Before sharing an agent trace
Review what your workflow records and exposes
BaristaLabs can help inspect one agent workflow’s logging, sharing process, and supplier questions without publishing sensitive traces.
Bring the workflow description and logging configuration. Do not send credentials or unredacted customer records through the contact form.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.
