Skip to main content
Small Business AI

OpenAI’s GPT-6 guide gives teams a reason to review their agent instructions

OpenAI’s October 2 guide recommends consistent prompts, skills, and repository instructions. Start by agreeing on the deliverable, autonomy, and definition of done.

Sean McLellan profile photo

Sean McLellan

Lead Architect & Founder

5 min read
Constructed guidance aligns a task prompt, reusable skill, and repository guidance on one assignment covering deliverable, autonomy, and definition of done.
Constructed diagramBaristaLabs suggested instruction review, informed by OpenAI’s October 2 guide. This is not product UI or an enforced control.

The instructions for an AI workflow can live in several places: a task prompt, a reusable skill, and a repository file explaining the project. Those sources need to agree about the work the agent is expected to deliver. A detailed prompt will not help much if the reusable instructions still describe a different output or stop at an earlier stage.

OpenAI’s October 2 guide to the GPT-6 family makes that consistency an explicit recommendation. Alongside model choice, cost management, and long-running work, it advises teams to align prompts, skills, and repository instructions about the deliverable, independent action, and what counts as done.

For a small team already using coding agents or preparing an automation pilot, that is useful maintenance work. Before adding another instruction file, review the ones the workflow already loads.

What OpenAI recommends

The guide says to start with a clear assignment: the desired result, its audience, relevant context and constraints, and the completion criteria. Its advice on reusable skills is to keep descriptions short and clear about when they should run, load supporting details when needed, and replace rigid recipes with guidance suited to the models in use.

For AGENTS.md, OpenAI recommends explaining when particular documents and tests matter and explicitly authorizing safe routine work, such as local tests with disposable data and no production access. It also recommends clear decision boundaries instead of blanket instructions to always ask for permission.

The completion advice is equally concrete: define “done” to include implementing the change, running it, inspecting the result, and fixing failures. Name the decisions that still need human review. These are recommendations for how to instruct an agent, not a guarantee that an agent will follow every boundary or complete every task correctly.

The guide covers much more than instruction design, including caching, model selection, and long-running tasks. The review below focuses on one part of that guidance and turns it into a practical team exercise. It is a BaristaLabs suggested method, not an OpenAI benchmark or a new product feature.

Compare the instructions for one real task

Choose a repeatable task that the team can test safely, such as a small website change on a preview branch. Gather the task prompt, the skills that normally load for it, and the repository instructions. Read them together rather than editing each in isolation.

Check three questions across those sources:

  • What is the deliverable? A proposed patch, a working preview, and a production release are different outcomes. Specify which one the assignment requires.
  • What may the agent decide? Name routine choices it can make, the resources it may use, and the changes that require a person’s decision.
  • What evidence ends the task? Identify the checks that establish completion and the situations that leave the task blocked or incomplete.

For a website preview, a team might authorize local edits and tests using sanitized data, require a browser check of the changed page, and reserve publishing for a person. The exact boundary depends on the project. Write the same boundary wherever the workflow needs it, and remove stale instructions that contradict it.

Avoid copying a long checklist into every skill. A reusable skill should explain when its specialist guidance is relevant; the task assignment should supply the current scope. Repository guidance can point to the project’s actual tests and constraints. Repetition makes maintenance harder when one copy changes and another does not.

Make completion visible

OpenAI suggests a short handoff covering what changed, what was checked, and what still needs attention. That is a useful reporting structure because it gives the reviewer something specific to inspect.

Suggested completion report separates completed work, tests and observed results, and untested areas or remaining review decisions.
Constructed diagramBaristaLabs reporting guidance. A report records evidence; it does not grant release permission.

Ask the agent to identify the changed files or output, the checks it actually ran, and their observed results. It should name any area it did not test. If a browser check could not run, the report should say that rather than treating a successful code test as proof of the rendered page.

Untested work should remain visible in the handoff. A reviewer can then decide whether the remaining uncertainty is acceptable for a preview, requires another check, or blocks release. Keeping that distinction explicit also makes it easier to compare repeated pilot tasks without relying on the agent’s confidence about its own output.

A completion report does not grant permission to release. The account permissions, deployment process, and human review rules still determine who may change production. Our guide to why system prompts are not an agent control plane covers the separate enforcement problem. Consistent instructions help communicate the intended workflow; they do not replace technical access controls.

Test the revised assignment before sharing it

Run the revised instructions on a sanitized example using the model and tools the team intends to use. Inspect the result and the handoff. Did the agent produce the requested deliverable? Did it perform the relevant checks? Did it stop at a decision reserved for a person? If it could not complete a check, did it identify that limitation accurately?

Keep the test narrow enough that a person can verify the answer. Record the instruction versions with the result so a later change can be evaluated against the same task. OpenAI recommends measuring task success, latency, and cost per successful task before deployment; instruction cleanup should be tested within that broader evaluation, not assumed to create a performance improvement on its own.

Start with one assignment your team already understands. Agree on the result, the agent’s discretion, and the evidence needed to accept the work. Then make the instructions consistent with that agreement and test them before applying them across the team.

Agent workflow preparation

Make one agent assignment ready to test

BaristaLabs can help review the task prompt, reusable instructions, permission boundaries, and completion evidence for one workflow your team wants to automate.

Bring a sanitized task description and the instructions your team uses. Do not send credentials, customer records, or confidential repository content through the form.

Turn this idea into a pilot

Which workflow should go first?

Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.

  • 3-5 minutes
  • Deterministic score
  • No sensitive data
Check workflow readiness

Practical AI Workflow Notes

Want more practical AI operations ideas?

Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.

A useful next step if you’re still exploring and not ready to request a 20-minute workflow assessment.

Occasional emails. Practical workflow guidance only. Unsubscribe anytime.