An AI pilot can produce useful evidence without earning expansion. A strong result on selected examples does not show how the workflow will behave with more users, sources, permissions, actions, or unusual cases. A weak result can also point to a fixable source or process problem instead of a reason to reject the full use case.
The post-pilot review should identify what the team tested, what passed, what failed, and what remains unknown. Then the pilot owner can choose one next state: expand, revise and retest, shadow longer, keep the workflow manual, or stop. Each decision needs a clear prerequisite, an owner, and a review date.
Restate the tested boundary before you judge the result
Start with the decision that the pilot was meant to support. The decision can be whether one workflow should receive more funding, broader access, a production build, or more permission. If the pilot started without a named decision, write the smallest decision that the collected evidence can support.
Next, record the tested boundary. The tested boundary is the exact workflow, people, sources, fields, permissions, actions, examples, dates, systems, and review conditions included in the pilot. It also names what the pilot did not test.
Keep this boundary beside every result. A pilot that drafted replies from approved documents under full human review did not test automatic sending. A pilot that used clean recent records did not test missing, old, or conflicting records. A test with one team did not establish how another team will review the same output.
The AI discovery and pilot comparison separates the purpose and outputs of each stage. It also states the limit directly: pilot evidence applies to the tested workflow and sample. It does not establish general return on investment, compliance, or readiness outside that boundary.
NIST's AI RMF Playbook Map guidance calls for a documented application scope based on the system's capability and context. It also calls for teams to compare expected benefits and costs with an appropriate benchmark. These points support a bounded review. NIST does not decide the next state for your pilot.
Review the evidence, including the work that people kept doing
Review representative cases, including outputs that did not appear in the final demonstration. Include cases that passed, needed edits, were rejected, failed, or stayed manual. Record important exceptions and untested conditions beside the sample.
Inspect the review work as part of the result. Record what evidence reviewers could see, which decisions they made, how long review took, and which corrections repeated. A correct draft can still be a poor workflow result when a person must search for the source, rebuild the reasoning, or correct the destination record.
Check the operating path as well. Record integration failures, permission denials, uncertain destination states, retries, and the result of any stop or rollback test. If the pilot never tested the stop path, state that limit before the team approves more permission.
The AI pilot proof guide provides detailed fields for the workflow baseline, source boundary, reviewer evidence, misses, exceptions, rollback, value signals, and owner decision. The first-pilot handoff guide explains how to preserve those materials for the next team. Use the evidence already collected; do not fill an empty field with a confident estimate.
Classify each material miss by its cause
A material miss is a result that changes the decision about value, review, safety, cost, or permission. Classify each material miss before you choose the next state. Different causes need different changes and different retests.
The classes below are BaristaLabs guidance. They are not NIST categories. Give each miss one primary cause and record contributing causes when the failure crossed more than one part of the workflow.
| Miss class | Evidence to inspect | What a change must address |
|---|---|---|
| Source or input | Missing, old, conflicting, incomplete, or incorrectly retrieved material | Repair, replace, restrict, or make the source visible; then rerun affected and ordinary cases |
| Process rule | An unclear policy, conflicting instruction, missing owner, or rule that people apply differently | Clarify the business rule and decision owner before changing the model or prompt |
| Model behavior | The approved sources and rules were available, but the model produced an unsupported, incomplete, or inconsistent result | Change the model, instructions, retrieval method, examples, or output constraint; then test the same failure class again |
| Integration | A source, tool, or destination failed, timed out, duplicated an action, or returned an uncertain state | Repair state handling, confirmation, retry, reconciliation, or rollback before another live action |
| Permission | The pilot read or attempted to change data, records, or actions outside the approved boundary | Reduce or correct access, verify the allowed action, and test the block as well as the permitted path |
| Review design | The reviewer lacked the source, policy, time, role, or choices needed for a sound decision | Change the review screen, queue, evidence, authority, or staffing and measure the resulting review work |
Miss class
Source or input
- Evidence to inspect
- Missing, old, conflicting, incomplete, or incorrectly retrieved material
- What a change must address
- Repair, replace, restrict, or make the source visible; then rerun affected and ordinary cases
Miss class
Process rule
- Evidence to inspect
- An unclear policy, conflicting instruction, missing owner, or rule that people apply differently
- What a change must address
- Clarify the business rule and decision owner before changing the model or prompt
Miss class
Model behavior
- Evidence to inspect
- The approved sources and rules were available, but the model produced an unsupported, incomplete, or inconsistent result
- What a change must address
- Change the model, instructions, retrieval method, examples, or output constraint; then test the same failure class again
Miss class
Integration
- Evidence to inspect
- A source, tool, or destination failed, timed out, duplicated an action, or returned an uncertain state
- What a change must address
- Repair state handling, confirmation, retry, reconciliation, or rollback before another live action
Miss class
Permission
- Evidence to inspect
- The pilot read or attempted to change data, records, or actions outside the approved boundary
- What a change must address
- Reduce or correct access, verify the allowed action, and test the block as well as the permitted path
Miss class
Review design
- Evidence to inspect
- The reviewer lacked the source, policy, time, role, or choices needed for a sound decision
- What a change must address
- Change the review screen, queue, evidence, authority, or staffing and measure the resulting review work
Do not treat every failure as a model problem. A new prompt does not repair an unavailable source, an unresolved refund rule, an uncertain write to a customer system, excessive access, or a reviewer who cannot see the evidence.
Use the pause-after-the-first-miss guide when one result could have changed a customer, financial, access, publishing, privacy, or record decision. Freeze the affected action while the owner inspects the cause and writes the restart condition.
Decide whether the gap is fixable and whether the business case remains useful
A fixable preparation gap has an identified cause, a feasible change, and a test that can show whether the change worked. Source cleanup, a clearer rule, corrected access, an integration repair, or a better review screen can meet this definition. The team still needs to decide whether the expected value justifies that work.
A gap changes the business case when the repair, review, exception work, maintenance, or failure cost removes the value that the pilot was meant to test. Failure cost is the customer, financial, legal, privacy, access, or operating consequence of a wrong, missing, or late result. A workflow can be technically repairable and still be a poor investment.
NIST Map 3.2 calls for teams to examine monetary and non-monetary costs from expected or realized errors in relation to the organization's risk tolerance. NIST Manage 2.1 also includes the resources needed to manage risk and viable non-AI alternatives. Apply that comparison to the local workflow, its manual baseline, and the evidence from the pilot.
Keep early value signals inside the same boundary. A shorter cycle-time sample, lower review time, cleaner intake, or fewer handoffs can guide the next test. It is not a general savings or production-readiness claim.
Choose one next state and name its prerequisite
The five states below are a BaristaLabs decision model. SBA's small-business AI guidance supports starting small and testing whether a tool adds value. It does not define these states or the evidence required for them.
| Next state | Evidence from the tested boundary | Prerequisite before the next review |
|---|---|---|
| Expand | The pilot met its agreed criteria on representative cases. Material misses are understood. Review work and failure costs remain acceptable. The stop or rollback path worked. | Name one scope increase, preserve the required review and evidence, assign its owner, and define the condition that pauses the new scope. |
| Revise and retest | A material miss has an identified and fixable cause. The expected value remains useful after the repair and retained human work. | Make the source, rule, model, integration, permission, or review change. Rerun the failed class and previously passing cases under a stated boundary. |
| Shadow longer | The result is promising, but the sample, period, reviewer load, exception coverage, or operating conditions are not representative enough for more permission. | Keep the current human path authoritative. Extend the sample across named dates, volume, users, and conditions, with a decision threshold and end date. |
| Keep manual | The workflow still needs human judgment, or the review and failure cost remove the value of automation. The current manual path remains the better operating choice. | Record the manual owner, the evidence that led to the decision, and any condition that would justify a later review. Do not leave a pilot integration active without an owner. |
| Stop | The pilot did not achieve its intended purpose, a viable non-AI option is better, or no feasible repair supports the business case. | Disengage the pilot, remove unneeded access, reconcile any destination state, preserve the handoff, and assign any remaining cleanup. |
Next state
Expand
- Evidence from the tested boundary
- The pilot met its agreed criteria on representative cases. Material misses are understood. Review work and failure costs remain acceptable. The stop or rollback path worked.
- Prerequisite before the next review
- Name one scope increase, preserve the required review and evidence, assign its owner, and define the condition that pauses the new scope.
Next state
Revise and retest
- Evidence from the tested boundary
- A material miss has an identified and fixable cause. The expected value remains useful after the repair and retained human work.
- Prerequisite before the next review
- Make the source, rule, model, integration, permission, or review change. Rerun the failed class and previously passing cases under a stated boundary.
Next state
Shadow longer
- Evidence from the tested boundary
- The result is promising, but the sample, period, reviewer load, exception coverage, or operating conditions are not representative enough for more permission.
- Prerequisite before the next review
- Keep the current human path authoritative. Extend the sample across named dates, volume, users, and conditions, with a decision threshold and end date.
Next state
Keep manual
- Evidence from the tested boundary
- The workflow still needs human judgment, or the review and failure cost remove the value of automation. The current manual path remains the better operating choice.
- Prerequisite before the next review
- Record the manual owner, the evidence that led to the decision, and any condition that would justify a later review. Do not leave a pilot integration active without an owner.
Next state
Stop
- Evidence from the tested boundary
- The pilot did not achieve its intended purpose, a viable non-AI option is better, or no feasible repair supports the business case.
- Prerequisite before the next review
- Disengage the pilot, remove unneeded access, reconcile any destination state, preserve the handoff, and assign any remaining cleanup.
Expansion should increase one named part of the boundary, such as a source, action, user group, or volume range. It does not convert the first pilot into proof for all conditions. Keep the current review point until new evidence supports a specific change.
Revision needs a cause-specific retest. If the team repairs a source, it must test missing and conflicting source cases. If it changes an integration, it must test uncertain destination state, retry, and rollback. A clean new demonstration is not a regression test.
Shadowing is useful when the main gap is insufficient evidence. In a shadow run, the system prepares or proposes work while people keep authority over the real action. Set an end date and decision threshold so that “shadow longer” does not become an indefinite state.
Keeping the workflow manual and stopping the pilot are different decisions. “Keep manual” selects the existing human process as the operating path. “Stop” ends the current pilot effort and its access or integrations. The same review can choose both, but the handoff should state each decision.
Assign the decision, evidence, owner, and review date
End the review with one recommendation. NIST's Manage guidance calls for a determination about whether the system should proceed. It also calls for documented response options, assigned responsibilities, and mechanisms to disengage or deactivate systems that do not match the intended use.
Record the decision in plain language:
Recommended next state:
Tested boundary:
Evidence supporting the decision:
Material misses by cause:
Untested conditions:
Prerequisite before the next review:
Decision owner:
Next review date:
Stop, rollback, or cleanup owner:
The decision owner must have authority over the workflow outcome. A developer can repair an integration but may not own the business rule or permission increase. Add technical, security, privacy, legal, or compliance owners when the prerequisite needs their decision.
A review date is still useful when the decision is to keep the work manual or stop. It can be a final closeout date for access removal, destination reconciliation, record retention, and handoff. Do not set a future restart date unless the evidence identifies a condition that justifies another review.
Preserve a clean handoff for every outcome
A stopped pilot can still reduce future uncertainty. Preserve the tested boundary, representative cases, reviewer decisions, miss classes, source and permission limits, integration notes, rollback result, and reason for the decision. Remove credentials and private data according to the approved retention rules.
Use the AI workflow controls guide to connect the decision to the workflow boundary, approval policy, review queue, run evidence, monitoring, escalation, and rollback. Select only the controls that the next state needs. A manual or stopped workflow still needs an owner for cleanup and any remaining records.
BaristaLabs uses AI consulting to help teams review one candidate workflow and the evidence for its next state. Your team still needs to agree on the cause of the misses and the next prerequisite. Review one pilot decision before you approve a larger scope.
Source note
The NIST AI RMF Playbook provides voluntary suggestions rather than a required post-pilot method. Its live Map and Manage pages state that AI RMF 1.0 is being updated and that the Playbook will be updated after the framework revision. NIST supports documented context and scope, expected benefits and failure costs, human oversight, assigned responsibility, response and recovery options, viable non-AI alternatives, and a decision about whether work should proceed. SBA supports starting small and testing whether an AI tool adds value.
The six miss classes and five next states in this article are BaristaLabs guidance. No cited source establishes a typical expansion rate, pass threshold, review-time limit, return on investment, recovery result, or production-readiness standard for an AI pilot.
Post-pilot review
Make the next pilot decision from the evidence
BaristaLabs can help your team review one tested workflow, classify the material misses, compare retained human work with expected value, and define the prerequisite for the next state.
Best fit when a pilot has produced evidence but the team has not agreed whether to expand, change, run longer, keep the workflow manual, or stop.
Turn this idea into a pilot
Which workflow should go first?
Use the readiness check to compare impact, effort, risk, owner, and next step before booking a call.
- 3-5 minutes
- Deterministic score
- No sensitive data
Practical AI Workflow Notes
Want more practical AI operations ideas?
Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.