Skip to main content

AI engagement stages

AI Discovery vs. AI Pilot: What Each Stage Should Produce

Discovery chooses and scopes the test. A pilot runs that bounded test and leaves evidence for the next decision. Your team may need discovery, a pilot, more preparation before either stage, or a decision not to build.

BaristaLabs publishes a 48-hour discovery model and a 3–6-week target for suitable scoped pilot deployments. The pilot window is not a blanket promise. Integrations, source quality, security review, acceptance criteria, and stakeholder availability can change the scope or make preparation the better next step.

Discovery and a pilot answer different questions

A discovery asks which workflow deserves a test, what boundary the test needs, and what evidence should determine the next decision. Its main deliverable is a recommendation and scope. It records what the team knows, what remains uncertain, and what should wait; it does not create operating evidence by itself.

A pilot applies that scope to representative examples or a controlled slice of real work. It shows what changed, what passed, what needed correction, what failed, what stayed manual, and how much review the workflow required. The result may be a working prototype, a narrow production asset, or a decision memo recommending revision, more preparation, or no further build.

Some teams move from discovery into a pilot. Others narrow the workflow, document the current process, resolve data access, choose an existing tool, keep the work manual, or stop. Those are valid outcomes when the evidence supports them.

Compare the purpose, work, and output

Comparison of AI discovery and a scoped AI pilot by decision, inputs, work, output, participation, timing, evidence, and valid outcomes.

48-hour AI discovery

Decision it supports
Should this workflow move into a pilot, under what boundary, and why?
What the buyer brings
Candidate workflows; current owners; recurring pain; systems and data; customer or staff consequences; review points; decision constraints.
Work performed
Compare candidates, trace the selected workflow, define data, action, and review boundaries, identify risks, and record deferred work.
Primary output
An opportunity map or roadmap with ranked candidates, a recommended first pilot, boundaries, risks, tool or architecture direction, deferred work, and the next decision.
Who participates
The business sponsor and workflow owner, with technical, security, privacy, or compliance input where the candidate requires it.
Published timing label
A 48-hour discovery for a suitable focused engagement.
Evidence produced
A documented recommendation based on available workflow facts, constraints, and assumptions. It does not establish production performance or ROI.
Valid outcomes
Pilot, narrow the workflow, prepare the process or data, choose an existing tool, keep the work manual, or stop.

Scoped AI pilot

Decision it supports
Should the tested workflow expand, revise, run under review for longer, stay manual, or stop?
What the buyer brings
One selected workflow; representative examples; approved sources; an accountable owner and reviewers; acceptance criteria; agreed access to systems or integrations.
Work performed
Build or configure the bounded test, run representative cases, collect reviewer decisions, record misses and exclusions, and test the rollback or stop path.
Primary output
A before/after workflow, source boundary, test cases, reviewer decisions, misses, exclusions, rollback path, prototype or decision memo, and an accountable next step.
Who participates
The workflow owner, representative users or reviewers, and the technical, security, privacy, or compliance stakeholders required by the agreed scope.
Published timing label
A 3–6-week target for suitable scoped pilot deployments; deeper integrations, source work, review, or security needs can change the plan.
Evidence produced
Bounded evidence from the tested workflow and sample. It does not establish universal ROI, compliance, or readiness outside the tested boundary.
Valid outcomes
Expand, revise, shadow longer, preserve human review, keep the work manual, or stop.

Timing begins with fit and scope. Treat both labels as delivery guidance for bounded work, not a promise that every workflow can or should move through the same sequence.

What a 48-hour discovery should produce

In BaristaLabs’ published model, discovery turns several possible AI ideas into a defensible next decision. The result should be useful to the business owner and to the people who would build, review, secure, or operate the pilot.

The roadmap should contain:

  • the candidate workflows considered and the reason for their ranking;
  • one recommended first pilot, or a recommendation to prepare the workflow before testing it;
  • the selected workflow’s owner, trigger, systems, sources, decisions, and important exceptions;
  • the allowed and excluded data, human review point, and action boundary;
  • the main implementation, security, privacy, vendor, and maintenance questions;
  • the acceptance evidence a pilot should collect;
  • the work deliberately deferred; and
  • the next decision, its owner, and the condition for revisiting it.

Discovery is still useful when the recommendation is to use a self-serve tool, run a small internal experiment, clean up the process, keep a consequential decision with a person, or avoid a custom build. Its job is to reduce uncertainty before the team commits more money or permission.

What a scoped pilot should produce

A pilot tests one bounded workflow. Before the work begins, the team should agree on the representative examples, allowed sources, review decisions, acceptance criteria, exclusions, and stop path. “The demo looked good” is too weak to support broader access or budget.

The pilot handoff should contain:

  • a readable view of the workflow before and after the test;
  • the sources and fields the test used, plus explicit exclusions;
  • representative test cases and the acceptance criteria applied to them;
  • reviewer decisions, edits, rejections, misses, and cases that stayed manual;
  • the stop and rollback path;
  • a prototype or narrow working asset when the evidence supports one, or a decision memo when it does not;
  • known failure modes, implementation assumptions, and untested conditions; and
  • the recommended next state, reason, prerequisites, owner, and review point.

The detailed AI pilot proof guide explains these fields. Early signals such as review load, cycle time, cleaner intake, or fewer handoffs should remain tied to the tested sample and baseline. They should not become a general savings or ROI promise without enough verified evidence.

Neither stage is ready when the decision has no owner or evidence

A short engagement cannot repair every missing prerequisite. Resolve these gaps before treating discovery or a pilot as the next purchase:

  • The workflow has no accountable owner.Name the person who can explain the current work, approve the test boundary, review the result, and own the next decision.
  • Representative examples are unavailable.Collect a small, privacy-safe sample that includes ordinary work, difficult cases, and known exceptions. A pilot built only around ideal examples cannot show how the real workflow behaves.
  • Sensitive data cannot be scoped.Identify the systems and data classes involved, what can stay excluded, and which security, privacy, legal, or compliance stakeholders must review the boundary.
  • The team cannot name the decision the work should support.Define whether the result will decide to proceed, narrow, revise, shadow longer, keep manual, choose another path, or stop.

Preparation is a valid next step. It is usually cheaper than forcing unclear ownership, unreviewable data, or a vague acceptance bar into a build.

Choose the next useful step

The workflow is still unclear

Use the workflow readiness assessment to score one recurring process by impact, effort, risk, systems, data sensitivity, approval needs, process clarity, and ownership.

Score one workflow

Several workflows look plausible

Review AI consulting when the team needs to rank candidates, define the first pilot boundary, record implementation risks, and decide what should wait.

Review AI consulting

A pilot proposal already exists

Use the AI pilot proof guide to inspect the workflow baseline, source boundary, test evidence, misses, rollback plan, and next-decision fields before approving more scope.

Review the pilot evidence

If the workflow is named but the implementation route remains unclear, use the AI implementation path decision matrix. If you want to compare a proposal with approved public work, browse the case studies without treating a different client’s result as a promise for your workflow.

Questions about discovery and pilots

Am I buying a plan or a working test?

A discovery primarily produces a recommendation and scope: which workflow to test, the boundary it needs, the risks and assumptions, and the next decision. A pilot runs the bounded test and should leave inspectable evidence. Depending on that evidence, its deliverable may be a prototype, a narrow working asset, or a decision memo recommending revision, preparation, or stopping.

Does every team need both discovery and a pilot?

No single sequence fits every workflow. Teams with several possible uses or unclear boundaries need discovery before committing to a test. A team with a well-defined workflow still needs its owner, sources, review point, acceptance criteria, and stop path confirmed before pilot work begins. Either review can show that preparation, an existing tool, manual work, or stopping is the better next step.

Does a successful pilot mean the workflow is ready for production?

No. A pilot supports a decision within its tested scope. Production readiness may require broader examples, stronger access controls, integration work, monitoring, user training, security or compliance review, and a clear support owner. The pilot handoff should state what it did not test or prove.

Are 48 hours and 3–6 weeks guaranteed timelines?

BaristaLabs publishes a 48-hour discovery model for suitable focused engagements and a 3–6-week target for suitable scoped pilot deployments. The pilot target is not universal. Integration depth, source quality, security review, acceptance criteria, and stakeholder availability can change the scope or make preparation the better next step.

Bring one candidate workflow when you are ready to discuss it

Contact is most useful after you can name the recurring workflow, current owner, systems and data involved, review point, consequence of a bad result, and decision you need the work to support. Incomplete answers are acceptable; a generic request to “add AI” still needs the readiness or discovery step first.

Discuss one candidate workflow

Best fit when you can name the workflow but need help deciding whether discovery, a scoped pilot, or more preparation should come next.