Journal

From AI pilot to production: a checklist before you scale

Review evaluation, human handoffs, operating costs and fallback before inviting more users.

An engineering team reviewing an AI pilot at a workstation.

A pilot can show that an idea is worth examining. Moving it into everyday use asks a different question: can the team operate the proposed workflow under the conditions its users will actually bring? A polished demonstration is only one piece of that discussion. Before expanding access, review the evidence, the responsibilities and the ways the system can stop or return work to a person.

Use this checklist as an agenda for a readiness meeting. Each item should produce a concrete answer, a named owner or a documented reason to delay the next step. The list is a general working aid, not a guarantee of suitability for every use case. A team should add checks that reflect the consequences of its own system and the needs of the people affected.

Check the scope you are approving

Write down who will use the system, what they will submit and what they will do with its output. Compare that description with the pilot. A trial involving a few trained colleagues may not support conclusions about a much larger group with different habits. If the proposed audience has changed, identify which assumptions need fresh evidence.

Keep the approved scope visible to users and operators. For an internal document assistant, that could include the supported document collection and the kinds of questions it is intended to address. A clear boundary gives support staff a way to distinguish a defect from a request for a new capability. Both deserve attention, but they lead to different decisions.

Inspect the evaluation set

Ask which examples were used to examine the pilot and why those examples were selected. The set should reflect the work expected at the next stage, including incomplete inputs, unfamiliar wording and cases where no useful answer is available. Record important gaps. A collection built only from successful demonstrations leaves the readiness meeting with little evidence about difficult cases.

Keep some examples separate from the material used while adjusting the system. When a team repeatedly tunes against the same questions, it may learn more about those questions than about the broader workflow. Record the version of the system evaluated and preserve enough context to reproduce the observations. Review examples of failure alongside any summary measures.

Confirm the human review path

Identify which outputs require review and who can approve them. A named reviewer needs clear criteria, enough context to check the answer and a practical way to report an issue. If the workflow produces more material than that person can examine, the operating plan needs adjustment. A human review step is a responsibility, not merely a label on a diagram.

Walk through an example where the reviewer disagrees with the output. Can they correct it, return it for more information or stop the action that follows? Confirm that the user interface and operating instructions support the intended choice. Also check what happens when the reviewer is unavailable. The answer should not depend on somebody noticing an unassigned queue by chance.

Trace inputs and information handling

List the information that enters the workflow, where it goes and which people or services can access it. Compare the actual configuration with the approved design. If the pilot used sample material but the next stage will use business records, review that change explicitly with the responsible people. Do not assume that a successful sample run resolves questions about real data.

Define what operational records are needed and what should be omitted from them. Support staff may need a request identifier and error details without needing the full source document. Agree on access and retention arrangements appropriate to the organisation. Where legal or specialist review is required, record the owner and the decision rather than turning this checklist into a claim of compliance.

Observe usage and operating costs

Assign responsibility for watching usage, request volume and the cost categories the team actually incurs. Include review and support effort in the discussion, even when those costs do not appear on a service invoice. A pilot with enthusiastic early users can produce a different pattern from a wider rollout. Treat the next stage as something to observe, not a simple multiplication of the demonstration.

Set a process for responding to unexpected usage. An operator should know which signals require investigation and who can restrict access or pause a feature. Check whether retries or repeated user submissions are counted in the team's estimates. The purpose is to make operating decisions with visible information, not to promise a particular economic outcome.

Rehearse the fallback

Choose a realistic interruption and walk through it. The source collection might be unavailable, a dependency might reject requests, or the system might return an output that cannot be checked. The user should know whether to retry, contact someone or use the previous process. The operator should know which action restores a workable path.

A fallback deserves its own owner and instructions. If returning to manual work requires access that users no longer have, it is not yet a usable alternative. Consider what happens to work already in progress when the system is paused. Record how participants will avoid losing track of those items and how they will be told that the normal process has resumed.

Review changes as part of operations

Identify which changes require another evaluation. A revised prompt, a different model configuration, a new document source or an altered integration can each change behavior. Keep a record that connects the change with the observations used to approve it. The team should be able to answer which version produced a reported output without reconstructing the project from memory.

Agree on who can introduce a change and who reviews it. A small team may combine roles, but it still benefits from making the decision explicit. Preserve a way to return to a previous configuration when appropriate. If a dependency cannot be pinned or rolled back, record that limitation and decide how it affects monitoring and rollout scope.

Prepare users and support staff

Give users a short explanation of the intended task, known limits and route for help. Show examples that make those limits understandable. Saying check the answer is less useful than explaining which source a person should consult and what to do when the two disagree. Training should match the actual interface and the decisions the user is expected to make.

Support staff need a way to collect useful reports without asking people to share unnecessary sensitive material. A report might include the approximate time, request identifier, expected behavior and a description of the issue. Decide who reviews those reports and how repeated patterns reach the product owner. An inbox alone does not define an operating process.

Make the readiness decision explicit

The NIST AI Risk Management Framework is a public reference for considering AI risk throughout design, use and evaluation. A team can consult it when organising a broader review. This checklist is narrower: it helps structure the practical conversation about the next operating stage, without assigning a certification or declaring a system universally ready.

Close the meeting with a written decision. State the approved audience and workflow, the remaining limitations, the people responsible and the date of the next review. If important answers are missing, narrow the proposed rollout or pause it. Run the checklist against one current pilot and attach evidence to each answer; the unresolved items will show what the team needs to learn next.