The pilot handled the selected cases, and invited testers found useful work for it. That result is enough to discuss a production release. Questions about ordinary inputs, operating cost and the authority to intervene remain open.
During a pilot, project staff can screen inputs, coach users and repair awkward output before it causes trouble. Production removes much of that shelter. Demand spikes, new staff arrive without the design rationale, and exceptions land with whoever is available.
The limits of the pilot result
A result often expands as it moves up an organisation. Acceptable answers on a reviewed sample become a claim about departmental automation. Encouraging comments from invited users turn into an adoption forecast. By the time the funding paper reaches its approver, it may describe a test that never took place.
Keep a test record with the sample, excluded cases and selection criteria. Name the people who prepared data or edited output. Include the comparison, test period and work performed outside the interface. The result may still be positive, but the record shows exactly how far it reaches.
Costs and responsibilities outside the model
The production service depends on data access, integrations and staff behaviour as much as model output. Someone must monitor performance and handle wrong answers. Mandatory human review belongs in the cost model. The approval paper should also state what happens if a model provider changes its price, policy or availability.
- Problem: the workflow failure costs enough to warrant changing it.
- Evaluation: the test set resembles production inputs and includes expensive failure cases.
- Staff work: review time, corrections, escalation and support appear in the operating plan.
- Cost: model use, infrastructure and maintenance remain affordable at forecast volume.
- Containment: staff can find a serious error before it passes into another process.
- Ownership: named people can change permissions, suspend use and retire the system.
Strong model performance does not settle the production case. Customers may withhold the required data, the workflow may have nowhere safe to send uncertain output, or review labour may consume the saving. The budget commits the organisation to those operating conditions as well as the model.
A smaller production commitment
The next release might cover one workflow, a volume cap or a limited group of accounts. Hiring can wait, and a vendor expansion can depend on a customer commitment. Choose a release that exposes production behaviour while reversal is still affordable.
Write each condition as an observable threshold with an owner and review date. “Quality should improve” gives the sponsor nothing to decide. A measured error range on representative cases, paired with a ceiling for correction time, can govern whether the next release occurs.
Approve the production exposure you can describe and control.
Scope of an independent production review
An Independent Initiative Review treats the proposed production release as one defined decision. It compares the claimed business result with the product mechanism, test evidence and full operating cost. If management and engineering read the same result differently, both readings remain visible in the decision record.
The opinion may support the proposed release, set conditions for a narrower one or call for another test. If correction time doubles when the project team steps away, the approval should already identify who measures it and which spending stops.
