A prototype should answer a bounded question. Can the system extract the required fields from the documents available? Will this team use a generated first draft? Does the interaction remove enough work to justify another test? A short build can produce a useful answer without carrying the whole product case.
The claim often widens after the demo. “The extraction worked on the test set” appears in the next presentation as “the product is feasible.” Adoption, reliability and margin enter the forecast even though the trial never measured them.
What happened inside the test
Prototype conditions are usually favourable by design. The team limits scope, prepares inputs and stays nearby to explain failures. Early users know they are testing unfinished work, so they tolerate delays and workarounds that ordinary users may reject.
The test record should list included cases, work performed outside the interface and any edited output. It should explain how quality was judged and where testing stopped. That boundary determines which funding claim the result can support.
Questions left for the product case
- Will intended users change their routine once the project team leaves?
- How does performance change on incomplete and unusual inputs?
- Where do unfinished cases go, and how much staff time do they consume?
- Who detects a quality change after the model or source data changes?
- Does the price or saving still work after support and review are included?
- Can the company reverse a consequential wrong action before it spreads?
The system’s role sets the evidence requirement. A specialist drafting aid can tolerate failures that would be unacceptable in a customer-facing process or a tool that moves money. The greater the consequence, the less a polished demonstration can establish on its own.
The next test and the next cheque
“Scale the prototype” gives no boundary for spending. The next budget should name the open question it will settle. That could mean a representative quality evaluation, a paid customer trial, a supervised live release or a short integration test against the dependency carrying the schedule.
If the prototype proved extraction quality, the next budget might measure correction time in a live queue. A department-wide rollout can wait for that figure.
