A useful AI pilot answers a narrow question: does this system help people perform a particular job under the company's actual conditions? A general chatbot demonstration cannot establish that. Use representative inputs, ordinary users and acceptance criteria agreed before configuration begins.

Choose an output that can be checked

A draft answer based on a product catalogue is easier to evaluate than “make the department more effective”. The first has a message, permitted sources and an output. The second has too many possible causes of success or failure. Initially avoid workflows where an incorrect response immediately triggers an irreversible action.

Consider a proposed incoming-enquiry workflow. The assistant identifies the topic, extracts items and drafts a response. The manager still sends it. This tests recognition and completeness without making autonomous communication part of the experiment.

Build an honest evaluation set

Include simple messages, ambiguous wording, poor attachments and requests outside the catalogue. Reserve some examples for final evaluation and keep them out of configuration work. Otherwise the team may learn to produce good answers for familiar documents without learning how the system handles new ones.

For each case, record expected facts, acceptable alternatives and mandatory stopping conditions. Refusing to quote without a current price list may be correct. Confidently inventing a price is a failure, even when the prose is excellent.

Measure the whole job

- Employee time spent preparing and reviewing the result. - The share accepted without substantial correction. - Missing facts, wrong sources and impermissible actions. - Waiting time under normal and peak workloads.

Decide what happens next

Record one of three outcomes: deploy the bounded workflow, fix an identified failure cause, or stop the experiment. Do not increase permissions just because the demonstration was attractive. Name a process owner and define incident handling. This fits NIST's approach to managing AI risks throughout development and operation, rather than treating launch as the end of evaluation. NIST · AI Risk Management Framework