Labelled pressure histories establish what normal demand and pump starts look like. The pilot then measures whether smaller, persistent deviations can be detected early without creating an unusable false-alarm rate.
The pilot defines usable camera positions, image quality and visible review criteria before evaluating detection performance. Uncertain or safety-relevant findings remain with a qualified human reviewer.
Representative requests are used to measure routing quality by category, language and difficult edge cases. Low-confidence items go to a review queue instead of being sent automatically.
A clear target schema and field-level evidence make extraction testable. Messy layouts, missing values and conflicting content are included so uncertain results can be flagged rather than silently accepted.
Retrieval is tested for source coverage, permissions, traceable context and the ability to abstain. The assistant should support a decision without presenting unsupported text as operational fact.
Historical data is split in time and compared with a simple baseline. Evaluation covers drift, false alarms and operational lead time before forecasting or anomaly detection is considered useful.