05Measurement-Before-Prediction Architecture

Listen to this framework

Why Industrial AI Fails on Accidental Data

0:00
05

Measurement-Before-Prediction Architecture

Developed for infrastructure and industrial AI deployments where prediction accuracy depends on a measurement schema designed before data collection begins.

Most infrastructure AI deployments fail not because the models are wrong but because the measurement framework was not designed before the data was collected. When the data schema, collection frequency, and quality standards are designed to serve a different purpose, AI predictions built on that data require months of recalibration after go-live. This framework reverses that sequence.

Reference diagramThe conventional collect-then-model sequence and the five-step sequence that replaces it. Step 3, the schema, is the gate: it is fixed before any collection begins.
Measurement-Before-Prediction Architecture, reference diagramThe conventional sequence collects whatever data is available, builds the model, then recalibrates for months after go-live. The measurement-before-prediction sequence starts from the operational decision, derives the data requirements from it, fixes the measurement schema before collection begins, validates against simulation, and designs operations so they produce the data that improves the model.The sequence this replacesCollect the data availablegathered for another purposeBuild the modelperforms well in testingRecalibrate for months after go-liveor recollect from scratchMeasurement before prediction: the sequence is not negotiable1Operational decisionWhat must an operatorknow, by when, at whatconfidence, to act?2Data requirementsVariables, frequency,precision and sourcereliability, from step 1.3Measurement schemaStructure, labelling,quality thresholds, audittrail. Fixed beforecollection starts.4Simulation validationGreenfield models testedon comparable sites;results are projections.5Improvement by designOperations produce thetraining data; retrain atreview intervals.the prediction defines the data, not the inventoryretrain and re-benchmark at review intervalsStep 3 is the gate.A schema designed after collection begins produces data quality problems that only recollection resolves. Fix it first, every time.
Step 1

Define what the prediction needs to produce

Start with the operational decision, not the data. What does an operator need to know, by when, with what confidence level, to take a meaningful action? This definition determines what the prediction must produce, and therefore what data the prediction requires as input.

Step 2

Work backwards to the data requirements

From the prediction output requirements, define the input data: variables, update frequency, precision, and source reliability. This becomes the data collection specification: not an inventory of available data, but a specification of required data derived from the prediction requirements.

Step 3

Design the measurement schema before collection begins

Define data structure, labelling conventions, quality thresholds, and audit trail requirements before any collection infrastructure is built. A measurement schema designed retrospectively around collected data produces persistent data quality problems that cannot be resolved without recollection.

Step 4

Validate predictions against simulation before go-live

For greenfield deployments with no historical data, validate initial prediction models against simulation data from comparable operational environments. Treat simulation-validated figures as projections, not production metrics, and design the live data collection to confirm or revise them from day one.

Step 5

Build model improvement into operational design

The data collected during operations is the asset that improves prediction accuracy over time. Design operational workflows to produce data that improves model accuracy as a by-product of normal operations. Define the review intervals at which prediction models are retrained and accuracy benchmarks are reassessed.

Application principle: The sequence is non-negotiable. Changing the measurement schema after data collection begins requires recollection or produces a persistent gap between historical and current data. Building prediction models on data collected for a different purpose produces models that perform well in testing and fail under operational conditions. Design the measurement framework first, every time.

Version
1.1
First published
5 June 2026
Last revised
12 September 2026

Reproduce the diagram and text with attribution for non-commercial use. For use in a tender, board paper or training material, get in touch; permission is normally a formality.

This is a personal site. The views, frameworks and publications here are my own analysis. They do not speak for Orion Five Engineering or any past employer or client, and they do not draw on the confidential information, data or proprietary methods of any of them.

The free tool scores you. The paid formats put the framework to work on your own programme, with me in the room.

Why does the measurement schema need to come before data collection?

Because changing the schema after collection begins requires recollection, or leaves a persistent gap between historical and current data. Working backwards from the prediction to the data requirements defines the schema correctly the first time.

What if we already have historical data collected for a different purpose?

Models built on data collected for a different purpose tend to perform well in testing and fail under operational conditions. That data can still inform the schema design, but it should not be treated as the collection specification for the new prediction target.

How does this apply to greenfield projects with no historical data?

Validate initial models against simulation data from comparable operational environments, treat those figures as projections rather than production metrics, and design live data collection from day one to confirm or revise them.

How often should prediction models be retrained?

Define the review intervals as part of the operational design itself. Build workflows so the data that improves model accuracy is produced as a by-product of normal operations, then reassess accuracy benchmarks on a fixed cadence.

Terence Kok
Before You Go

This one exists because I watched a greenfield airport project nearly build its prediction models on data collected for someone else's purpose, and I know how expensive that mistake is to unwind once concrete has been poured. It's not a glamorous framework. Nobody gets excited about a measurement schema. But the teams who slow down here are the same ones who aren't recollecting data eighteen months later. Patience at step three is what buys you speed everywhere after it.

Terence Kok