MEMO · TO readers evaluating a workshop or a speaker · RE Strategy
Why large enterprises are deploying AI at scale but measuring it incorrectly
Common measurement frameworks target the wrong layer of the technology stack, missing the business outcomes that actually matter.
Terence Kok, Enterprise AI Strategist, Author, Keynote Speaker
Capital expenditure on AI infrastructure across Southeast Asia, Greater China and the Indian subcontinent reached record levels in 2025, and much of it is being measured against the wrong yardstick. A model that is ninety-four percent accurate is a capability specification. It is not proof that the business it was deployed into has actually changed, and enterprises that conflate the two are optimising for a number that was never the point.
There are two distinct measurement layers, and most AI programmes only instrument one of them. The technical layer covers accuracy, F1 score, inference latency and model drift, useful for engineering teams, largely irrelevant to a board. The operational layer covers processing time per transaction, decision consistency, error and exception frequency, and cost per unit of output measured against a baseline set before deployment. Conflating the two produces decisions that optimise a model metric while the business outcome it was meant to move stays flat.
Three structural causes recur. AI sponsorship sits inside technology functions rather than the operating units that own the outcome. Vendor contracts are written around model performance benchmarks rather than business impact. And deployment speed keeps outpacing the measurement infrastructure needed to know whether any of it worked.
The corrective sequence is five steps: define the operational baseline before deployment, set operational KPIs with explicit attribution rules, instrument the workflow rather than only the model, connect model monitoring to operational alerting, and report the operational impact, not the model metrics, to executive governance on a fixed cadence. None of this replaces model-level monitoring. It sits above it, and it is the layer a board can actually act on.
Exhibit · Two measurement layers
Where enterprises are pointing the instrument
Most AI programmes only instrument one layer. The board can only act on the other.
| Layer | What it tracks | Who it's for |
|---|---|---|
| Technical | Accuracy, F1 score, inference latency, model drift | Engineering teams — largely irrelevant to a board |
| Operational | Processing time per transaction, decision consistency, error and exception frequency, cost per unit vs. baseline | Executive governance — the layer a board can actually act on |
- 01
Define the operational baseline
Before deployment, not after.
- 02
Set operational KPIs
With explicit attribution rules.
- 03
Instrument the workflow
Not only the model.
- 04
Connect model monitoring to operational alerting
So drift surfaces where it matters.
- 05
Report the operational impact to governance
On a fixed cadence, not the model metrics.
Reference
This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›
More in Strategy
- Is your business ready for AI? Start with these five questions
A diagnostic framework for identifying whether a business is actually positioned to operationalise AI tools, not just experiment with them. The same framework the Executive Programme is built on.
- The productivity gap: why companies must fix their systems before scaling AI
Most organisations use AI frequently but few have scaled it company-wide, and the gap comes down to architectural debt, not appetite.