Which operational decision arrives too late today?
An early warning creates value only when the operation knows who investigates and what response is permitted.
A manufacturer can meet its monthly production target and still be heading toward a serious problem.
One supplier has started shipping late. Demand for a product family is rising faster in two regions. A critical pump is drawing slightly more current and vibrating differently from its normal pattern. None of these signals alone proves that a disruption will occur. Together, they may be the first warning.
Traditional reporting tells leaders what has already happened. Artificial intelligence can add three forward-looking capabilities:
- Forecasting estimates what is likely to happen next.
- Anomaly detection identifies behaviour that differs from an established pattern.
- Agentic analysis investigates exceptions, gathers context, and recommends a response.
These capabilities are related but not interchangeable. A forecast is not a promise. An anomaly is not a diagnosis. A recommendation is not an authorised decision. Operational value appears when each output is connected to a clear human process.
Start with the decision, not the model
“Predict demand” is not yet a business use case. Ask what decision the prediction will change.
For example:
- How many units should be ordered for each location and week?
- Which orders need expediting today?
- Which production line needs inspection during the next planned stop?
- Which spare part should be positioned closer to a site?
- Which forecast disagreement requires a sales and operations review?
- Which exception is important enough to interrupt a planner?
The decision defines the time horizon, granularity, acceptable error, data, owner, and response. A daily forecast by product family may help capacity planning but be useless for replenishing individual stores. A five-minute equipment anomaly score may suit a continuously running compressor, while a daily result may suit a slow batch process.
The first design document should describe the decision in plain language: who makes it, what information they use now, what it costs to be late or wrong, and what action is available.
The three layers of operational data
- 01
What happened
The transactional record — orders, movements, downtime, completions.
- 02
What was happening around it
The context that explains the record: weather, promotions, staffing, supply.
- 03
What is happening now
Live signal from equipment and process, arriving fast enough to still matter.
Layer 01 alone produces rear-view reporting. It is layers 02 and 03 that turn a report into a warning with enough lead time to act.
Useful operational AI usually combines three kinds of data.
What happened
Orders, shipments, production, stock, downtime, work orders, returns, and actual demand provide the historical record.
What was happening around it
Promotions, prices, holidays, weather, supplier changes, maintenance events, production schedules, and stockouts explain why history may not repeat cleanly.
What is happening now
Current inventory, open orders, queue positions, equipment sensors, vibration, temperature, pressure, current draw, and recent quality results reveal the live state.
Data quality is not a preliminary IT concern that can be declared complete. A stockout can make sales look like low demand. A replaced sensor can shift readings without an equipment fault. A promotion entered under the wrong product code can distort future forecasts. Domain experts need to label these events and explain how the operation behaves.
Forecasting: turn history into a range for planning
A forecast estimates future values from historical patterns and relevant inputs. It may learn trend, seasonality, intermittent demand, promotional effects, and relationships across many items.
Amazon SageMaker Canvas allows users to create machine-learning predictions without writing code, including time-series forecasts. AWS lists business uses such as inventory planning, pricing and revenue, customer churn, and on-time delivery. For organisations that want an accessible custom model, Canvas provides a path to import data, build, evaluate, and generate predictions.
Amazon Quick Sight ML Insights offers forecasting and anomaly detection in a business intelligence environment. AWS documentation says its forecasting can automatically account for real-world features including seasonality, trends, outliers, and missing values. This can be useful when a team already works from dashboards and needs a forward view of a business metric.
For more specialised supply-chain planning, AWS announced the general availability of Amazon Connect Decisions in April 2026. Its Demand Intelligence capability combines forecasting tools and foundation models to create item-level forecasts, monitor actual demand against them, analyse exceptions, and support consensus planning through natural-language interaction. Its Supply Intelligence capability is designed to monitor inventory, suppliers, demand, and execution, group exceptions, and surface prioritised decisions.
Availability matters. At its April 2026 general-availability announcement, AWS listed Amazon Connect Decisions in US East (N. Virginia) and Europe (Ireland). A Singapore-based or globally regulated organisation should confirm current region availability, data requirements, and architecture before choosing it.
Forecast accuracy is not one universal percentage
A leader should ask how error is measured and at what level.
Common measures include:
- MAE, mean absolute error: the average size of the error, expressed in the same units as the forecast.
- WAPE, weighted absolute percentage error: total absolute error divided by total actual volume, useful for comparing performance across a portfolio.
- Bias: whether the operation persistently forecasts too high or too low. State the sign convention because organisations calculate it differently.
- Service-level outcomes: fill rate, on-time delivery, lost sales, or backlog.
- Economic outcomes: working capital, waste, expedited freight, overtime, and obsolescence.
A model can improve WAPE while making an important low-volume item worse. Review results by product, location, horizon, and business criticality. Compare against a simple baseline, such as last year, a moving average, or the current planning method. Complexity should earn its place.
Forecasts should also show uncertainty. A range makes the trade-off visible. A narrow range may support routine replenishment. A wide range may justify a scenario plan, supplier discussion, or delayed commitment.
Anomaly detection: find the unusual before it becomes obvious
Anomaly detection learns or defines normal behaviour and flags a meaningful departure. It is useful when fixed thresholds are too crude.
A vibration threshold might alert whenever a motor exceeds one value. A machine-learning model can consider several sensor signals and operating conditions together. It may notice that vibration, temperature, and current draw form an unusual combination even though each remains below its individual alarm level.
AWS IoT SiteWise is designed to collect, organise, and analyse industrial equipment data. Its native anomaly detection capability can train a custom model on selected equipment properties, learn normal operating conditions, process new time-series data on a schedule, and optionally use labelled periods of known failure. AWS positions it for fixed and stationary equipment such as pumps, compressors, motors, CNC machines, turbines, heat exchangers, boilers, and inverters.
AWS documentation also describes configurable inference schedules from high-frequency monitoring, as often as every five minutes, through lower-frequency daily checks, with operating windows that can avoid idle or planned maintenance periods. Models can be retrained, versioned, evaluated, and rolled back.
The output is an indication of abnormal behaviour, not a work order by itself. A high score may reflect a true developing failure, an unusual operating mode, maintenance activity, a bad sensor, or a changed process. The response should include the relevant signals and enough context for an engineer to investigate.
A worked example: one operation, connected decisions
Consider a company that manufactures and distributes packaged industrial materials through several plants and warehouses.
Demand changes
Forecasting identifies a likely increase for one product family in the north region. The model also shows a wide uncertainty range because a similar promotion has limited history.
The supply plan tightens
Available capacity can meet the midpoint forecast, but a supplier is already late on a key input. An exception-detection process raises the likely risk of a shortfall in week four.
Equipment behaviour shifts
At the plant that could provide extra volume, anomaly detection finds a change in a pump’s vibration and current pattern. It is not yet a failure, but the equipment supports a constrained line.
AI assembles context
An operations assistant gathers the demand range, current inventory, open purchase orders, supplier performance, line schedule, maintenance history, and sensor contribution diagnostics. It prepares three options:
- transfer finished stock from another region;
- bring forward a planned pump inspection and move production;
- maintain the plan but negotiate a supplier expedite and establish a customer-allocation trigger.
People decide
The planner, maintenance engineer, procurement lead, and commercial owner review the evidence and consequences. Approved changes flow through the existing planning, maintenance, and financial controls.
This is a better model of operational AI than an autonomous “control tower” that silently changes plans. AI compresses investigation and makes relationships visible. Accountable people choose the trade-off.
Predictive maintenance is a business process, not a sensor project
A predictive-maintenance initiative fails if it ends at an alert. The organisation needs to decide:
- who owns the alert;
- how quickly it must be reviewed;
- what evidence is shown;
- how it becomes an inspection or work order;
- whether spare parts and skilled labour are available;
- how the result is recorded;
- how confirmed faults and false alarms improve the model.
Start with equipment where the business case is real: high downtime cost, observable failure patterns, sufficient data, and an action that can be taken before failure. Not every asset needs machine learning. Calendar maintenance, usage thresholds, alarms, and operator rounds remain appropriate where they are effective.
Separate the anomaly system from safety controls. A model should not replace protective relays, emergency shutdown systems, regulated inspections, or operating procedures. It adds an early-warning layer.
Keep people in the exception loop without overwhelming them
One promise of AI is to surface “what matters.” That requires careful prioritisation.
Rank exceptions using business impact, urgency, confidence, customer consequence, safety, and reversibility. Group related alerts so planners do not receive twenty notifications caused by one supplier delay. Suppress known maintenance windows and identified bad sensors. Let users explain why they accepted, modified, or rejected a recommendation.
Monitor false positives, which waste attention, and false negatives, which create false assurance. The acceptable balance depends on consequence. Missing a minor efficiency opportunity is different from missing a developing fault in critical equipment.
Generative explanations can make forecasts and anomalies easier to understand, but they must point back to the data and rules used. A plausible story about causation is not proof. Distinguish “the pattern is associated with” from “this caused.”
A staged programme for operations AI
Stage 1: baseline the decision
Select one demand segment, exception type, or equipment class. Record current accuracy, delay, manual effort, downtime, stockout, waste, or service outcomes.
Stage 2: run in shadow mode
Generate forecasts or alerts without changing work. Compare them with actual events and expert decisions. Identify data gaps, bad labels, seasonal shifts, and operating modes.
Stage 3: provide decision support
Show the prediction, uncertainty, drivers, relevant history, and recommended investigation. Require people to record the disposition.
Stage 4: integrate the workflow
Create an inspection request, planning scenario, or supplier task only after defined confirmation. Keep system-of-record and approval rules intact.
Stage 5: automate narrow responses
Automate only low-risk, reversible actions with proven performance, such as requesting additional data, opening a draft work order, or refreshing a planning scenario. Maintain stop conditions and escalation.
Stage 6: manage it as a living system
Track performance drift, operational changes, sensor replacements, product launches, new suppliers, and user behaviour. Retrain or recalibrate under version control. Preserve the ability to revert to a stable model.
Safeguards for real operations
An operational AI control set should include:
- named business and technical owners;
- documented purpose, scope, users, and prohibited use;
- quality checks for source, time, units, missing values, and sensor health;
- access controls separating plants, suppliers, customers, and sensitive commercial data;
- model and rule versioning with change approval;
- shadow testing and back-testing against historical periods;
- explicit uncertainty and evidence;
- human approval for production, maintenance, purchasing, allocation, and safety decisions;
- stop conditions, fallback procedures, and rollback;
- monitoring for drift, false alerts, missed events, latency, and cost;
- cybersecurity review where cloud analytics connects to operational technology;
- separation from safety-instrumented and emergency systems.
AWS’s Responsible AI Lens recommends defining a narrow use case and considering controllability, privacy, security, safety, veracity, robustness, fairness, explainability, transparency, and governance throughout the lifecycle. In operations, robustness and controllability deserve particular attention because unusual conditions are exactly when the system will be needed.
Measure whether decisions improve
Track technical and business results together:
- forecast WAPE, MAE, and bias by horizon and segment;
- stockout, fill rate, backlog, and on-time delivery;
- inventory turns, working capital, waste, and expediting cost;
- anomaly precision, missed-event rate, and lead time before confirmed failure;
- unplanned downtime and maintenance plan compliance;
- mean time to acknowledge, investigate, and resolve an exception;
- percentage of recommendations accepted, modified, or rejected;
- false alerts per asset or planner;
- model performance after operating or product changes;
- cost per useful forecast, alert, or resolved exception.
Do not claim savings from a model demonstration. Measure the full process, including data preparation, review, false alerts, integration, and ongoing operation.
Early warning becomes valuable only when the business can respond
AI can help an operation look forward, but it cannot create capacity, parts, authority, or collaboration by itself. A forecast matters when it changes a plan. An anomaly matters when it leads to a timely inspection. An agentic recommendation matters when accountable people can see the evidence and act within a controlled process.
The objective is not a more impressive dashboard. It is more time to make a better decision.
- https://docs.aws.amazon.com/sagemaker/latest/dg/canvas.html
- https://docs.aws.amazon.com/sagemaker/latest/dg/canvas-build-model.html
- https://docs.aws.amazon.com/quicksight/latest/user/making-data-driven-decisions-with-ml-in-quicksight.html
- https://docs.aws.amazon.com/quicksight/latest/user/anomaly-detection.html
- https://aws.amazon.com/about-aws/whats-new/2026/04/amazon-connect-decisions-april/
- https://aws.amazon.com/products/connect/decisions/demand-intelligence/
- https://aws.amazon.com/products/connect/decisions/faqs/
- https://docs.aws.amazon.com/iot-sitewise/latest/userguide/sitewise-anomaly-detection.html
- https://docs.aws.amazon.com/iot-sitewise/latest/userguide/advanced-inference-configurations.html
- https://docs.aws.amazon.com/iot-sitewise/latest/userguide/adv-training-configs.html
- https://aws.amazon.com/iot-sitewise/
- https://docs.aws.amazon.com/prescriptive-guidance/latest/mes-on-aws/ai-ml.html
- https://docs.aws.amazon.com/wellarchitected/latest/responsible-ai-lens/responsible-ai-lens.html
This article is original editorial work for a business audience. Product facts were checked against the official sources above. Any performance figures are targets to validate against your own baseline, not vendor guarantees.
Bring us the operating need, risk, or opportunity. We will connect the people, security, AI and enterprise systems required to act.