Atul Iwale · Project 9 · Predictive maintenance with machine learning

Which machines will break down next?

Weekly telematics from 80 cranes, excavators, pumps, generators and hoists predict which machines are likely to break down in the next four weeks, so the plant team can inspect and repair them before they stop a site. XGBoost and a Support Vector Machine both run live in your browser.

XGBoostSVM (RBF kernel)Random Forest · Neural Network · KNN · Logistic Regression · Decision TreeSHAP explanationsCost-based alert threshold

This week's inspection list

#MachineXGBoost riskSVM riskAction

Risk = probability of a breakdown in the next 4 weeks. Rows are ranked by XGBoost. Select a machine to see its history, its top reasons, and a what-if.

How the models compare

Trained on Oct 2024 to Dec 2025, tuned on Jan to Mar 2026, tested once on Apr to Aug 2026. About one operating week in six is followed by a breakdown, so a random ranking scores a PR-AUC of about .

Ranking quality on the test months

PR-AUC rewards ranking the machines that do break down above those that don't. Recall at 50% precision is the share of breakdown weeks flagged when half the alerts are real.

Breakdowns caught when only K machines can be inspected per week

A breakdown counts as caught if the machine was on the list in any of the 4 weeks before it.

What it's worth: alert-threshold simulator

Move the threshold and see the result on the test months. Costs come from the workbook: a missed breakdown costs its actual repair plus replacement hire for the downtime; a caught one becomes a planned repair (35% of the repair cost plus 1 day of hire); every alert costs an inspection (Rs 6,000 plus half a day of hire).

Caught by root cause (at cost-optimal thresholds)

Sudden external damage gives no warning in the sensors. No model can catch it reliably, and none should claim to.

What drives the XGBoost risk

Mean absolute SHAP value on the test months. The model relies on the signals a plant engineer would check.

How it works

  • Data: 103 weeks of telematics (vibration, operating temperature, hydraulic pressure, fault codes, load, hours) plus monthly oil analysis, and 1,100+ maintenance events, on the same asset register as the original project.
  • Label: a breakdown in the next 4 weeks. Weeks when a machine is broken down or in transit are not scored.
  • Features: readings compared with each machine's own normal level (weeks t−16 to t−4), 6-week trends, service and repair history, age and class. They use past data only.
  • Models: seven classifiers tuned on validation months. The SVM is Platt-calibrated so its scores are probabilities.
  • In this page: the exported 400 XGBoost trees and SVM support vectors score every machine in JavaScript.

Honest limits

  • The telemetry is synthetic, built on realistic wear physics. Real fleets are noisier, so accuracy must be re-measured on real breakdown records.
  • Sudden damage (impacts, burst hoses) cannot be predicted from sensors.
  • The rupee savings depend on the inspection and planned-repair cost assumptions, and on five test months.
  • The models were trained up to Dec 2025. In use they should be retrained every quarter.
  • Machine drivers come from SHAP for the live week. The what-if sliders update the risk scores but not the drivers.