Atul Iwale · Project 1 · Cost & margin forecasting with machine learning

What will this project finally cost?

Part-way through a job, the standard earned-value formula assumes the rest will perform like the past. This model learned from 239 completed projects how cost really ends up. It forecasts the final cost and margin of each of the 30 portfolio projects, with a P10–P90 range and next-quarter spend. The XGBoost model runs live in your browser.

XGBoost + quantile rangesEarned value (EVM) baselinesRandom Forest · Ridge · Neural Network · LSTMHolt's exponential smoothingSHAP explanations

Portfolio: forecast margin at completion

ProjectStatusCompleteTender marginCPI-formula marginML forecast marginFinal margin

Forecasts are as of each project's latest cost report (August 2026 for live projects; the last month before completion for finished ones). Margin = (contract value − cost) ÷ contract value. Select a project to replay its forecasts month by month.

How accurate is it?

Every monthly forecast for the portfolio projects that have finished was scored against their true final cost. None of these projects were used in training.

Final-cost forecast error (MAPE) by stage of completion

Bias above 0 means forecasts too high, below 0 too low. The CPI × SPI formula is left out of the ranking: it overshoots badly early on.

Next-quarter spend: error on the test projects

WAPE = total absolute error ÷ total actual spend over 3-month windows. Monthly cost reports are noisy (accruals booked and reversed), so no method is exact month by month.

Why the CPI formula misleads

What moves the forecast (SHAP)

Mean absolute SHAP value on the test projects: how much each input shifts the XGBoost forecast of final cost ÷ budget.

How it works

  • Data: monthly cost reports (planned value, earned value, actual cost) in five categories for 239 completed projects and the 30 portfolio projects, plus a steel/cement price index.
  • Snapshots: each project at each month-end from 5% to 99% complete, using only data up to that month. The target is final cost ÷ budget.
  • Split by time: trained on projects started before July 2021, tuned on later ones, and tested on the portfolio.
  • Range: two quantile XGBoost models, widened on validation data (conformal calibration).
  • In this page: exported trees score each month in JavaScript.

Honest limits

  • The cost data is synthetic, built on realistic cost behaviour. Accuracy must be re-measured on a company's own completed projects.
  • The test set is projects, so the figures will move as more finish.
  • The LSTM and MLP did not beat the tree models: about 240 projects is little data for neural networks.
  • The model should be retrained each quarter as projects close.