Failure risk, calibrated triage bands and remaining tool life for a CNC milling cell, from six sensor readings.
Most predictive maintenance demos stop at a probability. A probability on its own is hard to act on: a shift lead needs to know whether to stop the spindle, how confident the number is, roughly how much tool life is left, and what in the reading caused the alert. forgewatch answers those four questions from the same six inputs.
Three things come out of one scoring call:
- a failure probability from a glass-box additive model, with the specific sensor readings that drove it
- a conformal triage band (clear, review, stop) that carries a measured coverage guarantee rather than a hand-picked cut point
- a tool-life budget from a survival model: how many more minutes of wear before cumulative failure risk crosses a chosen ceiling
The alert threshold is chosen by minimising the expected cost of running the cell rather than by maximising F1, under cost assumptions written down in one file and stress-tested in a sensitivity table. Whether that beats a naive 0.5 cut is answered over twenty independent splits rather than one, for a reason the results section explains.
All numbers below come from artifacts/metrics.json, written by a full training run. The held-out fold is 2,000 readings the model never saw during fitting or calibration.
Three candidates, ranked by cross-validated average precision on the training fold. Average precision rather than ROC-AUC, because a 3.39% positive rate flatters ROC and does not flatter AP.
| model | CV average precision | test ROC-AUC | test average precision | test Brier | test F1 at 0.16 |
|---|---|---|---|---|---|
| logistic regression | 0.4641 ± 0.0471 | 0.9386 | 0.5386 | 0.10997 | 0.1425 |
| gradient boosting | 0.8278 ± 0.0613 | 0.9744 | 0.8295 | 0.02149 | 0.5203 |
| explainable boosting (shipped) | 0.8773 ± 0.0654 | 0.9752 | 0.8903 | 0.00867 | 0.8296 |
The glass-box model wins outright, which is the convenient outcome and not the one I expected. Additive models usually trade a little accuracy for readability. Here the failure modes in the data are driven by a handful of physical quantities and one strong interaction, which is exactly the shape an additive model with pairwise terms fits well.
The two boosted models are nearly tied on ROC-AUC, 0.9752 against 0.9744, and that near-tie is exactly why ROC-AUC is the wrong metric here. On average precision the gap opens to 0.8903 against 0.8295, and on Brier score the glass-box model is two and a half times better. The linear baseline is not close on any of them: an ROC-AUC of 0.939 looks respectable until you see an average precision of 0.539 and a Brier score thirteen times worse than the winner's.
Calibration is the reason the cost model can be trusted. A threshold chosen on miscalibrated probabilities is a threshold chosen on nothing. The shipped model's Brier score of 0.00867 against a base rate of 0.0339 means the probabilities are close enough to real frequencies that expected-cost arithmetic on top of them means something.
The threshold comes from the calibration fold by minimising expected cost, then gets applied unchanged to the held-out fold.
| calibration fold (1,600) | held-out fold (2,000) | |
|---|---|---|
| threshold | 0.16 | 0.16 |
| true positives | 46 | 56 |
| false positives | 8 | 11 |
| false negatives | 8 | 12 |
| precision | 0.852 | 0.836 |
| recall | 0.852 | 0.824 |
| cost with no model | $685,800 | $863,600 |
| cost with alerts | $276,710 | $367,620 |
| saving | $409,090 | $495,980 |
That is $247.99 saved per reading scored, catching 56 of 68 failures at the price of 11 false alarms.
This is the part of the project I got wrong first, so it is worth showing the working.
An earlier draft answered this from a single held-out fold. On that fold the cost-optimised threshold lost to a plain 0.5 cut by $2,710, and I wrote it up as a null result. Then I changed the random seed for an unrelated reason and the same procedure beat 0.5 by $28,600. Same data, same code, same method: only the split differed. The first answer was not a finding, it was a sample of size one.
So the question gets answered over twenty independent splits. Each one re-splits from scratch, refits the model, reselects the threshold on its own calibration fold, and scores both that threshold and a fixed 0.5 on its own held-out fold.
| measure over 20 splits | value |
|---|---|
| splits where optimising beat 0.5 | 20 of 20 |
| mean advantage | $33,346 |
| median advantage | $31,295 |
| standard deviation | $16,827 |
| worst split | $3,225 |
| best split | $66,920 |
| threshold the sweep picked | 0.09 to 0.36, median 0.173 |
| mean gap to each fold's own optimum | $10,634 |
Cost-optimal thresholding wins, and it wins on every split. The effect is real and worth roughly $31,000 per 2,000 readings at the median, which is about 6% of what alerting is worth in the first place.
Two caveats sit alongside that, and both matter more than the headline.
The advantage is noisy. A standard deviation of $16,827 against a mean of $33,346 means any single fold's number carries about half its own size in noise. That is exactly how my first draft got the sign wrong, and it is why the single-fold figure is no longer quoted anywhere in this README.
And the threshold the sweep picks is itself unstable: across the twenty splits it landed anywhere from 0.09 to 0.36, a fourfold range, and sat on average $10,634 above what that fold's own optimum would have cost. So the procedure reliably finds a good region, and reliably fails to find the best point inside it. The basin is wide and shallow enough that most of the value comes from moving off 0.5 at all, not from where exactly you land.
The practical reading for a plant: run the sweep, because it pays every time. Do not treat the number it returns as precise, and do not re-tune it on every batch of new data, because the run-to-run variation in that number is larger than the difference it is chasing.
This whole study lives in src/forgewatch/stability.py and runs as part of every full training run. A test file runs it again at six splits and asserts the direction and rough size of each claim above, so a change that reverses any of them fails the suite rather than quietly outdating this section.
Every dollar figure above rests on these, and they are assumptions about one machining cell, not measurements:
| assumption | value | reasoning |
|---|---|---|
| unplanned stoppage | 4.5 h | mean time to diagnose, source a part and restart |
| cell hourly rate | $2,400 | lost contribution margin plus labour for one cell |
| extra unplanned cost | $1,900 | scrapped workpiece, expedited freight, callout |
| planned stop | 0.6 h | tool change slotted into a shift boundary |
| planned consumables | $180 | insert or cutter |
| intervention efficacy | 0.85 | share of flagged failures actually averted |
That puts an unplanned failure at $12,700 and a planned intervention at $1,620. Published surveys quote a median of roughly $125,000 per hour of unplanned downtime, but that is plant-wide across large manufacturers; applying it to a single cell would inflate these savings by more than an order of magnitude. I used the conservative number deliberately.
The two softest assumptions are how long an unplanned stop really lasts and how often an alert actually prevents the failure. Held-out fold savings across both:
| unplanned stop | efficacy 0.70 | efficacy 0.85 | efficacy 0.95 |
|---|---|---|---|
| 2.0 h ($6,700 per failure) | $154,100 | $210,380 | $247,900 |
| 4.5 h ($12,700 per failure) | $389,300 | $495,980 | $567,100 |
| 8.0 h ($21,100 per failure) | $718,580 | $895,820 | $1,013,980 |
The savings stay positive and substantial across the whole grid. Even at the pessimistic corner, a two-hour stoppage with only 70% of alerts landing, the fold still clears $154,100.
Split-conformal calibration on a fold the model never fit on, targeting 90% coverage:
| measure | value |
|---|---|
| requested confidence | 0.90 |
| measured coverage on held-out fold | 0.9195 |
| singleton sets | 92.10% |
| empty sets | 7.90% |
| two-label sets | 0.00% |
Coverage lands at 0.9195 against a request of 0.90, which is the guarantee doing what it says on data it has not seen, erring toward covering more than asked rather than less.
One honest note on the bands. MAPIE only offers the LAC conformity score for binary targets, and on this data LAC never produces a two-label set. So the review band, 158 of 2,000 readings, is made entirely of empty sets: readings where neither label cleared the calibrated score cut. That means review here reads as "this operating state does not resemble the calibration fold" rather than "the two labels are neck and neck". It is still the right routing, and arguably a more interesting signal, but it is not what a reader would assume from the band names alone.
A Cox proportional hazards model over the tool-wear axis, giving a wear budget rather than a yes-or-no call.
| covariate | hazard ratio | 95% interval | p |
|---|---|---|---|
| torque (Nm) | 1.0219 | 1.0148 – 1.0291 | <0.001 |
| mechanical power (W) | 1.0002 | 1.0001 – 1.0002 | <0.001 |
| heat rejected (K) | 0.8660 | 0.8107 – 0.9251 | <0.001 |
| spindle speed (rpm) | 1.0000 | 0.9996 – 1.0004 | 0.932 |
| workpiece grade | 0.9279 | 0.8400 – 1.0249 | 0.140 |
Concordance on the held-out fold is 0.7319, against 0.5 for random ordering.
Torque and mechanical power both raise the hazard, which matches the overstrain and power failure modes the data documents. Heat rejection cutting the hazard looks backwards until you notice what a large process-minus-air gap means physically: the cell is shedding heat rather than trapping it, and the heat dissipation failure mode fires when that gap closes. Spindle speed and workpiece grade do not reach significance once the others are in the model, and I have left them in rather than pruning to a prettier table.
A caveat that belongs next to every survival number here. The source readings are a cross-section, not per-tool trajectories: each row is one machining state at one wear value, not a tool followed from new to scrapped. Treating wear as the time axis, failures as events and healthy readings as right-censored is the standard approximation for this table and it produces curves that behave sensibly, but the output is a population statement about tools at that wear level, not a forecast for one specific spindle.
The spindle-speed and heat-rejection interaction leads, followed by each of them alone, then tool wear. Nothing in the top of that list is a proxy, an identifier or a leak, which is the main thing I wanted to confirm.
339 failures in 10,000 readings, a 3.39% base rate. The mode flags break down as heat dissipation 115, overstrain 98, power 95, tool wear 46, random 19.
Two quirks in the source table worth knowing about, both checked by the test suite rather than quietly smoothed over: 9 readings are labelled as failures with no mode flag set, and 18 readings carry the random-failure flag while the headline label stays 0. Neither is a bug in the loader. The headline label is what gets modelled.
data/ai4i2020.csv (10,000 readings)
|
load_raw + audit (DuckDB profile,
schema and range checks, fail loud)
|
feature_frame: drop UDI, Product ID and all
five mode flags <-- the one leakage guard
|
stratified 3-way split 64% fit / 16% calibrate / 20% test
|
+--------------------------+--------------------------+
| |
RISK HEAD TOOL-LIFE HEAD
| |
Pipeline([ survival_fit_frame
DerivedFeatures temp delta, power, |
wear x torque, speed/torque CoxPHFitter over
ColumnTransformer scale + one-hot the wear axis
Estimator LR | HGB | EBM ]) |
| wear budget at a
select on CV average precision chosen hazard cap
| |
+-----+---------------------+ |
| | |
cost sweep on the SplitConformalClassifier |
calibration fold on the calibration fold |
| | |
threshold 0.16 clear / review / stop |
| | |
(stability.py repeats | |
this over 20 splits) | |
| | |
+-----------+---------------+-------------------------------+
|
Bundle (joblib): pipeline + bander
+ tool-life model + threshold
|
+-------------+--------------+
| |
scoring.assess() artifacts/metrics.json
one reading in, assets/*.png
one Assessment out
|
Gradio console (app.py)
git clone https://github.com/ranjithguggilla/forgewatch.git
cd forgewatch
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
pip install -e .Python 3.10 or newer. The dataset ships in data/, so there is nothing to download.
For the console and the test suite:
pip install -r requirements-dev.txt
pip install -e ".[app]"python -m forgewatch.trainAbout 65 seconds on this machine. Most of that is the repeated-split threshold study, which refits the model twenty times; the candidate cross-validation accounts for most of the rest. It prints the candidate comparison, the chosen threshold and the tool-life concordance, then writes:
artifacts/metrics.json— every number quoted in this READMEartifacts/forgewatch.joblib— the scoring bundleassets/*.png— the charts above
A faster path for CI or a smoke check, which skips the glass-box candidate and the split study and drops to 3-fold cross-validation:
python -m forgewatch.train --quickOr keep the full run and shorten the study:
python -m forgewatch.train --stability-splits 5from forgewatch.scoring import Bundle, score_reading
bundle = Bundle.load()
result = score_reading(
bundle,
product_type="L", # workpiece grade: L, M or H
air_temp_k=300.1,
process_temp_k=310.2,
speed_rpm=1380,
torque_nm=62.0,
tool_wear_min=225,
)
print(result.band) # 'review'
print(result.failure_probability) # 0.96682
print(result.tool_life["remaining_min"]) # 0.0 - already past budget
print(result.explanation)For those inputs result.explanation prints:
Borderline. Have someone eyeball the cell this shift.
Failure probability 96.7% (conformal band: review).
What drove it:
- wear-torque load (overstrain proxy) at 1.395e+04 raises the risk score by 4.224
- tool wear at 225 raises the risk score by 3.156
- spindle speed at 1380 raises the risk score by 0.875
- speed-to-torque ratio at 22.26 raises the risk score by 0.689
The tool budget comes back as zero remaining minutes because this tool is at 225 minutes of wear against a budget that ran out at 139, with cumulative risk already at 0.61 against the 0.05 ceiling.
This reading is also a good illustration of the review band doing something a threshold cannot. The probability is 96.7%, far above the 0.16 alert threshold, so the alert flag is set. But the conformal set came back empty, meaning neither label cleared the calibrated cut, so the band is review rather than stop. The model is flagging this as an operating state unlike anything in its calibration fold. A shift lead gets told to go look, which is the right instruction for a confident score the model cannot vouch for.
python -m forgewatch.appOpens on port 7860 with sliders for the six inputs, four preset operating states, the triage card and the machine-readable payload side by side.
Everything tunable lives in src/forgewatch/config.py.
| setting | default | what it changes |
|---|---|---|
CostModel.unplanned_downtime_hours |
4.5 | length of an unplanned stop |
CostModel.cell_hourly_rate_usd |
2400.0 | value of an hour of cell time |
CostModel.intervention_efficacy |
0.85 | share of alerts that avert the failure |
CONFIDENCE_LEVEL |
0.9 | conformal coverage target |
SurvivalConfig.hazard_budget |
0.05 | cumulative risk ceiling for the tool budget |
TEST_SIZE / CALIBRATION_SIZE |
0.2 / 0.2 | fold sizes |
RANDOM_STATE |
613829 | seed for splits and models |
Changing a cost moves every dollar figure in metrics.json on the next training run. Nothing is hard-coded downstream.
artifacts/
forgewatch.joblib pipeline, conformal bander, tool-life model, threshold
metrics.json dataset profile, model comparison, operating point,
economics with sensitivity, conformal report,
hazard ratios, the repeated-split threshold study
with its per-split rows, chart index
assets/
data-profile.png class balance and failure mode flags
model-curves.png ROC and precision-recall for all three candidates
calibration.png reliability curves with Brier scores
cost-curve.png expected cost against threshold
confusion.png held-out confusion matrix at the chosen threshold
global-importance.png glass-box term importances
tool-survival.png Kaplan-Meier curves by workpiece grade
conformal-bands.png triage band distribution
threshold-stability.png advantage over a 0.5 cut across 20 splits
forgewatch/
src/forgewatch/
config.py paths, column names, cost and survival settings
data.py loading, DuckDB audit, quality checks, three-way split
features.py derived physical quantities as a stateless transformer
models.py the three candidate pipelines
economics.py cost model, threshold sweep, sensitivity grid
conformal.py split-conformal wrapper and the triage bands
survival.py Kaplan-Meier and Cox proportional hazards tool-life head
explain.py global importances and per-reading plain-language reasons
evaluate.py metrics and every chart
train.py the orchestration entry point
stability.py the repeated-split study behind the threshold section
scoring.py the bundle and the single scoring function
app.py Gradio console, binding straight to scoring.py
tests/ 98 tests
data/ the readings
assets/ charts
.github/workflows/ci.yml
Leakage. Three guards, all of them tested. The five per-mode flags and both identifiers never enter the feature frame, and test_feature_frame_drops_every_label_derived_column asserts on the exact surviving column set. Feature derivation is a stateless transformer placed first in the pipeline, so it cannot carry training statistics into inference; test_transformer_is_stateless_across_folds fits it on one slice and transforms another to prove it. Scaling and one-hot encoding sit inside the pipeline, so cross-validation refits them per fold.
Class imbalance. 339 positives in 10,000 rows. Handled with balanced class weights rather than synthetic oversampling. Oversampling would invent machining states that no cell would ever produce, and with a glass-box model those invented states end up visible in the shape functions. Model selection uses average precision, which does not reward a model for getting the 96.6% majority right.
Calibration fold. The conformal layer needs scores from data the model never fit on, otherwise the coverage guarantee quietly evaporates. That is why the split is three-way rather than the usual two. The same fold also picks the cost-optimal threshold, keeping the test fold genuinely untouched until the final evaluation.
Derived features. Four, all computable from one reading with no pooled statistics: process-minus-air temperature, mechanical power at the spindle from torque and angular velocity, the wear-times-torque product that the overstrain mode keys on, and the speed-to-torque ratio. The dataset's failure modes are generated from documented physical rules, so features aligned with that physics do well. That is worth saying plainly: performance here is an upper bound on what the same approach would deliver on raw shop-floor telemetry, where the rules are not known and the sensors are noisier.
Threshold. Chosen by minimising expected cost over a 197-point grid on the calibration fold, with ties broken toward the higher threshold to keep alert volume down. On the shipped split it lands at 0.16, well below the default 0.5, because a missed failure costs roughly eight times a false alarm. The repeated-split study in the results section measures what that selection is worth and how much the selected value moves between splits; the short version is that the procedure pays every time and the number it returns should not be read as precise.
Explainability. Per-reading contributions come from the additive model's term scores, so they are exact rather than an approximation of a black box. Pairwise interaction terms are excluded from the sentences on purpose: they appear in the global chart, but "spindle speed crossed with heat rejection" is not something anyone can check against the panel in front of them. When a non glass-box estimator is in the pipeline, which happens on the quick path, the reasons come back empty rather than raising mid-alert.
pytest
pytest --cov=forgewatch --cov-report=term-missing98 tests, 94% statement coverage. They are organised by the thing that could go wrong rather than by module:
test_data.py— schema and range validation, the DuckDB audit against pandas, split disjointness and stratification, reproducibility, and the two label quirks in the source tabletest_features.py— every derived formula checked by hand, input immutability, zero-torque division, and the stateless-across-folds prooftest_models.py— pipeline step ordering, that no label-derived column reaches the estimator, probability well-formedness, and an unseen workpiece grade degrading rather than raisingtest_economics.py— the savings formula worked out by hand, threshold extremes, monotonicity in efficacy, and the full sensitivity gridtest_conformal.py— coverage near the requested level, set-share arithmetic, that the stop band is riskier than the clear band, and that a tighter confidence level never narrows coveragetest_survival.py— monotone survival curves, concordance above chance, hazard-ratio interval ordering, torque raising the hazard, and a harder cut shortening the budgettest_scoring.py— input validation, the assessment contract, JSON round-tripping, bundle persistence, and a punishing cut scoring riskier than a gentle onetest_explain.py— label prettification, sorted contributions, narrative wording, and graceful degradationtest_app.py— every preset scores, the card renders, the interface buildstest_pipeline_end_to_end.py— a full quick training run, then asserting every chart lands on disk, the metrics sit in range, the operating point saves money and the saved bundle scores a fresh readingtest_threshold_finding.py— runs the repeated-split study for real against the shipped model, at six splits rather than twenty to keep the suite usable, and asserts the direction and rough magnitude of all three claims the README makes from it: that optimising beats 0.5 on every split, that the advantage is noisy enough to justify repeating, and that the selected threshold moves between splits without ever reaching 0.5. The bounds are deliberately loose, so they catch a change of direction rather than a change of decimal place
Lint and format are enforced by ruff with E, F, W, I, UP, B, C4, SIM, ARG, RET enabled.
AI4I 2020 Predictive Maintenance Dataset, 10,000 readings from a simulated milling machine with five documented failure modes. Donated to the UCI Machine Learning Repository by Stephan Matzka (Hochschule für Technik und Wirtschaft Berlin).
- Source: https://archive.ics.uci.edu/dataset/601/ai4i+2020+predictive+maintenance+dataset
- Vendored copy used here: https://github.com/lingchenlijing/AI4I-2020-Predictive-Maintenance-Dataset
- Licence: CC BY 4.0
The readings are synthetic but built to mirror real milling behaviour, including the correlations between tool wear, torque, spindle speed and heat rejection that drive the failure modes. That is a strength for reproducibility and a limitation for claims about live shop-floor data, noted above.
MIT. See LICENSE.








