Skip to content

About

Predictive maintenance for a CNC cell: glass-box failure model, conformal triage bands (MAPIE), survival-based tool-life budgets, cost-optimal thresholds.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

forgewatch

ci python license

Failure risk, calibrated triage bands and remaining tool life for a CNC milling cell, from six sensor readings.

Most predictive maintenance demos stop at a probability. A probability on its own is hard to act on: a shift lead needs to know whether to stop the spindle, how confident the number is, roughly how much tool life is left, and what in the reading caused the alert. forgewatch answers those four questions from the same six inputs.

Three things come out of one scoring call:

  • a failure probability from a glass-box additive model, with the specific sensor readings that drove it
  • a conformal triage band (clear, review, stop) that carries a measured coverage guarantee rather than a hand-picked cut point
  • a tool-life budget from a survival model: how many more minutes of wear before cumulative failure risk crosses a chosen ceiling

The alert threshold is chosen by minimising the expected cost of running the cell rather than by maximising F1, under cost assumptions written down in one file and stress-tested in a sensitivity table. Whether that beats a naive 0.5 cut is answered over twenty independent splits rather than one, for a reason the results section explains.

Results

All numbers below come from artifacts/metrics.json, written by a full training run. The held-out fold is 2,000 readings the model never saw during fitting or calibration.

Model selection

Three candidates, ranked by cross-validated average precision on the training fold. Average precision rather than ROC-AUC, because a 3.39% positive rate flatters ROC and does not flatter AP.

model CV average precision test ROC-AUC test average precision test Brier test F1 at 0.16
logistic regression 0.4641 ± 0.0471 0.9386 0.5386 0.10997 0.1425
gradient boosting 0.8278 ± 0.0613 0.9744 0.8295 0.02149 0.5203
explainable boosting (shipped) 0.8773 ± 0.0654 0.9752 0.8903 0.00867 0.8296

The glass-box model wins outright, which is the convenient outcome and not the one I expected. Additive models usually trade a little accuracy for readability. Here the failure modes in the data are driven by a handful of physical quantities and one strong interaction, which is exactly the shape an additive model with pairwise terms fits well.

The two boosted models are nearly tied on ROC-AUC, 0.9752 against 0.9744, and that near-tie is exactly why ROC-AUC is the wrong metric here. On average precision the gap opens to 0.8903 against 0.8295, and on Brier score the glass-box model is two and a half times better. The linear baseline is not close on any of them: an ROC-AUC of 0.939 looks respectable until you see an average precision of 0.539 and a Brier score thirteen times worse than the winner's.

ROC and precision-recall

Calibration

Calibration is the reason the cost model can be trusted. A threshold chosen on miscalibrated probabilities is a threshold chosen on nothing. The shipped model's Brier score of 0.00867 against a base rate of 0.0339 means the probabilities are close enough to real frequencies that expected-cost arithmetic on top of them means something.

Operating point and money

The threshold comes from the calibration fold by minimising expected cost, then gets applied unchanged to the held-out fold.

calibration fold (1,600) held-out fold (2,000)
threshold 0.16 0.16
true positives 46 56
false positives 8 11
false negatives 8 12
precision 0.852 0.836
recall 0.852 0.824
cost with no model $685,800 $863,600
cost with alerts $276,710 $367,620
saving $409,090 $495,980

That is $247.99 saved per reading scored, catching 56 of 68 failures at the price of 11 false alarms.

Cost against threshold

Confusion matrix

Does cost-optimal thresholding actually beat the default?

This is the part of the project I got wrong first, so it is worth showing the working.

An earlier draft answered this from a single held-out fold. On that fold the cost-optimised threshold lost to a plain 0.5 cut by $2,710, and I wrote it up as a null result. Then I changed the random seed for an unrelated reason and the same procedure beat 0.5 by $28,600. Same data, same code, same method: only the split differed. The first answer was not a finding, it was a sample of size one.

So the question gets answered over twenty independent splits. Each one re-splits from scratch, refits the model, reselects the threshold on its own calibration fold, and scores both that threshold and a fixed 0.5 on its own held-out fold.

measure over 20 splits value
splits where optimising beat 0.5 20 of 20
mean advantage $33,346
median advantage $31,295
standard deviation $16,827
worst split $3,225
best split $66,920
threshold the sweep picked 0.09 to 0.36, median 0.173
mean gap to each fold's own optimum $10,634

Threshold stability

Cost-optimal thresholding wins, and it wins on every split. The effect is real and worth roughly $31,000 per 2,000 readings at the median, which is about 6% of what alerting is worth in the first place.

Two caveats sit alongside that, and both matter more than the headline.

The advantage is noisy. A standard deviation of $16,827 against a mean of $33,346 means any single fold's number carries about half its own size in noise. That is exactly how my first draft got the sign wrong, and it is why the single-fold figure is no longer quoted anywhere in this README.

And the threshold the sweep picks is itself unstable: across the twenty splits it landed anywhere from 0.09 to 0.36, a fourfold range, and sat on average $10,634 above what that fold's own optimum would have cost. So the procedure reliably finds a good region, and reliably fails to find the best point inside it. The basin is wide and shallow enough that most of the value comes from moving off 0.5 at all, not from where exactly you land.

The practical reading for a plant: run the sweep, because it pays every time. Do not treat the number it returns as precise, and do not re-tune it on every batch of new data, because the run-to-run variation in that number is larger than the difference it is chasing.

This whole study lives in src/forgewatch/stability.py and runs as part of every full training run. A test file runs it again at six splits and asserts the direction and rough size of each claim above, so a change that reverses any of them fails the suite rather than quietly outdating this section.

Cost assumptions

Every dollar figure above rests on these, and they are assumptions about one machining cell, not measurements:

assumption value reasoning
unplanned stoppage 4.5 h mean time to diagnose, source a part and restart
cell hourly rate $2,400 lost contribution margin plus labour for one cell
extra unplanned cost $1,900 scrapped workpiece, expedited freight, callout
planned stop 0.6 h tool change slotted into a shift boundary
planned consumables $180 insert or cutter
intervention efficacy 0.85 share of flagged failures actually averted

That puts an unplanned failure at $12,700 and a planned intervention at $1,620. Published surveys quote a median of roughly $125,000 per hour of unplanned downtime, but that is plant-wide across large manufacturers; applying it to a single cell would inflate these savings by more than an order of magnitude. I used the conservative number deliberately.

Sensitivity

The two softest assumptions are how long an unplanned stop really lasts and how often an alert actually prevents the failure. Held-out fold savings across both:

unplanned stop efficacy 0.70 efficacy 0.85 efficacy 0.95
2.0 h ($6,700 per failure) $154,100 $210,380 $247,900
4.5 h ($12,700 per failure) $389,300 $495,980 $567,100
8.0 h ($21,100 per failure) $718,580 $895,820 $1,013,980

The savings stay positive and substantial across the whole grid. Even at the pessimistic corner, a two-hour stoppage with only 70% of alerts landing, the fold still clears $154,100.

Conformal triage

Split-conformal calibration on a fold the model never fit on, targeting 90% coverage:

measure value
requested confidence 0.90
measured coverage on held-out fold 0.9195
singleton sets 92.10%
empty sets 7.90%
two-label sets 0.00%

Conformal bands

Coverage lands at 0.9195 against a request of 0.90, which is the guarantee doing what it says on data it has not seen, erring toward covering more than asked rather than less.

One honest note on the bands. MAPIE only offers the LAC conformity score for binary targets, and on this data LAC never produces a two-label set. So the review band, 158 of 2,000 readings, is made entirely of empty sets: readings where neither label cleared the calibrated score cut. That means review here reads as "this operating state does not resemble the calibration fold" rather than "the two labels are neck and neck". It is still the right routing, and arguably a more interesting signal, but it is not what a reader would assume from the band names alone.

Tool life

A Cox proportional hazards model over the tool-wear axis, giving a wear budget rather than a yes-or-no call.

covariate hazard ratio 95% interval p
torque (Nm) 1.0219 1.0148 – 1.0291 <0.001
mechanical power (W) 1.0002 1.0001 – 1.0002 <0.001
heat rejected (K) 0.8660 0.8107 – 0.9251 <0.001
spindle speed (rpm) 1.0000 0.9996 – 1.0004 0.932
workpiece grade 0.9279 0.8400 – 1.0249 0.140

Concordance on the held-out fold is 0.7319, against 0.5 for random ordering.

Tool survival

Torque and mechanical power both raise the hazard, which matches the overstrain and power failure modes the data documents. Heat rejection cutting the hazard looks backwards until you notice what a large process-minus-air gap means physically: the cell is shedding heat rather than trapping it, and the heat dissipation failure mode fires when that gap closes. Spindle speed and workpiece grade do not reach significance once the others are in the model, and I have left them in rather than pruning to a prettier table.

A caveat that belongs next to every survival number here. The source readings are a cross-section, not per-tool trajectories: each row is one machining state at one wear value, not a tool followed from new to scrapped. Treating wear as the time axis, failures as events and healthy readings as right-censored is the standard approximation for this table and it produces curves that behave sensibly, but the output is a population statement about tools at that wear level, not a forecast for one specific spindle.

What the model leans on

Global importance

The spindle-speed and heat-rejection interaction leads, followed by each of them alone, then tool wear. Nothing in the top of that list is a proxy, an identifier or a leak, which is the main thing I wanted to confirm.

Data profile

Data profile

339 failures in 10,000 readings, a 3.39% base rate. The mode flags break down as heat dissipation 115, overstrain 98, power 95, tool wear 46, random 19.

Two quirks in the source table worth knowing about, both checked by the test suite rather than quietly smoothed over: 9 readings are labelled as failures with no mode flag set, and 18 readings carry the random-failure flag while the headline label stays 0. Neither is a bug in the loader. The headline label is what gets modelled.

How it fits together

                          data/ai4i2020.csv  (10,000 readings)
                                   |
                        load_raw + audit (DuckDB profile,
                        schema and range checks, fail loud)
                                   |
                    feature_frame: drop UDI, Product ID and all
                    five mode flags  <-- the one leakage guard
                                   |
             stratified 3-way split  64% fit / 16% calibrate / 20% test
                                   |
        +--------------------------+--------------------------+
        |                                                     |
  RISK HEAD                                            TOOL-LIFE HEAD
        |                                                     |
  Pipeline([                                          survival_fit_frame
    DerivedFeatures  temp delta, power,                       |
                     wear x torque, speed/torque        CoxPHFitter over
    ColumnTransformer  scale + one-hot                  the wear axis
    Estimator          LR | HGB | EBM ])                      |
        |                                              wear budget at a
  select on CV average precision                       chosen hazard cap
        |                                                     |
  +-----+---------------------+                               |
  |                           |                               |
cost sweep on the      SplitConformalClassifier               |
calibration fold       on the calibration fold                |
  |                           |                               |
threshold 0.16           clear / review / stop                |
  |                           |                               |
(stability.py repeats         |                               |
 this over 20 splits)         |                               |
  |                           |                               |
  +-----------+---------------+-------------------------------+
                              |
                    Bundle (joblib): pipeline + bander
                    + tool-life model + threshold
                              |
                +-------------+--------------+
                |                            |
      scoring.assess()              artifacts/metrics.json
      one reading in,               assets/*.png
      one Assessment out
                |
        Gradio console (app.py)

Installation

git clone https://github.com/ranjithguggilla/forgewatch.git
cd forgewatch

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate

pip install -r requirements.txt
pip install -e .

Python 3.10 or newer. The dataset ships in data/, so there is nothing to download.

For the console and the test suite:

pip install -r requirements-dev.txt
pip install -e ".[app]"

Usage

Train

python -m forgewatch.train

About 65 seconds on this machine. Most of that is the repeated-split threshold study, which refits the model twenty times; the candidate cross-validation accounts for most of the rest. It prints the candidate comparison, the chosen threshold and the tool-life concordance, then writes:

  • artifacts/metrics.json — every number quoted in this README
  • artifacts/forgewatch.joblib — the scoring bundle
  • assets/*.png — the charts above

A faster path for CI or a smoke check, which skips the glass-box candidate and the split study and drops to 3-fold cross-validation:

python -m forgewatch.train --quick

Or keep the full run and shorten the study:

python -m forgewatch.train --stability-splits 5

Score one reading in Python

from forgewatch.scoring import Bundle, score_reading

bundle = Bundle.load()
result = score_reading(
    bundle,
    product_type="L",      # workpiece grade: L, M or H
    air_temp_k=300.1,
    process_temp_k=310.2,
    speed_rpm=1380,
    torque_nm=62.0,
    tool_wear_min=225,
)

print(result.band)                      # 'review'
print(result.failure_probability)       # 0.96682
print(result.tool_life["remaining_min"])  # 0.0 - already past budget
print(result.explanation)

For those inputs result.explanation prints:

Borderline. Have someone eyeball the cell this shift.
Failure probability 96.7% (conformal band: review).

What drove it:
- wear-torque load (overstrain proxy) at 1.395e+04 raises the risk score by 4.224
- tool wear at 225 raises the risk score by 3.156
- spindle speed at 1380 raises the risk score by 0.875
- speed-to-torque ratio at 22.26 raises the risk score by 0.689

The tool budget comes back as zero remaining minutes because this tool is at 225 minutes of wear against a budget that ran out at 139, with cumulative risk already at 0.61 against the 0.05 ceiling.

This reading is also a good illustration of the review band doing something a threshold cannot. The probability is 96.7%, far above the 0.16 alert threshold, so the alert flag is set. But the conformal set came back empty, meaning neither label cleared the calibrated cut, so the band is review rather than stop. The model is flagging this as an operating state unlike anything in its calibration fold. A shift lead gets told to go look, which is the right instruction for a confident score the model cannot vouch for.

Run the console

python -m forgewatch.app

Opens on port 7860 with sliders for the six inputs, four preset operating states, the triage card and the machine-readable payload side by side.

Configuration

Everything tunable lives in src/forgewatch/config.py.

setting default what it changes
CostModel.unplanned_downtime_hours 4.5 length of an unplanned stop
CostModel.cell_hourly_rate_usd 2400.0 value of an hour of cell time
CostModel.intervention_efficacy 0.85 share of alerts that avert the failure
CONFIDENCE_LEVEL 0.9 conformal coverage target
SurvivalConfig.hazard_budget 0.05 cumulative risk ceiling for the tool budget
TEST_SIZE / CALIBRATION_SIZE 0.2 / 0.2 fold sizes
RANDOM_STATE 613829 seed for splits and models

Changing a cost moves every dollar figure in metrics.json on the next training run. Nothing is hard-coded downstream.

Output structure

artifacts/
  forgewatch.joblib      pipeline, conformal bander, tool-life model, threshold
  metrics.json           dataset profile, model comparison, operating point,
                         economics with sensitivity, conformal report,
                         hazard ratios, the repeated-split threshold study
                         with its per-split rows, chart index
assets/
  data-profile.png       class balance and failure mode flags
  model-curves.png       ROC and precision-recall for all three candidates
  calibration.png        reliability curves with Brier scores
  cost-curve.png         expected cost against threshold
  confusion.png          held-out confusion matrix at the chosen threshold
  global-importance.png  glass-box term importances
  tool-survival.png      Kaplan-Meier curves by workpiece grade
  conformal-bands.png    triage band distribution
  threshold-stability.png  advantage over a 0.5 cut across 20 splits

Project structure

forgewatch/
  src/forgewatch/
    config.py       paths, column names, cost and survival settings
    data.py         loading, DuckDB audit, quality checks, three-way split
    features.py     derived physical quantities as a stateless transformer
    models.py       the three candidate pipelines
    economics.py    cost model, threshold sweep, sensitivity grid
    conformal.py    split-conformal wrapper and the triage bands
    survival.py     Kaplan-Meier and Cox proportional hazards tool-life head
    explain.py      global importances and per-reading plain-language reasons
    evaluate.py     metrics and every chart
    train.py        the orchestration entry point
    stability.py    the repeated-split study behind the threshold section
    scoring.py      the bundle and the single scoring function
    app.py          Gradio console, binding straight to scoring.py
  tests/            98 tests
  data/             the readings
  assets/           charts
  .github/workflows/ci.yml

Technical details

Leakage. Three guards, all of them tested. The five per-mode flags and both identifiers never enter the feature frame, and test_feature_frame_drops_every_label_derived_column asserts on the exact surviving column set. Feature derivation is a stateless transformer placed first in the pipeline, so it cannot carry training statistics into inference; test_transformer_is_stateless_across_folds fits it on one slice and transforms another to prove it. Scaling and one-hot encoding sit inside the pipeline, so cross-validation refits them per fold.

Class imbalance. 339 positives in 10,000 rows. Handled with balanced class weights rather than synthetic oversampling. Oversampling would invent machining states that no cell would ever produce, and with a glass-box model those invented states end up visible in the shape functions. Model selection uses average precision, which does not reward a model for getting the 96.6% majority right.

Calibration fold. The conformal layer needs scores from data the model never fit on, otherwise the coverage guarantee quietly evaporates. That is why the split is three-way rather than the usual two. The same fold also picks the cost-optimal threshold, keeping the test fold genuinely untouched until the final evaluation.

Derived features. Four, all computable from one reading with no pooled statistics: process-minus-air temperature, mechanical power at the spindle from torque and angular velocity, the wear-times-torque product that the overstrain mode keys on, and the speed-to-torque ratio. The dataset's failure modes are generated from documented physical rules, so features aligned with that physics do well. That is worth saying plainly: performance here is an upper bound on what the same approach would deliver on raw shop-floor telemetry, where the rules are not known and the sensors are noisier.

Threshold. Chosen by minimising expected cost over a 197-point grid on the calibration fold, with ties broken toward the higher threshold to keep alert volume down. On the shipped split it lands at 0.16, well below the default 0.5, because a missed failure costs roughly eight times a false alarm. The repeated-split study in the results section measures what that selection is worth and how much the selected value moves between splits; the short version is that the procedure pays every time and the number it returns should not be read as precise.

Explainability. Per-reading contributions come from the additive model's term scores, so they are exact rather than an approximation of a black box. Pairwise interaction terms are excluded from the sentences on purpose: they appear in the global chart, but "spindle speed crossed with heat rejection" is not something anyone can check against the panel in front of them. When a non glass-box estimator is in the pipeline, which happens on the quick path, the reasons come back empty rather than raising mid-alert.

Testing

pytest
pytest --cov=forgewatch --cov-report=term-missing

98 tests, 94% statement coverage. They are organised by the thing that could go wrong rather than by module:

  • test_data.py — schema and range validation, the DuckDB audit against pandas, split disjointness and stratification, reproducibility, and the two label quirks in the source table
  • test_features.py — every derived formula checked by hand, input immutability, zero-torque division, and the stateless-across-folds proof
  • test_models.py — pipeline step ordering, that no label-derived column reaches the estimator, probability well-formedness, and an unseen workpiece grade degrading rather than raising
  • test_economics.py — the savings formula worked out by hand, threshold extremes, monotonicity in efficacy, and the full sensitivity grid
  • test_conformal.py — coverage near the requested level, set-share arithmetic, that the stop band is riskier than the clear band, and that a tighter confidence level never narrows coverage
  • test_survival.py — monotone survival curves, concordance above chance, hazard-ratio interval ordering, torque raising the hazard, and a harder cut shortening the budget
  • test_scoring.py — input validation, the assessment contract, JSON round-tripping, bundle persistence, and a punishing cut scoring riskier than a gentle one
  • test_explain.py — label prettification, sorted contributions, narrative wording, and graceful degradation
  • test_app.py — every preset scores, the card renders, the interface builds
  • test_pipeline_end_to_end.py — a full quick training run, then asserting every chart lands on disk, the metrics sit in range, the operating point saves money and the saved bundle scores a fresh reading
  • test_threshold_finding.py — runs the repeated-split study for real against the shipped model, at six splits rather than twenty to keep the suite usable, and asserts the direction and rough magnitude of all three claims the README makes from it: that optimising beats 0.5 on every split, that the advantage is noisy enough to justify repeating, and that the selected threshold moves between splits without ever reaching 0.5. The bounds are deliberately loose, so they catch a change of direction rather than a change of decimal place

Lint and format are enforced by ruff with E, F, W, I, UP, B, C4, SIM, ARG, RET enabled.

Dataset

AI4I 2020 Predictive Maintenance Dataset, 10,000 readings from a simulated milling machine with five documented failure modes. Donated to the UCI Machine Learning Repository by Stephan Matzka (Hochschule für Technik und Wirtschaft Berlin).

The readings are synthetic but built to mirror real milling behaviour, including the correlations between tool wear, torque, spindle speed and heat rejection that drive the failure modes. That is a strength for reproducibility and a limitation for claims about live shop-floor data, noted above.

License

MIT. See LICENSE.

About

Predictive maintenance for a CNC cell: glass-box failure model, conformal triage bands (MAPIE), survival-based tool-life budgets, cost-optimal thresholds.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages