This repo is responsible for training models simple models (im not expert but who really gives a damn).
This repo's models are used as inference package that another system can use for prediction.
Artifacts and the inference wheel built here get copied into the serving repo
(model-garden-interface).
| Path | Responsibility |
|---|---|
<model>/train/*.py |
quite obvious |
<model>/train/export.py |
Writes the versioned artifact folder (data files + manifest.json) super imp |
<model>/inference/*.py |
the other system actually loads and calls .predict() on. Does not import anything from train/ |
scripts/train_<model>.py |
Orchestration entrypoint: prep → export → smoke test. |
artifacts/<model>/<version>/ |
The physical output that gets copied to the other system |
pyproject.toml |
Defines the installable *.inference package(s) that get built into a wheel |
Every inference class, regardless of model type, implements:
class SomeModel:
def __init__(self, artifacts_path: str): ...
def predict(self, **kwargs) -> dict: ...And every export() function writes a manifest.json with at least:
{
"name": "model-name",
"version": "1.0.0",
"model_type": "...",
"package": "ml-garden-inference==1.0.0",
"entrypoint": "package.module:ClassName",
"input_schema": {},
"output_schema": {},
"config": {}
}This is the only thing linking train/ and inference/ — there's no direct
function call between them, just agreement on file names and manifest keys.
cd Simple-ml-model-garden/
python -m venv env && source env/bin/activate
pip install -e . # editable install, so scripts/ and notebooks/ can import the packagecd Simple-ml-model-garden/
python -m scripts.train_expense_categorizerThis should leave with some "things" built in the artifact dir:
artifacts/expense-categorizer/1.0.0/
├── embeddings.npy
├── transactions.parquet
└── manifest.json
pip install build
python -m buildProduces dist/ml_garden_inference-<version>-py3-none-any.whl. Sanity-check
what's actually inside it before shipping — you should only see */inference/*,
never */train/*:
unzip -l dist/ml_garden_inference-*.whl-
Create the folder pair:
<model_name>/ ├── train/ │ ├── __init__.py │ ├── ... # clean.py / features.py / embed.py — whatever prep it needs │ └── export.py └── inference/ ├── __init__.py └── <model_name>.py -
Write
train/export.py— fit/produce the model, write its data files (.npy+.parquet,.joblib,.pt, whatever fits the model type) toartifacts/<model-name>/<version>/, and writemanifest.jsonalongside them withentrypoint,input_schema,output_schema, and anyconfigthe inference class needs. -
Write
inference/<model_name>.py— a class with__init__(self, artifacts_path)(reads config frommanifest.json, loads the data files) andpredict(self, **kwargs) -> dict. It must not import anything fromtrain/. -
Write
scripts/train_<model_name>.py— prep data → callexport()→ smoke-test by instantiating the inference class against the artifact just written and calling.predict()once. -
Update
pyproject.toml— add the new inference package toinclude:[tool.setuptools.packages.find] include = [ "expense_categorizer.inference*", "late_fee_predictor.inference*", "<model_name>.inference*", # add this line ]
Add any new dependencies (e.g.
xgboost) todependencies. Then:pip install -e . -
Run it:
python -m scripts.train_<model_name>
-
Build and ship:
python -m build
Copy the new
artifacts/<model-name>/<version>/folder and the new/updated wheel fromdist/over to Repo 2 (model-garden-api/artifacts/...andmodel-garden-api/vendor/...), thenpip install -r requirements.txtthere. No code changes needed in Repo 2 — its registry discovers the new model from the manifest alone.