Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Simple ML Model Garden — Training Repo

This repo is responsible for training models simple models (im not expert but who really gives a damn).

This repo's models are used as inference package that another system can use for prediction.

Artifacts and the inference wheel built here get copied into the serving repo (model-garden-interface).


What each piece does

Path Responsibility
<model>/train/*.py quite obvious
<model>/train/export.py Writes the versioned artifact folder (data files + manifest.json) super imp
<model>/inference/*.py the other system actually loads and calls .predict() on. Does not import anything from train/
scripts/train_<model>.py Orchestration entrypoint: prep → export → smoke test.
artifacts/<model>/<version>/ The physical output that gets copied to the other system
pyproject.toml Defines the installable *.inference package(s) that get built into a wheel

The contract every model must follow

Every inference class, regardless of model type, implements:

class SomeModel:
    def __init__(self, artifacts_path: str): ...
    def predict(self, **kwargs) -> dict: ...

And every export() function writes a manifest.json with at least:

{
  "name": "model-name",
  "version": "1.0.0",
  "model_type": "...",
  "package": "ml-garden-inference==1.0.0",
  "entrypoint": "package.module:ClassName",
  "input_schema": {},
  "output_schema": {},
  "config": {}
}

This is the only thing linking train/ and inference/ — there's no direct function call between them, just agreement on file names and manifest keys.


One-time setup (probably not required in most cases but gonna keep this in here just in case)

cd Simple-ml-model-garden/
python -m venv env && source env/bin/activate
pip install -e .          # editable install, so scripts/ and notebooks/ can import the package

Training a model (existing model)

cd Simple-ml-model-garden/
python -m scripts.train_expense_categorizer

This should leave with some "things" built in the artifact dir:

artifacts/expense-categorizer/1.0.0/
├── embeddings.npy
├── transactions.parquet
└── manifest.json

Building the wheel to ship to the other sys

pip install build
python -m build

Produces dist/ml_garden_inference-<version>-py3-none-any.whl. Sanity-check what's actually inside it before shipping — you should only see */inference/*, never */train/*:

unzip -l dist/ml_garden_inference-*.whl

Refresher: adding a new model

  1. Create the folder pair:

    <model_name>/
    ├── train/
    │   ├── __init__.py
    │   ├── ...            # clean.py / features.py / embed.py — whatever prep it needs
    │   └── export.py
    └── inference/
        ├── __init__.py
        └── <model_name>.py
    
  2. Write train/export.py — fit/produce the model, write its data files (.npy+.parquet, .joblib, .pt, whatever fits the model type) to artifacts/<model-name>/<version>/, and write manifest.json alongside them with entrypoint, input_schema, output_schema, and any config the inference class needs.

  3. Write inference/<model_name>.py — a class with __init__(self, artifacts_path) (reads config from manifest.json, loads the data files) and predict(self, **kwargs) -> dict. It must not import anything from train/.

  4. Write scripts/train_<model_name>.py — prep data → call export() → smoke-test by instantiating the inference class against the artifact just written and calling .predict() once.

  5. Update pyproject.toml — add the new inference package to include:

    [tool.setuptools.packages.find]
    include = [
        "expense_categorizer.inference*",
        "late_fee_predictor.inference*",
        "<model_name>.inference*",       # add this line
    ]

    Add any new dependencies (e.g. xgboost) to dependencies. Then:

    pip install -e .
  6. Run it:

    python -m scripts.train_<model_name>
  7. Build and ship:

    python -m build

    Copy the new artifacts/<model-name>/<version>/ folder and the new/updated wheel from dist/ over to Repo 2 (model-garden-api/artifacts/... and model-garden-api/vendor/...), then pip install -r requirements.txt there. No code changes needed in Repo 2 — its registry discovers the new model from the manifest alone.


About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages