Skip to content
ranjithguggillaPublic

About

Crop recommendation engine (99.5% accuracy) with a plain-language soil-nutrient advisory layer. scikit-learn, Optuna tuning, Pandera validation, Gradio.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

fieldwise

ci python license

A crop recommendation engine that takes a soil test and a local climate reading and tells a grower which of 22 crops fits the field, along with the nutrient actions behind the call. The model is the easy part on data this clean; the part I spent most of my time on is making the output something an agronomist can actually act on rather than a bare label.

The recommender reaches 99.5% accuracy on a held-out test fold, but the headline number is not the interesting bit. A Gaussian naive Bayes classifier ties a tuned 450-tree random forest on cross-validation, so the model I ship is the simple one. What earns the project its keep is the advisory layer: it compares the field reading against per-crop agronomic targets and writes out, in plain language, where the soil is short and where fertilizer is being wasted.

what it does

You give it seven numbers a soil lab and a weather feed can supply: nitrogen, phosphorus, potassium, temperature, humidity, pH, and rainfall. It gives back:

  • the best-fit crop with a calibrated ranking of the alternatives
  • a plain-language reason, including a confidence reading so nobody over-trusts a single probability
  • nutrient-by-nutrient guidance: raise nitrogen by roughly this much, ease off potassium here, phosphorus is fine

results

All figures below come from a single training run on a stratified 80/20 split (1,760 train, 440 test rows) and are written to reports/metrics.json by the pipeline. Nothing here is typed by hand.

model comparison (5-fold cross-validated macro F1 on the training split)

model macro F1 std
gaussian naive bayes (shipped) 0.9949 0.0042
random forest, tuned with optuna 0.9949 —
random forest, default 0.9931 0.0053
extra trees 0.9926 0.0039
logistic regression 0.9678 0.0068
k-nearest neighbours 0.9650 0.0124

The forest and naive Bayes are a dead heat, and 40 rounds of Optuna search moved the forest from 0.9931 to 0.9949 without passing the far cheaper baseline. Given that, I ship naive Bayes: it trains in a blink and there is nothing to tune.

model comparison

held-out test performance (440 rows, shipped model)

metric value
accuracy 0.9955
macro F1 0.9954
weighted F1 0.9954
top-3 accuracy 1.0000

Twenty of the 22 crops land a perfect F1 on the test fold. The only real confusion is between rice (F1 0.947) and jute (0.952), which is honest: both want warm, humid, high-rainfall ground, so their nutrient profiles overlap. Top-3 accuracy is a clean 1.0, meaning the correct crop is always among the top three suggestions even when the top pick is wrong.

confusion matrix

what the model leans on

Permutation importance on the test fold, so it reflects the shipped model rather than any one estimator's internals:

feature importance
humidity 0.451
potassium (K) 0.434
rainfall 0.359
phosphorus (P) 0.181
nitrogen (N) 0.153
temperature 0.083
pH 0.052

feature importance

Humidity, potassium, and rainfall carry most of the signal, which matches how an agronomist would separate a paddy crop from an orchard. pH barely moves the needle on this dataset.

the data itself

class balance

The dataset is perfectly balanced at 100 samples per crop, so there is no imbalance to correct. I still use stratified splits and macro-averaged metrics throughout, because a recommender that quietly ignored the rarer crops in a real deployment would be worse than useless. An interactive view of the nutrient space is written to assets/nutrient_space.html.

business case

This is decision support, not a yield guarantee, so treat the following as an illustrative scenario with the assumptions stated openly.

Take a cooperative of 500 farms averaging 2 hectares each, so 1,000 hectares a season. Assume fertilizer runs about $120 per hectare and gross crop value about $900 per hectare (both vary widely by region and crop; plug in your own). The advisory flags over-application of nutrients: trimming even 10% of wasted fertilizer on the fields it flags saves roughly $12 per hectare, or about $12,000 a season across the cooperative. Separately, matching crop to soil and climate instead of planting by habit is where the larger number sits. A conservative 5% yield improvement on suitable ground is about $45 per hectare, or roughly $45,000 a season. The point of quoting both is that the input-cost saving is the safe, near-certain part, and the yield uplift is the upside that depends on how far current practice sits from the recommendation.

architecture

        soil test + climate reading
                    |
                    v
        +-----------------------+
        |   input validation    |   pandera schema + range guards
        +-----------------------+
                    |
                    v
        +-----------------------+
        |   scaler + model      |   one sklearn Pipeline, no leakage
        |   (GaussianNB)        |
        +-----------------------+
                    |
        ranked crop probabilities
                    |
                    v
        +-----------------------+
        |   advisory layer      |   per-crop agronomic targets
        |   nutrient-gap notes  |   -> plain-language actions
        +-----------------------+
                    |
                    v
        recommendation + reasons  --->  gradio app / json / library call

Preprocessing lives inside the pipeline, so the scaler is fit on training folds only and cross-validation numbers stay honest.

installation

git clone https://github.com/ranjithguggilla/fieldwise.git
cd fieldwise
python -m pip install -r requirements.txt
python -m pip install -e .

Python 3.10 or newer.

usage

Train the model end to end. This runs the EDA charts, cross-validates the model zoo, runs the Optuna search, fits the winner, evaluates on the held-out fold, and writes the model, metrics, and figures:

python -m fieldwise.train_pipeline

Score a field from Python:

from fieldwise.predict import score_field

rec = score_field({
    "N": 90, "P": 42, "K": 43,
    "temperature": 20.9, "humidity": 82,
    "ph": 6.5, "rainfall": 203,
})
print(rec.crop)        # 'rice'
print(rec.advisory)    # the full plain-language note

Launch the interface:

python -m fieldwise.app

The app is a thin wrapper over score_field; every slider maps to one input and the output box shows the advisory verbatim.

configuration

Everything tunable sits in src/fieldwise/config.py: the feature list, the plausible range for each input (used both as a validation guard and as the UI slider bounds), the random seed, the test-set fraction, and the number of cross-validation folds. The Optuna trial count is an argument to train_pipeline.run(n_trials=...) if you want a longer or shorter search.

output structure

reports/metrics.json          every number quoted above, machine readable
models/crop_recommender.joblib the fitted pipeline (rebuilt by training/CI)
assets/class_balance.png       samples per crop
assets/feature_correlation.png feature correlation heatmap
assets/model_comparison.png    cross-validated model scores
assets/confusion_matrix.png    test-fold confusion matrix
assets/feature_importance.png  permutation importance
assets/nutrient_space.html     interactive nutrient scatter

project structure

fieldwise/
├── src/fieldwise/
│   ├── config.py          paths, feature names, bounds, seeds
│   ├── data.py            load, validate (pandera), stratified split
│   ├── model.py           pipelines, model zoo, optuna tuning
│   ├── evaluate.py        metrics, confusion, permutation importance
│   ├── advisory.py        nutrient-gap notes and plain-language reasons
│   ├── predict.py         score_field: the one scoring surface
│   ├── plots.py           matplotlib/seaborn/plotly charts
│   ├── train_pipeline.py  end-to-end training run
│   └── app.py             gradio interface
├── tests/                 pytest suite (26 tests)
├── data/                  crop dataset + fertilizer targets
├── assets/                generated charts
├── reports/               metrics.json
└── .github/workflows/     ci

technical details

The seven inputs are all numeric, so the pipeline is a StandardScaler followed by the estimator. The scaler is inert for the tree models and necessary for logistic regression and k-nearest neighbours, and keeping one code path means the comparison is fair. Model selection happens strictly on cross-validated macro F1 over the training split; the test fold is touched exactly once, at the end, on the single chosen model.

A note on the shipped model's probabilities: naive Bayes tends to produce confident, extreme probabilities. I lean on the ranking and the top-3 behaviour rather than reading the raw percentage as a true likelihood, and the advisory phrases confidence in words for the same reason.

The advisory targets come from the fertilizer reference table shipped with the data, not from the model, so the nutrient guidance is grounded in published agronomic ranges rather than learned correlations.

testing

ruff check src tests
pytest -q

The suite has 26 tests covering the data schema (including that it rejects an impossible pH), the stratified split and absence of row overlap between partitions, the advisory logic across low, on-target, and excess nutrient cases, the scoring surface and its input-range warnings, the pipeline construction, and the evaluation metrics. Both lint and tests run in CI on every push and pull request, and the main branch additionally trains the model and uploads it with the metrics as a build artifact.

dataset

Crop recommendation dataset, originally published on Kaggle by Atharva Ingle (atharvaingle/crop-recommendation-dataset) and obtained here via the vendored copy in the Harvestify repository, along with its fertilizer target table. 2,200 rows, 22 crops, seven soil and climate features. Credit to the original author.

license

MIT. See LICENSE.

About

Crop recommendation engine (99.5% accuracy) with a plain-language soil-nutrient advisory layer. scikit-learn, Optuna tuning, Pandera validation, Gradio.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages