12

Good, damaged, clod, or wood: a detector that grades soybean grains one at a time

AgricultureVision

Two counting problems with the same shape. At the silo, a sample of grains is graded for damage and foreign material (clods of soil, fragments of wood).

At a glance
Client
Turing Agro product line - grain quality and yield-component estimation
Industry
Agriculture · Grains
Problem
Grain grading and pod counting are done by eye and by sample; the results are slow, subjective, and impossible to audit
Solution
Two object detectors trained from public datasets: a four-class grain detector (good, damaged, clod, wood) and a single-class pod detector, evaluated with per-class precision/recall, mAP and confusion matrices on verified non-overlapping test sets
Stack
Roboflow datasets (7,768 grain images; 2,689 pod images) · YOLO-family object detection · best/last checkpoint comparison · per-class metrics · normalized confusion matrices · augmented dataset expansion (~20k grains, ~7k pods)
Result
Grains: mAP50 0.912, F1 0.844. Pods: mAP50 0.907, F1 0.849
The challenge

Two counting problems with the same shape. At the silo, a sample of grains is graded for damage and foreign material (clods of soil, fragments of wood). In the field, pods per plant are a yield component estimated by counting. Both are repetitive, both are done by people with other things to do, and both produce numbers nobody can re-check afterward.

What they needed
  • A grain detector that separates good grain from damaged grain and from the two common contaminants
  • A pod detector that finds every pod on a photographed plant
  • Metrics per class, not a single accuracy number
  • Test sets guaranteed disjoint from training
Goals & success metrics
  • Four-class grain model with per-class precision and recall
  • Single-class pod model with mAP50 above 0.9
  • Confusion matrices on held-out photos
  • A plan to scale training to the augmented datasets
How we did it
  1. 01

    Grains: four classes, 7,768 images

    The dataset labels each grain as good, damaged, clod, or wood. Training was run to 185 epochs (15.6 h); the best checkpoint landed at epoch 84 (4.91 h). Checkpoint - Epochs - Time - Precision - Recall - F1 - mAP50 - mAP50-95; Best - 84 - 4.91 h - 0.836 - 0.836 - 0.844 - 0.912 - 0.781; Last - 185 - 15.6 h - 0.855 - 0.829 - 0.842 - 0.904 - 0.776

    Mosaic of training images with grain bounding boxes.
    Fig. 01. Training mosaic: every grain boxed and classed. Dense, overlapping, and small, which is what makes it a detection task rather than a classification one.
    Per-class precision, recall and mAP table for the grain model.
    Fig. 02. Per-class metrics. The contaminant classes are rarer and the numbers say so.
    Composite of test images with detections and the ground-truth legend.
    Fig. 03. Test photos with detections. Legend: ground truth, wood, good, damaged.
    Normalized confusion matrix for the grain detector.
    Fig. 04. Normalized confusion matrix on the test set. Most confusion is between damaged and good, which is also where human graders disagree.
  2. 02

    Pods: one class, 2,689 images

    The pod dataset has a single class. Training ran 154 epochs (2.54 h); best checkpoint at epoch 104 (1.16 h). It was verified that no test image appears in the training directory. Checkpoint - Epochs - Time - Precision - Recall - F1 - mAP50 - mAP50-95; Best - 104 - 1.16 h - 0.866 - 0.832 - 0.849 - 0.907 - 0.613; Last - 154 - 2.54 h - 0.879 - 0.823 - 0.850 - 0.899 - 0.605

    Mosaic of training images with pod bounding boxes.
    Fig. 05. Pod training mosaic. Pods overlap, hide behind leaves, and vary in orientation.
    Composite of test plant photos with pod detections.
    Fig. 06. Test photos with detected pods. The lower mAP50-95 reflects loose box fit on occluded pods, not missed pods.
    Normalized confusion matrix for the pod detector.
    Fig. 07. Confusion matrix on the test set. Pods are found; background is not mistaken for pods.
  3. 03

    Next step

    Both models are being retrained on the augmented versions of the datasets: roughly 20,000 grain images and 7,000 pod images.

What made it work
  • Best versus last checkpoint compared explicitly, so the shipped model is not just the longest-trained one
  • Per-class metrics and confusion matrices instead of a headline accuracy
  • Test/train overlap verified before reporting
  • Two problems with the same shape handled with the same pipeline
FAQ

Still have a question?

Ask us directly. A person reads it and gets back to you quickly.

Contact us