IP Library › Granted Patent US 12,632,957
Granted Patent B2
US 12,632,957 · App. 17/956,679 · Granted May 19, 2026

Methods and systems for use in processing images related to crops

Inventors: Robert Brauer (Lincoln, NE); Nima Hamidi Ghalehjegh (Chesterfield, MO)
Assignee: MONSANTO TECHNOLOGY LLC
G06T7/0012G06V10/454G06V10/764G06V10/82G06V20/17G06V20/188G06T2207/10032G06T2207/20024G06T2207/20081G06T2207/20084G06T2207/30188
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,957
App. No.
17/956,679
Granted
May 19, 2026
Kind
B2
Abstract

Systems and methods for processing image data associated with plots are provided. One example computer-implemented method includes accessing a data set including multiple images, a mask for each of the images, and classification data for each of the images, and inputting each of the images to a classifier of a model architecture. The method also includes, for each of the images input to the classifier, generating, by an encoder of the model architecture, a latent image from the input image; generating, by a decoder of the model architecture, an output mask from the latent image; determining, by the classifier, an output classification indicative of a type of crop in the image; comparing the output mask to the corresponding mask in the data set; comparing the output classification to the corresponding classification data in the data set; and modifying a parameter of the model architecture based on the comparisons.

Claims (85)

1 . A computer-implemented method for use in processing image data associated with one or more plots, the method comprising:

accessing, by a computing device, a training data set included in a data structure, the training data set including (i) multiple images, (ii) a mask for each of the multiple images, and (iii) classification data for each of the multiple images, wherein each of the multiple images is representative of a plot, wherein each of the masks corresponds to one of the multiple images and is indicative of non-crop attributes of the plot represented by the one of the multiple images, and wherein the classification data is indicative of a type of crop included in the plot represented by the one of the multiple images;

inputting, by the computing device, each of the multiple images to a model architecture of the computing device, wherein the model architecture includes an encoder connected to a classifier, a first decoder and a second decoder;

for each of the multiple images input to the model architecture:

generating, by the encoder of the model architecture of the computing device, a latent image from the input image;

generating, by the first decoder of the model architecture of the computing device, a first output mask, from the latent image from said encoder;

generating, by the second decoder of the model architecture of the computing device, a second output mask, from the latent image from said encoder, for the input image;

determining, by the classifier of the model architecture of the computing device, an output classification for the crop based on the latent image from said encoder, the output classification indicative of a type of the crop included in the plot represented by the input image;

comparing the first output mask generated by the first decoder to the mask in the training data set corresponding to the input image;

comparing the second output mask generated by the second decoder to the mask in the training data set corresponding to the input image;

comparing the output classification of the input image from the classifier to the classification data for the input image in the training data set; and

modifying, by the computing device, at least one parameter of the model architecture based on the comparisons; and

storing, by the computing device, the at least one parameter of the model architecture in a memory, whereby the model architecture is suited to generating masks, to distinguish between the crop and the non-crop attributes, for at least one subsequent production image of at least one production plot.

2 . The computer-implemented method of claim 1 , wherein generating the latent image includes generating the latent image through incremental convolutions of the image; and

wherein the model architecture defines a convolution neural network (CNN).

3 . The computer-implemented method of claim 1 , wherein generating the latent image includes reducing, by the encoder, a size of the input image by one quarter or less.

4 . The computer-implemented method of claim 1 , wherein comparing the first output mask generated by the first decoder to the mask in the training data set corresponding to the input image includes calculating a first loss indicative of a difference between a first output mask and the mask in the data structure corresponding to the image;

wherein comparing the second output mask generated by the second decoder to the mask in the training data set corresponding to the input image includes calculating a second loss indicative of a difference between a second output mask and the mask in the data structure corresponding to the image;

wherein comparing the output classification of the input image from the classifier to the classification data for the input image in the training data set includes calculating a third loss indicative of a difference between the output classification and the classification data for the input image in the training data set; and

wherein modifying the at least one parameter of the model architecture is based on the calculated first loss, calculated second loss, and calculated third loss.

5 . The computer-implemented method of claim 1 , further comprising, as part of a next iteration of training the model architecture after storing the at least one parameter of the model architecture:

generating, through the model architecture, multiple masks and associated classifier data;

defining a next training data set, which includes, for each of the multiple masks, an input image and the classifier data; and

filtering the next training data set based on greenness-based masks for the input images of the next training data set; and then

inputting, by the computing device, each of the multiple images of the next training data set to the model architecture of the computing device;

for each of the images input of the next training data set:

generating, by the encoder, a latent image from the input image;

generating, by the first decoder, a first output mask, from the latent image;

determining, by the classifier, an output classification for the crop based on the latent image, the output classification indicative of a type of the crop included in the plot represented by the input image;

comparing the first output mask generated by the first decoder to the mask in the data set corresponding to the input image;

comparing the output classification of the input image from the classifier to the classification data for the input image in the data set; and

modifying, by the computing device, the at least one parameter of the model architecture based on the comparisons; and

storing, by the computing device, the modified at least one parameter of the model architecture in the memory.

6 . The computer-implemented method of claim 5 , wherein filtering the next training data set includes filtering the next training data set based on (i) an intersection of the greenness-based mask and the mask for the input image and (ii) a union of the greenness-based mask and the mask for the input image.

7 . The computer-implemented method of claim 1 , further comprising:

generating a mask for the at least one production image of the at least one production plot; and

applying the generated mask for the at least one production image to the at least one production image, to eliminate non-crop attributes of the at least one production image.

8 . The computer-implemented method of claim 7 , further comprising determining phenotypic data from the at least one production image after application of the generated mask.

9 . The computer-implemented method of claim 8 , wherein the phenotypic data includes at least one of stand count, canopy coverage, and/or gap detection.

10 . The computer-implemented method of claim 9 , further comprising identifying non-crop vegetation based on a difference between the first output mask and a greenness-based mask for the input image.

11 . The computer-implemented method of claim 7 , further comprising generating a map representing one or more locations of the crop and/or the non-crop attributes in the at least one production plot, based on the generated mask for the at least one production image and location data associated with the at least one production image.

12 . A system for use in processing image data associated with one or more plots, the system comprising:

a memory including a model architecture, the model architecture including a classifier, an encoder, a first decoder, and a second decoder, wherein said encoder is connected to each of the classifier, the first decoder, and the second decoder; and

a computing device in communication with the memory, the computing device configured to:

access a data set included in a data structure, the data set including (i) multiple images, (ii) a mask for each of the multiple images, and (iii) classification data for each of the multiple images, wherein each of the multiple images is representative of a plot, wherein each of the masks corresponds to one of the multiple images and is indicative of non-crop attributes of the plot represented by the one of the multiple images, and wherein the classification data is indicative of a type of crop included in the plot represented by the one of the multiple images;

input each of the multiple images to the model architecture;

for each of the multiple images input:

(a) generate, via the encoder of the model architecture, a latent image from the input image;

(b) generate, via the first decoder of the model architecture, a first output mask, from the latent image and generate, via the second decoder of the model architecture, a second output mask, from the latent image;

(c) determine, via the classifier, an output classification for the crop based on the latent image, the output classification indicative of a type of the crop included in the plot represented by the input image;

(d) compare i) the first output mask generated by the first decoder to the mask in the data set corresponding to the input image and ii) the second output mask generated by the second decoder to the mask in the data set corresponding to the input image;

(e) compare the output classification of the input image from the classifier to the classification data for the input image in the data set; and

(f) modify at least one parameter of the model architecture based on (f) the comparisons; and

store the at least one parameter of the model architecture in the memory, whereby the model architecture is suited to generate masks for subsequent production images of production plots.

13 . The system of claim 12 , wherein the computing device is configured, in order to generate the latent image, to generate the latent image through incremental convolutions reducing the size of the image to ¼ or less of the input image; and

wherein the model architecture defines a convolution neural network (CNN).

14 . The system of claim 12 , wherein the classification data is indicative of either a first crop or a second crop, wherein the first decoder is specific to the first crop, and wherein the second decoder is specific to the second crop.

15 . The system of claim 12 , wherein the computing device is configured, in order to compare the first output mask generated by the first decoder to the mask in the data set corresponding to the input image, to calculate a first loss indicative of a difference between the first output mask and the mask in the data structure corresponding to the image;

wherein the computing device is configured, in order to compare the second output mask generated by the second decoder to the mask in the data set corresponding to the input image, to calculate a second loss indicative of a difference between the second output mask and the mask in the data structure corresponding to the image;

wherein the computing device is configured, in order to compare the output classification of the input image from the classifier to the classification data for the input image in the data set, to calculate a third loss indicative of a difference between the output classification and the classification data for the input image in the data set; and

wherein the computing device is configured, in order to modify the at least one parameter of the model architecture, to modify the at least one parameter of the model architecture based on the calculated first loss, the calculated second loss, and the calculated third loss.

16 . The system of claim 12 , wherein the computing device is further configured to:

generate a mask for at least one production image of at least one production plot;

apply the generated mask for the at least one production image to the at least one production image, to eliminate non-crop attributes of the at least one production image; and

determine phenotypic data from the image after application of the generated mask, wherein the phenotypic data includes at least one of stand count, canopy coverage, and/or gap detection.

17 . The system of claim 12 , wherein the computing device is further configured to:

generate a second test set of images of a second training data set; and

repeat steps (a)-(f) based on the images of the second training data set, to further modify the at least one parameter of the model architecture, thereby providing a second iteration of training for the model architecture.

18 . The system of claim 17 , wherein the computing device is further configured to filter the second training data set, prior to repeating steps (a)-(f), based on a greenness-based mask for the input images of the second training data set.

19 . A non-transitory computer-readable storage medium including executable instructions for processing image data, which when executed by at least one processor, cause the at least one processor to:

access a first training data set included in a data structure, the first training data set including (i) multiple images, (ii) a mask for each of the multiple images, and (iii) classification data for each of the multiple images, wherein each of the multiple images is representative of a plot, wherein each of the masks corresponds to one of the multiple images and is indicative of non-crop attributes of the plot represented by the one of the multiple images, and wherein the classification data is indicative of either a first crop or a second crop included in the plot represented by the one of the multiple images;

input each of the multiple images to a model architecture, which includes an encoder connected to a classifier, a first decoder and a second decoder;

for each of the multiple images input:

(a) generate, via the encoder of the model architecture, a latent image from the input image;

(b) generate, via the first decoder of the model architecture specific to the first crop, a first output mask, from the latent image;

(c) determine, via the classifier of the model architecture, an output classification for the crop based on the latent image, the output classification indicative of a type of the crop included in the plot represented by the input image;

(d) compare the first output mask generated by the first decoder to the mask in the first training data set corresponding to the input image;

(e) generate, via the second decoder of the model architecture specific to the second crop, a second output mask, from the latent image, for the input image, wherein the second crop is different than the first crop;

(f) compare the second output mask generated via the second decoder to the mask in the first training data set corresponding to the input image;

(g) compare the output classification of the input image from the classifier to the classification data for the input image in the first training data set; and

(h) modify at least one parameter of the model architecture based on the comparisons;

store the at least one parameter of the model architecture in memory;

generate a test set of images for a second training data set;

filter the second training data set based on a greenness-based mask for the set of test images of the second training data set; and

repeat steps (a)-(h) based on the test set of images of the filtered second training data set to further modify the at least one parameter of the model architecture, thereby providing a second iteration of training for the model architecture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2022
From: BRAUER, ROBERT; GHALEHJEGH, NIMA HAMIDI
To: MONSANTO TECHNOLOGY LLC
Reel/Frame 062214/0018 →
Continuity (2)
Provisional Application 63250629 · Sep 30, 2021
Related Publication 20230100004A1 · Mar 30, 2023
References Cited (34)
US 20120250962A1 · Scharf et al. · 2012 [cited by applicant]
US 20150071528A1 · Marchisio et al. · 2015 [cited by applicant]
US 20160171680A1 · Lobell · 2016 [cited by applicant]
US 20180189564A1 · Freitag et al. · 2018 [cited by applicant]
US 20180211156A1 · Guan et al. · 2018 [cited by applicant]
US 20200074605A1 · Goyal et al. · 2020 [cited by applicant]
US 20200125929A1 · Guo et al. · 2020 [cited by applicant]
US 20200342226A1 · Bengtson et al. · 2020 [cited by applicant]
US 20210012109A1 · Chou · 2021 [cited by examiner]
US 20210153500A1 · Kuenzi · 2021 [cited by applicant]
US 20210201024A1 · Lin et al. · 2021 [cited by applicant]
US 20210247313A1 · Ogawa · 2021 [cited by applicant]
US 20210286998A1 · Wilson · 2021 [cited by examiner]
US 20210312591A1 · Ren et al. · 2021 [cited by applicant]
US 20210397836A1 · Li et al. · 2021 [cited by applicant]
US 20220198221A1 · Avegliano et al. · 2022 [cited by applicant]
US 20220217894A1 · Guo et al. · 2022 [cited by applicant]
US 20220327815A1 · Picon Ruiz · 2022 [cited by examiner]
US 20220335715A1 · Geach et al. · 2022 [cited by applicant]
Machefer, M.; Lemarchand, F.; Bonnefond, V.; Hitchins, A.; Sidiropoulos, P. Mask R-CNN Refitting Strategy for Plant Counting and Sizing in UAV Imagery. Remote Sens. 2020, 12, 3015. https://doi.org/10.3390/rs12183015 (Ye… [cited by examiner]
Machefer, M.; Lemarchand, F.; Bonnefond, V.; Hitchins, A.; Sidiropoulos, P. Mask R-CNN Refitting Strategy for Plant Counting and Sizing in UAV Imagery. Remote Sens. 2020, 12, 3015. (Year: 2020). [cited by examiner]
P. Lottes, J. Behley, N. Chebrolu, A. Milioto and C. Stachniss, “Joint Stem Detection and Crop-Weed Classification for Plant-Specific Treatment in Precision Farming,” 2018 IEEE/RSJ International Conference on Intelligen… [cited by examiner]
Birla. “Single Image Super Resolution Using GANS—Keras,” In: medium.com, Nov. 8, 2018, 9 pgs. [cited by applicant]
Shendryk et al. “Integrating satellite imagery and environmental data to predict field-level cane and sugar yields in Australia using machine learning.” In: Field Crops Research, vol. 260. Jan. 1, 2021, 13 pgs. [cited by applicant]
Mai et al. “Batch Inverse-Variance Weighting: Deep Heteroscedastic Regression.” In: arXiv; Jul. 9, 2021, 15 pgs. [cited by applicant]
Kulkarni et al. “Semantic Segmentation of Medium-Resolution Satellite Imagery using Conditional Generative Adversarial Networks”; AI for Earth Sciences Workshop at NeurIP 2020; 7 pgs. [cited by applicant]
Isola et al. “Image-to-Image Translation with Conditional Adversarial Networks”; Berkeley AI Research (BAIR) Laboratory, UC Berkeley, Nov. 26, 2018; 17 pgs. [cited by applicant]
Kwak et al. “Two-stage Deep Learning Model with LSTM-based Autoencoder and CNN for Crop Classification Using Multi-temporal Remote Sensing Images.” In: Korean Journal of Remote Sensing, vol. 37, No. 4, 2021; Aug. 23, 20… [cited by applicant]
Tenreiro et al. “Using NDV1 for the assessment of canopy in agricultural crops within modelling research.” In: Computers and Electronics in Agriculture 182 (2021), Feb. 22, 2021;12 pages. [cited by applicant]
Ellinger, Understanding Spatial Resolution with Drones, TLT Photography, 2017 (Year: 2017), 8 pages. [cited by applicant]
Roman et al., Noise Estimation for Generative Diffusion Models, arXiv, Sep. 12, 2021 (Year: 2021), 11 pages. [cited by applicant]
Gandikota et al., RTC-GAN Real-Time Classification of Satellite Imagery Using Deep Generative Adversarial Networks With Infused Spectral Information, IEEE, 2020 (Year: 2020), 4 pages. [cited by applicant]
Jiang et al., GAN-Based Multi-Level Mapping Network for Satellite Imagery Super-Resolution, IEEE, 2019 (Year: 2019), 6 pages. [cited by applicant]
Liu et al., PSGAN a Generative Adversarial Network for Remote Sensing Image Pan-Sharpening, IEEE, 2018 (Year: 2018), 5 pages. [cited by applicant]