Apparatus and method for crop yield prediction
Aspects of the subject disclosure may include, for example, a device comprising: a processing system including a processor; and a memory that stores executable instructions that, when executed by the processing system, perform operations, the operations comprising: identifying an occurrence of one or multiple phenology stages of a crop, resulting in identified occurrences; optimizing, based upon the identified occurrences, a yield model, wherein the yield model produces, after the optimizing, a first predicted yield for a first region; and generating a second predicted yield based upon the first predicted yield, wherein the second predicted yield covers a second region that is smaller than the first region. Additional embodiments are disclosed.
1 . A device for creating a model capable of predicting and/or estimating crop yield and for predicting and/or estimating crop yield, the device comprising:
a processing system including a processor; and
a memory that stores executable instructions that, when executed by the processing system, perform operations, the operations comprising:
defining a set of at least one geographic region;
obtaining a set of environmental, weather, soil, and satellite-derived remote sensing variables in the set of at least one geographic region during a first set of one or more time windows, wherein the set of variables is associated with a crop and wherein the obtaining of the set of variables at least partially involves converting the remote sensing variables to a geographic projection using nearest neighbor resampling;
training a yield model, wherein the set of variables is a first set of inputs to the yield model, and wherein the training utilizes machine learning and comprises:
generating a gap-free, smoothed time series of a metric for a plurality of dates by applying a Savitzky-Golay filter to the geographic projection for each geographic region of the set of at least one geographic region;
determining a date when the metric exhibits a particular characteristic;
applying a plurality of time shifts to the time series on the determined date to estimate an occurrence of a phenology stage of the crop;
calculating a fraction of pixels that have achieved a threshold of the phenology stage based, at least in part, on each time shift of the plurality of time shifts;
conducting regressions using the calculated fractions to predict yield data;
comparing the predicted yield data to actual yield data;
selecting at least one time shift from the plurality of time shifts based on an error of the at least one time shift, the error being extracted from the conducted regressions and wherein the first set of one or more time windows is defined based on the at least one time shift;
performing additional regressions using least absolute shrinkage and selection operator (LASSO), wherein inputs to the additional regressions comprise the remote sensing variables, the geographic projection, the time series, or the at least one time shift;
predicting, after the training, a first predicted yield for a first region via the trained yield model;
identifying a unit of land, and identifying an occurrence of a second phenology stage of the crop for the unit of land, resulting in a second occurrence;
obtaining a second set of environmental, weather, soil, and satellite-derived remote sensing variables for the unit of land during a second set of one or more time windows defined based on the second occurrence; and
generating a second predicted yield using the trained yield model, and using the second set of environmental, weather, soil, and satellite-derived remote sensing variables as a second set of inputs to the trained yield model, wherein the second predicted yield covers a second region that is smaller than the first region;
wherein the yield model is capable of modeling one or more crop yields at a pixel scale.
2 . The device of claim 1 , wherein the occurrence of the phenology stage comprises multiple phenology stages of the crop.
3 . The device of claim 1 , wherein the satellite data is real-time satellite data.
4 . The device of claim 1 , wherein:
the first region is on a scale of a country, state, county, or user-defined region; and
the second region is on a scale of a farm, field, or portion of a field.
5 . The device of claim 1 , wherein the first predicted yield is based on an aggregation of a plurality of models of yields at a pixel scale or at a field scale within the first region.
6 . The device of claim 1 , wherein the crop is corn, soybean, wheat, rice, cotton, or other row crops.
7 . The device of claim 1 , wherein the set of at least one geographic region is on a scale of a country, state, county, single farm, single field, or user-defined region.
8 . The device of claim 1 , wherein the metric is green chlorophyll vegetation index and the particular characteristic is a maximum.
9 . The device of claim 1 , wherein the operations further comprise:
generating a yield map for the first region based on the first predicted yield for the first region; and
generating a second yield map for the second region based on the second predicted yield for the second region.
10 . A method for use with predicting and/or estimating crop yield, the method comprising:
obtaining satellite data associated with a crop in a first set of one or more geographical regions;
obtaining additional data associated with the crop in a second set of one or more geographical regions, wherein the second set of one or more geographical regions at least partially overlaps the first set of one or more geographical regions;
obtaining actual yield information regarding the crop in a third set of one or more geographical regions, wherein the third set of one or more geographical regions at least partially overlaps the first set of one or more geographical regions or the second set of one or more geographical regions;
extracting, calculating, or deriving a first set of variables from the satellite data and/or the additional data over at least partially overlapping geographical regions from the first set, second set, and/or third set by sampling from the satellite data in the at least partially overlapping geographical regions and/or sampling from the additional data in the at least partially overlapping geographical regions, wherein the sampling is performed by iterating over pixels of the satellite data or additional data and selecting pixels that either partially or fully overlap the first set of one or more geographical regions, the second set of one or more geographical regions, and/or the third set of one or more geographical regions, and wherein the extracting, calculating, or deriving of the first set of variables at least partially involves converting the satellite data to a geographic projection using nearest neighbor resampling;
training a yield model, resulting in a trained yield model, wherein the yield model receives the first set of variables as input, and wherein a goal of the trained yield model is to minimize a difference between predicted yield information output by the yield model regarding the crop in the third set of one or more geographical regions and the actual yield information regarding the crop in the third set of one or more geographical regions, and wherein the training utilizes machine learning and comprises:
generating a gap-free, smoothed time series of a metric for a plurality of dates by applying a Savitzky-Golay filter to the geographic projection for each geographical region of the first set of one or more geographical regions;
determining a date when the metric exhibits a particular characteristic;
applying a plurality of time shifts to the time series on the determined date to estimate an occurrence of a phenology stage of the crop;
calculating a fraction of pixels that have achieved a threshold of the phenology stage based, at least in part, on each time shift of the plurality of time shifts;
conducting regressions using the calculated fractions to predict yield data;
comparing the predicted yield data to actual yield data;
selecting at least one time shift from the plurality of time shifts based on an error of the at least one time shift, the error being extracted from the conducted regressions;
performing additional regressions using least absolute shrinkage and selection operator (LASSO), wherein inputs to the additional regressions comprise the satellite data, the geographic projection, the time series, or the at least one time shift;
defining a fourth set of one or more geographical regions wherein the fourth set of one or more geographical regions is smaller than the first, second, or third sets of one or more geographical regions in area;
extracting, calculating, or deriving a second set of variables from satellite data and/or additional data over the fourth set of one or more geographical regions, by sampling from the satellite data and/or the additional data over the fourth set of one or more geographical regions; and
predicting yield information regarding the crop in the fourth set of one or more geographical regions using the trained yield model.
11 . The method of claim 10 , wherein the phenology stage comprises a reproductive period.
12 . The method of claim 10 , wherein the additional data associated with the crop in the second set of one or more geographical regions comprises climate data, weather data, environmental data, soil data, surface reflectance data, crop biomass data, water stress data, vapor pressure deficit, and/or vegetation indices.
13 . The method of claim 10 , wherein each of the first, second, third, and fourth sets of one or more geographical regions is/are on a scale of a country, state, county, single farm, single field, or user-defined region.
14 . The method of claim 10 , wherein the metric is green chlorophyll vegetation index and the particular characteristic is a maximum.
15 . The method of claim 10 , further comprising generating a yield map for the predicted yield information regarding the crop in the fourth set of one or more geographical regions.
16 . A method of training and/or applying a machine-learning model capable of predicting and/or estimating crop yield, the method comprising:
training a machine-learning model via the steps comprising:
obtaining, from a dataset, actual yield data associated with a crop;
obtaining, from a second dataset, satellite data associated with the crop;
converting the satellite data to a geographic projection using nearest neighbor resampling;
generating a gap-free, smoothed time series of a metric for a plurality of dates by applying a Savitzky-Golay filter to the geographic projection;
determining a date when the metric exhibits a particular characteristic;
applying a plurality of time shifts to the time series on the determined date to estimate an occurrence of a phenology stage of the crop;
calculating a fraction of pixels that have achieved a threshold of the phenology stage based, at least in part, on each time shift of the plurality of time shifts;
conducting regressions using the calculated fractions to predict yield data;
comparing the predicted yield data to actual yield data;
selecting at least one time shift from the plurality of time shifts based on an error of the at least one time shift, the error being extracted from the conducted regressions;
performing additional regressions using least absolute shrinkage and selection operator (LASSO), wherein inputs to the additional regressions comprise the satellite data, the geographic projection, the time series, or the at least one time shift.
17 . The method of claim 16 , wherein the metric is green chlorophyll vegetation index and the particular characteristic is a maximum.
18 . The method of claim 16 , further comprising generating a yield map showing a predicted yield in a region wherein the predicted yield is based on the predicted yield data.