IP Library Patent Application 18374568
Patent Application
App. No. 18/374,568

FUSING MULTIMODAL ENVIRONMENTAL DATA FOR AGRICULTURAL INFERENCE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/374,568
Abstract

Implementations are disclosed for fusing multiple modalities of data into a multimodal feature embedding and then processing the multimodal feature embedding using various downstream processes for training and/or inference purposes. In various implementations, multiple different modalities of agricultural data about an agricultural parcel may be obtained. Each modality of agricultural data may be processed based on a respective modality-specific encoder to generate a respective modality-specific embedding. The plurality of modality-specific embeddings may be processed based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel. In some implementations, the multimodal feature embedding may be processed using downstream computer process(es) to generate agricultural prediction(s) about the agricultural parcel. Additionally or alternatively, the multimodal feature embedding may be used to train the multimodal fusion model and/or the modality specific encoder(s).

Claims (37)

1 . A method implemented using one or more processors, comprising:

obtaining multiple different modalities of agricultural data about an agricultural parcel;

processing each modality of agricultural data based on a respective modality-specific encoder to generate a respective modality-specific embedding, wherein the respective modality-specific encoder is pre-trained for that modality using masked autoencoding;

processing the plurality of modality-specific embeddings based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel;

processing the multimodal feature embedding using one or more downstream computer processes to generate one or more agricultural predictions about the agricultural parcel; and

causing one or more computing devices to render output that includes one or more of the agricultural predictions.

2 . The method of claim 1 , wherein the multimodal fusion machine learning model is jointly trained with at least some of the modality-specific encoders.

3 . The method of claim 2 , wherein the multimodal machine learning model is jointly trained using masked autoencoding.

4 . The method of claim 1 , wherein the multimodal fusion machine learning model comprises a transformer.

5 . The method of claim 1 , wherein the multiple different modalities of data include at least one modality that comprises agricultural time series data about the agricultural parcel.

6 . The method of claim 5 , wherein the agricultural time series data about the agricultural parcel comprises soil moisture data.

7 . The method of claim 5 , wherein the agricultural time series data about the agricultural parcel comprises climate data.

8 . The method of claim 1 , wherein the multiple different modalities of data include at least one modality that comprises tabular data about the agricultural parcel.

9 . The method of claim 8 , wherein the tabular data comprises soil properties of the agricultural parcel.

10 . The method of claim 1 , wherein the multiple different modalities of data include at least one modality that comprises satellite or aerial imagery of the agricultural parcel.

11 . The method of claim 1 , wherein one or more of the downstream computer processes comprises identifying one or more reference multimodal feature embeddings that are sufficiently proximate to the multimodal feature embedding in embedding space, wherein the one or more reference multimodal feature embeddings were generated by processing multiple different modalities of agricultural data about one or more reference agricultural parcels.

12 . The method of claim 11 , wherein the output comprises a recommendation of a suitable crop for the agricultural parcel, wherein the suitable crop is selected based on having been grown in one or more of the identified reference agricultural parcels.

13 . The method of claim 1 , wherein one or more of the downstream computer processes comprises processing the multimodal feature embedding using a downstream machine learning model to perform multi-crop yield forecasting for the agricultural parcel.

14 . A system comprising one or more processors and memory storing instructions that, in response to execution by the one or more processors, cause the one or more processors to:

obtain multiple different modalities of agricultural data about an agricultural parcel;

process each modality of agricultural data based on a respective modality-specific encoder to generate a respective modality-specific embedding, wherein the respective modality-specific encoder is pre-trained for that modality using masked autoencoding;

process the plurality of modality-specific embeddings based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel;

process the multimodal feature embedding using one or more downstream computer processes to generate one or more agricultural predictions about the agricultural parcel; and

cause one or more computing devices to render output that includes one or more of the agricultural predictions.

15 . The system of claim 1 , wherein the multimodal fusion machine learning model is jointly trained with at least some of the modality-specific encoders.

16 . The system of claim 15 , wherein the multimodal machine learning model is jointly trained using masked autoencoding.

17 . The system of claim 14 , wherein the multimodal fusion machine learning model comprises a transformer.

18 . The system of claim 14 , wherein the multiple different modalities of data include at least one modality that comprises agricultural time series data about the agricultural parcel.

19 . A method implemented using one or more processors, comprising:

obtaining multiple different modalities of agricultural data about an agricultural parcel;

masking one or more of the different modalities of agricultural data;

processing the remaining modalities of agricultural data based on respective modality-specific encoders to generate respective modality-specific embeddings, wherein the respective modality-specific encoders are pre-trained for the respective modality using masked autoencoding;

processing the plurality of modality-specific embeddings based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel;

processing the multimodal feature embedding using one or more downstream computer processes to generate one or more agricultural predictions about the agricultural parcel;

comparing the one or more agricultural predictions to one or more ground truth observations; and

training the multimodal fusion machine learning model based on the comparing.

20 . The method of claim 19 , further comprising jointly training one or more of the modality-specific encoders based on the comparing.

Assignments (2)
MERGER Recorded Jun 26, 2024
From: MINERAL EARTH SCIENCES LLC
To: DEERE & CO.
Reel/Frame 068055/0420 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2023
From: ZHANG, YAWEN; CHEN, KEZHEN; RAO, JINMENG; GUO, XIAOYUAN; YANG, JIE; OUTON, LUIS PAZOS
To: MINERAL EARTH SCIENCES LLC
Reel/Frame 065270/0178 →