Image-derived text delivery location descriptions
A computer-implemented method includes obtaining an aerial image representing an object in an environment and providing the aerial image as input to a machine learning model. Based on the aerial image, and using the machine learning model, a textual description of a location of the object in the environment is generated and the textual description of the location of the object is outputted.
1 . A computer-implemented method comprising:
obtaining an aerial image representing an object in an environment;
generating, using a semantic model and based on the aerial image, a semantic map that represents, for each respective visual feature of a plurality of visual features in the aerial image, a corresponding classification of the respective visual feature;
providing the aerial image and the semantic map as input to a machine learning model;
generating, using the machine learning model and based on the aerial image and the semantic map, a textual description of a location of the object in the environment; and
outputting the textual description of the location of the object.
2 . The computer-implemented method of claim 1 , wherein the object comprises a package.
3 . The computer-implemented method of claim 2 , wherein the package has been delivered to the environment by an unmanned aerial vehicle, and wherein the aerial image has been captured by the unmanned aerial vehicle.
4 . The computer-implemented method of claim 1 , wherein the machine learning model has been trained using a plurality of training samples, wherein each respective training sample of the plurality of training samples comprises (i) a corresponding aerial image of a corresponding training environment and (ii) a corresponding textual description of a location of a training object located in the corresponding training environment.
5 . The computer-implemented method of claim 1 , wherein the machine learning model is configured to generate textual descriptions that anonymize visual information contained in aerial images, and wherein the textual description anonymizes at least some visual information contained in the aerial image.
6 . The computer-implemented method of claim 1 , wherein the textual description does not refer to objects outside of a designated boundary within the training environment.
7 . The computer-implemented method of claim 1 , wherein the aerial image comprises a composite aerial image, and wherein obtaining the aerial image comprises:
obtaining a plurality of aerial images of the environment, wherein at least some aerial images of the plurality of aerial images represent the object, and wherein the plurality of aerial images represents the environment from different points of view; and
determining the composite aerial image by combining image data from the plurality of aerial images.
8 . The computer-implemented method of claim 1 , wherein the aerial image comprises a plurality of aerial images, and wherein the semantic model is configured to generate the semantic map based on semantic information from the plurality of aerial images.
9 . The computer-implemented method of claim 1 , further comprising:
determining, based on satellite-based navigation data associated with the aerial image, an estimated location of the object in the environment; and
providing the estimated location of the object as input to the machine learning model, wherein the machine learning model is configured to generate the textual description further based on the estimated location of the object.
10 . The computer-implemented method of claim 9 , further comprising:
determining, based on a comparison between the estimated location of the object in the environment and a predicted location of the object in the environment, an error measurement; and
providing the error measurement as input to the machine learning model, wherein the machine learning model is configured to generate the textual description further based on the error measurement.
11 . The computer-implemented method of claim 1 , further comprising:
providing, as input to the machine learning model, a time at which the aerial image was captured, wherein the machine learning model is configured to generate the textual description further based on the time at which the aerial image was captured.
12 . The computer-implemented method of claim 1 , further comprising:
providing, as input to the machine learning model, a representation of an altitude at which the aerial image was captured, wherein the machine learning model is configured to generate the textual description further based on the representation of the altitude at which the aerial image was captured.
13 . The computer-implement method of claim 1 , further comprising:
transmitting, to a client device, the textual description;
receiving, from the client device, a response to the transmitted textual description; and
updating, based on the response, a status associated with the object.
14 . The computer-implement method of claim 13 , further comprising:
modifying, based on the response, the textual description; and
transmitting, to the client device, the modified textual description.
15 . The computer-implemented method of claim 1 , wherein the textual description of the location of the object describes the location of the object relative to another object in the environment.
16 . The computer-implemented method of claim 1 , wherein the textual description of the location of the object includes a cardinal direction.
17 . The computer-implemented method of claim 1 , wherein the semantic map facilitates anonymization of visual information in the aerial image by (i) including a classification of a first visual feature of the aerial image and (ii) masking a second visual feature of the aerial image.
18 . A system comprising:
a processor; and
a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations comprising:
obtaining an aerial image representing an object in an environment;
generating, using a semantic model and based on the aerial image, a semantic map that represents, for each respective visual feature of a plurality of visual features in the aerial image, a corresponding classification of the respective visual feature;
providing the aerial image and the semantic map as input to a machine learning model;
generating, using the machine learning model and based on the aerial image and the semantic map, a textual description of a location of the object in the environment; and
outputting the textual description of the location of the object.
19 . The system of claim 18 , further comprising:
an unmanned aerial vehicle, wherein the aerial image is captured by the unmanned aerial vehicle and wherein the operations further comprise:
delivering the object to the environment.
20 . A non-transitory computer readable medium comprising program instructions executable by one or more processors to perform operations, the operations comprising:
obtaining an aerial image representing an object in an environment;
generating, using a semantic model and based on the aerial image, a semantic map that represents, for each respective visual feature of a plurality of visual features in the aerial image, a corresponding classification of the respective visual feature;
providing the aerial image and the semantic map as input to a machine learning model;
generating, using the machine learning model and based on the aerial image and the semantic map, a textual description of a location of the object in the environment; and
outputting the textual description of the location of the object.