Automated generation and use of building videos with accompanying narration from analysis of acquired images and other building information
Techniques are described for using computing devices to perform automated operations for automatically generating information about attributes of buildings from automated analysis of building information that includes floor plans and acquired building images and to subsequently using the generated building information in one or more further automated manners. In some situations, such automated generation of building information includes automatically determining objects in a building and other attributes of the building, and automatically generating descriptions about the determined building attributes. Information about such determined attributes and generated descriptions may be used in various automated manners, including for updating and/or validating information in existing building descriptions, for determining matching buildings that have similarities to indicated building descriptions or other specified criteria, for controlling device navigation (e.g., autonomous vehicles), for display on client devices in corresponding graphical user interfaces, etc.
1 . A computer-implemented method comprising:
obtaining, by one or more computing devices, data about a house with multiple rooms, including a plurality of images acquired at the house, and a floor plan for the house that includes a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms;
generating, by the one or more computing devices, description information for the house based on the obtained data, including:
analyzing, by the one or more computing devices and using one or more trained first neural networks models, the plurality of images to identify multiple objects inside the house and to determine attributes of the multiple objects;
analyzing, by the one or more computing devices and using one or more trained second neural network models, the floor plan to determine further attributes of the house each corresponding to a characteristic of the room layout;
generating, by the one or more computing devices and using one or more trained language models, textual descriptions of each of the attributes and further attributes, and combining the generated textual descriptions to create the attribute description information for the house;
generating, by the one or more computing devices, a video that describes the house using narration based on the generated description information, including:
determining, by the one or more computing devices and using one or more trained third neural network models, a group of multiple images to use in a determined sequence for the video that are a subset of the plurality of images, wherein the multiple images include at least one image in each of the multiple rooms and further include one or more panorama images;
generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the multiple images, at least some visual data of that image to include at a position in the visual portion corresponding to the determined sequence, and further including inserting additional visual data in the visual portion to provide one or more transitions between selected visual data of adjacent images in the determined sequence, wherein the selecting of the at least some visual data of each of the one or more panorama images includes selecting multiple subsets of the visual data of that panorama image that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image; and
generating, by the one or more computing devices, an audio portion of the video, including:
for each of the multiple images, determining one or more of the objects that are visible in the image, and using the textual descriptions for one or more of the attributes of those one or more objects to include audible narrated information in the audio portion that is about those one or more attributes and that occurs concurrent with selected visual data of that image in the visual portion;
adding, for each of the one or more transitions, additional audible narrated information in the audio portion about that transition and that occurs concurrent with the additional visual data in the visual portion for that transition; and
adding, concurrent with at least one of a beginning or an ending of the visual portion of the video, further audible narrated information in the audio portion based on the textual descriptions for each of the further attributes corresponding to a characteristic of the room layout;
receiving, by the one or more computing devices, one or more search criteria; and
presenting, by the one or more computing devices and in response to determining that the house matches the one or more search criteria based at least in part on the narration for the video, search results that indicate the house and include the generated video that describes the house.
2 . The computer-implemented method of claim 1 wherein the analyzing of the plurality of images further includes determining positions of the multiple objects within the multiple rooms, wherein the selecting of the at least some visual data for the multiple images includes selecting visual data to show the determined positions of the multiple objects, wherein including of audible narrated information in the audio portion about attributes of objects further includes indicating those objects and the determined positions of those objects, wherein the search criteria further include indications of one or more positions of the one or more types of objects, wherein the determining that the house matches the one or more search criteria is further based on the determined positions of one or more identified objects of the one or more types, and wherein the presenting of the search results further includes transmitting, by the one or more computing devices over one or more computer networks to a client device from which the search criteria are received, the search results to cause the client device to display the visual portion of the generated video and to audibly play the audio portion of the generated video.
3 . The computer-implemented method of claim 1 wherein the identified multiple objects in the house include at least appliances and fixtures and structural elements, wherein the determined attributes of the multiple objects include colors and types of surface materials, wherein the determined further attributes of the house include both objective attributes about the house that are able to be independently verified and subjective attributes for the house that are predicted by the one or more trained second neural network models, wherein the search criteria include indications of one or more colors and one or more types of surface materials and one or more types of objects, and wherein the determining that the house matches the one or more search criteria is based on one or more of the identified multiple objects and on one or more of the determined attributes of the multiple objects and on one or more subjective attributes of the determined further attributes.
4 . A computer-implemented method comprising:
obtaining, by one or more computing devices, data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building;
generating, by the one or more computing devices and based on the obtained data, a video for the indicated building that describes at least some of the multiple rooms, including:
determining, by the one or more computing devices, multiple attributes for the indicated building based on objects in the indicated building and visible characteristics of the objects, including analyzing the plurality of images to identify the objects and to determine the visible characteristics;
selecting, by the one or more computing devices, a group of at least two images to use for the video that are a subset of the plurality of images, and determining a sequence in which to include visual data of the at least two images in the video, wherein the at least two images include at least one image in each of the at least some rooms and further include one or more panorama images;
generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the at least two images, at least some visual data of that image to include at a position in the visual portion based on the determined sequence, wherein the selecting of the at least some visual data of each of the one or more panorama images includes selecting multiple subsets of the visual data of that panorama image that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image;
generating, by the one or more computing devices, and for each of the at least two images, a textual description of at least one of the multiple attributes that is visible in the visual data of that image;
generating, by the one or more computing devices, an audio portion of the video, including using, for each of the at least two images, the textual description of the at least one attribute visible in the visual data of that image to produce audible narrated information in the audio portion that is about that at least one attribute and occurs concurrently with selected visual data of that image in the visual portion; and
presenting, by the one or more computing devices, at least some of the generated video about the indicated building.
5 . The computer-implemented method of claim 4 wherein the analyzing of the plurality of images to identify the objects and to determine the visible characteristics includes using one or more trained first neural networks, wherein the selecting of the group of at least two images and the determining of the sequence includes using one or more trained second neural networks that further reject at least one of the plurality of images from inclusion in the subset of the at least two images, wherein the generating of the textual description of the at least one attribute for each of the at least two images includes using one or more trained language models, and wherein the method further comprises, before the generating of the video, training the one or more first neural networks to identify objects in images and determine visual characteristics of those objects, training the one or more second neural networks to select images to include in videos and to determine sequences of those images, and training the one or more language models to generate textual descriptions of attributes of buildings.
6 . The computer-implemented method of claim 4 further comprising receiving one or more search criteria, and determining that the indicated building matches the one or more search criteria based at least in part on the narrated information for the generated video, and wherein the presenting of the at least some of the generated video includes transmitting, by the one or more computing devices and over one or more computer networks to one or more client devices, search results that include the generated video for presentation on the one or more client devices.
7 . The computer-implemented method of claim 4 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional audible narrated information in the audio portion based on additional textual description that is generated to describe the one or more building attributes.
8 . A system comprising:
one or more hardware processors of one or more computing devices; and
one or more memories with stored instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform automated operations including at least:
obtaining data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building, and information about multiple attributes for the indicated building that are based at least in part on objects in the indicated building;
generating, for each of at least some of the multiple attributes, a textual description of that attribute;
generating a video for the indicated building that is based on the obtained data and describes at least some of the multiple rooms, including:
selecting a group of at least two images having visual data of the at least some rooms;
generating a visual portion of the video, including selecting, for each of the at least two images, at least some visual data of that image to include in the visual portion; and
generating an audio portion of the video, including, for each of the at least two images, and for at least one attribute of the at least some attributes that is visible in that image, using the generated textual description of that at least one attribute to produce narrated information in the audio portion that is about that at least one attribute and is synchronized with selected visual data of that image in the visual portion; and
providing information about the indicated building that includes the generated video.
9 . The system of claim 8 wherein the at least one computing device includes a server computing device and wherein the one or more computing devices further include a client computing device of a user, and wherein the stored instructions include software instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform further automated operations including:
receiving, by the server computing device, one or more search criteria from the client computing device;
determining, by the server computing device, search results for the search criteria that include the indicated building based at least in part on the generated video;
performing, by the server computing device, the providing of the information about the indicated building by transmitting the information about the indicated building over one or more computer networks to the client computing device, the transmitted information including the determined search results; and
receiving, by the client computing device, the transmitted information including the determined search results, and displaying the determined search results on the client computing device to enable presentation of the generated video on the client computing device.
10 . The system of claim 8 wherein the obtaining of the information about the multiple attributes includes analyzing visual data of the plurality of images to identify the objects in the indicated building, and wherein the at least some of the multiple attributes are at least some of the objects.
11 . The system of claim 8 wherein the selecting of the group of at least two images includes selecting a subset of the plurality of images to include in the group based at least in part on images of the subset being acquired in the at least some rooms, and further includes excluding at least one of the plurality of images from the subset.
12 . The system of claim 8 wherein one or more images of the at least two images are panorama images, and wherein the selecting of at least some visual data of each of the at least two images includes, for each of the one or more images, selecting multiple subsets of the visual data of that image that are to be shown in succession and that correspond to at least one of panning or tilting within that image.
13 . The system of claim 8 wherein the selecting of at least some visual data of each of the at least two images includes, for one of the at least two images, performing zooming within that one image to show information corresponding to one or more of the at least one attributes visible in that one image.
14 . The system of claim 8 wherein the generating of the visual portion of the video further includes adding further visual data in the visual portion to provide one or more transitions between selected visual data of the at least two images, and wherein the generating of the audio portion of the video further includes, for each of the one or more transitions, producing additional narrated information in the audio portion that describes that transition and that is synchronized with further visual data for that transition.
15 . The system of claim 8 wherein the at least two images of the group include images in all of the multiple rooms and the generated video further describes all of the multiple rooms, wherein the automated operations further include, after the generating of the video, revising the video to satisfy one or more indicated criteria by removing some of the visual portion and audio portion, and wherein the providing of the information about the indicated building includes providing the revised video.
16 . The system of claim 15 wherein the generating of the video further includes generating multiple video segments within the video that each corresponds to at least one of one or more of the multiple rooms or one or more of the objects, wherein the indicated criteria include at least one of an indication of a video length or an indication corresponding to at least one room of the multiple rooms or an indication corresponding to at least one object of the objects, and wherein the revising of the video includes removing at least one of the multiple video segments.
17 . The system of claim 15 wherein the indicated criteria are specific to an indicated recipient, wherein the revising of the video is performed to personalize the revised video for the indicated recipient, and wherein the providing of the revised video includes presenting the revised video to an indicated recipient.
18 . The system of claim 8 wherein the generating of the video further includes generating multiple videos that each corresponds to at least one of one or more of the multiple rooms or one or more of the objects, and wherein the providing of the information about the indicated building includes selecting and providing one of the multiple videos that satisfies one or more indicated criteria.
19 . The system of claim 8 wherein the automated operations further include receiving one or more criteria specific to an indicated user, wherein the generating of the video is further performed to personalize the generated video for the indicated user by satisfying the one or more criteria, and wherein the providing of the information about the indicated building includes providing the video to the indicated user.
20 . The system of claim 8 wherein the generating of the textual description of each of the at least some attributes using one or more trained language models, wherein the selecting of the group of at least two images includes using one or more trained neural networks that further reject at least one of the plurality of images from inclusion in the group, and wherein the automated operations further include, before the generating of the video, training the one or more language models to generate textual descriptions of attributes of buildings, and training the one or more neural networks to select images to include in videos.
21 . The system of claim 8 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional narrated information in the audio portion that is generated to describe the one or more building attributes.
22 . The system of claim 8 wherein the obtained data includes a floor plan of the indicated building, wherein the at least some attributes include at least one of one or more subjective attributes generated from analyzing of the floor plan that include at least one of an open floor plan or an accessible floor plan or a non-standard floor plan, or of one or more global attributes generated from the analyzing of the floor plan and is associated with all of the indicated building, or of one or more local attributes generated from analyzing of the plurality of images and each associated with one of the multiple rooms.
23 . The system of claim 8 wherein the objects include at least appliances and fixtures and structural elements that are determined from analyzing of the plurality of images, and wherein the at least some attributes include colors and types of surface materials for the objects that are determined from the analyzing of the plurality of images.
24 . The system of claim 8 wherein the obtained data further includes additional building information including at least one of a textual description of the building, or labels associated with the objects, or labels associated with the rooms, or descriptive textual annotations associated with the objects, or descriptive textual annotations associated with the rooms, or a group of inter-connections that link at least some of the plurality of images, and wherein the automated operations further include analyzing the additional building information to determine some or all of the at least some attributes.
25 . The system of claim 8 wherein the generating of the audio portion of the video includes using one or more language models that are trained to use, as input, information about the at least some attributes, and about locations in the building corresponding to the at least some attributes, and about timing and/or a sequence for the at least some attributes, wherein the one or more language models include at least one of a Vision and Language Model (VLM) that is trained using image/caption tuples, or a Knowledge Enhanced Natural Language Generation (VENLG) model that is trained using one or more defined knowledge sources, or a language model that uses a knowledge graph in which nodes represent entities and edges represent predicate relationships.
26 . A non-transitory computer-readable medium having stored contents that cause one or more computing devices to perform automated operations, the automated operations including at least:
obtaining, by the one or more computing devices, data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building;
generating, by the one or more computing devices and based on the obtained data, a video for the indicated building, including:
determining, by the one or more computing devices, multiple attributes for the indicated building that include objects in the indicated building, including analyzing the plurality of images to identify the objects;
selecting, by the one or more computing devices, a group of one or more images to use for the video that are a subset of the plurality of images and that include at least one panorama image;
generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the one or more images, at least some visual data of that image to include in the visual portion, including selecting multiple subsets of the visual data of each of the at least one panorama images that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image;
generating, by the one or more computing devices, textual descriptions of two or more attributes of the multiple attributes that are visible in the visual data of the one or more images;
generating, by the one or more computing devices, an audio portion of the video, including using the generated textual descriptions of the two or more attributes to produce audible narrated information in the audio portion that is about the two or more attributes and accompanies selected visual data in the visual portion in which the two or more attributes are visible; and
providing, by the one or more computing devices, the generated video for the indicated building.
27 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include software instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform further automated operations including:
receiving, by the one or more computing devices, one or more search criteria from a client computing device;
determining, by the one or more computing devices, search results for the search criteria that include the indicated building based at least in part on the generated video; and
performing, by the one or more computing devices, the providing of the generated video as part of transmitting, over one or more computer networks to the client computing device, the determined search results to enable presentation of the generated video on the client computing device.
28 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include software instructions that, when executed, cause the one or more computing devices to acquire the plurality of images at a plurality of acquisition locations inside the indicated building, wherein the selecting of the group of one or more images to use for the video includes selecting two or more images to be used in a determined sequence in the video, and wherein the generating of the audio portion of the video further includes, for each of the two or more images, and for each of at least one of the two or more attributes that is visible in that image, using a generated textual description of that attribute to produce a portion of the audible narrated information that occurs concurrently with selected visual data of that image in the visual portion.
29 . The non-transitory computer-readable medium of claim 28 wherein the generating of the visual portion of the video further includes adding further visual data in the visual portion to provide one or more transitions between selected visual data of adjacent images in the determined sequence, and wherein the generating of the audio portion of the video further includes, for each of the one or more transitions, producing additional audible narrated information in the audio portion that describes that transition and that occurs concurrently with further visual data for that transition.
30 . The non-transitory computer-readable medium of claim 26 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional audible narrated information in the audio portion that is generated to describe the one or more building attributes.
31 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include one or more data structures, the one or more data structures including at least one of one or more first trained machine learning models used for the analyzing of the plurality of images to identify the objects, or of one or more second trained machine learning models used for the selecting of the group of one or more images to use for the video, or of one or more trained language models used for the generating of the textual descriptions of the two or more attributes.