IP Library Granted Patent US 12,633,119
Granted Patent B1
US 12,633,119 · App. 17/892,427 · Granted May 19, 2026

Automated generation and use of building videos with accompanying narration from analysis of acquired images and other building information

Inventors: Eric M. Penner (Centennial, CO); Ivaylo Boyadzhiev (Seattle, WA); Sing Bing Kang (Redmond, WA)
Assignee: MFTB Holdco, Inc.
G06V20/36G06T3/4038G06V10/25G06V10/70G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,119
App. No.
17/892,427
Granted
May 19, 2026
Kind
B1
Abstract

Techniques are described for using computing devices to perform automated operations for automatically generating information about attributes of buildings from automated analysis of building information that includes floor plans and acquired building images and to subsequently using the generated building information in one or more further automated manners. In some situations, such automated generation of building information includes automatically determining objects in a building and other attributes of the building, and automatically generating descriptions about the determined building attributes. Information about such determined attributes and generated descriptions may be used in various automated manners, including for updating and/or validating information in existing building descriptions, for determining matching buildings that have similarities to indicated building descriptions or other specified criteria, for controlling device navigation (e.g., autonomous vehicles), for display on client devices in corresponding graphical user interfaces, etc.

Claims (77)

1 . A computer-implemented method comprising:

obtaining, by one or more computing devices, data about a house with multiple rooms, including a plurality of images acquired at the house, and a floor plan for the house that includes a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms;

generating, by the one or more computing devices, description information for the house based on the obtained data, including:

analyzing, by the one or more computing devices and using one or more trained first neural networks models, the plurality of images to identify multiple objects inside the house and to determine attributes of the multiple objects;

analyzing, by the one or more computing devices and using one or more trained second neural network models, the floor plan to determine further attributes of the house each corresponding to a characteristic of the room layout;

generating, by the one or more computing devices and using one or more trained language models, textual descriptions of each of the attributes and further attributes, and combining the generated textual descriptions to create the attribute description information for the house;

generating, by the one or more computing devices, a video that describes the house using narration based on the generated description information, including:

determining, by the one or more computing devices and using one or more trained third neural network models, a group of multiple images to use in a determined sequence for the video that are a subset of the plurality of images, wherein the multiple images include at least one image in each of the multiple rooms and further include one or more panorama images;

generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the multiple images, at least some visual data of that image to include at a position in the visual portion corresponding to the determined sequence, and further including inserting additional visual data in the visual portion to provide one or more transitions between selected visual data of adjacent images in the determined sequence, wherein the selecting of the at least some visual data of each of the one or more panorama images includes selecting multiple subsets of the visual data of that panorama image that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image; and

generating, by the one or more computing devices, an audio portion of the video, including:

for each of the multiple images, determining one or more of the objects that are visible in the image, and using the textual descriptions for one or more of the attributes of those one or more objects to include audible narrated information in the audio portion that is about those one or more attributes and that occurs concurrent with selected visual data of that image in the visual portion;

adding, for each of the one or more transitions, additional audible narrated information in the audio portion about that transition and that occurs concurrent with the additional visual data in the visual portion for that transition; and

adding, concurrent with at least one of a beginning or an ending of the visual portion of the video, further audible narrated information in the audio portion based on the textual descriptions for each of the further attributes corresponding to a characteristic of the room layout;

receiving, by the one or more computing devices, one or more search criteria; and

presenting, by the one or more computing devices and in response to determining that the house matches the one or more search criteria based at least in part on the narration for the video, search results that indicate the house and include the generated video that describes the house.

2 . The computer-implemented method of claim 1 wherein the analyzing of the plurality of images further includes determining positions of the multiple objects within the multiple rooms, wherein the selecting of the at least some visual data for the multiple images includes selecting visual data to show the determined positions of the multiple objects, wherein including of audible narrated information in the audio portion about attributes of objects further includes indicating those objects and the determined positions of those objects, wherein the search criteria further include indications of one or more positions of the one or more types of objects, wherein the determining that the house matches the one or more search criteria is further based on the determined positions of one or more identified objects of the one or more types, and wherein the presenting of the search results further includes transmitting, by the one or more computing devices over one or more computer networks to a client device from which the search criteria are received, the search results to cause the client device to display the visual portion of the generated video and to audibly play the audio portion of the generated video.

3 . The computer-implemented method of claim 1 wherein the identified multiple objects in the house include at least appliances and fixtures and structural elements, wherein the determined attributes of the multiple objects include colors and types of surface materials, wherein the determined further attributes of the house include both objective attributes about the house that are able to be independently verified and subjective attributes for the house that are predicted by the one or more trained second neural network models, wherein the search criteria include indications of one or more colors and one or more types of surface materials and one or more types of objects, and wherein the determining that the house matches the one or more search criteria is based on one or more of the identified multiple objects and on one or more of the determined attributes of the multiple objects and on one or more subjective attributes of the determined further attributes.

4 . A computer-implemented method comprising:

obtaining, by one or more computing devices, data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building;

generating, by the one or more computing devices and based on the obtained data, a video for the indicated building that describes at least some of the multiple rooms, including:

determining, by the one or more computing devices, multiple attributes for the indicated building based on objects in the indicated building and visible characteristics of the objects, including analyzing the plurality of images to identify the objects and to determine the visible characteristics;

selecting, by the one or more computing devices, a group of at least two images to use for the video that are a subset of the plurality of images, and determining a sequence in which to include visual data of the at least two images in the video, wherein the at least two images include at least one image in each of the at least some rooms and further include one or more panorama images;

generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the at least two images, at least some visual data of that image to include at a position in the visual portion based on the determined sequence, wherein the selecting of the at least some visual data of each of the one or more panorama images includes selecting multiple subsets of the visual data of that panorama image that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image;

generating, by the one or more computing devices, and for each of the at least two images, a textual description of at least one of the multiple attributes that is visible in the visual data of that image;

generating, by the one or more computing devices, an audio portion of the video, including using, for each of the at least two images, the textual description of the at least one attribute visible in the visual data of that image to produce audible narrated information in the audio portion that is about that at least one attribute and occurs concurrently with selected visual data of that image in the visual portion; and

presenting, by the one or more computing devices, at least some of the generated video about the indicated building.

5 . The computer-implemented method of claim 4 wherein the analyzing of the plurality of images to identify the objects and to determine the visible characteristics includes using one or more trained first neural networks, wherein the selecting of the group of at least two images and the determining of the sequence includes using one or more trained second neural networks that further reject at least one of the plurality of images from inclusion in the subset of the at least two images, wherein the generating of the textual description of the at least one attribute for each of the at least two images includes using one or more trained language models, and wherein the method further comprises, before the generating of the video, training the one or more first neural networks to identify objects in images and determine visual characteristics of those objects, training the one or more second neural networks to select images to include in videos and to determine sequences of those images, and training the one or more language models to generate textual descriptions of attributes of buildings.

6 . The computer-implemented method of claim 4 further comprising receiving one or more search criteria, and determining that the indicated building matches the one or more search criteria based at least in part on the narrated information for the generated video, and wherein the presenting of the at least some of the generated video includes transmitting, by the one or more computing devices and over one or more computer networks to one or more client devices, search results that include the generated video for presentation on the one or more client devices.

7 . The computer-implemented method of claim 4 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional audible narrated information in the audio portion based on additional textual description that is generated to describe the one or more building attributes.

8 . A system comprising:

one or more hardware processors of one or more computing devices; and

one or more memories with stored instructions that, when executed by at least one of the one or more hardware processors, cause at least one of the one or more computing devices to perform automated operations including at least:

obtaining data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building, and information about multiple attributes for the indicated building that are based at least in part on objects in the indicated building;

generating, for each of at least some of the multiple attributes, a textual description of that attribute;

generating a video for the indicated building that is based on the obtained data and describes at least some of the multiple rooms, including:

selecting a group of at least two images having visual data of the at least some rooms;

generating a visual portion of the video, including selecting, for each of the at least two images, at least some visual data of that image to include in the visual portion; and

generating an audio portion of the video, including, for each of the at least two images, and for at least one attribute of the at least some attributes that is visible in that image, using the generated textual description of that at least one attribute to produce narrated information in the audio portion that is about that at least one attribute and is synchronized with selected visual data of that image in the visual portion; and

providing information about the indicated building that includes the generated video.

9 . The system of claim 8 wherein the at least one computing device includes a server computing device and wherein the one or more computing devices further include a client computing device of a user, and wherein the stored instructions include software instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform further automated operations including:

receiving, by the server computing device, one or more search criteria from the client computing device;

determining, by the server computing device, search results for the search criteria that include the indicated building based at least in part on the generated video;

performing, by the server computing device, the providing of the information about the indicated building by transmitting the information about the indicated building over one or more computer networks to the client computing device, the transmitted information including the determined search results; and

receiving, by the client computing device, the transmitted information including the determined search results, and displaying the determined search results on the client computing device to enable presentation of the generated video on the client computing device.

10 . The system of claim 8 wherein the obtaining of the information about the multiple attributes includes analyzing visual data of the plurality of images to identify the objects in the indicated building, and wherein the at least some of the multiple attributes are at least some of the objects.

11 . The system of claim 8 wherein the selecting of the group of at least two images includes selecting a subset of the plurality of images to include in the group based at least in part on images of the subset being acquired in the at least some rooms, and further includes excluding at least one of the plurality of images from the subset.

12 . The system of claim 8 wherein one or more images of the at least two images are panorama images, and wherein the selecting of at least some visual data of each of the at least two images includes, for each of the one or more images, selecting multiple subsets of the visual data of that image that are to be shown in succession and that correspond to at least one of panning or tilting within that image.

13 . The system of claim 8 wherein the selecting of at least some visual data of each of the at least two images includes, for one of the at least two images, performing zooming within that one image to show information corresponding to one or more of the at least one attributes visible in that one image.

14 . The system of claim 8 wherein the generating of the visual portion of the video further includes adding further visual data in the visual portion to provide one or more transitions between selected visual data of the at least two images, and wherein the generating of the audio portion of the video further includes, for each of the one or more transitions, producing additional narrated information in the audio portion that describes that transition and that is synchronized with further visual data for that transition.

15 . The system of claim 8 wherein the at least two images of the group include images in all of the multiple rooms and the generated video further describes all of the multiple rooms, wherein the automated operations further include, after the generating of the video, revising the video to satisfy one or more indicated criteria by removing some of the visual portion and audio portion, and wherein the providing of the information about the indicated building includes providing the revised video.

16 . The system of claim 15 wherein the generating of the video further includes generating multiple video segments within the video that each corresponds to at least one of one or more of the multiple rooms or one or more of the objects, wherein the indicated criteria include at least one of an indication of a video length or an indication corresponding to at least one room of the multiple rooms or an indication corresponding to at least one object of the objects, and wherein the revising of the video includes removing at least one of the multiple video segments.

17 . The system of claim 15 wherein the indicated criteria are specific to an indicated recipient, wherein the revising of the video is performed to personalize the revised video for the indicated recipient, and wherein the providing of the revised video includes presenting the revised video to an indicated recipient.

18 . The system of claim 8 wherein the generating of the video further includes generating multiple videos that each corresponds to at least one of one or more of the multiple rooms or one or more of the objects, and wherein the providing of the information about the indicated building includes selecting and providing one of the multiple videos that satisfies one or more indicated criteria.

19 . The system of claim 8 wherein the automated operations further include receiving one or more criteria specific to an indicated user, wherein the generating of the video is further performed to personalize the generated video for the indicated user by satisfying the one or more criteria, and wherein the providing of the information about the indicated building includes providing the video to the indicated user.

20 . The system of claim 8 wherein the generating of the textual description of each of the at least some attributes using one or more trained language models, wherein the selecting of the group of at least two images includes using one or more trained neural networks that further reject at least one of the plurality of images from inclusion in the group, and wherein the automated operations further include, before the generating of the video, training the one or more language models to generate textual descriptions of attributes of buildings, and training the one or more neural networks to select images to include in videos.

21 . The system of claim 8 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional narrated information in the audio portion that is generated to describe the one or more building attributes.

22 . The system of claim 8 wherein the obtained data includes a floor plan of the indicated building, wherein the at least some attributes include at least one of one or more subjective attributes generated from analyzing of the floor plan that include at least one of an open floor plan or an accessible floor plan or a non-standard floor plan, or of one or more global attributes generated from the analyzing of the floor plan and is associated with all of the indicated building, or of one or more local attributes generated from analyzing of the plurality of images and each associated with one of the multiple rooms.

23 . The system of claim 8 wherein the objects include at least appliances and fixtures and structural elements that are determined from analyzing of the plurality of images, and wherein the at least some attributes include colors and types of surface materials for the objects that are determined from the analyzing of the plurality of images.

24 . The system of claim 8 wherein the obtained data further includes additional building information including at least one of a textual description of the building, or labels associated with the objects, or labels associated with the rooms, or descriptive textual annotations associated with the objects, or descriptive textual annotations associated with the rooms, or a group of inter-connections that link at least some of the plurality of images, and wherein the automated operations further include analyzing the additional building information to determine some or all of the at least some attributes.

25 . The system of claim 8 wherein the generating of the audio portion of the video includes using one or more language models that are trained to use, as input, information about the at least some attributes, and about locations in the building corresponding to the at least some attributes, and about timing and/or a sequence for the at least some attributes, wherein the one or more language models include at least one of a Vision and Language Model (VLM) that is trained using image/caption tuples, or a Knowledge Enhanced Natural Language Generation (VENLG) model that is trained using one or more defined knowledge sources, or a language model that uses a knowledge graph in which nodes represent entities and edges represent predicate relationships.

26 . A non-transitory computer-readable medium having stored contents that cause one or more computing devices to perform automated operations, the automated operations including at least:

obtaining, by the one or more computing devices, data for an indicated building with multiple rooms, including a plurality of images acquired at the indicated building;

generating, by the one or more computing devices and based on the obtained data, a video for the indicated building, including:

determining, by the one or more computing devices, multiple attributes for the indicated building that include objects in the indicated building, including analyzing the plurality of images to identify the objects;

selecting, by the one or more computing devices, a group of one or more images to use for the video that are a subset of the plurality of images and that include at least one panorama image;

generating, by the one or more computing devices, a visual portion of the video, including selecting, for each of the one or more images, at least some visual data of that image to include in the visual portion, including selecting multiple subsets of the visual data of each of the at least one panorama images that are to be shown in succession and that correspond to at least one of panning or tilting within that panorama image;

generating, by the one or more computing devices, textual descriptions of two or more attributes of the multiple attributes that are visible in the visual data of the one or more images;

generating, by the one or more computing devices, an audio portion of the video, including using the generated textual descriptions of the two or more attributes to produce audible narrated information in the audio portion that is about the two or more attributes and accompanies selected visual data in the visual portion in which the two or more attributes are visible; and

providing, by the one or more computing devices, the generated video for the indicated building.

27 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include software instructions that, when executed by the one or more computing devices, cause the one or more computing devices to perform further automated operations including:

receiving, by the one or more computing devices, one or more search criteria from a client computing device;

determining, by the one or more computing devices, search results for the search criteria that include the indicated building based at least in part on the generated video; and

performing, by the one or more computing devices, the providing of the generated video as part of transmitting, over one or more computer networks to the client computing device, the determined search results to enable presentation of the generated video on the client computing device.

28 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include software instructions that, when executed, cause the one or more computing devices to acquire the plurality of images at a plurality of acquisition locations inside the indicated building, wherein the selecting of the group of one or more images to use for the video includes selecting two or more images to be used in a determined sequence in the video, and wherein the generating of the audio portion of the video further includes, for each of the two or more images, and for each of at least one of the two or more attributes that is visible in that image, using a generated textual description of that attribute to produce a portion of the audible narrated information that occurs concurrently with selected visual data of that image in the visual portion.

29 . The non-transitory computer-readable medium of claim 28 wherein the generating of the visual portion of the video further includes adding further visual data in the visual portion to provide one or more transitions between selected visual data of adjacent images in the determined sequence, and wherein the generating of the audio portion of the video further includes, for each of the one or more transitions, producing additional audible narrated information in the audio portion that describes that transition and that occurs concurrently with further visual data for that transition.

30 . The non-transitory computer-readable medium of claim 26 wherein the obtained data further includes a floor plan for the indicated building indicating a room layout with at least two-dimensional room shapes and relative positions of the multiple rooms, wherein the multiple attributes for the indicated building further include one or more building attributes that are identified from analyzing the floor plan and that each corresponds to a characteristic of the room layout, and wherein the generating of the audio portion of the video further includes producing additional audible narrated information in the audio portion that is generated to describe the one or more building attributes.

31 . The non-transitory computer-readable medium of claim 26 wherein the stored contents include one or more data structures, the one or more data structures including at least one of one or more first trained machine learning models used for the analyzing of the plurality of images to identify the objects, or of one or more second trained machine learning models used for the selecting of the group of one or more images to use for the video, or of one or more trained language models used for the generating of the textual descriptions of the two or more attributes.

Assignments (4)
MERGER AND CHANGE OF NAME Recorded Dec 13, 2022
From: PUSH SUB I, INC.; MFTB HOLDCO, INC.
To: MFTB HOLDCO, INC.
Reel/Frame 062117/0936 →
ENTITY CONVERSION Recorded Dec 12, 2022
From: ZILLOW, INC.
To: ZILLOW, LLC
Reel/Frame 062116/0084 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2022
From: ZILLOW, LLC
To: PUSH SUB I, INC.
Reel/Frame 062116/0094 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2022
From: PENNER, ERIC M.; BOYADZHIEV, IVAYLO; KANG, SING BING
To: ZILLOW, INC.
Reel/Frame 060986/0441 →
References Cited (278)
US 5140352A · Moore et al. · 1992 [cited by applicant]
US 6031540A · Golin et al. · 2000 [cited by applicant]
US 6141034A · McCutchen · 2000 [cited by applicant]
US 6317166B1 · McCutchen · 2001 [cited by applicant]
US 6320584B1 · Golin et al. · 2001 [cited by applicant]
US 6323858B1 · Gilbert et al. · 2001 [cited by applicant]
US 6337683B1 · Gilbert et al. · 2002 [cited by applicant]
US 6654019B2 · Gilbert et al. · 2003 [cited by applicant]
US 6683608B2 · Golin et al. · 2004 [cited by applicant]
US 6690374B2 · Park et al. · 2004 [cited by applicant]
US 6731305B1 · Park et al. · 2004 [cited by applicant]
US 6738073B2 · Park et al. · 2004 [cited by applicant]
US 7050085B1 · Park et al. · 2006 [cited by applicant]
US 7129971B2 · McCutchen · 2006 [cited by applicant]
US 7196722B2 · White et al. · 2007 [cited by applicant]
US 7525567B2 · McCutchen · 2009 [cited by applicant]
US 7620909B2 · Park et al. · 2009 [cited by applicant]
US 7627235B2 · McCutchen et al. · 2009 [cited by applicant]
US 7782319B2 · Ghosh et al. · 2010 [cited by applicant]
US 7791638B2 · McCutchen · 2010 [cited by applicant]
US 7909241B2 · Stone et al. · 2011 [cited by applicant]
US 7973838B2 · McCutchen · 2011 [cited by applicant]
US 8072455B2 · Temesvari et al. · 2011 [cited by applicant]
US 8094182B2 · Park et al. · 2012 [cited by applicant]
US RE43786E · Cooper · 2012 [cited by applicant]
US 8463020B1 · Schuckmann et al. · 2013 [cited by applicant]
US 8517256B2 · Stone et al. · 2013 [cited by applicant]
US 8520060B2 · Zomet et al. · 2013 [cited by applicant]
US 8523066B2 · Stone et al. · 2013 [cited by applicant]
US 8523067B2 · Stone et al. · 2013 [cited by applicant]
US 8528816B2 · Stone et al. · 2013 [cited by applicant]
US 8540153B2 · Stone et al. · 2013 [cited by applicant]
US 8594428B2 · Aharoni et al. · 2013 [cited by applicant]
US 8654180B2 · Zomet et al. · 2014 [cited by applicant]
US 8666815B1 · Chau · 2014 [cited by applicant]
US 8699005B2 · Likholyot · 2014 [cited by applicant]
US 8705892B2 · Aguilera et al. · 2014 [cited by applicant]
US RE44924E · Cooper et al. · 2014 [cited by applicant]
US 8854684B2 · Zomet · 2014 [cited by applicant]
US 8861840B2 · Bell et al. · 2014 [cited by applicant]
US 8861841B2 · Bell et al. · 2014 [cited by applicant]
US 8879828B2 · Bell et al. · 2014 [cited by applicant]
US 8953871B2 · Zomet · 2015 [cited by applicant]
US 8989440B2 · Klusza et al. · 2015 [cited by applicant]
US 8996336B2 · Malka et al. · 2015 [cited by applicant]
US 9021947B2 · Landa · 2015 [cited by applicant]
US 9026947B2 · Lee et al. · 2015 [cited by applicant]
US 9035968B2 · Zomet · 2015 [cited by applicant]
US 9041796B2 · Malka et al. · 2015 [cited by applicant]
US 9071714B2 · Zomet · 2015 [cited by applicant]
US 9129438B2 · Aarts et al. · 2015 [cited by applicant]
US 9151608B2 · Malka et al. · 2015 [cited by applicant]
US 9165410B1 · Bell et al. · 2015 [cited by applicant]
US 9171405B1 · Bell et al. · 2015 [cited by applicant]
US 9324190B2 · Bell et al. · 2016 [cited by applicant]
US 9361717B2 · Zomet · 2016 [cited by applicant]
US 9396586B2 · Bell et al. · 2016 [cited by applicant]
US 9438759B2 · Zomet · 2016 [cited by applicant]
US 9438775B2 · Powers et al. · 2016 [cited by applicant]
US 9489775B1 · Bell et al. · 2016 [cited by applicant]
US 9495783B1 · Samarasekera et al. · 2016 [cited by applicant]
US 9576401B2 · Zomet · 2017 [cited by applicant]
US 9619933B2 · Spinella-Marno et al. · 2017 [cited by applicant]
US 9635252B2 · Accardo et al. · 2017 [cited by applicant]
US 9641702B2 · Bin-Nun et al. · 2017 [cited by applicant]
US 9760994B1 · Bell et al. · 2017 [cited by applicant]
US 9786097B2 · Bell et al. · 2017 [cited by applicant]
US 9787904B2 · Birkler et al. · 2017 [cited by applicant]
US 9836885B1 · Eraker et al. · 2017 [cited by applicant]
US 9852351B2 · Aguilera Perez et al. · 2017 [cited by applicant]
US 9953111B2 · Bell et al. · 2018 [cited by applicant]
US 9953430B1 · Zakhor · 2018 [cited by applicant]
US 9990760B2 · Aguilera Perez et al. · 2018 [cited by applicant]
US 9990767B1 · Sheffield et al. · 2018 [cited by applicant]
US 10026224B2 · Bell et al. · 2018 [cited by applicant]
US 10030979B2 · Bjorke et al. · 2018 [cited by applicant]
US 10055876B2 · Ford et al. · 2018 [cited by applicant]
US 10068344B2 · Jovanovic et al. · 2018 [cited by applicant]
US 10083522B2 · Jovanovic et al. · 2018 [cited by applicant]
US 10102639B2 · Bell et al. · 2018 [cited by applicant]
US 10102673B2 · Eraker et al. · 2018 [cited by applicant]
US 10120397B1 · Zakhor et al. · 2018 [cited by applicant]
US 10122997B1 · Sheffield et al. · 2018 [cited by applicant]
US 10127718B2 · Zakhor et al. · 2018 [cited by applicant]
US 10127722B2 · Shakib et al. · 2018 [cited by applicant]
US 10139985B2 · Mildrew et al. · 2018 [cited by applicant]
US 10163261B2 · Bell et al. · 2018 [cited by applicant]
US 10163271B1 · Powers et al. · 2018 [cited by applicant]
US 10181215B2 · Sedeffow · 2019 [cited by applicant]
US 10192115B1 · Sheffield et al. · 2019 [cited by applicant]
US 10204185B2 · Mrowca et al. · 2019 [cited by applicant]
US 10210285B2 · Wong et al. · 2019 [cited by applicant]
US 10235797B1 · Sheffield et al. · 2019 [cited by applicant]
US 10242400B1 · Eraker et al. · 2019 [cited by applicant]
US 10339716B1 · Powers et al. · 2019 [cited by applicant]
US 10366531B2 · Sheffield · 2019 [cited by applicant]
US 10375306B2 · Shan et al. · 2019 [cited by applicant]
US 10395435B2 · Powers et al. · 2019 [cited by applicant]
US 10530997B2 · Shan et al. · 2020 [cited by applicant]
US 10643386B2 · Li et al. · 2020 [cited by applicant]
US 10708507B1 · Dawson et al. · 2020 [cited by applicant]
US 10809066B2 · Colburn et al. · 2020 [cited by applicant]
US 10825247B1 · Vincent et al. · 2020 [cited by applicant]
US 10834317B2 · Shan et al. · 2020 [cited by applicant]
US 11030709B2 · McLinden et al. · 2021 [cited by applicant]
US 11057561B2 · Shan et al. · 2021 [cited by applicant]
US 11164361B2 · Moulon et al. · 2021 [cited by applicant]
US 11164368B2 · Vincent et al. · 2021 [cited by applicant]
US 11165959B2 · Shan et al. · 2021 [cited by applicant]
US 11217019B2 · Li et al. · 2022 [cited by applicant]
US 11238652B2 · Impas et al. · 2022 [cited by applicant]
US 11243656B2 · Li et al. · 2022 [cited by applicant]
US 11252329B1 · Cier et al. · 2022 [cited by applicant]
US 11284006B2 · Dawson et al. · 2022 [cited by applicant]
US 11405549B2 · Cier et al. · 2022 [cited by applicant]
US 11405558B2 · Dawson et al. · 2022 [cited by applicant]
US 11408738B2 · Colburn et al. · 2022 [cited by applicant]
US 20060256109A1 · Acker et al. · 2006 [cited by applicant]
US 20100232709A1 · Zhang et al. · 2010 [cited by applicant]
US 20120075414A1 · Park et al. · 2012 [cited by applicant]
US 20120162253A1 · Collins · 2012 [cited by examiner]
US 20120293613A1 · Powers et al. · 2012 [cited by applicant]
US 20130050407A1 · Brinda et al. · 2013 [cited by applicant]
US 20130342533A1 · Bell et al. · 2013 [cited by applicant]
US 20140043436A1 · Bell et al. · 2014 [cited by applicant]
US 20140044343A1 · Bell et al. · 2014 [cited by applicant]
US 20140044344A1 · Bell et al. · 2014 [cited by applicant]
US 20140125658A1 · Bell et al. · 2014 [cited by applicant]
US 20140125767A1 · Bell et al. · 2014 [cited by applicant]
US 20140125768A1 · Bell et al. · 2014 [cited by applicant]
US 20140125769A1 · Bell et al. · 2014 [cited by applicant]
US 20140125770A1 · Bell et al. · 2014 [cited by applicant]
US 20140236482A1 · Dorum et al. · 2014 [cited by applicant]
US 20140267631A1 · Powers et al. · 2014 [cited by applicant]
US 20140307100A1 · Myllykoski et al. · 2014 [cited by applicant]
US 20140320674A1 · Kuang · 2014 [cited by applicant]
US 20150116691A1 · Likholyot · 2015 [cited by applicant]
US 20150189165A1 · Milosevski et al. · 2015 [cited by applicant]
US 20150262421A1 · Bell et al. · 2015 [cited by applicant]
US 20150269785A1 · Bell et al. · 2015 [cited by applicant]
US 20150302636A1 · Arnoldus et al. · 2015 [cited by applicant]
US 20150310596A1 · Sheridan et al. · 2015 [cited by applicant]
US 20150332464A1 · O'Keefe et al. · 2015 [cited by applicant]
US 20160055268A1 · Bell et al. · 2016 [cited by applicant]
US 20160134860A1 · Jovanovic et al. · 2016 [cited by applicant]
US 20160140676A1 · Fritze et al. · 2016 [cited by applicant]
US 20160217225A1 · Bell et al. · 2016 [cited by applicant]
US 20160260250A1 · Jovanovic et al. · 2016 [cited by applicant]
US 20160286119A1 · Rondinelli · 2016 [cited by applicant]
US 20160300385A1 · Bell et al. · 2016 [cited by applicant]
US 20170034430A1 · Fu et al. · 2017 [cited by applicant]
US 20170067739A1 · Siercks et al. · 2017 [cited by applicant]
US 20170194768A1 · Powers et al. · 2017 [cited by applicant]
US 20170195654A1 · Powers et al. · 2017 [cited by applicant]
US 20170263050A1 · Ha et al. · 2017 [cited by applicant]
US 20170324941A1 · Birkler · 2017 [cited by applicant]
US 20170330273A1 · Holt et al. · 2017 [cited by applicant]
US 20170337737A1 · Edwards et al. · 2017 [cited by applicant]
US 20180007340A1 · Stachowski · 2018 [cited by applicant]
US 20180025536A1 · Bell et al. · 2018 [cited by applicant]
US 20180075168A1 · Tiwari et al. · 2018 [cited by applicant]
US 20180139431A1 · Simek et al. · 2018 [cited by applicant]
US 20180143023A1 · Bjorke et al. · 2018 [cited by applicant]
US 20180143756A1 · Mildrew et al. · 2018 [cited by applicant]
US 20180144487A1 · Bell et al. · 2018 [cited by applicant]
US 20180144535A1 · Ford et al. · 2018 [cited by applicant]
US 20180144547A1 · Shakib et al. · 2018 [cited by applicant]
US 20180144555A1 · Ford et al. · 2018 [cited by applicant]
US 20180146121A1 · Hensler et al. · 2018 [cited by applicant]
US 20180146193A1 · Safreed et al. · 2018 [cited by applicant]
US 20180146212A1 · Hensler et al. · 2018 [cited by applicant]
US 20180165871A1 · Mrowca · 2018 [cited by applicant]
US 20180203955A1 · Bell et al. · 2018 [cited by applicant]
US 20180241985A1 · O'Keefe et al. · 2018 [cited by applicant]
US 20180293793A1 · Bell et al. · 2018 [cited by applicant]
US 20180300936A1 · Ford et al. · 2018 [cited by applicant]
US 20180306588A1 · Bjorke et al. · 2018 [cited by applicant]
US 20180348854A1 · Powers et al. · 2018 [cited by applicant]
US 20180365496A1 · Hovden et al. · 2018 [cited by applicant]
US 20190012833A1 · Eraker et al. · 2019 [cited by applicant]
US 20190026956A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190026957A1 · Gausebeck · 2019 [cited by applicant]
US 20190026958A1 · Gausebeck et al. · 2019 [cited by applicant]
US 20190035165A1 · Gausebeck · 2019 [cited by applicant]
US 20190041972A1 · Bae · 2019 [cited by applicant]
US 20190050137A1 · Mildrew et al. · 2019 [cited by applicant]
US 20190051050A1 · Bell et al. · 2019 [cited by applicant]
US 20190051054A1 · Jovanovic et al. · 2019 [cited by applicant]
US 20190087067A1 · Hovden et al. · 2019 [cited by applicant]
US 20190122422A1 · Sheffield et al. · 2019 [cited by applicant]
US 20190164335A1 · Sheffield et al. · 2019 [cited by applicant]
US 20190180104A1 · Sheffield et al. · 2019 [cited by applicant]
US 20190251645A1 · Winans · 2019 [cited by applicant]
US 20190287164A1 · Eraker et al. · 2019 [cited by applicant]
US 20200336675A1 · Dawson et al. · 2020 [cited by applicant]
US 20200389602A1 · Dawson et al. · 2020 [cited by applicant]
US 20200408532A1 · Colburn et al. · 2020 [cited by applicant]
US 20210044760A1 · Dawson et al. · 2021 [cited by applicant]
US 20210073449A1 · Segev et al. · 2021 [cited by applicant]
US 20210117071A1 · Gharpuray · 2021 [cited by applicant]
US 20210142564A1 · Impas · 2021 [cited by examiner]
US 20210150088A1 · Gallo et al. · 2021 [cited by applicant]
US 20210377442A1 · Boyadzhiev et al. · 2021 [cited by applicant]
US 20210385378A1 · Cier et al. · 2021 [cited by applicant]
US 20220003555A1 · Colburn et al. · 2022 [cited by applicant]
US 20220028156A1 · Boyadzhiev et al. · 2022 [cited by applicant]
US 20220028159A1 · Vincent et al. · 2022 [cited by applicant]
US 20220076019A1 · Moulon et al. · 2022 [cited by applicant]
US 20220092227A1 · Yin et al. · 2022 [cited by applicant]
US 20220114291A1 · Li et al. · 2022 [cited by applicant]
US 20220164493A1 · Li et al. · 2022 [cited by applicant]
US 20220189122A1 · Li et al. · 2022 [cited by applicant]
US 20230384924A1 · Al-Sharieh · 2023 [cited by examiner]
EP 2413097A2 · 2012 [cited by applicant]
EP 2505961A2 · 2012 [cited by applicant]
EP 2506170A2 · 2012 [cited by applicant]
KR 101770648B1 · 2017 [cited by applicant]
KR 101930796B1 · 2018 [cited by applicant]
WO 2005091894A2 · 2005 [cited by applicant]
WO 2016154306A1 · 2016 [cited by applicant]
WO 2018204279A1 · 2018 [cited by applicant]
WO 2019083832A1 · 2019 [cited by applicant]
WO 2019104049A1 · 2019 [cited by applicant]
WO 2019118599A2 · 2019 [cited by applicant]
WO 2020076880A1 · 2020 [cited by applicant]
Perez-Martin, J., Bustos, B., Guimaraes, S.J.F., Sipiran, I., Pérez, J. and Said, G.C., 2022. A comprehensive review of the video-to-text problem. Artificial Intelligence Review, pp. 1-75. [cited by examiner]
Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y. and Liu, W., 2021. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv … [cited by examiner]
Pintore et al., “State-of-the-Art in Automatic 3D Reconstruction of Structured Indoor Environments”, Eurographics 2020 39:2 pp. 667-699, Jul. 13, 2020, 33 pages. [cited by applicant]
Kara Van Horn, “How to Create a Teaser Video From Matterport”, accessed Aug. 1, 2024 from https://support.homevisit.com/hc/en-us/articles/115005302363-How-to-Create-a-Teaser-Video-from-Matterport, 4 pages. [cited by applicant]
CubiCasa | From video to floor plan in under 5 minutes, retrieved on Mar. 26, 2019, from https://www.cubi.casa/, 6 pages. [cited by applicant]
CubiCasa FAQ & Manual, retrieved on Mar. 26, 2019, from https://www.cubi.casa/faq/, 5 pages. [cited by applicant]
Cupix Home, retrieved on Mar. 26, 2019, from https://www.cupix.com/, 1 page. [cited by applicant]
Cupix—FAQ, retrieved on Mar. 26, 2019, from https://www.cupix.com/faq.html, 3 pages. [cited by applicant]
IGUIDE: 3D Virtual Tours, retrieved on Mar. 26, 2019, from https://goiguide.com/, 6 pages. [cited by applicant]
immoviewer.com | Automated Video Creation & Simple Affordable 3D 360 Tours, retrieved on Mar. 26, 2019, from https://www.immoviewer.com/, 5 pages. [cited by applicant]
MagicPlan | #1 Floor Plan App, Construction & Surveying Samples, retrieved on Mar. 26, 2019, from https://www.magicplan.app/, 9 pages. [cited by applicant]
EyeSpy360 Virtual Tours | Virtual Tour with any 360 camera, retrieved on Mar. 27, 2019, from https://www.eyespy360.com/en-us/, 15 pages. [cited by applicant]
Indoor Reality, retrieved on Mar. 27, 2019, from https://www.indoorreality.com/, 9 pages. [cited by applicant]
InsideMaps, retrieved on Mar. 27, 2019, from https://www.insidemaps.com/, 7 pages. [cited by applicant]
IStaging | Augmented & Virtual Reality Platform For Business, retrieved on Mar. 27, 2019, from https://www.istaging.com/en/, 7 pages. [cited by applicant]
Metareal, retrieved on Mar. 27, 2019, from https://www.metareal.com/, 4 pages. [cited by applicant]
PLNAR—The AR 3D Measuring / Modeling Platform, retrieved on Mar. 27, 2019, from https://www.plnar.co, 6 pages. [cited by applicant]
YouVR Global, retrieved on Mar. 27, 2019, from https://global.youvr.io/, 9 pages. [cited by applicant]
GeoCV, retrieved on Mar. 28, 2019, from https://geocv.com/, 4 pages. [cited by applicant]
Biersdorfer, J.D., “How to Make a 3-D Model of Your Home Renovation Vision,” in The New York Times, Feb. 13, 2019, retrieved Mar. 28, 2019, 6 pages. [cited by applicant]
Chen et al. “Rise of the indoor crowd: Reconstruction of building interior view via mobile crowdsourcing.” In: Proceedings of the 13th ACM Conference on Embedded Networked Sensor Systems. Nov. 4, 2015, 13 pages. [cited by applicant]
Immersive 3D for the Real World, retrieved from https://matterport.com/, on Mar. 27, 2017, 5 pages. [cited by applicant]
Learn About Our Complete 3D System, retrieved from https://matterport.com/how-it-works/, on Mar. 27, 2017, 6 pages. [cited by applicant]
Surefield FAQ, retrieved from https://surefield.com/faq, on Mar. 27, 2017, 1 page. [cited by applicant]
Why Surefield, retrieved from https://surefield.com/why-surefield, on Mar. 27, 2017, 7 pages. [cited by applicant]
Schneider, V., “Create immersive photo experiences with Google Photo Sphere,” retrieved from http://geojoumalism.org/2015/02/create-immersive-photo-experiences-with-google-photo-sphere/, on Mar. 27, 2017, 7 pages. [cited by applicant]
Tango (platform), Wikipedia, retrieved from https://en.wikipedia.org/wiki/Tango_(platform), on Jun. 12, 2018, 6 pages. [cited by applicant]
Zou et al. “LayoutNet: Reconstructing the 3D Room Layout from a Single RGB Image” in arXiv:1803.08999, submitted Mar. 23, 2018, 9 pages. [cited by applicant]
Lee et al. “RoomNet: End-to-End Room Layout Estimation” in arXiv:1703.00241v2, submitted Aug. 7, 2017, 10 pages. [cited by applicant]
Time-of-flight camera, Wikipedia, retrieved from https://en.wikipedia.org/wiki/Time-of-flight_camera, on Aug. 30, 2018, 8 pages. [cited by applicant]
Magicplan—Android Apps on Go . . . , retrieved from https://play.google.com/store/apps/details?id=com.sensopia.magicplan, on Feb. 21, 2018, 5 pages. [cited by applicant]
Pintore et al., “AtlantaNet: Inferring the 3D Indoor Layout from a Single 360 Image beyond the Manhattan World Assumption”, ECCV 2020, 16 pages. [cited by applicant]
Cowles, Jeremy, “Differentiable Rendering”, Aug. 19, 2018, accessed Dec. 7, 2020 at https://towardsdatascience.com/differentiable-rendering-d00a4b0f14be, 3 pages. [cited by applicant]
Yang et al., “DuLa-Net: a Dual-Projection Network for Estimating Room Layouts from a Single RGB Panorama”, in arXiv:1811.11977[cs.v2], submitted Apr. 2, 2019, 14 pages. [cited by applicant]
Sun et al., “HoHoNet: 360 Indoor Holistic Understanding with Latent Horizontal Features”, in arXiv:2011.11498[cs.v2], submitted Nov. 24, 2020, 15 pages. [cited by applicant]
Nguyen-Phuoc et al., “RenderNet: a deep convolutional network for differentiable rendering from 3D shapes”, in arXiv:1806.06575[cs.v3], submitted Apr. 1, 2019, 17 pages. [cited by applicant]
Convolutional neural network, Wikipedia, retrieved from https://en.wikipedia.org/wiki/Convolutional_neural_network, on Dec. 7, 2020, 25 pages. [cited by applicant]
Hamilton et al., “Inductive Representation Learning on Large Graphs”, in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 19 pages. [cited by applicant]
Kipf et al., “Variational Graph Auto-Encoders”, in arXiv:1611.07308v1 [stat.ML], submitted Nov. 21, 2016, 3 pages. [cited by applicant]
Cao et al., “MolGAN: an Implicit Generative Model for Small Molecular Graphs”, in arXiv:1805.11973v1 [stat.ML], submitted May 30, 2018, 11 pages. [cited by applicant]
Chen et al., “Intelligent Home 3D: Automatic 3D-House Design from Linguistic Descriptions Only”, in arXiv:2003.00397v1 [cs.CV], submitted Mar. 1, 2020, 14 pages. [cited by applicant]
Cucurull et al., “Context-Aware Visual Compatibility Prediction”, in arXiv:1902.03646v2 [cs.CV], submitted Feb. 12, 2019, 10 pages. [cited by applicant]
Fan et al., “Labeled Graph Generative Adversarial Networks”, in arXiv:1906.03220v1 [cs.LG], submitted Jun. 7, 2019, 14 pages. [cited by applicant]
Gong et al., “Exploiting Edge Features in Graph Neural Networks”, in arXiv:1809.02709v2 [cs.LG], submitted Jan. 28, 2019, 10 pages. [cited by applicant]
Genghis Goodman, “A Machine Learning Approach to Artificial Floorplan Generation”, University of Kentucky Theses and Dissertations-Computer Science, 2019, accessible at https://uknowledge.uky.edu/cs_etds/89, 40 pages. [cited by applicant]
Grover et al., “node2vec: Scalable Feature Learning for Networks”, in arXiv:1607.00653v1 [cs.SI], submitted Jul. 3, 2016, 10 pages. [cited by applicant]
Nauata et al., “House-GAN: Relational Generative Adversarial Networks for Graph-constrained House Layout Generation”, in arXiv:2003.06988v1 [cs.CV], submitted Mar. 16, 2020, 17 pages. [cited by applicant]
Kang et al., “A Review of Techniques for 3D Reconstruction of Indoor Environments”, in ISPRS International Journal of Geo-Information 2020, May 19, 2020, 31 pages. [cited by applicant]
Kipf et al., “Semi-Supervised Classification With Graph Convolutional Networks”, in arXiv:1609.02907v4 [cs.LG], submitted Feb. 22, 2017, 14 pages. [cited by applicant]
Li et al., “Graph Matching Networks for Learning the Similarity of Graph Structured Objects”, in Proceedings of the 36th International Conference on Machine Learning (PMLR 97), 2019, 18 pages. [cited by applicant]
Liu et al., “Hyperbolic Graph Neural Networks”, in 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 12 pages. [cited by applicant]
Merrell et al., “Computer-Generated Residential Building Layouts”, in ACM Transactions on Graphics, Dec. 2010, 13 pages. [cited by applicant]
Lando Loic, “The 9 Best AI Video Generators (Text-to-Video)”, accessed May 23, 2022 at https://www.makeuseof.com/best-ai-video-generators-text-to-video/, Nov. 19, 2021, 8 pages. [cited by applicant]