IP Library Granted Patent US 12,445,717
Granted Patent B2
US 12,445,717 · App. 17/163,043 · Granted Oct 14, 2025

Techniques for enhanced image capture using a computer-vision network

Inventors: William Castillo (Belmont, CA); Brandon Scott (New York, NY); Alrik Firl (San Francisco, CA); David Royston Cutts (San Francisco, CA); Jonathan Mark Igner (San Francisco, CA); Dario Rethage (Kendall Park, NJ); Domenico Curro (San Francisco, CA); Giridhar Murali (Sunnyvale, CA); Panfeng Li (San Francisco, CA)
Assignee: Hover Inc.
H04N23/64G06T7/11G06T7/12G06T7/174G06T7/277G06T7/74G06T17/00G06V10/26G06V10/44H04N23/635G06F3/167G06T15/00G06T2207/20072G06T2207/20084G06T2210/00G06V30/19013G06V30/19107G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,717
App. No.
17/163,043
Granted
Oct 14, 2025
Kind
B2
Abstract

Disclosed are techniques for enhancing two-dimensional (2D) image capture of subjects (e.g., a physical structure, such as a residential building) to maximize the feature correspondences available for three-dimensional (3D) model reconstruction. More specifically, disclosed is a computer-vision network configured to provide viewfinder interfaces and analyses to guide the improved capture of an intended subject for specified purposes. Additionally, the computer-vision network can be configured to generate a metric representing a quality of feature correspondences between images of a complete set of images used for reconstructing a 3D model of a physical structure. The computer-vision network can also be configured to generate feedback at or before image capture time to guide improvements to the quality of feature correspondences between a pair of images.

Claims (50)

1. A computer-implemented method, comprising:

capturing a set of pixels representing a scene visible to an image capturing device including a display, the set of pixels including a plurality of border pixels, and each border pixel of the plurality of border pixels being located at or within a defined range of a boundary of the set of pixels;

detecting a physical structure depicted within the set of pixels, the physical structure being represented by a subset of the set of pixels;

generating a segmentation mask associated with the physical structure depicted within the set of pixels, the segmentation mask including one or more segmentation pixels, wherein the segmentation mask comprises an irregular shape that conforms to contours of the subset of the set of pixels;

determining a pixel value for each border pixel of the plurality of border pixels;

generating an indicator based on the pixel value of one or more border pixels of the plurality of border pixels by:

determining a number of the one or more border pixels with pixel values indicating overlap with the one or more segmentation pixels; and

generating the indicator based on whether the number of the one or more border pixels with pixel values indicating overlap with the one or more segmentation pixels exceeds a threshold percentage of the plurality of border pixels; and

presenting the indicator, the indicator representing an instruction for framing the physical structure within the display.

2. The computer-implemented method of claim 1 , wherein determining the pixel value for each border pixel further comprises:

detecting that the one or more border pixels of the plurality of border pixels includes a segmentation pixel of the one or more segmentation pixels, and wherein the plurality of border pixels includes:

one or more left edge border pixels located at a left edge of the set of pixels;

one or more or more top edge border pixels located at a top edge of the set of pixels;

one or more right edge border pixels located at a right edge of the set of pixels; and

one or more bottom edge border pixels located at a bottom edge of the set of pixels.

3. The computer-implemented method of claim 2 , wherein:

when a left edge border pixel of the one or more left edge border pixels includes a segmentation pixel, the instruction represented by the indicator instructs a user viewing the display to move the image capturing device in a leftward direction;

when a top edge border pixel of the one or more top edge border pixels includes a segmentation pixel, the instruction represented by the indicator instructs the user viewing the display to move the image capturing device in an upward direction;

when a right edge border pixel of the one or more right edge border pixels includes a segmentation pixel, the instruction represented by the indicator instructs the user viewing the display to move the image capturing device in a rightward direction; and

when a bottom edge border pixel of the one or more bottom edge border pixels includes a segmentation pixel, the instruction represented by the indicator instructs the user viewing the display to move the image capturing device in a downward direction.

4. The computer-implemented method of claim 2 , wherein:

when each of a left edge border pixel, a top edge border pixel, a right edge border pixel, and a bottom edge border pixel includes a segmentation pixel, the instruction represented by the indicator instructs a user viewing the display to move backward.

5. The computer-implemented method of claim 2 , wherein when none of the one or more left edge border pixels, the one or more top edge border pixels, the one or more right edge border pixels, and the one or more bottom edge border pixels includes a segmentation pixel, the instruction represented by the indicator instructs a user viewing the display to zoom in to frame the physical structure.

6. The computer-implemented method of claim 1 , wherein the plurality of border pixels comprises a border having a pixel width determined by a width of the set of pixels.

7. The computer-implemented method of claim 1 , wherein presenting the indicator comprises:

displaying the indicator on the display of the image capturing device; or

audibly presenting the indicator to a user operating the image capturing device.

8. The computer-implemented method of claim 1 , wherein the border pixels comprise fields storing image information and a field to store the pixel value indicating whether a border pixel intersects with the segmentation mask.

9. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause a data processing apparatus to perform operation including:

capturing a set of pixels representing a scene visible to an image capturing device including a display, the set of pixels including a plurality of border pixels, and each border pixel of the plurality of border pixels being located at or within a defined range of a boundary of the set of pixels;

detecting a physical structure depicted within the set of pixels, the physical structure being represented by a subset of the set of pixels;

generating a segmentation mask associated with the physical structure depicted within the set of pixels, the segmentation mask including one or more segmentation pixels, wherein the segmentation mask comprises an irregular shape that conforms to contours of the subset of the set of pixels;

determining a pixel value for each border pixel of the plurality of border pixels;

generating an indicator based on the pixel value of one or more border pixels of the plurality of border pixels by:

determining a number of the one or more border pixels with pixel values indicating overlap with the one or more segmentation pixels; and

generating the indicator based on whether the number of the one or more border pixels with pixel values indicating overlap with the one or more segmentation pixels exceeds a threshold percentage of the plurality of border pixels; and

presenting the indicator, the indicator representing an instruction for framing the physical structure within the display.

10. The computer-program product of claim 9 , wherein the threshold percentage is a function of a related pixel dimension of the segmentation mask.

11. The computer-program product of claim 9 , wherein the threshold percentage comprises a predetermined number of consecutive border pixels in the plurality of border pixels that overlap with the one or more segmentation pixels.

12. The computer-program product of claim 9 , wherein the defined range of the boundary of the set of pixels comprises between 2 and 10 pixels from the boundary of the set of pixels, such that the plurality of border pixels have a width of between 2 and 10 pixels around the boundary of the set of pixels.

13. The computer-program product of claim 9 , wherein the segmentation mask is generated by a classifier trained to identify physical structures in sets of pixels.

14. The computer-program product of claim 9 , further comprising:

generating a bounding box around the segmentation mask, such that all of the one or more segmentation pixels fit inside of the bounding box, and the bounding box includes a buffer region such that the bounding box does not tangentially touch any of the one or more segmentation pixels; and

presenting the bounding box and the set of pixels with the indicator.

15. The computer-program product of claim 9 , further comprising:

denoising the segmentation mask by smoothing a boundary of the segmentation mask.

16. The computer-program product of claim 15 , wherein

denoising the segmentation mask comprises combining a plurality of segmentation masks from a plurality of frames of the scene that include the physical structure.

17. The computer-program product of claim 16 , wherein combining the plurality of segmentation masks from the plurality of frames comprises pixel voting using pixels from the plurality of frames.

18. The computer-program product of claim 17 , wherein the pixel voting is weighted based on a motion of the image capturing device when capturing each of the plurality of frames.

Assignments (7)
TERMINATION OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 6, 2022
From: SILICON VALLEY BANK
To: HOVER INC.
Reel/Frame 061622/0741 →
TERMINATION OF INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 6, 2022
From: SILICON VALLEY BANK, AS AGENT
To: HOVER INC.
Reel/Frame 061622/0761 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2022
From: LI, PANFENG
To: HOVER INC.
Reel/Frame 061273/0846 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: MURALI, GIRIDHAR
To: HOVER INC.
Reel/Frame 058724/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2021
From: CASTILLO, WILLIAM; SCOTT, BRANDON; FIRL, ALRIK; CUTTS, DAVID ROYSTON; IGNER, JONATHAN MARK; RETHAGE, DARIO; CURRO, DOMENICO
To: HOVER INC.
Reel/Frame 056640/0947 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 28, 2021
From: HOVER INC.
To: SILICON VALLEY BANK
Reel/Frame 056423/0199 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 28, 2021
From: HOVER INC.
To: SILICON VALLEY BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 056423/0222 →
Continuity (4)
Provisional Application 63140716 · Jan 22, 2021
Provisional Application 63059093 · Jul 30, 2020
Provisional Application 62968977 · Jan 31, 2020
Related Publication 20210243362A1 · Aug 5, 2021
References Cited (45)
US 9159164B2 · Ciarcia · 2015 [cited by applicant]
US 9531957B1 · George et al. · 2016 [cited by applicant]
US 9716826B2 · Wu et al. · 2017 [cited by applicant]
US 9721177B2 · Lee et al. · 2017 [cited by applicant]
US 9826145B2 · Malgimani et al. · 2017 [cited by applicant]
US 10038838B2 · Castillo et al. · 2018 [cited by applicant]
US 10382673B2 · Upendran et al. · 2019 [cited by applicant]
US 20080089577A1 · Wang · 2008 [cited by applicant]
US 20110157406A1 · Tauchi · 2011 [cited by applicant]
US 20110285874A1 · Showering · 2011 [cited by examiner]
US 20120105647A1 · Yoshizumi · 2012 [cited by examiner]
US 20130129156A1 · Wang · 2013 [cited by examiner]
US 20170094184A1 · Gao · 2017 [cited by examiner]
US 20170208245A1 · Castillo et al. · 2017 [cited by applicant]
US 20170310884A1 · Li · 2017 [cited by examiner]
US 20180198976A1 · Upendran · 2018 [cited by examiner]
US 20180198978A1 · Cho et al. · 2018 [cited by applicant]
US 20180220066A1 · Kitamura · 2018 [cited by examiner]
US 20180300855A1 · Tang et al. · 2018 [cited by applicant]
US 20190174056A1 · Jung et al. · 2019 [cited by applicant]
US 20190236394A1 · Price · 2019 [cited by examiner]
US 20190311202A1 · Lee et al. · 2019 [cited by applicant]
US 20190385026A1 · Richeimer · 2019 [cited by examiner]
US 20200202533A1 · Cohen et al. · 2020 [cited by applicant]
US 20210004648A1 · Ghosh et al. · 2021 [cited by applicant]
EP 1981268B1 · 2017 [cited by applicant]
Badrinarayanan et al., “SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation”, Dec. 2017, IEEE, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, No. 12, p. 2481-2495. … [cited by examiner]
PCT/US2021/015850, “International Search Report and Written Opinion”, Jul. 12, 2021, 21 pages. [cited by applicant]
Hegarty, et al., “SAMATS—Triangle Grouping and Structure Recovery for 3D Building Modeling and Visualization”, Technological University Dublin, Digital Media Centre, Jan. 12, 2005, 15 pages. [cited by applicant]
Wirtz, et al., “Semiautomatic Generation of Semantic Building Models From Image Series”, Proceedings of SPIE—The International Society for Optical Engineering, Feb. 2011, 8 pages. [cited by applicant]
PCT/US2021/015850, “Invitation to Pay Additional Fees and, Where Applicable, Protest Fee”, May 19, 2021, 5 pages. [cited by applicant]
PCT/US2021/015850, “International Preliminary Report on Patentability”, Aug. 11, 2022, 16 pages. [cited by applicant]
U.S. Appl. No. 17/163,105, “Notice of Allowance”, Jul. 26, 2023, 10 pages. [cited by applicant]
Bischke, et al., “Multi-Task Learning for Segmentation of Building Footprints with Deep Neural Networks”, IEEE, 2019, pp. 1480-1484. [cited by applicant]
Brostow, et al., “Semantic Object Classes in Video: A High-Definition Ground Truth Database”, Pattern Recognition Letters, vol. 30, No. 2, Jan. 15, 2009, pp. 88-97. [cited by applicant]
CA3163137, “Office Action”, Aug. 8, 2023, 4 pages. [cited by applicant]
Huang, et al., “Building Extraction from Multi-Source Remote Sensing Images Via Deep Deconvolution Neural Networks”, IEEE, Jul. 2016, pp. 1835-1838. [cited by applicant]
Kisantal, et al., “Augmentation for Small Object Detection”, Available Online at: https://arxiv.org/pdf/1902.07296.pdf, Feb. 19, 2019, pp. 1-15. [cited by applicant]
Wu, et al., “Automatic Building Segmentation of Aerial Imagery using Multi-Constraint Fully Convolutional Networks”, MDPI, Remote Sens., vol. 10, No. 3, Mar. 6, 2018, pp. 1-18. [cited by applicant]
Zhou, et al., “Building Segmentation from Airborne Vhr Images using Mask R-Cnn”, ISPRS, Jun. 2019, pp. 155-161. [cited by applicant]
U.S. Appl. No. 18/495,336, “Notice of Allowance”, May 30, 2024, 11 pages. [cited by applicant]
CA3163137, “Office Action”, May 15, 2024, 4 pages. [cited by applicant]
Tarantino, et al., “Extracting Buildings from True Color Stereo Aerial Images Using a Decision Making Strategy”, Remote Sensing, vol. 3, Issue 8, Jul. 25, 2011, pp. 1553-1567. [cited by applicant]
Wen, et al., “Automatic Building Extraction from Google Earth Images under Complex Backgrounds Based on Deep Instance Segmentation Network”, Sensors, vol. 19, No. 2, Jan. 15, 2019, pp. 1-16. [cited by applicant]
EP24188777.7, “Extended European Search Report”, Oct. 2, 2024, 10 pages. [cited by applicant]