IP Library Granted Patent US 11,734,827
Granted Patent B2
US 11,734,827 · App. 17/317,755 · Granted Aug 22, 2023

User guided iterative frame and scene segmentation via network overtraining

Inventor: Gary Bradski (Palo Alto, CA)
Assignee: Matterport, Inc.
G06T7/11G06F3/04883G06N3/08G06T7/187G06T11/203
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,734,827
App. No.
17/317,755
Granted
Aug 22, 2023
Kind
B2
Abstract

Systems and methods for user guided iterative frame and scene segmentation are disclosed herein. The systems and methods can rely on overtraining a segmentation network on a frame. A disclosed method includes selecting a frame from a scene and generating a frame segmentation using the frame and a segmentation network. The method also includes displaying the frame and frame segmentation overlain on the frame, receiving a correction input on the frame, and training the segmentation network using the correction input. The method includes overtraining the segmentation network for the scene by iterating the above steps on the same frame or a series of frames from the scene.

Claims (125)

1. A computer-implemented method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input directed to the frame;

training the segmentation network using the correction input;

iterating through the generating segmentation step, displaying segmentation step, receiving correction input step, and training the segmentation step until a number of iterations reaches a predetermined threshold, the predetermined threshold being determined based on a statistical variation within the scene;

generating, after training the segmentation network using the correction input, a revised frame segmentation using: (i) the frame; and (ii) the segmentation network; and

displaying on the device: (i) the frame; and (ii) the revised frame segmentation overlaid on the frame.

2. The computer-implemented method of claim 1 , further comprising:

scanning the scene for a set of statistical variation points using a frame selector,

wherein the selecting the frame step is conducted: (i) using the set of statistical variation points; and (ii) so that the frame is from a frame located between two statistical variation points in the set of statistical variation points.

3. The computer-implemented method of claim 1 , further comprising:

dilating the first frame segmentation to form an outer boundary of a tri-map;

eroding the first frame segmentation to form an inner boundary of the tri-map; and

presenting a region of the frame for receiving the correction input using the tri-map.

4. The computer-implemented method of claim 1 , wherein:

training the segmentation network using the correction input includes tagging a portion of the frame identified by the correction input with a tag for the segmentation target; and

the portion of the frame and the tag are used as a supervisor in a training routine for the segmentation network.

5. The computer-implemented method of claim 1 , further comprising:

selecting a second frame from the scene;

generating a second frame segmentation using: (i) the second frame; and (ii) the segmentation network;

displaying on the device: (i) the second frame; and (ii) the second frame segmentation overlain on the second frame;

receiving a frame skip input; and

selecting a third frame from the scene, wherein the selecting the third frame step is conducted as a response to the receiving the frame skip input.

6. The computer-implemented method of claim 1 , further comprising:

displaying a prompt to select a segmentation target, wherein the correction input is provided in response to the prompt, wherein training the segmentation using the correction input includes tagging a portion of the frame identified by the correction input with a tag for the segmentation target, and wherein the portion of the frame and the tag are used as a supervisor in a training routine for the segmentation network.

7. A computer-implemented method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input; and

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene, wherein the training step, the generating the revised frame segmentation step, and the displaying the frame and revised frame segmentation step are all conducted as a single response to the receiving the correction input step until a number of iterations reaches a predetermined threshold, the predetermined threshold being determined based on a statistical variation within the scene.

8. The computer-implemented method of claim 7 , further comprising:

generating, after training the segmentation network using the correction input, a second frame segmentation using: (i) a second frame; and (ii) the segmentation network, wherein the second frame is from the scene, and wherein the generating the second frame segmentation step is conducted as part of the single response.

9. The computer-implemented method of claim 7 , further comprising:

displaying a prompt to select a segmentation target, wherein the correction input is provided in response to the prompt, wherein training the segmentation using the correction input includes tagging a portion of the frame identified by the correction input with a tag for the segmentation target, and wherein the portion of the frame and the tag are used as a supervisor in a training routine for the segmentation network.

10. A computer-implemented method comprising:

selecting a frame from a scene;

scanning the scene for a set of statistical variation points, wherein the selecting the frame step is conducted: (i) using the set of statistical variation points; and (ii) so that the frame is from a frame located between two statistical variation points in the set of statistical variation points;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input; and

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene.

11. A computer-implemented method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input;

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene;

dilating the first frame segmentation to form an outer boundary of a tri-map;

eroding the first frame segmentation to form an inner boundary of the tri-map; and

presenting a region of the frame for receiving the correction input using the tri-map.

12. A computer-implemented method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input;

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene; and

receiving a frame skip input instead of an additional correction input during an iteration of the training the segmentation network for the scene step, wherein the training the segmentation network using the correction input step is skipped during the iteration.

13. A device comprising:

a display;

a frame selector instantiated on the device, wherein the frame selector is programmed to select a frame from a scene;

a segmentation editor instantiated on the device, wherein the segmentation editor is programmed to, in response to the frame selector selecting the frame, display on the display: (i) the frame; and (ii) a frame segmentation overlaid on the frame;

a correction interface configured to receive a correction input directed to the frame;

wherein the device is programmed to:

provide the correction input to a trainer for a segmentation network;

receive a revised frame segmentation from the segmentation network after the trainer has applied the correction input to the segmentation network; and

display the revised frame segmentation overlaid on the frame; and

iterate through the receive the correction input, and training the corrected input until a number of iterations reaches a predetermined threshold, the predetermined threshold being determined based on a statistical variation within the scene.

14. The device of claim 13 , wherein the frame selector is programmed to:

scan the scene for a set of statistical variation points, wherein the selecting the frame step is conducted: (i) using the set of statistical variation points; and (ii) so that the frame is from a frame located between two statistical variation points in the set of statistical variation points.

15. The device of claim 13 , wherein:

the correction interface includes a prompt to select a segmentation target;

the correction input is provided in response to the prompt;

the trainer is a supervised learning trainer; and

a portion of the frame identified by the correction input and a segmentation target type form a supervisor for the trainer.

16. A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method, the method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input directed to the frame;

training the segmentation network using the correction input;

iterating through the generating segmentation step, displaying segmentation step, receiving correction input step, and training the segmentation step until a number of iterations reaches a predetermined threshold, the predetermined threshold being determined based on a statistical variation within the scene;

generating, after training the segmentation network using the correction input, a revised frame segmentation using: (i) the frame; and (ii) the segmentation network; and

displaying on the device: (i) the frame; and (ii) the revised frame segmentation overlaid on the frame.

17. The non-transitory computer-readable medium of claim 16 , the executable instructions that are executable by the one or more processors to further:

scanning the scene for a set of statistical variation points using a frame selector, wherein the selecting the frame step is conducted: (i) using the set of statistical variation points; and (ii) so that the frame is from a frame located between two statistical variation points in the set of statistical variation points.

18. The non-transitory computer-readable medium of claim 16 , the executable instructions that are executable by the one or more processors to further:

dilating the first frame segmentation to form an outer boundary of a tri-map;

eroding the first frame segmentation to form an inner boundary of the tri-map; and

presenting a region of the frame for receiving the correction input using the tri-map.

19. The non-transitory computer-readable medium of claim 16 , the executable instructions that are executable by the one or more processors to further:

training the segmentation network using the correction input includes tagging a portion of the frame identified by the correction input with a tag for the segmentation target; and

the portion of the frame and the tag are used as a supervisor in a training routine for the segmentation network.

20. A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method, the method comprising:

selecting a frame from a scene;

scanning the scene for a set of statistical variation points, wherein the selecting the frame step is conducted: (i) using the set of statistical variation points; and (ii) so that the frame is from a frame located between two statistical variation points in the set of statistical variation points;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input; and

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene.

21. A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method, the method comprising:

selecting a frame from a scene;

dilating the first frame segmentation to form an outer boundary of a tri-map;

eroding the first frame segmentation to form an inner boundary of the tri-map;

presenting a region of the frame for receiving the correction input using the tri-map;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input; and

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene.

22. A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method, the method comprising:

selecting a frame from a scene;

generating a first frame segmentation using: (i) the frame; and (ii) a segmentation network;

displaying on a device: (i) the frame; and (ii) the first frame segmentation overlaid on the frame;

receiving a correction input on the frame;

training the segmentation network using the correction input;

training the segmentation network for the scene by iterating the selecting, generating, displaying, receiving, and training steps using one of: the same frame; or a series of frames from the scene; and

receiving a frame skip input instead of an additional correction input during an iteration of the training the segmentation network for the scene step, wherein the training the segmentation network using the correction input step is skipped during the iteration.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2025
From: MATTERPORT, LLC
To: COSTAR REALTY INFORMATION, INC.
Reel/Frame 072938/0375 →
MERGER AND CHANGE OF NAME Recorded Sep 10, 2025
From: MATTERPORT, INC.; MATRIX MERGER SUB II LLC
To: MATTERPORT, LLC
Reel/Frame 072827/0337 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNORS PREVIOUSLY RECORDED AT REEL: 057679 FRAME: 0638. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: BRADSKI, GARY
To: ARRAIY, INC.
Reel/Frame 064159/0421 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2021
From: BRADSKI, GARY; TOBIN, MARK
To: ARRAIY, INC.
Reel/Frame 057679/0638 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2021
From: ARRAIY, INC.
To: MATTERPORT, INC.
Reel/Frame 057688/0932 →
Continuity (2)
Continuation 16411739 · May 14, 2019
Related Publication 20210264609A1 · Aug 26, 2021
Cited By (2)
US 12,424,002 US 12,456,100