IP Library Granted Patent US 11,080,861
Granted Patent B2
US 11,080,861 · App. 16/412,031 · Granted Aug 3, 2021

Scene segmentation using model subtraction

Inventors: Gary Bradski (Palo Alto, CA); Ethan Rublee (Mountain View, CA)
Assignee: Matterport, Inc.
G06T7/194G06K9/6257G06T7/155G06T7/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,080,861
App. No.
16/412,031
Granted
Aug 3, 2021
Kind
B2
Abstract

Systems and methods for frame and scene segmentation are disclosed herein. A disclosed method includes providing a frame of a scene. The scene includes a scene background. The method also includes providing a model of the scene background. The method also includes determining a frame background using the model and subtracting the frame background from the frame to obtain an approximate segmentation. The method also includes training a segmentation network using the approximate segmentation.

Claims (56)

1. A computer-implemented method comprising:

providing a frame of a scene, the scene including a scene background, the scene background being a chroma screen;

providing a model of the scene background based on fitting an intensity plane to the chroma screen in the frame;

determining a frame background using the model based on mixing a chroma of the chroma screen with the intensity plane;

subtracting the frame background from the frame to obtain an approximate segmentation; and

training a segmentation network using the approximate segmentation.

2. The computer-implemented method of claim 1 , further comprising:

segmenting, after training the segmentation network, a second frame from the scene using the segmentation network.

3. The computer-implemented method of claim 1 , wherein training the segmentation network using the approximate segmentation comprises:

tagging a first portion of the frame, included in the approximate segmentation, with a subject tag;

tagging a second portion of the frame, excluded from the approximate segmentation, with a background tag;

generating a segmentation inference from the segmentation network using the frame; and

evaluating the segmentation inference with at least one of the subject tag and the background tag.

4. The computer-implemented method of claim 3 , wherein training the segmentation network using the approximate segmentation comprises:

morphologically dilating the approximate segmentation to define the second portion of the frame; and

morphologically eroding the approximate segmentation to define the first portion of the frame.

5. The computer-implemented method of claim 1 , wherein:

providing the model of the scene background comprises capturing a three-dimensional model of the scene background; and

determining the frame background from the model comprises generating the frame background using the three-dimensional model and a camera pose of the frame.

6. The computer-implemented method of claim 1 , wherein determining the frame background from the model comprises:

registering the model with a set of trackable tags in the scene background;

deriving a camera pose for the frame using the set of trackable tags; and

deriving the frame background using the model and the camera pose of the frame.

7. The computer-implemented method of claim 1 , wherein:

providing the model of the scene background comprises: (i) capturing a clean plate scene background; (ii) selecting a set of key frames from the clean plate scene background; and (iii) training, using the clean plate scene background, an optical flow network; and

determining the frame background from the model comprises inferring the frame background using the optical flow network.

8. The computer-implemented method of claim 7 , wherein:

training the optical flow network includes synthesizing occlusions on the clean plate scene background.

9. A non-transitory computer-readable medium storing instructions to execute a computer-implemented method, the method comprising:

providing a frame of a scene, the scene including a scene background, the scene background being a chroma screen;

providing a model of the scene background based on fitting an intensity plane to the chroma screen in the frame;

determining a frame background from the model based on mixing a chroma of the chroma screen with the intensity plane;

subtracting the frame background from the frame to obtain a frame foreground; and

training a segmentation network using the frame foreground.

10. The non-transitory computer-readable medium of claim 9 , the method further comprising:

segmenting, after training the segmentation network, a second frame from the scene using the segmentation network.

11. The non-transitory computer-readable medium of claim 9 , wherein training the segmentation network using the frame foreground comprises:

tagging a first portion of the frame, included in the frame foreground, with a subject tag;

tagging a second portion of the frame, excluded from the frame foreground, with a background tag;

generating a segmentation inference from the segmentation network using the frame; and

evaluating the segmentation inference with at least one of the subject tag and the background tag.

12. The non-transitory computer-readable medium of claim 11 , wherein training the segmentation network using the frame foreground comprises:

morphologically dilating the frame foreground to define the second portion of the frame; and

morphologically eroding the frame foreground to define the first portion of the frame.

13. The non-transitory computer-readable medium of claim 9 , wherein:

providing the model of the scene background comprises capturing a three-dimensional model of the scene background; and

determining the frame background from the model comprises generating the frame background using the three-dimensional model and a camera pose of the frame.

14. The non-transitory computer-readable medium of claim 9 , wherein determining the frame background from the model comprises:

registering the model with a set of trackable tags in the scene background;

deriving a camera pose for the frame using the set of trackable tags; and

deriving the frame background using the model and the camera pose of the frame.

15. The non-transitory computer-readable medium of claim 9 , wherein:

providing the model of the scene background comprises: (i) capturing a clean plate scene background; (ii) selecting a set of key frames from the clean plate scene background; and (iii) training, using the clean plate scene background, an optical flow network; and

determining the frame background from the model comprises inferring the frame background using the optical flow network.

16. The non-transitory computer-readable medium of claim 15 , wherein:

training the optical flow network includes synthesizing occlusions on the clean plate scene background.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2025
From: MATTERPORT, LLC
To: COSTAR REALTY INFORMATION, INC.
Reel/Frame 072938/0375 →
MERGER AND CHANGE OF NAME Recorded Sep 10, 2025
From: MATTERPORT, INC.; MATRIX MERGER SUB II LLC
To: MATTERPORT, LLC
Reel/Frame 072827/0337 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: ARRAIY, INC.
To: MATTERPORT, INC.
Reel/Frame 051396/0020 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2019
From: BRADSKI, GARY; RUBLEE, ETHAN
To: ARRAIY, INC.
Reel/Frame 050352/0509 →
Continuity (1)
Related Publication 20200364877A1 · Nov 19, 2020
Cited By (1)
US 12,217,311