IP Library Granted Patent US 10,565,729
Granted Patent B2
US 10,565,729 · App. 15/971,997 · Granted Feb 18, 2020

Optimizations for dynamic object instance detection, segmentation, and structure mapping

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,565,729
App. No.
15/971,997
Granted
Feb 18, 2020
Kind
B2
Abstract

In one embodiment, a method includes a system accessing an image and generating a feature map using a first neural network. The system identifies a plurality of regions of interest in the feature map. A plurality of regional feature maps may be generated for the plurality of regions of interest, respectively. Using a second neural network, the system may detect at least one regional feature map in the plurality of regional feature maps that corresponds to a person depicted in the image, and generate a target region definition associated with a location of the person using the regional feature map. Based on the target region definition associated with the location of the person, a target regional feature map may be generated by sampling the feature map for the image. The system may process the target regional feature map to generate a keypoint mask and an instance segmentation mask.

Claims (53)

1. A method comprising, by a computing system:

accessing an image;

generating a feature map for the image using a first neural network;

identifying a plurality of regions of interest in the feature map;

generating a plurality of regional feature maps for the plurality of regions of interest, respectively, by sampling the feature map for the image;

processing the plurality of regional feature maps using a second neural network to:

detect at least one regional feature map in the plurality of regional feature maps that corresponds to a person depicted in the image; and

generate a target region definition associated with a location of the person using the regional feature map;

generating, based on the target region definition associated with the location of the person, a target regional feature map by sampling the feature map for the image; and

generating:

a keypoint mask associated with the person by processing the target regional feature map using a third neural network; or

an instance segmentation mask associated with the person by processing the target regional feature map using a fourth neural network.

2. The method of claim 1 , wherein the instance segmentation mask and keypoint mask are both generated and are being generated concurrently.

3. The method of claim 1 , wherein the first neural network comprises four or fewer convolutional layers.

4. The method of claim 3 , wherein each of the convolutional layers uses a kernel size of 3×3 or less.

5. The method of claim 1 , wherein the first neural network comprises a total of one pooling layer.

6. The method of claim 1 , wherein the first neural network comprises three or fewer inception modules.

7. The method of claim 6 , wherein each of the inception modules performs convolutional operations with kernel sizes of 5×5 or less.

8. The method of claim 1 , wherein each of the second neural network, third neural network, and fourth neural network is configured to process an input regional feature map using a total of one inception module.

9. A system comprising: one or more processors and one or more computer-readable non-transitory storage media coupled to one or more of the processors, the one or more computer-readable non-transitory storage media comprising instructions operable when executed by one or more of the processors to cause the system to perform operations comprising:

accessing an image;

generating a feature map for the image using a first neural network;

identifying a plurality of regions of interest in the feature map;

generating a plurality of regional feature maps for the plurality of regions of interest, respectively, by sampling the feature map for the image;

processing the plurality of regional feature maps using a second neural network to:

detect at least one regional feature map in the plurality of regional feature maps that corresponds to a person depicted in the image; and

generate a target region definition associated with a location of the person using the regional feature map;

generating, based on the target region definition associated with the location of the person, a target regional feature map by sampling the feature map for the image; and

generating:

a keypoint mask associated with the person by processing the target regional feature map using a third neural network; or

an instance segmentation mask associated with the person by processing the target regional feature map using a fourth neural network.

10. The system of claim 9 , wherein the instance segmentation mask and keypoint mask are both generated and are being generated concurrently.

11. The system of claim 9 , wherein the first neural network comprises four or fewer convolutional layers.

12. The system of claim 11 , wherein each of the convolutional layers uses a kernel size of 3×3 or less.

13. The system of claim 9 , wherein the first neural network comprises a total of one pooling layer.

14. The system of claim 9 , wherein the first neural network comprises three or fewer inception modules.

15. One or more computer-readable non-transitory storage media embodying software that is operable when executed to cause one or more processors to perform operations comprising:

accessing an image;

generating a feature map for the image using a first neural network;

identifying a plurality of regions of interest in the feature map;

generating a plurality of regional feature maps for the plurality of regions of interest, respectively, by sampling the feature map for the image;

processing the plurality of regional feature maps using a second neural network to:

detect at least one regional feature map in the plurality of regional feature maps that corresponds to a person depicted in the image; and

generate a target region definition associated with a location of the person using the regional feature map;

generating, based on the target region definition associated with the location of the person, a target regional feature map by sampling the feature map for the image; and

generating:

a keypoint mask associated with the person by processing the target regional feature map using a third neural network; or

an instance segmentation mask associated with the person by processing the target regional feature map using a fourth neural network.

16. The media of claim 15 , wherein the instance segmentation mask and keypoint mask are both generated and are being generated concurrently.

17. The media of claim 15 , wherein the first neural network comprises four or fewer convolutional layers.

18. The media of claim 17 , wherein each of the convolutional layers uses a kernel size of 3×3 or less.

19. The media of claim 15 , wherein the first neural network comprises a total of one pooling layer.

20. The media of claim 15 , wherein the first neural network comprises three or fewer inception modules.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2018
From: VAJDA, PETER; ZHANG, PEIZHAO; YANG, FEI; WANG, YANGHAN
To: FACEBOOK, INC.
Reel/Frame 046359/0609 →