IP Library Granted Patent US 10,496,895
Granted Patent B2
US 10,496,895 · App. 15/853,290 · Granted Dec 3, 2019

Generating refined object proposals using deep-learning models

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,496,895
App. No.
15/853,290
Granted
Dec 3, 2019
Kind
B2
Abstract

In one embodiment a plurality of patches of an image are processed, using a first set of layers of a convolutional neural network, to output a plurality of object proposals associated with the plurality of patches of the image. Each patch includes one or more pixels of the image. Each object proposal includes a prediction as to a location of an object in the respective patch. Using a second set of layers of the convolutional neural network, the plurality of object proposals outputted by the first set of layers are processed to generate a plurality of refined object proposals. Each refined object proposal includes pixel-level information for the respective patch of the image. The first layer in the second set of layers of the convolutional neural network takes as input the plurality of object proposals outputted by the first set of layers. Each layer after the first layer in the second set of layers takes as input the output of a preceding layer in the second set of layers combined with the output of a respective layer of the first set of layers.

Claims (31)

1. A method comprising, for each patch of a plurality of patches of an image:

processing, by one or more computing devices, using a first set of layers of a convolutional neural network, the patch to output one or more coarse object proposals associated with the patch, wherein the patch comprises one or more pixels of the image, and wherein each coarse object proposal of the one or more coarse object proposals comprises a prediction as to a location of an object in the patch; and

processing, by one or more of the computing devices, using a second set of layers of the convolutional neural network, the one or more coarse object proposals outputted by the first set of layers of the convolutional neural network, to generate one or more refined object proposals, each refined object proposal of the one or more refined object proposals comprising pixel-level information for the patch, wherein:

a first layer in the second set of layers of the convolutional neural network takes as input the one or more coarse object proposals outputted by the first set of layers, and

each layer after the first layer in the second set of layers takes as input an output of a preceding layer in the second set of layers combined with a respective output of a corresponding layer of the first set of layers.

2. The method of claim 1 , wherein each layer in the first set of layers corresponds to a respective layer in the second set of layers.

3. The method of claim 1 , wherein the one or more coarse object proposals comprise object-level information.

4. The method of claim 1 , wherein the pixel-level information for the patch comprises indications of edges of at least one object in the patch.

5. The method of claim 1 , further comprising generating identifications of one of more objects in the image based on a subset of coarse object proposals.

6. The method of claim 1 , further comprising combining, during a refinement stage, the output of the preceding layer in the second set of layers with the respective output of the corresponding layer in the first set of layers.

7. The method of claim 1 , wherein each refined object proposal of the one or more refined object proposals and the image are of a same resolution.

8. One or more computer-readable non-transitory storage media embodying software that is operable when executed to, for each patch of a plurality of patches of an image:

process, using a first set of layers of a convolutional neural network, the patch to output one or more coarse object proposals associated with the patch, wherein the patch comprises one or more pixels of the image, and wherein each coarse object proposal of the one or more coarse object proposals comprises a prediction as to a location of an object in the patch; and

process, using a second set of layers of the convolutional neural network, the one or more coarse object proposals outputted by the first set of layers of the convolutional neural network, to generate one or more refined object proposals, each refined object proposal of the one or more refined object proposals comprising pixel-level information for the patch, wherein:

a first layer in the second set of layers of the convolutional neural network takes as input the one or more coarse object proposals outputted by the first set of layers, and

each layer after the first layer in the second set of layers takes as input an output of a preceding layer in the second set of layers combined with a respective output of a corresponding layer of the first set of layers.

9. The media of claim 8 , wherein each layer in the first set of layers corresponds to a respective layer in the second set of layers.

10. The media of claim 8 , wherein the one or more coarse object proposals comprise object-level information.

11. The media of claim 8 , wherein the pixel-level information for the patch comprises indications of edges of at least one object in the patch.

12. The media of claim 8 , wherein the software is further operable when executed to generate identifications of one of more objects in the image based on a subset of coarse object proposals.

13. The media of claim 8 , wherein the software is further operable when executed to combine, during a refinement stage, the output of the preceding layer in the second set of layers with the respective output of the corresponding layer in the first set of layers.

14. A system comprising: one or more processors; and a memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to, for each patch of a plurality of patches of an image:

process, using a first set of layers of a convolutional neural network, the patch to output one or more coarse object proposals associated with the patch, wherein the patch comprises one or more pixels of the image, and wherein each coarse object proposal of the one or more coarse object proposals comprises a prediction as to a location of an object in the patch; and

process, using a second set of layers of the convolutional neural network, the one or more coarse object proposals outputted by the first set of layers of the convolutional neural network, to generate one or more refined object proposals, each refined object proposal of the one or more refined object proposals comprising pixel-level information for the patch, wherein:

a first layer in the second set of layers of the convolutional neural network takes as input the one or more coarse object proposals outputted by the first set of layers, and

each layer after the first layer in the second set of layers takes as input an output of a preceding layer in the second set of layers combined with a respective output of a corresponding layer of the first set of layers.

15. The system of claim 14 , wherein each layer in the first set of layers corresponds to a respective layer in the second set of layers.

16. The system of claim 14 , wherein the one or more coarse object proposals comprise object-level information.

17. The system of claim 14 , wherein the pixel-level information for the patch comprises indications of edges of at least one object in the patch.

18. The system of claim 14 , wherein the processors are further operable when executing the instructions to generate identifications of one of more objects in the image based on a subset of coarse object proposals.

19. The system of claim 14 , wherein the processors are further operable when executing the instructions to combine, during a refinement stage, the output of the preceding layer in the second set of layers with the respective output of a corresponding layer in the first set of layers.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2019
From: PINHEIRO, PEDRO HENRIQUE OLIVEIRA; COLLOBERT, RONAN STÉFAN; DOLLAR, PIOTR
To: FACEBOOK, INC.
Reel/Frame 050324/0685 →