IP Library › Granted Patent US 11,430,084
Granted Patent B2
US 11,430,084 · App. 16/121,978 · Granted Aug 30, 2022

Systems and methods for saliency-based sampling layer for neural networks

Inventors: Simon A. I. Stent (Cambridge, MA); Adrià Recasens (Cambridge, MA); Antonio Torralba (Cambridge, MA); Petr Kellnhofer (Cambridge, MA); Wojciech Matusik (Cambridge, MA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; MASSACHUSETTS INSTITUTE OF TECHNOLOGY
G06T3/40G06F3/013G06N3/08G06V10/462
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,430,084
App. No.
16/121,978
Granted
Aug 30, 2022
Kind
B2
Abstract

A method includes receiving, with a computing device, an image, identifying one or more salient features in the image, and generating a saliency map of the image including the one or more salient features. The method further includes sampling the image based on the saliency map such that the one or more salient features are sampled at a first density of sampling and at least one portion of the image other than the one or more salient features are sampled at a second density of sampling, where the first density of sampling is greater than the second density of sampling, and storing the sampled image in a non-transitory computer readable memory.

Claims (77)

1. A method comprising:

receiving, with a computing device, an image;

implementing, with the computing device, a neural network comprising a saliency-based distortion layer trained to:

learn one or more salient features in the image corresponding to a predefined task,

generate a saliency map of the image including the learned one or more salient features, and

sample the image based on the saliency map such that the one or more salient features are sampled at a first density of sampling and at least one portion of the image other than the one or more salient features are sampled at a second density of sampling, wherein the first density of sampling is greater than the second density of sampling, thereby creating a distorted sampled image where more salient features are enlarged compared to less salient features that are less salient than the more salient features; and

storing the distorted sampled image in a non-transitory computer readable memory.

2. The method of claim 1 , further comprising:

generating a sampling grid based on the saliency map, wherein the sampling grid defines a density of sampling for portions of the image; and

sampling the image based on the sampling grid.

3. The method of claim 1 , further comprising:

retrieving the distorted sampled image from the non-transitory computer readable memory; and

inputting the distorted sampled image into a task network for the predefined task, the task network configured to estimate a gaze position of a creature within the image.

4. The method of claim 1 , further comprising:

retrieving the distorted sampled image from the non-transitory computer readable memory; and

inputting the distorted sampled image into a task network for the predefined task, the task network configured to determine a classification of an object within the image.

5. The method of claim 1 , wherein the step of generating the saliency map of the image implements a saliency network trained to learn the one or more salient features for the predefined task of determining a gaze position of a creature within the image.

6. The method of claim 1 , wherein the step of generating the saliency map of the image implements a saliency network trained to learn the one or more salient features for the predefined task of determining a classification of an object within the image.

7. The method of claim 1 , further comprising:

down-sampling the image from an original image; and

storing both the image and the original image in the non-transitory computer readable memory.

8. The method of claim 2 , wherein the step of sampling the image based on the sampling grid further comprises:

retrieving an original image that the image was generated from; and

sampling the original image based on the sampling grid.

9. A computer-implemented system comprising:

a saliency network;

a grid generator;

a sampler; and

a processor and a non-transitory computer readable memory storing computer readable instructions that, when executed by the processor, cause the processor to:

receive an image;

implement a neural network comprising a saliency-based distortion layer trained to:

learn one or more salient features within the image corresponding to a predefined task,

generate, with the saliency network, a saliency map of the image including the learned one or more salient features,

generate, with the grid generator, a sampling grid based on the saliency map, wherein the sampling grid defines a density of sampling for portions of the image such that the one or more salient features are sampled at a first density of sampling and at least one portion of the image other than the one or more salient features are sampled at a second density of sampling, and wherein the first density of sampling is greater than the second density of sampling, and

sample, with the sampler, the image based on the sampling grid, thereby generating a distorted sampled image where more salient features are enlarged compared to less salient features that are less salient than the more salient features; and

store the distorted sampled image in the non-transitory computer readable memory.

10. The computer-implemented system of claim 9 , further comprising a display communicatively coupled to the processor and the non-transitory computer readable memory, wherein the computer readable instructions, when executed by the processor, further causes the processor to:

configure the distorted sampled image for display; and

display the distorted sampled image on the display.

11. The computer-implemented system of claim 9 , wherein the computer readable instructions, when executed by the processor, further causes the processor to:

retrieve the distorted sampled image from the non-transitory computer readable memory;

and

input the distorted sampled image into a task network for the predefined task, the task network configured to estimate a gaze position of a creature within the image.

12. The computer-implemented system of claim 9 , wherein the computer readable instructions, when executed by the processor, further causes the processor to:

retrieve the distorted sampled image from the non-transitory computer readable memory; and

input the distorted sampled image into a task network for the predefined task, the task network configured to determine a classification of an object within the image.

13. The computer-implemented system of claim 9 , wherein the saliency network is trained to learn the one or more salient features for the predefined task of determining a gaze position of a creature within the image.

14. The computer-implemented system of claim 9 , wherein the saliency network is trained to learn the one or more salient features for the predefined task of determining a classification of an object within the image.

15. The computer-implemented system of claim 9 , wherein the computer readable instructions, when executed by the processor, further causes the processor to:

down-sample the image from an original image; and

store both the image and the original image in the non-transitory computer readable memory.

16. The computer-implemented system of claim 9 , wherein when the step of sampling the image, with the sampler, based on the sampling grid is executed by the processor, the computer readable instructions cause the processor to:

retrieve an original image that the image was generated from; and

sample the original image based on the sampling grid.

17. A system comprising:

a camera configured to capture a higher resolution image;

a computing device having a processor and a non-transitory computer readable memory communicatively coupled to the camera; and

a computer readable instruction set that, when executed by the processor, causes the processor to:

receive the higher resolution image from the camera;

down-sample the higher resolution image to a lower resolution image, wherein a resolution of the lower resolution image is lower than a resolution of the higher resolution image;

implement, with the computing device, a neural network comprising a saliency-based distortion layer trained to:

learn one or more salient features within the lower resolution image corresponding to a predefined task,

generate, with a saliency network, a saliency map of the lower resolution image including the learned one or more salient features,

generate, with a grid generator, a sampling grid based on the saliency map, wherein the sampling grid defines a density of sampling for portions of the lower resolution image such that the one or more salient features are sampled at a first density of sampling and at least one portion of the lower resolution image other than the one or more salient features are sampled at a second density of sampling, and wherein the first density of sampling is greater than the second density of sampling, and

sample, with a sampler, the higher resolution image based on the sampling grid, thereby generating a distorted sampled image where more salient features are enlarged compared to less salient features that are less salient than the more salient features; and

store the distorted sampled image in the non-transitory computer readable memory.

18. The system of claim 17 , further comprising a display communicatively coupled to the computing device, wherein the computer readable instruction set, when executed by the processor, further causes the processor to:

configure the distorted sampled image for display; and

display the distorted sampled image on the display.

19. The system of claim 17 , wherein the saliency network is trained to identify the one or more salient features for a task of determining a gaze position of a creature captured by the camera, and wherein the computer readable instruction set, when executed by the processor, further causes the processor to:

retrieve the distorted sampled image from the non-transitory computer readable memory;

and

input the distorted sampled image into a task network for the predefined task, the task network configured to estimate the gaze position of the creature.

20. The system of claim 17 , wherein the saliency network is trained to identify the one or more salient features for a task of determining a classification of an object captured by the camera, and wherein the non-transitory computer readable memory, when executed by the processor, further causes the processor to:

retrieve the distorted sampled image from the non-transitory computer readable memory; and

input the distorted sampled image into a task network for the predefined task, the task network configured to determine the classification of the object.

21. The method of claim 1 , wherein the first density of sampling and the second density of sampling each define a percentage or amount of pixels sampled from a portion of the image during sampling.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2022
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 061472/0408 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2018
From: STENT, SIMON A. I.
To: TOYOTA RESEARCH INSTITUTE, INC.
Reel/Frame 046791/0756 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2018
From: RECASENS, ADRIA; TORRALBA, ANTONIO; KELLNHOFER, PETR; MATUSIK, WOJCIECH
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 046791/0772 →
Continuity (1)
Related Publication 20200074589A1 · Mar 5, 2020
Cited By (1)
US 12,632,941