IP Library Granted Patent US 10,185,895
Granted Patent B1
US 10,185,895 · App. 15/467,546 · Granted Jan 22, 2019

Systems and methods for classifying activities captured within images

Inventors: Daniel Tse (San Mateo, CA); Desmond Chik (San Mateo, CA); Guanhang Wu (Hunan, CN)
Assignee: GoPro, Inc.
G06K9/6267G06K9/00744G06T5/20G06K2009/00738G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,185,895
App. No.
15/467,546
Granted
Jan 22, 2019
Kind
B1
Abstract

An image including a visual capture of a scene may be accessed. The image may be processed through a convolutional neural network. The convolutional neural network may generate a set of two-dimensional feature maps based on the image. The set of two-dimensional feature maps may be processed through a contextual long short-term memory unit. The contextual long short-term memory unit may generate a set of two-dimensional outputs based on the set of two-dimensional feature maps. A set of attention-masks for the image may be generated based on the set of two-dimensional outputs and the set of two-dimensional feature maps. The set of attention-masks may define dimensional portions of the image. The scene may be classified based on the two-dimensional outputs.

Claims (35)

1. A system for classifying activities captured within images, the system comprising:

one or more physical processors configured by machine-readable instructions to:

access an image, the image including a visual capture of a scene;

process the image through a convolutional neural network, the convolutional neural network generating a set of two-dimensional feature maps based on the image;

process the set of two-dimensional feature maps through a contextual long short-term memory unit, the contextual long short-term memory unit generating a set of two-dimensional outputs based on the set of two-dimensional feature maps, wherein the contextual long short-term memory unit includes a loss function characterized by a non-overlapping loss, an entropy loss, and a cross-entropy loss and the non-overlapping loss, the entropy loss, and the cross-entropy loss are combined into the loss function through a linear combination with a first hyper parameter for the non-overlapping loss, a second hyper parameter for the entropy loss, and a third hyper parameter for the cross-entropy loss;

generate a set of attention-masks for the image based on the set of two-dimensional outputs and the set of two-dimensional feature maps, the set of attention-masks defining dimensional portions of the image; and

classify the scene based on the set of two-dimensional outputs.

2. The system of claim 1 , wherein the convolutional neural network includes a plurality of convolution layers, and the set of two-dimensional feature maps is generated by a last convolution layer in the convolutional neural network.

3. The system of claim 2 , wherein the set of two-dimensional feature maps is obtained from the convolutional neural network before the set of two-dimensional feature maps is flattened.

4. The system of claim 1 , wherein the set of two-dimensional outputs is used to visualize the dimensional portions of the image.

5. The system of claim 1 , wherein the set of two-dimensional outputs is used to constrain the dimensional portions of the image.

6. The system of claim 1 , wherein the classification of the scene is performed by a fully connected layer that takes as input the set of two-dimensional outputs.

7. The system of claim 1 , wherein the loss function discourages the set of attention masks defining a same dimensional portion of the image across multiple time-steps.

8. A method for classifying activities captured within images, the method comprising:

accessing an image, the image including a visual capture of a scene;

processing the image through a convolutional neural network, the convolutional neural network generating a set of two-dimensional feature maps based on the image;

processing the set of two-dimensional feature maps through a contextual long short-term memory unit, the contextual long short-term memory unit generating a set of two-dimensional outputs based on the set of two-dimensional feature maps, wherein the contextual long short-term memory unit includes a loss function characterized by a non-overlapping loss, an entropy loss, and a cross-entropy loss and the non-overlapping loss, the entropy loss, and the cross-entropy loss are combined into the loss function through a linear combination with a first hyper parameter for the non-overlapping loss, a second hyper parameter for the entropy loss, and a third hyper parameter for the cross-entropy loss;

generating a set of attention-masks for the image based on the set of two-dimensional outputs and the set of two-dimensional feature maps, the set of attention-masks defining dimensional portions of the image; and

classifying the scene based on the set of two-dimensional outputs.

9. The method of claim 8 , wherein the convolutional neural network includes a plurality of convolution layers, and the set of two-dimensional feature maps is generated by a last convolution layer in the convolutional neural network.

10. The method of claim 9 , wherein the set of two-dimensional feature maps is obtained from the convolutional neural network before the set of two-dimensional feature maps is flattened.

11. The method of claim 8 , wherein the set of two-dimensional outputs is used to visualize the dimensional portions of the image.

12. The method of claim 8 , wherein the set of two-dimensional outputs is used to constrain the dimensional portions of the image.

13. The method of claim 8 , wherein the classification of the scene is performed by a fully connected layer that takes as input the set of two-dimensional outputs.

14. The method of claim 8 , wherein the loss function discourages the set of attention masks defining a same dimensional portion of the image across multiple time-steps.

15. A system for classifying activities captured within images, the system comprising:

one or more physical processors configured by machine-readable instructions to:

access an image, the image including a visual capture of a scene;

process the image through a convolutional neural network, the convolutional neural network generating a set of two-dimensional feature maps based on the image;

process the set of two-dimensional feature maps through a contextual long short-term memory unit, the contextual long short-term memory unit generating a set of two-dimensional outputs based on the set of two-dimensional feature maps, wherein:

the contextual long short-term memory unit includes a loss function characterized by a non-overlapping loss, an entropy loss, and a cross-entropy loss; and

the non-overlapping loss, the entropy loss, and the cross-entropy loss are combined into the loss function through a linear combination with a first hyper parameter for the non-overlapping loss, a second hyper parameter for the entropy loss, and a third hyper parameter for the cross-entropy loss;

generate a set of attention-masks for the image based on the set of two-dimensional outputs and the set of two-dimensional feature maps, the set of attention-masks defining dimensional portions of the image, wherein the loss function discourages the set of attention masks defining a same dimensional portion of the image across multiple time-steps; and

classify the scene based on the set of two-dimensional outputs.

16. The system of claim 15 , wherein the convolutional neural network includes a plurality of convolution layers, and the set of two-dimensional feature maps is generated by a last convolution layer in the convolutional neural network and is obtained from the convolutional neural network before the set of two-dimensional feature maps is flattened.

Assignments (7)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SCHEDULE TO REMOVE APPLICATION 15387383 AND REPLACE WITH 15385383 PREVIOUSLY RECORDED ON REEL 042665 FRAME 0065. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Oct 23, 2019
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 050808/0824 →
SECURITY INTEREST Recorded Mar 5, 2019
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 048508/0728 →
SECURITY INTEREST Recorded Jun 1, 2017
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 042665/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2017
From: TSE, DANIEL; CHIK, DESMOND; WU, GUANHANG
To: GOPRO, INC.
Reel/Frame 041704/0897 →
Cited By (1)
US 12,468,952