IP Library Granted Patent US 10,536,700
Granted Patent B1
US 10,536,700 · App. 15/594,380 · Granted Jan 14, 2020

Systems and methods for encoding videos based on visuals captured within the videos

Inventor: Sandeep Doshi (Sunnyvale, CA)
Assignee: GoPro, Inc.
H04N19/126G06K9/6218G06K9/6269H04N7/183H04N19/105H04N19/122H04N19/159H04N19/177
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,536,700
App. No.
15/594,380
Granted
Jan 14, 2020
Kind
B1
Abstract

Video information defining video content to be encoded may be obtained. Scene composition information for the video content may be obtained. The scene composition information may be determined by a convolutional neural network based on visuals represented within the video content. The video content may be encoded based on the scene composition information. The encoding of the video content may generate encoded video information defining the encoded video content.

Claims (42)

1. A system that encodes videos based on visuals captured within the videos, the system comprising:

one or more physical processors configured by machine-readable instructions to:

obtain video information, the video information defining video content to be encoded;

obtain scene composition information for the video content, the scene composition information determined by a convolutional neural network based on visuals represented within the video content, wherein the determination of the scene composition information includes object identification of a first object within the video content or face identification of a first face within the video content and depth identification of the first object or the first face within the video content, the depth identification including identification of whether the first object or the first face is located within a foreground or a background of the video content; and

encode the video content based on the scene composition information, the encoding of the video content generating encoded video information defining the encoded video content;

wherein:

the video content includes frames;

the determination of the scene composition information is performed for groupings of frames, wherein the determination of the scene composition information is performed for a first grouping of frames including a first number of frames and for a second grouping of frames including a second number of frames; and

number of frames within individual groupings of frames is determined based on detection of speed of activity within the frames such that the first grouping of frames includes the first number of frames based on detection of fast activity and the second grouping of frames includes the second number of frames based on detection of slow activity, the second number of frames being greater than the first number of frames.

2. The system of claim 1 , wherein the video information includes raw video information generated from an image sensor's capture of the video content.

3. The system of claim 1 , wherein the determination of the scene composition information includes analysis of a lower fidelity version of the video content.

4. The system of claim 1 , wherein the determination of the scene composition information includes analysis of the video content in a single color channel.

5. The system of claim 1 , wherein the determination of the scene composition information is performed by dedicated hardware.

6. The system of claim 1 , wherein the determination of the scene composition information further includes one or more of activity identification and luminance identification.

7. The system of claim 1 , wherein the one or more physical processors are, to encode the video content based on the scene composition information, further configured by the machine-readable instructions to set at least one of a quantization parameter, a block mode type selection, a block size selection, a transform size selection, an intra-frame bit distribution, or grouping of pictures setting based on the scene composition information.

8. The system of claim 1 , wherein the one or more physical processors are located in an image capture device.

9. A method for encoding videos based on visuals captured within the videos, the method comprising:

obtaining video information, the video information defining video content to be encoded;

obtaining scene composition information for the video content, the scene composition information determined by a convolutional neural network based on visuals represented within the video content, wherein the determination of the scene composition information includes object identification of a first object within the video content or face identification of a first face within the video content and depth identification of the first object or the first face within the video content, the depth identification including identification of whether the first object or the first face is located within a foreground or a background of the video content; and

encoding the video content based on the scene composition information, the encoding of the video content generating encoded video information defining the encoded video content;

wherein:

the video content includes frames;

determining the scene composition information is performed for groupings of frames, wherein the determination of the scene composition information is performed for a first grouping of frames including a first number of frames and for a second grouping of frames including a second number of frames; and

number of frames within individual groupings of frames is determined based on detection of speed of activity within the frames such that the first grouping of frames includes the first number of frames based on detection of fast activity and the second grouping of frames includes the second number of frames based on detection of slow activity, the second number of frames being greater than the first number of frames.

10. The method of claim 9 , wherein the video information includes raw video information generated from an image sensor's capture of the video content.

11. The method of claim 9 , wherein determining the scene composition information includes analyzing a lower fidelity version of the video content.

12. The method of claim 9 , wherein determining the scene composition information includes analyzing the video content in a single color channel.

13. The method of claim 9 , wherein determining the scene composition information is performed by dedicated hardware.

14. The method of claim 9 , wherein determining the scene composition information further includes identifying one or more of an activity and luminance.

15. The method of claim 9 , wherein encoding the video content based on the scene composition information includes setting at least one of a quantization parameter, a block mode type selection, a block size selection, a transform size selection, an intra-frame bit distribution, or grouping of pictures setting based on the scene composition information.

16. The method of claim 9 , wherein encoding the video content is performed by one or more processors of an image capture device.

17. A system that encodes videos based on visuals captured within the videos, the system comprising:

one or more physical processors configured by machine-readable instructions to:

obtain video information, the video information defining video content to be encoded;

obtain scene composition information for the video content, the scene composition information determined by a convolutional neural network based on visuals represented within the video content, wherein:

the determination of the scene composition information is performed by dedicated hardware; and

the determination of the scene composition information includes object identification of a first object within the video content or face identification of a first face within the video content and depth identification of the first object or the first face within the video content, the depth identification including identification of whether the first object or the first face is located within a foreground or a background of the video content; and

encode the video content based on the scene composition information, the encoding of the video content generating encoded video information defining the encoded video content, wherein encoding the video content based on the scene composition information includes setting at least one of a quantization parameter, a block mode type selection, a block size selection, a transform size selection, an intra-frame bit distribution, or grouping of pictures setting based on the scene composition information

wherein:

the video content includes frames;

the determination of the scene composition information is performed for groupings of frames, wherein the determination of the scene composition information is performed for a first grouping of frames including a first number of frames and for a second grouping of frames including a second number of frames; and

number of frames within individual groupings of frames is determined based on detection of speed of activity within the frames such that the first grouping of frames includes the first number of frames based on detection of fast activity and the second grouping of frames includes the second number of frames based on detection of slow activity, the second number of frames being greater than the first number of frames.

Assignments (7)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
SECURITY INTEREST Recorded Oct 19, 2020
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 054113/0594 →
SECURITY INTEREST Recorded Jul 31, 2017
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 043380/0163 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SPELLING OF INVENTOR'S NAME PREVIOUSLY RECORDED ON REEL 042364 FRAME 0347. ASSIGNOR(S) HEREBY CONFIRMS THE INVENTOR'S NAME SHOULD BE SPELLED SANDEEP DOSHI. Recorded May 30, 2017
From: DOSHI, SANDEEP
To: GOPRO, INC.
Reel/Frame 042623/0853 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2017
From: DOSHI, SANEEP
To: GOPRO, INC.
Reel/Frame 042364/0347 →