IP Library › Granted Patent US 11,477,533
Granted Patent B2
US 11,477,533 · App. 17/063,445 · Granted Oct 18, 2022

Automated video cropping

Inventors: Apurvakumar Dilipkumar Kansara (San Jose, CA); Sanford Holsapple (Sherman Oaks, CA); Arica Westadt (Los Angeles, CA); Kunal Bisla (Pleasanton, CA); Sameer Shah (Fremont, CA)
Assignee: Netflix, Inc.
H04N21/4728G06V20/46G06V20/49H04N21/4318H04N21/440272H04N21/4854H04N21/4858
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,477,533
App. No.
17/063,445
Granted
Oct 18, 2022
Kind
B2
Abstract

The disclosed computer-implemented method may include receiving, as an input, segmented video scenes, where each video scene includes a specified length of video content. The method may further include scanning the video scenes to identify objects within the video scene and also determining a relative importance value for the identified objects. The relative importance value may include an indication of which objects are to be included in a cropped version of the video scene. The method may also include generating a video crop that is to be applied to the video scene such that the resulting cropped version of the video scene includes those identified objects that are to be included based on the relative importance value. The method may also include applying the generated video crop to the video scene to produce the cropped version of the video scene. Various other methods, systems, and computer-readable media are also disclosed.

Claims (46)

1. A computer-implemented method comprising:

receiving, as an input, one or more segmented video scenes, each video scene comprising a specified length of video content;

scanning at least one of the video scenes to identify one or more objects within the video scene;

determining a relative importance value for one or more of the identified objects within the video scene;

determining which of the one or more identified objects are to be included in a cropped version of the video scene based on the determined relative importance value;

based on the determination, generating a video crop that is to be applied to the video scene, such that the resulting cropped version of the video scene includes those identified objects that are to be included in the cropped version of the video scene; and

applying the generated video crop to the video scene to produce the cropped version of the video scene, wherein the generated crop is specific to a size of a display screen.

2. The computer-implemented method of claim 1 , wherein the generated video crop is configured to generate a plurality of different aspect ratios for the cropped version of the video scene.

3. The computer-implemented method of claim 1 , wherein the generated video crop is configured to generate a plurality of different shapes for the cropped version of the video scene.

4. The computer-implemented method of claim 1 , wherein determining the relative importance value for one or more of the identified objects within the video scene includes at least one of:

determining which of the one or more identified objects a viewer is most likely to want to see; or

determining which of the one or more identified objects are to be included in a specific aspect ratio.

5. The computer-implemented method of claim 4 , further comprising:

determining that at least two objects in the video scene have a sufficient relative importance value to be included in the resulting cropped version of the video scene;

determining that the cropped version of the video scene has insufficient space to include each of the at least two objects;

determining prioritization values for the at least two objects; and

applying the generated video crop based on the prioritization values, such that the object with the highest prioritization value is included in the cropped version of the video scene.

6. The computer-implemented method of claim 1 , wherein determining the relative importance value for one or more of the identified objects within the video scene includes determining a frequency of occurrence of the one or more identified objects within the video scene.

7. The computer-implemented method of claim 1 , wherein determining the relative importance value for one or more of the identified objects within the video scene includes measuring an amount of movement of the one or more identified objects within the video scene.

8. The computer-implemented method of claim 1 , wherein determining the relative importance value for one or more of the identified objects within the video scene includes measuring an amount of blurring associated with each of the one or more identified objects in the video scene.

9. The computer-implemented method of claim 1 , wherein determining a relative importance value for one or more of the identified objects within the video scene includes, as a determining factor, the size of the display screen.

10. The computer-implemented method of claim 1 , wherein determining which of the one or more identified objects are to be included in a cropped version of the video scene is performed by a neural network.

11. A system comprising:

at least one physical processor; and

physical memory comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to:

receive, as an input, one or more segmented video scenes, each video scene comprising a specified length of video content;

scan at least one of the video scenes to identify one or more objects within the video scene;

determine a relative importance value for one or more of the identified objects within the video scene;

determine which of the one or more identified objects are to be included in a cropped version of the video scene based on the determined relative importance value;

based on the determination, generate a video crop that is to be applied to the video scene, such that the resulting cropped version of the video scene includes those identified objects that are to be included in the cropped version of the video scene; and

apply the generated video crop to the video scene to produce the cropped version of the video scene, wherein the generated crop is specific to a size of a display screen.

12. The system of claim 11 , further comprising determining a semantic context for one or more of the identified objects in the video scene.

13. The system of claim 12 , wherein the determined semantic context is implemented when determining a relative importance value for the one or more identified objects in the video scene.

14. The system of claim 11 , further comprising tracking which video crops were generated and applied to one or more of the video scenes.

15. The system of claim 14 , further comprising comparing at least one cropped version of the video scene to a user-cropped version of the same video scene to identify one or more differences in cropping.

16. The system of claim 15 , wherein the at least one physical processor automatically alters how the video crop is generated based on the identified differences in cropping.

17. The system of claim 11 , further comprising encoding the cropped version of the video scene according to a specified encoding format.

18. A non-transitory computer-readable medium comprising one or more computer-executable instructions that, when executed by at least one processor of a computing device, cause the computing device to:

receive, as an input, one or more segmented video scenes, each video scene comprising a specified length of video content;

scan at least one of the video scenes to identify one or more objects within the video scene;

determine a relative importance value for one or more of the identified objects within the video scene;

determine which of the one or more identified objects are to be included in a cropped version of the video scene based on the determined relative importance value;

based on the determination, generate a video crop that is to be applied to the video scene, such that the resulting cropped version of the video scene includes those identified objects that are to be included in the cropped version of the video scene; and

apply the generated video crop to the video scene to produce the cropped version of the video scene, wherein the generated crop is specific to a size of a display screen.

19. The non-transitory computer-readable medium of claim 18 , wherein the generated video crop is configured to generate a plurality of different aspect ratios for the cropped version of the video scene.

20. The non-transitory computer-readable medium of claim 18 , wherein the generated video crop is configured to generate a plurality of different shapes for the cropped version of the video scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2020
From: KANSARA, APURVAKUMAR DILIPKUMAR; HOLSAPPLE, SANFORD; WESTADT, ARICA; BISLA, KUNAL; SHAH, SAMEER
To: NETFLIX, INC
Reel/Frame 053991/0612 →
Continuity (2)
Continuation 16457586 · Jun 28, 2019
Related Publication 20210021900A1 · Jan 21, 2021