IP Library Granted Patent US 10,419,773
Granted Patent B1
US 10,419,773 · App. 15/933,222 · Granted Sep 17, 2019

Hybrid learning for adaptive video grouping and compression

Inventors: Hai Wei (Seattle, WA); Yang Yang (Issaquah, WA); Lei Li (Kirkland, WA); Amarsingh Buckthasingh Winston (Seattle, WA); Avisar Ten-Ami (Mercer Island, WA)
Assignee: Amazon Technologies, Inc.
H04N19/46G06N20/00H04N19/132H04N19/14H04N19/169
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,419,773
App. No.
15/933,222
Granted
Sep 17, 2019
Kind
B1
Abstract

Methods and apparatus are described in which both supervised and unsupervised machine learning are used to classify video content for compression using encoding profiles that are optimized for each type of video content.

Claims (54)

1. A computer program product, comprising one or more non-transitory computer-readable media having computer program instructions stored therein, the computer program instructions being configured such that, when executed by one or more computing devices, the computer program instructions cause the one or more computing devices to:

generate a supervised training feature set for each of a plurality of training video samples;

train a classifier to classify each of the training video samples into a corresponding one of a plurality of semantic classes based on the corresponding supervised training feature set;

generate an unsupervised training feature set for each of the plurality of training video samples;

cluster the training video samples within each of the semantic classes based on the corresponding unsupervised training feature sets thereby generating a plurality of clusters of video samples within each of the semantic classes;

determine an encoding profile for each of the clusters of video samples;

generate a semantic classification feature set based on a first video sample;

classify the first video sample into a first semantic class of the plurality of semantic classes based on the semantic classification feature set;

generate an encoding-complexity feature set based on the first video sample;

assign the first video sample to a first cluster of the plurality of clusters of video samples associated with the first semantic class based on the encoding-complexity feature set; and

encode the first video sample using the encoding profile corresponding to the first cluster of video samples.

2. The computer program product of claim 1 , wherein the computer program instructions are configured to cause the one or more computing devices to determine the encoding profile for each of the clusters of video samples by performing a rate-distortion optimization based on encoded versions of the corresponding training video samples.

3. The computer program product of claim 1 , wherein the computer program instructions are configured to cause the one or more computing devices to generate the semantic classification feature set by extracting spatial features and temporal features from the first video sample, and wherein the computer program instructions are configured to cause the one or more computing devices to classify the first video sample into the first semantic class by determining that a feature vector including values representing the spatial and temporal features of the first video sample corresponds to the first semantic class.

4. The computer program product of claim 1 wherein the computer program instructions are configured to cause the one or more computing devices to generate the encoding-complexity feature set by:

encoding the first video sample using a pre-encode encoding profile; and

extracting encoding-complexity features from encoding log data generated as a result of the encoding of the first video sample.

5. A computer-implemented method, comprising:

generating a semantic classification feature set based on a first video sample;

classifying the first video sample into a first semantic class of a plurality of semantic classes based on the semantic classification feature set, each of the semantic classes having a plurality of clusters of video samples associated therewith;

generating an encoding-complexity feature set based on the first video sample;

assigning the first video sample to a first cluster of video samples associated with the first semantic class based on the encoding-complexity feature set; and

encoding the first video sample using a first encoding profile corresponding to the first cluster of video samples.

6. The method of claim 5 , wherein generating the semantic classification feature set includes extracting spatial features and temporal features from the first video sample.

7. The method of claim 6 , wherein classifying the first video sample into the first semantic class includes determining that a feature vector including values representing the spatial and temporal features of the first video sample corresponds to the first semantic class.

8. The method of claim 5 , wherein generating the encoding-complexity feature set includes:

encoding the first video sample using a pre-encode encoding profile; and

extracting encoding-complexity features from encoding log data generated as a result of the encoding of the first video sample.

9. The method of claim 5 , wherein classifying the first video sample is done using a classifier, the method further comprising:

generating a supervised training feature set for each of a plurality of training video samples, each of the training video samples having a known correspondence to one of a plurality of semantic classes; and

training the classifier to classify each of the training video samples into a corresponding one of the semantic classes based on the corresponding supervised training feature set.

10. The method of claim 9 , further comprising:

generating an unsupervised training feature set for each of the training video samples; and

clustering the training video samples associated with each of the semantic classes based on the corresponding unsupervised training feature sets thereby generating the plurality of clusters of video samples associated with each of the semantic classes.

11. The method of claim 10 , further comprising modifying the clusters associated with the semantic classes based on run-time classification of new video samples.

12. The method of claim 10 , further comprising determining the first encoding profile for the first cluster of video samples by performing a rate-distortion optimization based on encoded versions of the training video samples included in the first cluster of video samples.

13. A system, comprising one or more computing devices configured to:

generate a semantic classification feature set based on a first video sample;

classify the first video sample into a first semantic class of a plurality of semantic classes based on the semantic classification feature set, each of the semantic classes having a plurality of clusters of video samples associated therewith;

generate an encoding-complexity feature set based on the first video sample;

assign the first video sample to a first cluster of video samples associated with the first semantic class based on the encoding-complexity feature set; and

encode the first video sample using a first encoding profile corresponding to the first cluster of video samples.

14. The system of claim 13 , wherein the one or more computing devices are configured to generate the semantic classification feature set by extracting spatial features and temporal features from the first video sample.

15. The system of claim 14 , wherein the one or more computing devices are configured to classify the first video sample into the first semantic class by determining that a feature vector including values representing the spatial and temporal features of the first video sample corresponds to the first semantic class.

16. The system of claim 13 , wherein the one or more computing devices are configured to generate the encoding-complexity feature set by:

encoding the first video sample using a pre-encode encoding profile; and

extracting encoding-complexity features from encoding log data generated as a result of the encoding of the first video sample.

17. The system of claim 13 , wherein the one or more computing devices are configured to classify the first video sample using a classifier, and wherein the one or more computing devices are further configured to:

generate a supervised training feature set for each of a plurality of training video samples, each of the training video samples having a known correspondence to one of a plurality of semantic classes; and

train the classifier to classify each of the training video samples into a corresponding one of the semantic classes based on the corresponding supervised training feature set.

18. The system of claim 17 , wherein the one or more computing devices are further configured to:

generate an unsupervised training feature set for each of the training video samples; and

cluster the training video samples associated with each of the semantic classes based on the corresponding unsupervised training feature sets thereby generating the plurality of clusters of video samples associated with each of the semantic classes.

19. The system of claim 18 , wherein the one or more computing devices are further configured to modify the clusters associated with the semantic classes based on run-time classification of new video samples.

20. The system of claim 18 , wherein the one or more computing devices are further configured to determine the first encoding profile for the first cluster of video samples by performing a rate-distortion optimization based on encoded versions of the training video samples included in the first cluster of video samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2018
From: WEI, HAI; YANG, YANG; LI, LEI; WINSTON, AMARSINGH BUCKTHASINGH; TEN-AMI, AVISAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 045626/0131 →
Cited By (7)
US 12,299,298 US 12,363,328 US 12,501,061 US 12,501,084 US 12,513,303 US 12,634,489 US 12,725,635