IP Library › Granted Patent US 11,544,562
Granted Patent B2
US 11,544,562 · App. 16/875,887 · Granted Jan 3, 2023

Perceived media object quality prediction using adversarial annotations for training and multiple-algorithm scores as input

Inventors: Thomas Sydney Austin Wallis (Tuebingen, DE); Luitpold Staudigl (Bonn, DE); Muhammad Bilal Javed (Berlin, DE); Pablo Barbachano (Berlin, DE); Mike Mueller (Tuebingen, DE)
Assignee: Amazon Technologies, Inc.
G06N3/08G06F30/27G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,562
App. No.
16/875,887
Filed
May 15, 2020
Granted
Jan 3, 2023
Kind
B2
Examiner
DINH, PAUL
Art Unit
2851
USPC
706/20
Abstract

Respective labels indicative of compression-related quality degradation for a set of media object tuples which meet a divergence criterion are obtained; each tuple comprises a reference media object and a pair of corresponding compressed media object versions. Pairs of training records for a machine learning model are generated using the labeled media object tuples and multiple perceptual quality algorithms, with each training record comprising respective perceived quality degradation scores generated by each of the multiple algorithms for a given compressed media object of a tuple. A machine learning model is trained, using the record pairs, to predict quality degradation scores for compressed media objects.

Claims (63)

1. A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices cause the one or more computing devices to:

identify a first plurality of image tuples which satisfy an algorithm-to-algorithm divergence threshold, wherein individual ones of the image tuples comprise a reference image, a first compressed version of the reference image, and a second compressed version of the reference image, and wherein, with respect to a given image tuple, a difference between (a) a first quality degradation score produced by a first perceptual quality algorithm of a first set of perceptual quality algorithms for one or more of the compressed versions relative to the reference image, and (b) a second quality degradation score produced by a second perceptual quality algorithm of the first set for the one or more of the compressed versions relative to the reference image exceeds the divergence threshold;

obtain respective labels from a group of one or more annotators for individual ones of the first plurality of image tuples, wherein a label for a given image tuple indicates which compressed version of the given image tuple is perceived to be more similar to the reference image of the given image tuple;

without utilizing an annotator, automatically generate labels for individual ones of a second plurality of image tuples using quality degradation scores produced by a second set of perceptual quality algorithms;

store a labeled image data set comprising at least some image tuples of the first and second pluralities of image tuples and their respective labels;

generate, using a third set of perceptual quality algorithms, a plurality of pairs of training records for at least a first machine learning model, wherein an individual pair of training records comprises:

a first record which includes (a) a plurality of quality degradation scores for a first compressed image of the labeled image data set, wherein individual ones of the quality degradation scores are obtained using respective perceptual quality algorithms of the third set, and (b) the particular label which was stored in the labeled image data set for the image tuple of which the first compressed version is a member; and

a second record which includes (a) a plurality of quality degradation scores for a second compressed image of the labeled image data set, wherein individual ones of the quality degradation scores are obtained using the respective perceptual quality algorithms of the third set, and (b) the particular label;

train the first machine learning model using the plurality of pairs of training records to predict, for a post-training input record comprising a plurality of quality degradation scores for a particular compressed version of an image, an output quality degradation score for the particular compressed version; and

utilize the output quality degradation score to identify an image for presentation to a viewer.

2. The system as recited in claim 1 , wherein the first machine learning model comprises a symmetric neural network with at least one fully-connected layer and at least one softmax layer.

3. The system as recited in claim 1 , wherein the one or more annotators comprise a plurality of annotators, and wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

analyze the inter-annotator consistency of labels produced by the plurality of annotators; and

exclude, from the labeled image data set, at least one image tuple based at least on part on results of the inter-annotator consistency analysis.

4. The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

initiate the identification of the first plurality of image tuples in response to one or more requests obtained via a programmatic interface of a network-accessible service of a provider network.

5. The system as recited in claim 1 , wherein a first the third set of perceptual quality algorithms comprises a particular algorithm in which a final quality degradation score is obtained from a plurality of intermediate scores, wherein the plurality of quality degradation scores included for the first compressed image in the first training record comprises an intermediate score of the plurality of intermediate scores.

6. A method, comprising:

performing, at one or more computing devices:

obtaining respective labels from a group of one or more annotators for individual ones of the first plurality of image tuples which satisfy a first divergence criterion, wherein individual ones of the image tuples comprise a reference image, a first compressed version of the reference image, and a second compressed version of the reference image, and wherein a label for a given image tuple indicates which compressed version of the given image tuple is perceived to be more similar to the reference image of the given image tuple;

storing an labeled image data set comprising at least some image tuples of the first plurality of image tuples and their respective labels;

generating, using a first set of perceptual quality algorithms, a plurality of pairs of training records for at least a first machine learning model, wherein an individual pair of training records comprises:

a first record which includes a plurality of quality degradation scores for a first compressed version of a particular reference image of the labeled image data set, wherein individual ones of the quality degradation scores are obtained using respective perceptual quality algorithms of the first set; and

a second record which includes a plurality of quality degradation scores for a second compressed version of the particular reference image, wherein individual ones of the quality degradation scores are obtained using the respective perceptual quality algorithms of the first set; and

training the first machine learning model using the plurality of pairs of training records to predict, for a post-training input record comprising a plurality of quality degradation scores for a particular compressed version of an image, a quality degradation score for the particular compressed version.

7. The method as recited in claim 6 , further comprising performing, at one or more computing devices:

without utilizing an annotator, automatically generating labels for individual ones of a second plurality of image tuples using quality degradation scores generated by a second set of perceptual quality algorithms; and

storing the second plurality of image tuples and their respective labels as part of the labeled image data set.

8. The method as recited in claim 6 , further comprising performing, at one or more computing devices:

determining a difference, with respect to a particular image tuple of a collection of image tuples, between (a) a first quality degradation score generated by a first perceptual quality algorithm for a compressed image of the image tuple and (b) a second quality degradation score generated by a second perceptual quality algorithm for the compressed image; and

evaluating the first divergence criterion with respect to the particular image tuple, wherein said evaluating comprises comparing the difference to a threshold.

9. The method as recited in claim 6 , wherein the group of one or more annotators comprises a plurality of annotators, the method further comprising performing, at one or more computing devices:

computing, for individual image tuples of the first plurality of image tuples, a measure of inter-annotator consistency; and

excluding, from the labeled image data set, at least one image tuple whose inter-annotator consistency measure is below a threshold.

10. The method as recited in claim 6 , wherein the group of one or more annotators comprises a plurality of annotators, the method further comprising performing, at one or more computing devices:

excluding, from the labeled image data set, at least one image tuple for which a label was generated by a particular annotator selected based on an analysis of inter-annotator consistency.

11. The method as recited in claim 6 , further comprising performing, at one or more computing devices:

training, using additional training records for which labels were generated automatically without using annotators, a second machine learning model to predict quality degradation scores, wherein training the first machine learning model using the plurality of pairs of training records comprises modifying the second machine learning model using the plurality of pairs of training records.

12. The method as recited in claim 6 , wherein the first machine learning model comprises a neural network-based model.

13. The method as recited in claim 6 , further comprising:

obtaining, at a network-accessible service of a provider network, one or more programmatic requests to train a machine learning model to predict perceived image quality degradation scores, wherein the first machine learning model is trained in response to the one or more programmatic requests.

14. The method as recited in claim 6 , further comprising:

obtaining, from a trained version of the first machine learning model, a first set of quality degradation scores for compressed images produced using a first set of hyper-parameters of a compression algorithm, and a second set of quality degradation scores for compressed images produced using a second set of hyper-parameters of the compression algorithm; and

causing, based at least in part on a comparison of the first and second sets of quality degradation scores, the first set of hyper-parameters to be employed for presenting a set of images.

15. The method as recited in claim 6 , further comprising:

obtaining respective resource consumption metrics of a plurality of perceptual quality algorithms; and

including, in the first set of perceptual quality algorithms, a first perceptual quality algorithm of the plurality of perceptual quality algorithms based at least in part on a comparison of a resource consumption metric of the first perceptual quality algorithm with a corresponding resource consumption metric of a second perceptual quality algorithm of the plurality of perceptual quality algorithms; and

excluding, from the first set of perceptual quality algorithms, the second perceptual quality algorithm based at least in part on the comparison.

16. One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:

obtain respective labels from a group of one or more annotators for individual ones of the first plurality of media object tuples which satisfy a first divergence criterion, wherein individual ones of the media object tuples comprise a reference media object, a first compressed version of the media object, and a second compressed version of the reference media object, and wherein a label for a given media object tuple indicates which compressed version of the given media object tuple is perceived to be more similar to the reference media object of the given media object tuple;

generate, using a first set of perceptual quality algorithms, a plurality of pairs of training records for at least a first machine learning model, wherein an individual pair of training records comprises:

a first record which includes a plurality of quality degradation scores for a first compressed version of a particular reference media object of a labeled media object data set, wherein individual ones of the quality degradation scores are obtained using respective perceptual quality algorithms of the first set, and wherein the labeled media object data set comprises at least some media object tuples of the first plurality of media object tuples and their respective labels; and

a second record which includes a plurality of quality degradation scores for a second compressed version of the particular reference media object, wherein individual ones of the quality degradation scores are obtained using the respective perceptual quality algorithms of the first set; and

train the first machine learning model using the plurality of pairs of training records to predict, for a post-training input record comprising a plurality of quality degradation scores for a particular compressed version of a media object, a quality degradation score for the particular compressed version.

17. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

automatically generate labels for individual ones of a second plurality of media object tuples using quality degradation scores generated by a second set of perceptual quality algorithms; and

include the second plurality of media object tuples and their respective labels as part of the labeled media object data set.

18. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the first set of perceptual quality algorithms comprises one or more of: (a) an algorithm which utilizes multi-scale decomposition to generate predicted perceived quality degradation scores, (b) an algorithm in which physical image differences are weighted at least according to assumptions about contrast sensitivity, or (c) an algorithm which measures phase coherence in spatial filters.

19. The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the first machine learning model comprises a neural network-based model.

20. The one or more non-transitory computer-accessible storage media as recited in claim 16 , storing further program instructions that when executed on or across the one or more processors further cause the one or more processors to:

identify, using a second set of perceptual quality algorithms, the media object tuples which satisfy the first divergence criterion, wherein the at least one algorithm of the second set is not in the first set of perceptual quality algorithms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2020
From: WALLIS, THOMAS SYDNEY AUSTIN; STAUDIGL, LUITPOLD; JAVED, MUHAMMAD BILAL; BARBACHANO, PABLO; MUELLER, MIKE
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053523/0333 →
Continuity (1)
Related Publication 20210357745A1 · Nov 18, 2021