IP Library Granted Patent US 12,273,521
Granted Patent B2
US 12,273,521 · App. 17/862,571 · Granted Apr 8, 2025

Obtaining video quality scores from inconsistent training quality scores

Inventors: Yilin Wang (Sunnyvale, CA); Balineedu Adsumilli (Sunnyvale, CA)
Assignee: GOOGLE LLC
H04N19/13G06F18/214G06N20/00G06T7/0002G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,273,521
App. No.
17/862,571
Granted
Apr 8, 2025
Kind
B2
Abstract

A training dataset that includes a first dataset and a second dataset is received. The first dataset includes a first subset of first videos corresponding to a first context and respective first ground truth quality scores of the first videos, and the second dataset includes a second subset of second videos corresponding to a second context and respective second ground truth quality scores of the second videos. A machine learning model is trained to predict the respective first ground truth quality scores and the respective second ground truth quality scores. Training the model includes training it to obtain a global quality score for one of the videos; and training it to map the global quality score to context-dependent predicted quality scores. The context-dependent predicted quality scores include a first context-dependent predicted quality score corresponding to the first context and a second context-dependent predicted quality score corresponding to the second context.

Claims (55)

1. A method, comprising:

receiving a training dataset,

wherein the training dataset comprises a first dataset and a second dataset,

wherein the first dataset comprises a first subset of first videos corresponding to a first context and respective first ground truth quality scores of the first videos, and

wherein the second dataset comprises a second subset of second videos corresponding to a second context that is different from the first context and respective second ground truth quality scores of the second videos; and

training a machine learning model to predict the respective first ground truth quality scores and the respective second ground truth quality scores, wherein training the machine learning model comprises:

training the machine learning model to obtain a global quality score for one of the videos of the training dataset, wherein the global quality score is context independent; and

training the machine learning model to map the global quality score to context-dependent predicted quality scores, wherein the context-dependent predicted quality scores comprising a first context-dependent predicted quality score corresponding to the first context and a second context-dependent predicted quality score corresponding to the second context.

2. The method of claim 1 , wherein the machine learning model is trained to map the global quality score to the context-dependent predicted quality scores using piecewise functions, wherein each piece of the piecewise functions corresponds to a context.

3. The method of claim 2 , wherein at least one of the piecewise functions is a monotonic function.

4. The method of claim 2 , wherein the at least one of the piecewise functions is a linear function.

5. The method of claim 4 , wherein all coefficients of the linear function are constrained during the training to be non-negative values.

6. The method of claim 1 , further comprising:

obtaining, for an input video, a plurality of predicted context-specific quality scores from the machine learning model; and

selecting one of the plurality of the predicted context-specific quality scores as a predicted quality score for the input video.

7. The method of claim 6 , wherein selecting the one of the plurality of the predicted context-specific quality scores as the predicted quality score for the input video comprises:

selecting the one of the plurality of the predicted context-specific quality scores based on a context of the input video.

8. The method of claim 7 , wherein the context of the input video is obtained using a context predictor machine learning model that is trained to receive input videos and output respective context predictions.

9. A device, comprising:

a processor that is configured to:

receive an input video;

obtain, from a machine learning model, a plurality of context-dependent predicted quality scores for the input video, wherein to obtain the plurality of context-dependent predicted quality scores comprises to:

obtain a global quality score for the input video, wherein the global quality score is context independent; and

map the global quality score to the plurality of context-dependent predicted quality scores; and

select one of the plurality of context-dependent predicted quality scores as a predicted quality score for the input video.

10. The device of claim 9 , wherein to select the one of the plurality of context-dependent predicted quality scores as the predicted quality score for the input video comprises to:

select the one of the plurality of context-dependent predicted quality scores based on a context of the input video.

11. The device of claim 10 , wherein the context of the input video is obtained using a context predictor machine learning model that is trained to receive input videos and output respective context predictions.

12. The device of claim 9 , wherein to select the one of the plurality of context-dependent predicted quality scores as the predicted quality score for the input video comprises to:

select the one of the plurality of context-dependent predicted quality scores corresponding to a maximum value amongst the plurality of context-dependent predicted quality scores.

13. The device of claim 9 , wherein the processor is further configured to:

identify a plurality of videos responsive to a request;

obtain, from the machine learning model, respective quality scores for the plurality of videos; and

rank the plurality of videos based on the respective quality scores.

14. The device of claim 9 , wherein the processor is further configured to:

encode a source video using a first value of an encoding parameter to obtain the input video; and

determine whether to re-encode the source video using a second value of the encoding parameter that is different from the first value based on the predicted quality score.

15. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

receiving an input video;

obtaining, from a machine learning model, a plurality of context-dependent predicted quality scores for the input video, wherein obtaining the plurality of context-dependent predicted quality scores comprises:

obtaining a global quality score for the input video, wherein the global quality score is context independent; and

mapping the global quality score to the plurality of context-dependent predicted quality scores; and

selecting one of the plurality of context-dependent predicted quality scores as a predicted quality score for the input video.

16. The non-transitory computer readable medium of claim 15 , wherein selecting the one of the plurality of context-dependent predicted quality scores as the predicted quality score for the input video comprises:

selecting the one of the plurality of context-dependent predicted quality scores based on a context of the input video.

17. The non-transitory computer readable medium of claim 16 , wherein the context of the input video is obtained using a context predictor machine learning model that is trained to receive input videos and output respective context predictions.

18. The non-transitory computer readable medium of claim 15 , wherein selecting the one of the plurality of context-dependent predicted quality scores as the predicted quality score for the input video comprises:

selecting the one of the plurality of context-dependent predicted quality scores corresponding to a maximum value amongst the plurality of context-dependent predicted quality scores.

19. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

identifying a plurality of videos responsive to a request;

obtaining, from the machine learning model, respective quality scores for the plurality of videos; and

ranking the plurality of videos based on the respective quality scores.

20. The non-transitory computer readable medium of claim 15 , wherein the operations further comprise:

encoding a source video using a first value of an encoding parameter to obtain the input video; and

determining whether to re-encode the source video using a second value of the encoding parameter that is different from the first value based on the predicted quality score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2022
From: WANG, YILIN; ADSUMILLI, BALINEEDU
To: GOOGLE LLC
Reel/Frame 060503/0100 →
Continuity (1)
Related Publication 20240022726A1 · Jan 18, 2024
References Cited (14)
US 20090067726A1 · Erol · 2009 [cited by examiner]
US 20140169662A1 · Liu · 2014 [cited by examiner]
US 20170076318A1 · Goswami · 2017 [cited by examiner]
US 20170372155A1 · Odry · 2017 [cited by examiner]
US 20190147305A1 · Lu · 2019 [cited by examiner]
US 20200204804A1 · Chen · 2020 [cited by examiner]
US 20210089963A1 · Baek · 2021 [cited by examiner]
US 20220076078A1 · Vdovjak · 2022 [cited by examiner]
US 20220237417A1 · Singhal · 2022 [cited by examiner]
US 20230138016A1 · Kaiser · 2023 [cited by examiner]
Bosse et al. “A deep neural network for image quality assessment.” 2016 IEEE international conference on image processing (ICIP) IEEE, 2016. (Year: 2016). [cited by examiner]
Jiang et al. “Deep optimization model for screen content image quality assessment using neural networks.” arXiv preprint arXiv: 1903.00705 (2019). (Year: 2019). [cited by examiner]
Kuzovkin et al. “Context-aware clustering and assessment of photo collections.” Proceedings of the symposium on Computational Aesthetics. 2017. (Year: 2017). [cited by examiner]
Lu et al. “Blind Surveillance Image Quality Assessment via Deep Neural Network Combined with the Visual Saliency.” arXiv preprint arXiv:2206.04318 (Jun. 2022). (Year: 2022). [cited by examiner]