IP Library Granted Patent US 12,417,522
Granted Patent B2
US 12,417,522 · App. 17/486,359 · Granted Sep 16, 2025

Method for constructing a perceptual metric for judging video quality

Inventors: Troy Chinen (Fremont, CA); Alex Sukhanov (Sunnyvale, CA); Eirikur Thor Agustsson (Zurich, CH); George Dan Toderici (Mountain View, CA)
Assignee: GOOGLE LLC
G06T7/0002G06F18/24G06N20/00H04N19/23G06T2207/10016G06T2207/20081G06T2207/30168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,522
App. No.
17/486,359
Granted
Sep 16, 2025
Kind
B2
Abstract

An example computer-implemented method for determining a perceptual quality of a subject video content item is provided. The example method can include inputting a subject frame set from the subject video content item into a first machine-learned model. The example method can also include generating, using the first machine-learned model, a feature based at least in part on the subject frame set. The example method can also include outputting, using a second machine-learned model, a score indicating the perceptual quality of the subject video content item based at least in part on the feature.

Claims (65)

1. A computer-implemented method for determining a perceptual quality of a subject video content item that is based on an encoding of a reference video content item, the method comprising:

generating, by one or more computing devices, subject flow data based at least in part on a frame sequence of the subject video content item;

generating, by the one or more computing devices, reference flow data based at least in part on a frame sequence of the reference video content item;

processing, by the one or more computing devices, the subject flow data and the reference flow data using a first machine-learned model to compare the subject flow data and the reference flow data, wherein processing the subject flow data and the reference flow data using the first machine-learned model comprises processing the subject flow data and the reference flow data using an adapted image classifier model;

generating, by the one or more computing devices and using the first machine-learned model, a temporal feature based at least in part on the subject flow data and the reference flow data; and

outputting, by the one or more computing devices using a second machine-learned model, a score indicating the perceptual quality of the subject video content item based at least in part on the temporal feature.

2. The computer-implemented method of claim 1 , wherein generating the subject flow data comprises:

determining, by the one or more computing devices, a difference between frames in the frame sequence of the subject video content item.

3. The computer-implemented method of claim 1 , wherein the adapted image classifier model comprises one or more weights re-trained for determining the perceptual quality of the subject video content item.

4. The computer-implemented method of claim 1 , wherein the first machine-learned model comprises weights pre-trained on an image classification task.

5. The computer-implemented method of claim 1 , comprising:

generating, by the one or more computing devices, a spatial feature based at least in part on a second subject frame set from the subject video content item; and

inputting, by the one or more computing devices, the spatial feature into the second machine-learned model;

wherein the second machine-learned model generates the score based on the temporal feature and the spatial feature.

6. The computer-implemented method of claim 1 , comprising:

for each respective frame set of a plurality of frame sets of the subject video content item, determining a respective temporal feature by comparison to a corresponding frame set of the reference video content item;

aggregating, by the one or more computing devices, the plurality of temporal features to obtain an aggregate temporal feature; and

inputting, by the one or more computing devices, the aggregate temporal feature into the second machine-learned model to generate the score.

7. The computer-implemented method of claim 1 , comprising:

updating, by the one or more computing devices, one or more parameters of a video encoder based at least in part on the score.

8. The computer-implemented method of claim 7 , comprising:

encoding, by the one or more computing devices, a reference video content item based at least in part on the updated one or more parameters.

9. The computer-implemented method of claim 1 , comprising:

updating, by the one or more computing devices, one or more parameters of a video encoder and a video decoder based at least in part on the score, wherein the video decoder is configured to decode video content encoded by the video encoder.

10. The computer-implemented method of claim 1 , wherein the adapted image classified model comprises a convolutional neural network.

11. The computer-implemented method of claim 1 , wherein processing the subject flow data and the reference flow data using the first machine-learned model comprises:

outputting, from the adapted image classifier model, a layer output of a layer of the adapted image classifier model.

12. The computer-implemented method of claim 1 , wherein part of a pre-trained image classifier model is used as the adapted image classifier model.

13. A computing system for determining a perceptual quality of a subject video content item that is based on an encoding of a reference video content item, comprising:

one or more processors; and

one or more memory devices storing computer-readable instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

generating subject flow data based at least in part on a frame sequence of the subject video content item;

generating reference flow data based at least in part on a frame sequence of the reference video content item;

processing the subject flow data and the reference flow data using a first machine-learned model to compare the subject flow data and the reference flow data, wherein processing the subject flow data and the reference flow data using the first machine-learned model comprises processing the subject flow data and the reference flow data using an adapted image classifier model;

generating, using the first machine-learned model, a temporal feature based at least in part on the subject flow data and the reference flow data; and

outputting, using a second machine-learned model, a score indicating the perceptual quality of the subject video content item based at least in part on the temporal feature.

14. The computing system of claim 13 , wherein the adapted image classifier model comprises one or more weights re-trained for determining the perceptual quality of the subject video content item.

15. The computing system of claim 13 , wherein the first machine-learned model comprises weights pre-trained on an image classification task.

16. The computing system of claim 13 , wherein the operations further comprise:

generating, by the one or more computing devices, a spatial feature based at least in part on a second subject frame set from the subject video content item; and

inputting, by the one or more computing devices, the spatial feature into the second machine-learned model;

wherein the second machine-learned model generates the score based on the temporal feature and the spatial feature.

17. The computing system of claim 13 , wherein the operations further comprise:

for each respective frame set of a plurality of frame sets of the subject video content item, determining a respective temporal feature by comparison to a corresponding frame set of the reference video content item;

aggregating, by the one or more computing devices, the plurality of temporal features to obtain an aggregate temporal feature; and

inputting the aggregate temporal feature into the second machine-learned model to generate the score.

18. The computing system of claim 13 , wherein the operations further comprise:

determining an objective function based at least in part on the score; and

updating one or more parameters of a video encoder based at least in part on the objective function.

19. The computing system of claim 18 , wherein the operations further comprise:

encoding the reference video content item based at least in part on the updated one or more parameters.

20. The computing system of claim 13 , wherein the operations further comprise:

updating one or more parameters of a video encoder and a video decoder based at least in part on the score, wherein the video decoder is configured to decode video content encoded by the video encoder.

21. A computing system for determining a perceptual quality of a subject video content item, comprising:

one or more processors; and

one or more memory devices storing computer-readable instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

inputting a subject frame set from the subject video content item into a first machine-learned model, wherein the first machine-learned model comprises an adapted image classifier model, wherein the adapted image classifier model comprises one or more weights re-trained for determining the perceptual quality of the subject video content item;

generating, using the first machine-learned model, a feature based at least in part on the subject frame set; and

outputting, using a second machine-learned model, a score indicating the perceptual quality of the subject video content item based at least in part on the feature.

22. One or more memory devices storing:

an encoder configured to compress video content based at least in part on one or more parameters, the one or more parameters trained based at least in part on a score indicating the perceptual quality of decoded video content, the score generated using a second machine-learned model based at least in part on a temporal feature generated using a first-machine-learned model based at least in part on processing flow data and reference flow data for the decoded video content using an adapted image classifier model, the flow data generated based at least in part on a frame sequence of the decoded video content and the reference flow data generated based at least in part on a frame sequence of reference video content, the decoded video content based on an encoding of the reference video content; and

computer-readable instructions that, when implemented, cause one or more processors to generate, using the encoder, an encoded data stream based at least in part on an input video content item.

23. One or more memory devices storing:

a decoder configured to decompress video content based at least in part on one or more parameters, the one or more parameters trained based at least in part on a score indicating the perceptual quality of decoded video content, the score generated using a second machine-learned model based at least in part on a temporal feature generated using a first-machine-learned model based at least in part on processing flow data and reference flow data for the decoded video content using an adapted image classifier model, the flow data generated based at least in part on a frame sequence of the decoded video content and the reference flow data generated based at least in part on a frame sequence of reference video content, the decoded video content based on an encoding of the reference video content; and

computer-readable instructions that, when implemented, cause one or more processors to generate, using the decoder, a decoded video data stream based at least in part on an input encoded data stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2022
From: CHINEN, TROY; SUKHANOV, ALEX; AGUSTSSON, EIRIKUR THOR; TODERICI, GEORGE DAN
To: GOOGLE LLC
Reel/Frame 059033/0307 →
Continuity (1)
Related Publication 20230099526A1 · Mar 30, 2023
References Cited (9)
US 20160034786A1 · Suri · 2016 [cited by examiner]
US 20190246111A1 · Li · 2019 [cited by examiner]
US 20210099715A1 · Topiwala · 2021 [cited by examiner]
US 20220051382A1 · Chen · 2022 [cited by examiner]
CN 111062297A · 2020 [cited by examiner]
Somraj et al., Understanding the Perceived Quality of Video Predictions, arXiv:2005.00356, Jun. 22 (Year: 2021). [cited by examiner]
Bampis et al., “Spatio Temporal Feature Integration and Model Fusion for Full Reference Video Quality Assessment”, Apr. 13, 2018, arXiv:1804.04813v1, arXiv.org, 12 pages. [cited by applicant]
Zhang et al., “Enhancing VMAF through New Feature Integration and Model Combination”, Mar. 10, 2021, arXiv:2103.06338v1, arXiv.org; 5 pages. [cited by applicant]
Chinen et al., “Towards a Semantic Perceptual Image Metric”, arXiv:1808.00447v1, Aug. 1, 2018, 5 pages. [cited by applicant]