IP Library › Granted Patent US 12,739,450
Granted Patent B2
US 12,739,450 · App. 19/009,366 · Granted Sep 15, 2026

Methods, systems, and media for determining perceptual quality indicators of video content items

Inventors: Yilin Wang (San Jose, CA); Balineedu Adsumilli (Foster City, CA); Junjie Ke (Stanford, CA); Hossein Talebi (San Francisco, CA); Joong Yim (Mountain View, CA); Neil Birkbeck (Santa Cruz, CA); Peyman Milanfar (Menlo Park, CA); Feng Yang (Sunnyvale, CA)
Assignee: Google LLC
H04N21/23418H04N19/154H04N21/4668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,739,450
App. No.
19/009,366
Granted
Sep 15, 2026
Kind
B2
Abstract

Techniques for determining perceptual quality indicators of video content items are provided. In some embodiments, a system including one or more processors executes instructions to: receive a video content item comprising a plurality of frames; determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame; determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame; generate a quality level for each frame based on at least the content quality indicator and the video distortion indicator for the frame; and output an indication of quality for the video content item based on the quality level for each frame.

Claims (62)

1 . A method comprising:

receiving, by a computing system, a video content item comprising a plurality of frames;

determining, by the computing system and using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;

determining, by the computing system and using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

generating, by the computing system, an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and

modifying, by the computing system and based on the indication of quality, the video content item or a transmission strategy associated with the video content item by at least:

adjusting a transmission bitrate for the video content item based on the indication of quality;

adjusting a transmission resolution for the video content item based on the indication of quality;

generating a compressed version of the video content item based on the indication of quality;

applying distortion correction to the video content item based on the indication of quality; or

transcoding the video content item from a first format to a second format based on the indication of quality.

2 . The method of claim 1 , further comprising:

determining, by the computing system and using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and

generating, by the computing system, a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.

3 . The method of claim 2 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.

4 . The method of claim 1 , further comprising:

determining, by the computing system and using the first subnetwork of the deep neural network, predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item; and

outputting, by the computing system, the predicted content labels.

5 . The method of claim 1 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item.

6 . The method of claim 2 , wherein generating the indication of quality comprises averaging the quality level for each frame of the plurality of frames.

7 . The method of claim 1 , further comprising causing, by the computing system, a video recommendation to be presented based on the indication of quality for the video content item.

8 . The method of claim 7 , wherein the video recommendation includes a recommendation to further compress the video content item based on the indication of quality for the video content item.

9 . The method of claim 7 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item.

10 . A system comprising:

a memory that stores instructions; and

one or more processors that execute the instructions to:

receive a video content item comprising a plurality of frames;

determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;

determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

generate an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and

modify, based on the indication of quality, the video content item or a transmission strategy associated with the video content item by executing the instructions to at least:

adjust a transmission bitrate for the video content item based on the indication of quality;

adjust a transmission resolution for the video content item based on the indication of quality;

generate a compressed version of the video content item based on the indication of quality;

apply distortion correction to the video content item based on the indication of quality; or

transcode the video content item from a first format to a second format based on the indication of quality.

11 . The system of claim 10 , wherein the one or more processors execute the instructions to:

determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and

generate a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.

12 . The system of claim 11 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.

13 . The system of claim 10 , wherein the one or more processors execute the instructions to:

determine, using the first subnetwork of the deep neural network, predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item; and

output the predicted content labels.

14 . The system of claim 10 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item.

15 . The system of claim 11 , wherein, to generate the indication of quality, the one or more processors execute the instructions to average the quality level for each frame of the plurality of frames.

16 . The system of claim 10 , wherein the one or more processors execute the instructions to cause a video recommendation to be presented based on the indication of quality for the video content item.

17 . The system of claim 16 , wherein the video recommendation includes a recommendation to further compress the video content item based on the indication of quality for the video content item.

18 . The system of claim 16 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item.

19 . Non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

receive a video content item comprising a plurality of frames;

determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator corresponds to one or more semantic content indicators for each frame of the plurality of frames of the video content item;

determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

generate an indication of quality for the video content item based on at least the video distortion indicator for one or more frames of the plurality of frames and the content quality indicator for one or more frames of the plurality of frames; and

modify, based on the indication of quality, the video content item or a transmission strategy associated with the video content item by executing the instructions to at least:

adjust a transmission bitrate for the video content item based on the indication of quality;

adjust a transmission resolution for the video content item based on the indication of quality;

generate a compressed version of the video content item based on the indication of quality;

apply distortion correction to the video content item based on the indication of quality; or

transcode the video content item from a first format to a second format based on the indication of quality.

20 . The non-transitory computer-readable storage media of claim 19 , wherein the instructions, when executed by the one or more processors, cause the one or more processors to:

determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; and

generate a quality level for each frame of the plurality of frames of the video content item based on at least the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for the frame, wherein generating the indication of quality is based on the quality level for each frame of the plurality of frames.

Continuity (3)
Continuation 18021636 · Jun 8, 2022
Provisional Application 63210003 · Jun 12, 2021
Related Publication 20250220251A1 · Jul 3, 2025
References Cited (48)
US 9741107B2 · Xu et al. · 2017 [cited by applicant]
US 11568637B2 · Zhou et al. · 2023 [cited by applicant]
US 11729387B1 · Khsib et al. · 2023 [cited by applicant]
US 11856203B1 · Tschannen et al. · 2023 [cited by applicant]
US 11895330B2 · Zhang et al. · 2024 [cited by applicant]
US 20170347159A1 · Baik et al. · 2017 [cited by applicant]
US 20200275016A1 · Citerin et al. · 2020 [cited by applicant]
US 20210158008A1 · Zhou et al. · 2021 [cited by applicant]
US 20220101629A1 · Liu et al. · 2022 [cited by applicant]
US 20230319327A1 · Wang et al. · 2023 [cited by applicant]
US 20230412808A1 · Holland et al. · 2023 [cited by applicant]
US 20240037802A1 · Solovyev et al. · 2024 [cited by applicant]
US 20240056570A1 · Li et al. · 2024 [cited by applicant]
CN 107454446A · 2017 [cited by applicant]
CN 108377387 · 2018 [cited by applicant]
CN 109215028A · 2019 [cited by applicant]
CN 109961434 · 2019 [cited by applicant]
CN 110853032A · 2020 [cited by applicant]
CN 111784694A · 2020 [cited by applicant]
WO 2020080698 · 2020 [cited by applicant]
WO 2020134926 · 2020 [cited by applicant]
Notice of Intent to Grant from counterpart Japanese Application No. 2023-558593 dated Mar. 21, 2025, 5 pp. [cited by applicant]
Notice of Intent to Grant from counterpart Korean Application No. 10-2023-7029021 dated Apr. 7, 2025, 10 pp. [cited by applicant]
Abu-El-Haija et al., “Youtube-8M: A Large-Scale Video Classification Benchmark”, in Google Research, Sep. 27, 2016, pp. 1-10. [cited by applicant]
Bosse et al., “Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment”, IEEE Transactions on Image Processing, vol. 27, No. 1, IEEE, Jan. 2018, pp. 206-219, Retrieved from the Internet on Dec.… [cited by applicant]
Choudhury, A., “Robust HDR Image Quality Assessment Using Combination of Quality Metrics”, in Multimedia Tools and Applications, vol. 79, May 31, 2020, pp. 22843-22867. [cited by applicant]
Fang, Y., et al., “Perceptual Quality Assessment of Smartphone Photography”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, Jun. 13-19, 2020, pp. 1-10. [cited by applicant]
Guo, Y., et al., “Deep Learning for Visual Understanding: A Review”, In Neurocomputing, vol. 187, Apr. 26, 2016 , pp. 27-48. [cited by applicant]
International Search Report and Written Opinion dated Sep. 5, 2022 in International Patent Application No. PCT/US2022/032668. [cited by applicant]
Li et al., “Quality Assessment of In-the-Wild Video”, in Proceedings of the ACM International Conference on Multimedia, Nice, France, Oct. 21-25, 2019, pp. 1-9. [cited by applicant]
Li, Y., et al., “User-generated Video Quality Assessment: A Subjective and Objective Study”, In Journal of Latex Class Files, vol. 14, No. 8, May 18, 2020, pp. 1-11. [cited by applicant]
Lin et al., “Kadid-10k: A Large-Scale Artificially Distorted IQA Database”, in the Proceedings of the International Conference on Quality of Multimedia Experience, Berlin, Germany, Jun. 5-7, 2019, pp. 1-3. [cited by applicant]
Ma et al., “End-to-End Blind Image Quality Assessment Using Deep Neural Networks”, in Transactions on Image Processing, vol. 27, No. 3, Mar. 2018, pp. 1202-1213. [cited by applicant]
Office Action, and translation thereof, from counterpart Japanese Application No. 2023-558593 dated Nov. 5, 2024, 6 pp. [cited by applicant]
Prosecution History from U.S. Appl. No. 18/021,636, dated May 7, 2024 through Dec. 18, 2024, 34 pp. [cited by applicant]
Response to Extended Search Report dated Aug. 18, 2023, from counterpart European Application No. 22743969.2 filed Feb. 14, 2024, 11 pp. [cited by applicant]
Russakovsky et al., “Imagenet Large Scale Visual Recognition Challenge”, in International Journal of Computer Vision, vol. 115, No. 3., Jan. 2015, pp. 1-43. [cited by applicant]
Stroud et al., “D3D: Distilled 3D Networks for Video Action Recognition”, in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Aspen, CO, USA, Mar. 1-5, 2020, pp. 1-10. [cited by applicant]
Van den Oord et al., “Representation Learning with Contrastive Predictive Coding”, Cornell University, Jan. 22, 2019, pp. 1-13. [cited by applicant]
Wang, Y., et al., “Video Transcoding Optimization Based on Input Perceptual Quality”, In Applications of Digital Image Processing XLIII, vol. 11510, Sep. 29, 2020, pp. 1-11. [cited by applicant]
Zadtootaghaj et al., “Demi: Deep Video Quality Estimation Model Using Perceptual Video Quality Dimensions”, Oct. 23, 2024, 9 pp. [cited by applicant]
Communication pursuant to Article 94(3) EPC from counterpart European Application No. 22743969.2 dated Jun. 6, 2025, 4 pp. [cited by applicant]
Response to Communication pursuant to Article 94(3) EPC dated Jun. 6, 2025, from counterpart European Application No. 22743969.2 filed Oct. 14, 2025, 8 pp. [cited by applicant]
Response to Office Action dated Jun. 11, 2025, from counterpart Indian Application No. 202347088663 filed Oct. 7, 2025, 18 pp. [cited by applicant]
Response to Office Action, and translation thereof, dated Nov. 5, 2024, from counterpart Japanese Application No. 2023-558593 filed Feb. 5, 2025, 12 pp. [cited by applicant]
First Examination Report from counterpart Indian Application No. 202347088663 dated Jun. 11, 2025, 6 pp. [cited by applicant]
First Office Action and Search Report, and translation thereof, from counterpart Chinese Application No. 202280015231.2 dated Apr. 3, 2026, 16 pp. [cited by applicant]
Saman et al., “DEMI: Deep Video Quality Estimation Model using Perceptual Video Quality Dimensions”, 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP), Dec 16, 2020, 6 pp. [cited by applicant]