IP Library › Granted Patent US 12,206,914
Granted Patent B2
US 12,206,914 · App. 18/021,636 · Granted Jan 21, 2025

Methods, systems, and media for determining perceptual quality indicators of video content items

Inventors: Yilin Wang (San Jose, CA); Balineedu Adsumilli (Foster City, CA); Junjie Ke (Stanford, CA); Hossein Talebi (San Francisco, CA); Joong Yim (Mountain View, CA); Neil Birkbeck (Santa Cruz, CA); Peyman Milanfar (Menlo Park, CA); Feng Yang (Sunnyvale, CA)
Assignee: Google LLC
H04N21/23418H04N19/154H04N21/4668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,206,914
App. No.
18/021,636
Granted
Jan 21, 2025
Kind
B2
Abstract

Methods, systems, and media for determining perceptual quality indicators of video content items are provided. In some embodiments, the method comprises: receiving a video content item; extracting a plurality of frames from the video content item; determining, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item; determining, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item; determining, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item; generating a quality level for each frame of the plurality of frames of the video content item that concatenates the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for that frame of the video content item; generating an overall quality level for video content item by aggregating the quality level of each frame of the plurality of frames; and causing a video recommendation to be presented based on the overall quality level of the video content item.

Claims (46)

1. A method for video quality assessment, the method comprising:

receiving a video content item;

extracting a plurality of frames from the video content item;

determining, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator represents semantic-level embeddings for each frame of the plurality of frames of the video content item;

determining, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

determining, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item;

generating a quality level for each frame of the plurality of frames of the video content item that concatenates the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for that frame of the video content item;

generating an overall quality level for video content item by aggregating the quality level of each frame of the plurality of frames; and

causing a video recommendation to be presented based on the overall quality level of the video content item.

2. The method of claim 1 , wherein the first subnetwork of the deep neural network further outputs predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item.

3. The method of claim 1 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item and wherein the second subnetwork of the deep neural network further outputs detected distortion types that describe distortion detected in each frame of the plurality of frames of the video content item.

4. The method of claim 1 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.

5. The method of claim 1 , wherein the overall quality level is generated using a convolutional neural network that outputs a respective quality level for each frame of the plurality of frames of the video content item and that averages the respective quality levels.

6. The method of claim 1 , wherein the video recommendation includes a recommendation to further compress the video content item based on the overall quality level.

7. The method of claim 1 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item based on the quality level associated with that frame of the video content item.

8. A system for video quality assessment, the system comprising:

a hardware processor that is configured to:

receive a video content item;

extract a plurality of frames from the video content item;

determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator represents semantic-level embeddings for each frame of the plurality of frames of the video content item;

determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item;

generate a quality level for each frame of the plurality of frames of the video content item that concatenates the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for that frame of the video content item;

generate an overall quality level for video content item by aggregating the quality level of each frame of the plurality of frames; and

cause a video recommendation to be presented based on the overall quality level of the video content item.

9. The system of claim 8 , wherein the first subnetwork of the deep neural network further outputs predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item.

10. The system of claim 8 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item and wherein the second subnetwork of the deep neural network further outputs detected distortion types that describe distortion detected in each frame of the plurality of frames of the video content item.

11. The system of claim 8 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.

12. The system of claim 8 , wherein the overall quality level is generated using a convolutional neural network that outputs a respective quality level for each frame of the plurality of frames of the video content item and that averages the respective quality levels.

13. The system of claim 8 , wherein the video recommendation includes a recommendation to further compress the video content item based on the overall quality level.

14. The system of claim 8 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item based on the quality level associated with that frame of the video content item.

15. A non-transitory computer-readable medium containing computer executable instructions that, when executed by a processor, cause the processor to:

receive a video content item;

extract a plurality of frames from the video content item;

determine, using a first subnetwork of a deep neural network, a content quality indicator for each frame of the plurality of frames of the video content item, wherein the content quality indicator represents semantic-level embeddings for each frame of the plurality of frames of the video content item;

determine, using a second subnetwork of the deep neural network, a video distortion indicator for each frame of the plurality of frames of the video content item, wherein the video distortion indicator indicates a quality of the frame based on distortions contained within the frame;

determine, using a third subnetwork of the deep neural network, a compression sensitivity indicator for each frame of the plurality of frames of the video content item;

generate a quality level for each frame of the plurality of frames of the video content item that concatenates the content quality indicator, the video distortion indicator, and the compression sensitivity indicator for that frame of the video content item;

generate an overall quality level for video content item by aggregating the quality level of each frame of the plurality of frames; and

cause a video recommendation to be presented based on the overall quality level of the video content item.

16. The non-transitory computer-readable medium of claim 15 , wherein the first subnetwork of the deep neural network further outputs predicted content labels that describe content appearing in each frame of the plurality of frames of the video content item.

17. The non-transitory computer-readable medium of claim 15 , wherein the video distortion indicator further represents distortion-sensitive embeddings for each frame of the plurality of frames of the video content item and wherein the second subnetwork of the deep neural network further outputs detected distortion types that describe distortion detected in each frame of the plurality of frames of the video content item.

18. The non-transitory computer-readable medium of claim 15 , wherein the compression sensitivity indicator represents compression-sensitive embedding features for each frame of the plurality of frames of the video content item and wherein the third subnetwork of the deep neural network further outputs a compression level score.

19. The non-transitory computer-readable medium of claim 15 , wherein the overall quality level is generated using a convolutional neural network that outputs a respective quality level for each frame of the plurality of frames of the video content item and that averages the respective quality levels.

20. The non-transitory computer-readable medium of claim 15 , wherein the video recommendation includes a recommendation to further compress the video content item based on the overall quality level.

21. The non-transitory computer-readable medium of claim 15 , wherein the video recommendation includes a recommendation to an uploader of the video content item to modify a portion of the video content item based on the quality level associated with that frame of the video content item.

Continuity (2)
Provisional Application 63210003 · Jun 12, 2021
Related Publication 20230319327A1 · Oct 5, 2023
References Cited (31)
US 9741107B2 · Xu et al. · 2017 [cited by applicant]
US 11729387B1 · Khsib · 2023 [cited by examiner]
US 11856203B1 · Tschannen · 2023 [cited by examiner]
US 11895330B2 · Zhang · 2024 [cited by examiner]
US 20200275016A1 · Citerin et al. · 2020 [cited by applicant]
US 20210158008A1 · Zhou et al. · 2021 [cited by applicant]
US 20220101629A1 · Liu et al. · 2022 [cited by applicant]
US 20230412808A1 · Holland · 2023 [cited by examiner]
US 20240037802A1 · Solovyev · 2024 [cited by examiner]
US 20240056570A1 · Li · 2024 [cited by examiner]
CN 108377387 · 2018 [cited by applicant]
CN 109961434 · 2019 [cited by applicant]
WO WO2020080698 · 2020 [cited by applicant]
WO WO2020134926 · 2020 [cited by applicant]
Abu-El-Haija et al., “Youtube-8M: A Large-Scale Video Classification Benchmark”, in Google Research, Sep. 27, 2016, pp. 1-10. [cited by applicant]
Choudhury, A., “Robust HDR Image Quality Assessment Using Combination of Quality Metrics”, in Multimedia Tools and Applications, vol. 79, May 31, 2020, pp. 22843-22867. [cited by applicant]
Fang, Y., et al., “Perceptual Quality Assessment of Smartphone Photography”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, Jun. 13-19, 2020, pp. 1-10. [cited by applicant]
Guo, Y., et al., “Deep Learning For Visual Understanding: A Review”, In Neurocomputing, vol. 187, Apr. 26, 2016, pp. 27-48. [cited by applicant]
International Search Report and Written Opinion dated Sep. 5, 2022 in International Patent Application No. PCT/US2022/032668. [cited by applicant]
Li et al., “Quality Assessment of In-the-Wild Video”, in Proceedings of the ACM International Conference on Multimedia, Nice, France, Oct. 21-25, 2019, pp. 1-9. [cited by applicant]
Li, Y., et al., “User-generated Video Quality Assessment: A Subjective and Objective Study”, In Journal of Latex Class Files, vol. 14, No. 8, May 18, 2020, pp. 1-11. [cited by applicant]
Lin et al., “Kadid-10k: A Large-Scale Artificially Distorted IQA Database”, in the Proceedings of the International Conference on Quality of Multimedia Experience, Berlin, Germany, Jun. 5-7, 2019, pp. 1-3. [cited by applicant]
Ma et al., “End-to-End Blind Image Quality Assessment Using Deep Neural Networks”, in Transactions on Image Processing, vol. 27, No. 3, Mar. 2018, pp. 1202-1213. [cited by applicant]
Russakovsky et al., “Imagenet Large Scale Visual Recognition Challenge”, in International Journal of Computer Vision, vol. 115, No. 3., Jan. 2015, pp. 1-43. [cited by applicant]
Stroud et al., “D3D: Distilled 3D Networks for Video Action Recognition”, in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Aspen, CO, USA, Mar. 1-5, 2020, pp. 1-10. [cited by applicant]
van den Oord et al., “Representation Learning with Contrastive Predictive Coding”, Cornell University, Jan. 22, 2019, pp. 1-13. [cited by applicant]
Wang, Y., et al., “Video Transcoding Optimization Based on Input Perceptual Quality”, In Applications of Digital Image Processing XLIII, vol. 11510, Sep. 29, 2020, pp. 1-11. [cited by applicant]
Response to Extended Search Report dated Aug. 18, 2023, from counterpart European Application No. 22743969.2 filed Feb. 14, 2024, 11 pp. [cited by applicant]
Bosse et al., “Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment”, IEEE Transactions on Image Processing, vol. 27, No. 1, IEEE, Jan. 2018, pp. 206-219, Retrieved from the Internet on Dec.… [cited by applicant]
Office Action, and translation thereof, from counterpart Japanese Application No. 2023-558593 dated Nov. 5, 2024, 6 pp. [cited by applicant]
Zadtootaghaj et al., “DEMI: Deep Video Quality Estimation Model Using Perceptual Video Quality Dimensions”, Oct. 23, 2024, 9 pp. [cited by applicant]