IP Library Granted Patent US 12,445,656
Granted Patent B2
US 12,445,656 · App. 17/698,116 · Granted Oct 14, 2025

Guided restoration of video data using neural networks

Inventors: Debargha Mukherjee (Cupertino, CA); Urvang Joshi (Mountain View, CA); Yue Chen (Kirkland, WA); Sarah Parker (San Francisco, CA)
Assignee: GOOGLE LLC
H04N19/82G06F18/214G06N3/045G06N3/08G06N20/20G06T3/40G06T5/50G06T5/60G06T9/002H04N19/176H04N19/70G06T2207/20081G06T2207/20084H04N19/117H04N19/17
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,656
App. No.
17/698,116
Granted
Oct 14, 2025
Kind
B2
Abstract

Guided restoration is used to restore video data degraded from a video frame. The video frame is divided into restoration units (RUs) which each correspond to one or more blocks of the video frame. Restoration schemes are selected for each RU. The restoration schemes may indicate to use one of a plurality of neural networks trained for the guided restoration. Alternatively, the restoration schemes may indicate to use a neural network and a filter-based restoration tool. The video frame is then restored by processing each RU according to the respective selected restoration scheme. During encoding, the restored video frame is encoded to an output bitstream, and the use of the selected restoration schemes may be signaled within the output bitstream. During decoding, the restored video frame is output to an output video stream.

Claims (50)

1. A method, comprising:

producing, during a frame decoding process, a reconstructed video frame including degraded video data by dequantizing, inverse transforming, predicting, and loop filtering encoded video data associated with an encoded video frame;

performing a guided restoration process after the loop filtering and during the frame decoding process, wherein performing the guided restoration process includes:

determining a first restoration scheme for restoring a first portion of the degraded video data based on information associated with the first portion, wherein the first restoration scheme indicates to use a first neural network trained on a first value range corresponding to one of a first quantization parameter range or a first non-zero quantized transform coefficient range;

determining a second restoration scheme for restoring a second portion of the degraded video data based on information associated with the second portion, wherein the second restoration scheme indicates to use a second neural network trained on a second value range corresponding to one of a second quantization parameter range or a second non-zero quantized transform coefficient range; and

producing a restored video frame by processing the first portion using the first neural network and by processing the second portion processing the second portion using the second neural network; and

outputting the restored video frame for storage or display.

2. The method of claim 1 , wherein the first restoration scheme is determined for the first portion based on the information associated with the first portion corresponding to one or more quantization parameters within the first quantization parameter range, and wherein the second restoration scheme is determined for the second portion based on the information associated with the second portion corresponding to one or more quantization parameters within the second quantization parameter range.

3. The method of claim 1 , wherein the first restoration scheme is determined for the first portion based on the information associated with the first portion corresponding to one or more non-zero quantized transform coefficients within the first non-zero quantized transform coefficient range, and wherein the second restoration scheme is determined for the second portion based on the information associated with the second portion corresponding to one or more non-zero quantized transform coefficients within the second non-zero quantized transform coefficient range.

4. The method of claim 1 , comprising:

training the first neural network using a first training data set including a number of first data samples and the second neural network using a second training data set including a number of second data samples, wherein each data sample of the first data samples and of the second data samples includes a degraded block and a corresponding original block.

5. The method of claim 1 , comprising:

decoding one or more syntax elements indicative of the first restoration scheme and the second restoration scheme from a bitstream including the encoded video frame, wherein the first restoration scheme and the second restoration scheme are determined based on the one or more syntax elements.

6. The method of claim 1 , wherein the first portion corresponds to a first restoration unit of the reconstructed video frame and the second portion corresponds to a second restoration unit of the reconstructed video frame.

7. The method of claim 6 , comprising:

dividing the reconstructed video frame into restoration units including the first restoration unit and the second restoration unit.

8. An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

produce a reconstructed video frame by dequantizing, inverse transforming, predicting, and loop filtering encoded video data associated with an encoded video frame;

perform, after the loop filtering, a guided restoration process to:

divide the reconstructed video frame into multiple portions including a first portion and a second portion;

determine a first restoration scheme for restoring degraded video data associated with the first portion, wherein the first restoration scheme indicates to use a first neural network trained on a first value range corresponding to one of a first quantization parameter range or a first non-zero quantized transform coefficient range;

determine a second restoration scheme for restoring degraded video data associated with the second portion, wherein the second restoration scheme indicates to use a second neural network trained on a second value range corresponding to one of a second quantization parameter range or a second non-zero quantized transform coefficient range; and

produce a restored video frame by processing the first portion using the first neural network and the second portion using the second neural network; and

output the restored video frame for storage or display.

9. The apparatus of claim 8 , wherein the instructions include instructions to:

decode one or more syntax elements indicative of the first restoration scheme and the second restoration scheme from a bitstream from which the encoded video frame is decoded.

10. The apparatus of claim 8 , wherein the multiple portions are restoration units each corresponding to one or more blocks of the reconstructed video frame.

11. The apparatus of claim 10 , wherein sizes of the restoration units are based on one of a configuration of a decoder, a non-block partition of the reconstructed video frame, or a size of a largest block within the reconstructed video frame.

12. The apparatus of claim 8 , wherein the processor is configured to execute the instructions to:

train the first neural network using a first training data set including a number of first data samples and the second neural network using a second training data set including a number of second data samples.

13. The apparatus of claim 12 , wherein each data sample of the first data samples and of the second data samples includes a degraded block and a corresponding original block.

14. A method, comprising:

producing a reconstructed video frame by dequantizing, inverse transforming, predicting, and loop filtering encoded video data associated with an encoded video frame;

performing, after the loop filtering, a guided restoration process including:

determining restoration schemes for restoring degraded video data of separate portions of the reconstructed video frame based on information associated with the separate portions, wherein each restoration scheme indicates to use a different neural network trained on a different value range each corresponding to one of a quantization range or a non-zero quantized transform coefficient range; and

producing a restored video frame by processing each portion of the separate portions using a neural network associated with a respective one of the restoration schemes determined for the portion; and

outputting the restored video frame for storage or display.

15. The method of claim 14 , wherein determining the restoration schemes comprises:

determining a first restoration scheme for a first portion of the separate portions based on the information associated with the first portion being within a first value range of a first neural network of the neural networks; and

determining a second restoration scheme for a second portion of the separate portions based on the information associated with the second portion being within a second value range of a second neural network of the neural networks.

16. The method of claim 14 , wherein determining the restoration schemes comprises:

decoding one or more syntax elements indicative of the restoration schemes from a bitstream that includes the encoded video frame.

17. The method of claim 14 , wherein the separate portions correspond to different restoration units of the reconstructed video frame.

18. The method of claim 17 , comprising:

dividing the reconstructed video frame into the different restoration units.

19. The method of claim 17 , wherein sizes of the different restoration units are based on one of a configuration of a decoder, a non-block partition of the reconstructed video frame, or a size of a largest block within the reconstructed video frame.

20. The method of claim 14 , comprising:

training each different neural network using a training data set including a number of data samples including degraded blocks and corresponding original blocks.

Continuity (3)
Continuation 16515226 · Jul 18, 2019
Provisional Application 62778266 · Dec 11, 2018
Related Publication 20220207654A1 · Jun 30, 2022
References Cited (34)
US 8819525B1 · Holmer · 2014 [cited by examiner]
US 10652565B1 · Zhang · 2020 [cited by examiner]
US 11095887B2 · Kim · 2021 [cited by examiner]
US 20180295320A1 · Breternitz · 2018 [cited by examiner]
US 20180349759A1 · Isogawa · 2018 [cited by examiner]
US 20190373276A1 · Hu · 2019 [cited by examiner]
US 20200244997A1 · Galpin · 2020 [cited by examiner]
US 20210099710A1 · Salehifar · 2021 [cited by examiner]
WO 2017036370A1 · 2017 [cited by applicant]
Bankoski et al., “VP8 Data Format and Decoding Guide”, Independent Submission RFC 6389, Nov. 2011, 305 pp. [cited by applicant]
Bankoski et al., “VP8 Data Format and Decoding Guide draft-bankoski-vp8-bitstream-02”, Network Working Group, Internet-Draft, May 18, 2011, 288 pp. [cited by applicant]
“Introduction to Video Coding Part 1: Transform Coding”, Mozilla, Mar. 2012, 171 pp. [cited by applicant]
“Overview VP7 Data Format and Decoder”, Version 1.5, On2 Technologies, Inc., Mar. 28, 2005, 65 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Amendment 2: New profiles for professional applications, International Telecommunication Union, Apr. 2007, 75 … [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Amendment 1: Support of additional colour spaces and r… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 3, International Telecommunication Union, Mar.… [cited by applicant]
Bankoski, et al., “Technical Overview of VP8, An Open Source Video Codec for the Web”, Jul. 11, 2011, 6 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Coding of moving video: Implementors Guide for H.264: Advanced video coding for generic audiovisual services, International Telecommunication Union, Jul. 30, 2010, 15 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, International Telecommunication Union, Version 11, Mar… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, International Telecommunication Union, Version 12, Mar… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 8, International Telecommunication Union, Nov.… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 1, International Telecommunication Union, May … [cited by applicant]
“VP8 Data Format and Decoding Guide, WebM Project”, Google On2, Dec. 1, 2010, 103 pp. [cited by applicant]
“VP6 Bitstream and Decoder Specification”, Version 1.02, On2 Technologies, Inc., Aug. 17, 2006, 88 pp. [cited by applicant]
“VP6 Bitstream and Decoder Specification”, Version 1.03, On2 Technologies, Inc., Oct. 29, 2007, 95 pp. [cited by applicant]
Rivaz et al.; “AV1 Bitstream & Decoding Process Specification”; Jan. 2019; pp. 1-669. [cited by applicant]
Lu et al. “Deep Kalman Filtering Network for Video Compression Artifact Reduction”; ECCV 2018; 17 pages. [cited by applicant]
Yu et al.; “Deep Convolution Networks for Compression Artifacts Reduction”; Aug. 2016; pp. 1-13. [cited by applicant]
International Search Report and Written Opinion of International Application No. PCT/US2019/059019 dated Feb. 17, 2020. [cited by applicant]
Zhang et al; “Residual Highway Convolutional Neural Networks for In-Loop Filtering in HEVC” IEEE Transactions on Image Processing; IEEE Services Center; vol. 27, No. 8, Aug. 2018; pp. 3827-3841. [cited by applicant]
Wang et al; “AHG9: Dense Residual Convolution Neural Network Based In-Loop Filter”; JVET Meeting; Jul. 14, 2018; 6 Pages. [cited by applicant]
Park et al; CNN-Based in-loop filtering for coding efficency improvement; 2016 IEEE 12 Image, Video and Multidimentional signal processing workshop; Jul. 11, 2016; pp. 1-5. [cited by applicant]
Jia Chuanmin et al; “Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video Coding” IEEE Transactions on Image Processing; vol. 28, No. 7; Jul. 1, 2019; pp. 3343-3356. [cited by applicant]
Anonymous; “An Intuitive Explanation of Convolutional Neural Networks—The Data Science Blog”; May 29, 2017; p. 13. [cited by applicant]