IP Library Granted Patent US 12,445,657
Granted Patent B1
US 12,445,657 · App. 18/383,798 · Granted Oct 14, 2025

Deriving in-loop filter parameters for video coding

Inventors: Yi Guo (Hangzhou, CN); Zhichu He (Hangzhou, CN); Rui Li (Hangzhou, CN); Bo Ling (Saratoga, CA); Jing Wu (Hangzhou, CN); Minxia Yang (Hangzhou, CN); Shiyan Zhang (Hangzhou, CN); Yichen Zhang (Hangzhou, CN)
Assignee: Zoom Communications, Inc.
H04N19/82H04N19/117H04N19/503
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,445,657
App. No.
18/383,798
Granted
Oct 14, 2025
Kind
B1
Abstract

Deriving in-loop filter parameters via training for video encoding is provided. A video encoder performs inter prediction for a frame in a set of frames of the video to generate prediction residuals for the frame. The inter prediction for the frame is performed based on a reconstructed frame in the set of frames filtered using an in-loop filter. The value of a parameter of the in-loop filter is determined by determining, for each candidate in-loop filter parameter value, a visual quality metric for a set of training frames in training video sequences filtered by the in-loop filter. The candidate in-loop filter parameter value that corresponds the highest visual quality metric can be selected as the value of the parameter of the in-loop filter. The video encoder further encodes the prediction residues of the frame and the parameter of the in-loop filter into a bitstream representing the video.

Claims (43)

1. A method for encoding a video, the method comprising:

accessing a plurality of frames of the video;

performing inter prediction for a frame in the plurality of frames to generate prediction residuals for the frame, wherein the inter prediction for the frame is performed based on a reconstructed frame in the plurality of frames filtered using an in-loop filter, wherein determining a value of a parameter of the in-loop filter comprises:

for each candidate in-loop filter parameter value among a plurality of candidate in-loop filter parameter values, determining a visual quality metric for a plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value, and

selecting a candidate in-loop filter parameter value among the plurality of candidate in-loop filter parameter values that corresponds a visual quality metric higher than another visual quality metric as the value of the parameter of the in-loop filter; and

encoding the prediction residues of the frame and the parameter of the in-loop filter into a bitstream representing the video.

2. The method of claim 1 , wherein determining the visual quality metric for the plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value comprises:

for each training frame in the plurality of training frames, determining a frame-level visual quality metric for the training frame filtered by the in-loop filter with the value of the candidate in-loop filter parameter; and

combining the frame-level visual quality metrics for the respective training frames to generate the visual quality metric.

3. The method of claim 1 , wherein the value of the parameter of the in-loop filter is determined for a corresponding quantization parameter used in encoding the frame.

4. The method of claim 1 , wherein the reconstructed frame is an intra-coded frame (I-frame) and the plurality of training frames comprise I-frames.

5. The method of claim 1 , wherein the reconstructed frame filtered using the in-loop filter is a predicted frame (P-frame), and wherein determining the value of the parameter of the in-loop filter further comprises determining a value of the parameter of the in-loop filter for P-frames based on the selected candidate in-loop filter parameter value.

6. The method of claim 1 , wherein the value of the parameter of the in-loop filter for a Y component of the video is determined separately from the value of the parameter for a U component of the video or a V component of the video.

7. The method of claim 1 , wherein the visual quality metric comprises a peak signal to noise ratio (PSNR).

8. A system comprising:

a processor; and

at least one memory device including instructions that are executable by the processor to cause the processor to:

access a plurality of frames of a video;

perform inter prediction for a frame in the plurality of frames to generate prediction residuals for the frame, wherein the inter prediction for the frame is performed based on a reconstructed frame in the plurality of frames filtered using an in-loop filter, wherein determining a value of a parameter of the in-loop filter comprises:

for each candidate in-loop filter parameter value among a plurality of candidate in-loop filter parameter values, determine a visual quality metric for a plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value, and

select a candidate in-loop filter parameter value among the plurality of candidate in-loop filter parameter values that corresponds a visual quality metric higher than another visual quality metric as the value of the parameter of the in-loop filter; and

encode the prediction residues of the frame and the parameter of the in-loop filter into a bitstream representing the video.

9. The system of claim 8 , wherein determining the visual quality metric for the plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value comprises:

for each training frame in the plurality of training frames, determining a frame-level visual quality metric for the training frame filtered by the in-loop filter with the value of the candidate in-loop filter parameter; and

combining the frame-level visual quality metrics for the respective training frames to generate the visual quality metric.

10. The system of claim 8 , wherein the value of the parameter of the in-loop filter is determined for a corresponding quantization parameter used in encoding the frame.

11. The system of claim 8 , wherein the reconstructed frame is an intra-coded frame (I-frame) and the plurality of training frames comprise I-frames.

12. The system of claim 8 , wherein the reconstructed frame filtered using the in-loop filter is a predicted frame (P-frame), and wherein determining the value of the parameter of the in-loop filter further comprises determining a value of the parameter of the in-loop filter for P-frames based on the selected candidate in-loop filter parameter value.

13. The system of claim 8 , wherein the value of the parameter of the in-loop filter for a Y component of the video is determined separately from the value of the parameter for a U component of the video or a V component of the video.

14. The system of claim 8 , wherein the visual quality metric comprises a peak signal to noise ratio (PSNR).

15. A non-transitory computer-readable medium comprising program code that is executable by one or more processors to cause the one or more processors to:

access a plurality of frames of a video;

perform inter prediction for a frame in the plurality of frames to generate prediction residuals for the frame, wherein the inter prediction for the frame is performed based on a reconstructed frame in the plurality of frames filtered using an in-loop filter, wherein determining a value of a parameter of the in-loop filter comprises:

for each candidate in-loop filter parameter value among a plurality of candidate in-loop filter parameter values, determine a visual quality metric for a plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value, and

select a candidate in-loop filter parameter value among the plurality of candidate in-loop filter parameter values that corresponds a visual quality metric higher than another visual quality metric as the value of the parameter of the in-loop filter; and

encode the prediction residues of the frame and the parameter of the in-loop filter into a bitstream representing the video.

16. The non-transitory computer-readable medium of claim 15 , wherein determining the visual quality metric for the plurality of training frames in one or more training video sequences filtered by the in-loop filter with the candidate in-loop filter parameter value comprises:

for each training frame in the plurality of training frames, determining a frame-level visual quality metric for the training frame filtered by the in-loop filter with the value of the candidate in-loop filter parameter; and

combining the frame-level visual quality metrics for the respective training frames to generate the visual quality metric.

17. The non-transitory computer-readable medium of claim 15 , wherein the value of the parameter of the in-loop filter is determined for a corresponding quantization parameter used in encoding the frame.

18. The non-transitory computer-readable medium of claim 15 , wherein the reconstructed frame is an intra-coded frame (I-frame) and the plurality of training frames comprise I-frames.

19. The non-transitory computer-readable medium of claim 15 , wherein the reconstructed frame filtered using the in-loop filter is a predicted frame (P-frame), and wherein determining the value of the parameter of the in-loop filter further comprises determining a value of the parameter of the in-loop filter for P-frames based on the selected candidate in-loop filter parameter value.

20. The non-transitory computer-readable medium of claim 15 , wherein the value of the parameter of the in-loop filter for a Y component of the video is determined separately from the value of the parameter for a U component of the video or a V component of the video.

Assignments (2)
CHANGE OF NAME Recorded Sep 17, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 072914/0370 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2024
From: LI, RUI; ZHANG, SHIYAN; GUO, YI; LING, BO; HE, ZHICHU; WU, JING; ZHANG, YICHEN; YANG, MINXIA
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 067480/0664 →
Continuity (1)
Provisional Application 63528230 · Jul 21, 2023
References Cited (11)
US 20140192869A1 · Laroche · 2014 [cited by examiner]
US 20190007680A1 · Chen · 2019 [cited by examiner]
US 20190052877A1 · Zhang · 2019 [cited by examiner]
US 20190141339A1 · Madajczak · 2019 [cited by examiner]
US 20190320196A1 · Yu · 2019 [cited by examiner]
US 20200351533A1 · Bampis · 2020 [cited by examiner]
US 20210409744A1 · Kawamura · 2021 [cited by examiner]
US 20230412800A1 · Wu · 2023 [cited by examiner]
“Alliance for Open Media”, Retrieved from internet on Oct. 17, 2023 from: https://aomedia.googlesource.com/aom, 12 pages. [cited by applicant]
Lei, et al., “GPGPU Implementation of VP9 In-Loop Deblocking Filter and Improvements for AV1 Codec”, ICIP 2017, 2017, pp. 925-929. [cited by applicant]
Zimichev, “BD-rate: one name—two metrics. AOM vs. the World.”, Retrieved from internet on Oct. 17, 2023 from: https://vicuesoft.com/blog/titles/bd_rate_one_name_two_metrics/, 9 pages. [cited by applicant]