IP Library › Granted Patent US 12,726,618
Granted Patent B2
US 12,726,618 · App. 18/925,857 · Granted Sep 1, 2026

Multimedia data processing method and apparatus, device, storage medium, and program product

Inventor: Liqiang Wang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
H04N19/117H04N19/176H04N19/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,726,618
App. No.
18/925,857
Granted
Sep 1, 2026
Kind
B2
Abstract

A multimedia data processing method, executed by a computer device, includes: determining a first associated image block associated with a target image block to be filtered in multimedia data; acquiring target coding and decoding information associated with the target image block; and filtering the target image block by an image filter based on a neural network according to the first associated image block and the target coding and decoding information to obtain a filtered image block corresponding to the target image block.

Claims (88)

1 . A multimedia data processing method executed by a computer device, comprising:

determining a first associated image block associated with a target image block to be filtered in multimedia data;

acquiring target coding and decoding information associated with the target image block; and

filtering the target image block by:

inputting the target image block, the first associated image block, and the target coding and decoding information into an image filter based on a neural network, wherein the neural network is trained to minimize a loss function measuring a difference between a true value comprising an original image block and a predicted filtered image block output by the image filter based on: (i) an input target image block, (ii) an input associated image block, and (iii) input target coding and decoding information,

fusing the first associated image block with first coding and decoding information corresponding to the first associated image block by an information fusion layer of the image filter to obtain target fusion data corresponding to the first associated image block, and

filtering the target image block according to the target fusion data by a filtering layer of the image filter to obtain a filtered image block corresponding to the target image block.

2 . The method according to claim 1 , wherein the filtering the target image block comprises:

fusing the first associated image block with first coding and decoding information corresponding to the first associated image block by an information fusion layer of the image filter to obtain target fusion data corresponding to the first associated image block; and

filtering the target image block according to the target fusion data and second coding and decoding information corresponding to the target image block by a filtering layer of the image filter to obtain the filtered image block.

3 . The method according to claim 2 , wherein the second coding and decoding information comprises at least one of: a sequence level quantization parameter, a slice level quantization parameter, a slice level coding type, a block coding type, filtering intensity image block information corresponding to the target image block, predicted image block information corresponding to the target image block, or partitioned image block information of the target image block.

4 . The method according to claim 3 , wherein image block information corresponding to the target image block comprises at least one of first color component information, second color component information, or third color component information,

wherein the first color component information indicates brightness of an image block,

wherein the second color component information and the third color component information indicate chrominance of the image block, and

wherein the image block information comprises the filtering intensity image block information, the partitioned image block information, or the predicted image block information.

5 . The method according to claim 1 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1,

wherein the image filter comprises M information fusion layers,

wherein one information fusion layer corresponds to one associated image block,

wherein the fusing the first associated image block comprises:

performing a convolution operation on a first associated image block i and second coding and decoding information corresponding to the first associated image block i by an information fusion layer i of the image filter to obtain first fusion data corresponding to the first associated image block i; and

determining second fusion data corresponding to M first associated image blocks respectively as the target fusion data based on the second fusion data being acquired,

wherein the first associated image block i belongs to M first associated image blocks,

wherein i is a positive integer less than or equal to M, and

wherein the information fusion layer i is a first information fusion layer corresponding to the first associated image block i in the M information fusion layers.

6 . The method according to claim 1 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1, and

wherein the fusing the first associated image block comprises performing a convolution operation on M first associated image blocks and M coding and decoding information by the information fusion layer to obtain the target fusion data.

7 . The method according to claim 1 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1,

wherein the image filter comprises M information fusion layers,

wherein one information fusion layer corresponds to one associated image block,

wherein the fusing the first associated image block comprises:

performing a point multiplication operation on a first associated image block i and second coding and decoding information corresponding to the first associated image block i by an information fusion layer i of the image filter to obtain first fusion data corresponding to the first associated image block i; and

performing a convolution operation on second fusion data corresponding to M first associated image blocks respectively to obtain the target fusion data based on the second fusion data being acquired,

wherein the first associated image block i belongs to M first associated image blocks,

wherein i is a positive integer less than or equal to M, and

wherein the information fusion layer i is a first information fusion layer corresponding to the first associated image block i in the M information fusion layers.

8 . The method according to claim 1 , wherein the target coding and decoding information indicates a degree of influence of the first associated image block on the target image block, and

wherein the target coding and decoding information comprises at least one of an influence factor of the first associated image block on the target image block, a reference direction of the first associated image block relative to the target image block, partitioned image block information of the first associated image block, a quantization parameter corresponding to the first associated image block, reconstructed image block information corresponding to the first associated image block, filtering intensity image block information corresponding to the first associated image block, or predicted image block information corresponding to the first associated image block, and

wherein the influence factor is determined according to an image distance between the first associated image block and the target image block.

9 . The method according to claim 8 , wherein the method further comprises:

determining the image distance between the first associated image block and the target image block according to a picture order count (POC) corresponding to the first associated image block and a POC corresponding to the target image block based on a slice in which the target image block is located being a non-full intraframe coding slice; and

determining a preset distance value as the image distance between the first associated image block and the target image block based on the slice in which the target image block is located being a full intraframe coding slice.

10 . A multimedia data processing apparatus, comprising:

at least one memory configured to store computer program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

determination code configured to cause at least one of the at least one processor to determine a first associated image block associated with a target image block to be filtered in multimedia data;

acquisition code configured to cause at least one of the at least one processor to acquire target coding and decoding information of the first associated image block associated with the target image block; and

first filtering code configured to cause at least one of the at least one processor to filter the target image block by:

inputting the target image block, the first associated image block, and the target coding and decoding information into an image filter based on a neural network, wherein the neural network is trained to minimize a loss function measuring a difference between a true value comprising an original image block and a predicted filtered image block output by the image filter based on: (i) an input target image block, (ii) an input associated image block, and (iii) input target coding and decoding information,

fusing the first associated image block with first coding and decoding information corresponding to the first associated image block by an information fusion layer of the image filter to obtain target fusion data corresponding to the first associated image block, and

filtering the target image block according to the target fusion data by a filtering layer of the image filter to obtain a filtered image block corresponding to the target image block.

11 . The apparatus according to claim 10 , wherein the first filtering code comprises first fusing code and second filtering code:

wherein the first fusing code is configured to cause at least one of the at least one processor to fuse the first associated image block with first coding and decoding information corresponding to the first associated image block by an information fusion layer of the image filter to obtain target fusion data corresponding to the first associated image block; and

wherein the second filtering code is configured to cause at least one of the at least one processor to filter the target image block according to the target fusion data and second coding and decoding information corresponding to the target image block by a filtering layer of the image filter to obtain the filtered image block.

12 . The apparatus according to claim 11 , wherein the second coding and decoding information comprises at least one of: a sequence level quantization parameter, a slice level quantization parameter, a slice level coding type, a block coding type, filtering intensity image block information corresponding to the target image block, predicted image block information corresponding to the target image block, or partitioned image block information of the target image block.

13 . The apparatus according to claim 12 , wherein image block information corresponding to the target image block comprises at least one of first color component information, second color component information, or third color component information,

wherein the first color component information indicates brightness of an image block,

wherein the second color component information and the third color component information indicate chrominance of the image block, and

wherein the image block information comprises the filtering intensity image block information, the partitioned image block information, or the predicted image block information.

14 . The apparatus according to claim 10 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1,

wherein the image filter comprises M information fusion layers,

wherein one information fusion layer corresponds to one associated image block,

wherein the first fusing code is configured to cause at least one of the at least one processor to:

perform a convolution operation on a first associated image block i and second coding and decoding information corresponding to the first associated image block i by an information fusion layer i of the image filter to obtain first fusion data corresponding to the first associated image block i; and

determine second fusion data corresponding to M first associated image blocks respectively as the target fusion data based on the second fusion data being acquired,

wherein the first associated image block i belongs to M first associated image blocks,

wherein i is a positive integer less than or equal to M, and

wherein the information fusion layer i is a first information fusion layer corresponding to the first associated image block i in the M information fusion layers.

15 . The apparatus according to claim 10 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1, and

wherein the first fusing code is configured to cause at least one of the at least one processor to perform a convolution operation on M first associated image blocks and M coding and decoding information by the information fusion layer to obtain the target fusion data.

16 . The apparatus according to claim 10 , wherein a number of the first associated image block is M, and M is an integer greater than or equal to 1,

wherein the image filter comprises M information fusion layers,

wherein one information fusion layer corresponds to one associated image block,

wherein the first fusing code is configured to cause at least one of the at least one processor to:

perform a point multiplication operation on a first associated image block i and second coding and decoding information corresponding to the first associated image block i by an information fusion layer i of the image filter to obtain first fusion data corresponding to the first associated image block i; and

perform a convolution operation on second fusion data corresponding to M first associated image blocks respectively to obtain the target fusion data based on the second fusion data being acquired,

wherein the first associated image block i belongs to M first associated image blocks,

wherein i is a positive integer less than or equal to M, and

wherein the information fusion layer i is a first information fusion layer corresponding to the first associated image block i in the M information fusion layers.

17 . The apparatus according to claim 10 , wherein the target coding and decoding information indicates a degree of influence of the first associated image block on the target image block, and

wherein the target coding and decoding information comprises at least one of an influence factor of the first associated image block on the target image block, a reference direction of the first associated image block relative to the target image block, partitioned image block information of the first associated image block, a quantization parameter corresponding to the first associated image block, reconstructed image block information corresponding to the first associated image block, filtering intensity image block information corresponding to the first associated image block, or predicted image block information corresponding to the first associated image block, and

wherein the influence factor is determined according to an image distance between the first associated image block and the target image block.

18 . A non-transitory computer-readable storage medium, storing computer code which, when executed by at least one processor, causes the at least one processor to at least:

determine a first associated image block associated with a target image block to be filtered in multimedia data;

acquire target coding and decoding information of the first associated image block associated with the target image block; and

filter the target image block by:

inputting the target image block, the first associated image block, and the target coding and decoding information into an image filter based on a neural network, wherein the neural network is trained to minimize a loss function measuring a difference between a true value comprising an original image block and a predicted filtered image block output by the image filter based on: (i) an input target image block, (ii) an input associated image block, and (iii) input target coding and decoding information,

fusing the first associated image block with first coding and decoding information corresponding to the first associated image block by an information fusion layer of the image filter to obtain target fusion data corresponding to the first associated image block, and

filtering the target image block according to the target fusion data by a filtering layer of the image filter to obtain a filtered image block corresponding to the target image block.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2024
From: WANG, LIQIANG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 069010/0318 →
Priority Claims (1)
CN 202211138569.5 · Sep 19, 2022 · national
Continuity (2)
Continuation PCTCN2023106277 · Jul 7, 2023
Related Publication 20250055990A1 · Feb 13, 2025
References Cited (20)
US 7630435B2 · Chen · 2009 [cited by examiner]
US 8149909B1 · Garbacea · 2012 [cited by examiner]
US 11265540B2 · Na · 2022 [cited by examiner]
US 11792438B2 · Li · 2023 [cited by examiner]
US 12126799B2 · Galpin · 2024 [cited by examiner]
US 20130188689A1 · Garbacea · 2013 [cited by examiner]
US 20190045224A1 · Huang · 2019 [cited by examiner]
US 20210021823A1 · Na · 2021 [cited by examiner]
US 20210037251A1 · Zhang · 2021 [cited by examiner]
US 20210409685A1 · Gao · 2021 [cited by examiner]
US 20220109890A1 · Li · 2022 [cited by examiner]
US 20220215593A1 · Wang et al. · 2022 [cited by applicant]
US 20230385643A1 · Rahman · 2023 [cited by examiner]
CN 113709504A · 2021 [cited by applicant]
CN 114157869A · 2022 [cited by applicant]
JP 2020198463A · 2020 [cited by applicant]
Liqiang Wang, et al., “EE1-1.4; Neural network based in-loop filter with 2 models”, (JVET-AA0087-v3), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29 27th Meeting, by teleconference, Jul. 13-… [cited by applicant]
International Search Report for PCT/CN2023/106277 dated Nov. 6, 2023. [cited by applicant]
Written Opinion for PCT/CN2023/106277 dated Nov. 6, 2023. [cited by applicant]
Communication issued Oct. 7, 2025 in JP Application No. 2024-563485. [cited by applicant]