IP Library Granted Patent US 12,413,721
Granted Patent B2
US 12,413,721 · App. 18/656,413 · Granted Sep 9, 2025

Content-adaptive online training method and apparatus for post-filtering

Inventors: Ding Ding (Washington, DC); Roman Chernyak (Santa Clara, CA); Shan Liu (San Jose, CA)
Assignee: TENCENT AMERICA LLC
H04N19/117H04N19/136H04N19/176H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,413,721
App. No.
18/656,413
Granted
Sep 9, 2025
Kind
B2
Abstract

Aspects of the disclosure provide methods and apparatuses, for video decoding and encoding. The apparatus includes processing circuitry configured to receive an image/video comprising one or more blocks and metadata for a machine task associated with the image/video. The metadata specifies neural network post-filtering characteristics for machine consumption. The processing circuitry decodes a first post-filtering parameter in the image/video corresponding to the one or more blocks to be reconstructed. The first post-filtering parameter applies to a block in the one or more blocks and has been updated by a post-filtering module in a post-filtering neural network (NN) that is trained based on a training dataset and the metadata. The processing circuitry determines the post-filtering NN in a video decoder corresponding to the one or more blocks based on the first post-filtering parameter, and decodes the block based on the determined post-filtering NN corresponding to the block and the metadata.

Claims (38)

1. A method for video decoding in a video decoder, comprising:

receiving one of an image and a video comprising one or more blocks and metadata for a machine task associated with the one of the image and the video, the metadata specifying neural network post-filtering characteristics for machine consumption;

decoding a first post-filtering parameter in the one of the image and the video corresponding to the one or more blocks to be reconstructed, the first post-filtering parameter applying to at least one of the one or more blocks, and the first post-filtering parameter having been updated by a post-filtering module in a post-filtering neural network (NN) that is trained based on a training dataset and the metadata for the machine task associated with the one of the image and the video;

determining the post-filtering NN in the video decoder corresponding to the one or more blocks based on the first post-filtering parameter; and

decoding the one or more blocks based on the determined post-filtering NN corresponding to the one or more blocks and the metadata for the machine task associated with the one of the image and the video.

2. The method of claim 1 , further comprising:

receiving the metadata for the machine task associated with the one of the image and the video in a Supplemental Enhancement Information (SEI) message.

3. The method of claim 2 , wherein the SEI message comprises a neural-network post-filter (NNPF) characteristics (NNPFC) SEI message, the NNPFC SEI message specifying characteristics of the post-filtering NN to be used to filter pictures after decoding for the machine task.

4. The method of claim 2 , wherein the SEI message comprises a neural-network post-filter activation (NNPFA) SEI message, the NNPFA SEI message specifying whether to enable a post-filter referenced for the one of the image and the video.

5. The method of claim 2 , wherein the SEI message comprises an annotated regions SEI message to signal parameters that identify annotated regions using bounding boxes representing a size and a location of a detected object in the one of the image and the video.

6. The method of claim 5 , wherein the bounding boxes representing the size and the location of the detected object comprises a rectangular bounding box of the detected object in the one of the image and the video.

7. The method of claim 2 , wherein the SEI message comprises an object mask information (OMI) SEI message to signal information indicating a shape of a detected object in the one of the image and the video.

8. The method of claim 2 , wherein the SEI message comprises an encoder optimization information (EOI) SEI message to indicate whether the one of the image and the video has been optimized for machine analysis and which types of optimizations have been applied in pre-processing or encoding.

9. A method for video encoding in a video encoder, comprising:

determining a first post-filtering parameter corresponding to one or more blocks to be encoded, the first post-filtering parameter applying to at least one of the one or more blocks, and the first post-filtering parameter having been updated by a post-filtering module in a post-filtering neural network (NN) that is trained based on a training dataset and metadata for a machine task associated with one of an image and a video that includes the one or more blocks, the metadata specifying neural network post-filtering characteristics for machine consumption; and

encoding the first post-filtering parameter and the metadata for the machine task in the one of the image and the video.

10. The method of claim 9 , further comprising signaling the metadata for the machine task associated with the one of the image and the video in a Supplemental Enhancement Information (SEI) message.

11. The method of claim 10 , wherein the SEI message comprises a neural-network post-filter (NNPF) characteristics (NNPFC) SEI message, the NNPFC SEI message specifying characteristics of the post-filtering NN to be used to filter pictures after decoding for the machine task.

12. The method of claim 10 , wherein the SEI message comprises a neural-network post-filter activation (NNPFA) SEI message, the NNPFA SEI message specifying whether to enable a post-filter referenced for the one of the image and the video.

13. The method of claim 10 , wherein the SEI message comprises an annotated regions SEI message to signal parameters that identify annotated regions using bounding boxes representing a size and a location of a detected object in the one of the image and the video.

14. The method of claim 13 , wherein the bounding boxes representing the size and the location of the detected object comprises a rectangular bounding box of the detected object in the one of the image and the video.

15. The method of claim 10 , wherein the SEI message comprises an object mask information (OMI) SEI message to signal information indicating a shape of a detected object in the one of the image and the video.

16. The method of claim 10 , wherein the SEI message comprises an encoder optimization information (EOI) SEI message to indicate whether the one of the image and the video has been optimized for machine analysis and which types of optimizations have been applied in pre-processing or encoding.

17. A method of processing visual media data, the method comprising:

processing a bitstream that includes the visual media data according to a format rule, wherein

the bitstream includes metadata and a first post-filtering parameter; and

the format rule specifies that

one of an image and a video includes one or more blocks and the metadata for a machine task associated with the one of the image and the video,

the metadata specifies neural network post-filtering characteristics for machine consumption;

a first post-filtering parameter in the one of the image and the video corresponding to the one or more blocks to be reconstructed is decoded,

the first post-filtering parameter applies to at least one of the one or more blocks,

the first post-filtering parameter is updated by a post-filtering module in a post-filtering neural network (NN) that is trained based on a training dataset and the metadata for the machine task associated with the one of the image and the video;

the post-filtering NN in a video decoder corresponding to the one or more blocks is determined based on the first post-filtering parameter; and

the one or more blocks is decoded based on the determined post-filtering NN corresponding to the one or more blocks and the metadata for the machine task associated with the one of the image and the video.

18. The method of claim 17 , wherein the format rule specifies that:

the metadata for the machine task associated with the one of the image and the video is received in a Supplemental Enhancement Information (SEI) message.

19. The method of claim 18 , wherein the SEI message comprises a neural-network post-filter (NNPF) characteristics (NNPFC) SEI message, the NNPFC SEI message specifying characteristics of the post-filtering NN to be used to filter pictures after decoding for the machine task.

20. The method of claim 18 , wherein the SEI message comprises a neural-network post-filter activation (NNPFA) SEI message, the NNPFA SEI message specifying whether to enable a post-filter referenced for the one of the image and the video.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2025
From: DING, DING; CHERNYAK, ROMAN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 071330/0524 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: DING, DING; CHERNYAK, ROMAN; LIU, SHAN
To: TENCENT AMERICA LLC
Reel/Frame 069543/0101 →
Continuity (4)
Continuation In Part 17749641 · May 20, 2022
Provisional Application 63641809 · May 2, 2024
Provisional Application 63194057 · May 27, 2021
Related Publication 20240291980A1 · Aug 29, 2024
References Cited (20)
US 20220385896A1 · Ding et al. · 2022 [cited by applicant]
EP 3941066A1 · 2022 [cited by applicant]
JP 2020010331A · 2020 [cited by applicant]
JP 2021072540A · 2021 [cited by applicant]
WO 2010143427A1 · 2010 [cited by applicant]
WO 2020192034A1 · 2020 [cited by applicant]
WO 2023196217A1 · 2023 [cited by applicant]
Office Action received for Japanese Patent Application No. 2022-570193, mailed on Aug. 5, 2024, 8 pages (4 pages of English Translation and 4 pages of Original Document). [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2024/059638, mailed on Feb. 7, 2025, 11 pages. [cited by applicant]
Ballé J, Minnen D, Singh S, et al. Variational image compression with a scale hyperprior[J]. arXiv preprint arXiv:1802.01436, 2018, pp. 1-23. [cited by applicant]
Cheng Z, Sun H, Takeuchi M, et al. Learned image compression with discretized gaussian mixture likelihoods and attention modules, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020: … [cited by applicant]
D. Liu, Z. Chen, S. Liu, F. Wu, “Deep Learning-Based Technology in Responses to the Joint Call for Proposals on Video Compression with Capability beyond HEVC”, IEEE Transactions on Circuits and Systems for Video Technol… [cited by applicant]
D. Minnen, J. Ballé, G. Toderici, “Joint Autoregressive and Hierarchical Priors for Learned Image Compression”, Proceedings of the 32nd International Conference on Neural Information Processing Systems, Dec. 2018, pp. 1… [cited by applicant]
Extended European Search Report received for European Patent Application No. 22790197.2, mailed on Aug. 7, 2023, 9 pages. [cited by applicant]
Lam Y et al: “AHG11: Content-adaptive neural network post-processing filter”, 22. JVET Meeting; Apr. 20, 2021-Apr. 28, 2021; Teleconference; (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11 and ITU-T SG.16 ),… [cited by applicant]
Office Action received for Japanese Patent Application No. 2022-570193, mailed on Feb. 5, 2024, 13 pages (6 pages of English Translation and 7 pages of Original Document). [cited by applicant]
S. Liu, L. Wang, P. Wu, H. Yang, “JVET AHG report 9: Neural Networks in Video Coding (AHG9)”, ISO/IEC JTC1/SC29/WG11 JVET-J0009, pp. 1-3. [cited by applicant]
S. Liu, X. Zhang, S. Lei, “Rectangular partitioning for Intra prediction in HEVC”, Visual Communications and Image Processing (VCIP), IEEE, Jan. 2012, pp. 1-6. [cited by applicant]
Toderici G, Vincent D, Johnston N, et al. “Full resolution image compression with recurrent neural networks”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017: 5306-5314. [cited by applicant]
Y. Li, S. Liu, K. Kawamura, “Methodology and reporting template for neural network coding tool testing”, ISO/IEC JTC1/SC29/WG11 JVET-M1006, pp. 1-4. [cited by applicant]