IP Library › Granted Patent US 12,641,266
Granted Patent B2
US 12,641,266 · App. 18/733,379 · Granted May 26, 2026

Method for neural network-based video encoding and decoding, and video encoding apparatus

Inventors: Seung Eon Kim (Suwon-si, KR); Yeongwoong Kim (Suwon-si, KR); Hui Yong Kim (Daejeon, KR); Won Hee Lee (Suwon-si, KR); Suyong Bahk (Yongin-si, KR); Young Hun Sung (Suwon-si, KR); Dokwan Oh (Suwon-si, KR)
Assignees: SAMSUNG ELECTRONICS CO., LTD.; UNIVERSITY-INDUSTRY COOPERATION GROUP OF KYUNG HEE UNIVERSITY
H04N19/42H04N19/105H04N19/124H04N19/13H04N19/139H04N19/31H04N19/61H04N19/80H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,641,266
App. No.
18/733,379
Granted
May 26, 2026
Kind
B2
Abstract

There is provided a method for neural network-based video encoding. The method includes estimating a motion vector between an input image and a reference image based on a temporal layer of the input image, transforming the motion vector into a latent representation, scaling the latent representation of the motion vector based on the temporal layer of the input image and obtaining a temporal context of the input image based on the scaled latent representation of the motion vector and the reference image.

Claims (30)

1 . A method for neural network-based video encoding, the method comprising:

estimating a motion vector between an input image and a reference image based on a temporal layer of the input image;

transforming the motion vector into a latent representation;

scaling the latent representation of the motion vector based on the temporal layer of the input image; and

obtaining a temporal context of the input image based on the scaled latent representation of the motion vector and the reference image.

2 . The method of claim 1 , wherein the reference image comprises a first bidirectional reference image temporally before the input image and a second bidirectional reference image temporally after the input image.

3 . The method of claim 1 , wherein the scaling based on the temporal layer comprises scaling the latent representation of the motion vector by dividing the latent representation of the motion vector into quantization step determining parameters defined for the temporal layer.

4 . The method of claim 1 , further comprising:

performing entropy encoding and entropy decoding on the latent representation of the motion vector;

rescaling the scaled latent representation of the motion vector based on the temporal layer; and

reconstructing motion vectors based on the rescaled latent representation of the motion vector,

wherein the obtaining of the temporal context comprises obtaining the temporal context based on the reconstructed motion vectors and the reference image.

5 . The method of claim 4 , wherein the rescaling based on the temporal layer comprises multiplying the scaled latent representation of the motion vector by the quantization step determining parameters defined for the temporal layer.

6 . The method of claim 4 , wherein the obtaining of the temporal context comprises:

outputting a reference feature map by inputting the reference image into a feature extraction neural network,

performing bilinear warping on the reference feature map based on the reconstructed motion vectors to output a warped reference feature map,

inputting the warped reference feature map to a post-processing neural network, and

inputting an output of the post-processing neural network to a context fusion network to output the temporal context.

7 . A method for neural network-based video encoding, the method comprising:

estimating a motion vector between an input image and a reference image based on a temporal layer of the input image;

transforming the motion vector into a latent representation;

obtaining a temporal context of the input image based on the latent representation of the motion vector and the reference image; and

performing a smoothing operation on the temporal context based on a smoothing object comprising at least one of the reference image, the input image, the motion vector, an input in the obtaining of the temporal context, an output in the obtaining of the temporal context, or an input or output of a sub-process in the obtaining of the temporal context.

8 . The method of claim 7 , wherein the reference image comprises a first bidirectional reference image temporally before the input image and a second bidirectional reference image temporally after the input image.

9 . An electronic device comprising:

a memory configured to store one or more instructions and a reference image; and

a processor configured to execute the one or more instructions to: estimate a motion vector between an input image and the reference image based on the reference image, the input image, and a temporal layer of the input image;

transform the motion vector into a latent representation;

scale the latent representation of the motion vector based on the temporal layer of the input image; and

obtain a temporal context of the input image based on the scaled latent representation of the motion vector and the reference image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: KIM, SEUNG EON; KIM, YEONGWOONG; KIM, HUI YONG; LEE, WON HEE; BAHK, SUYONG; SUNG, YOUNG HUN; OH, DOKWAN
To: SAMSUNG ELECTRONICS CO., LTD.; UNIVERSITY-INDUSTRY COOPERATION GROUP OF KYUNG HEE UNIVERSITY
Reel/Frame 067618/0214 →
Priority Claims (2)
KR 10-2023-0102601 · Aug 7, 2023 · national
KR 10-2023-0160501 · Nov 20, 2023 · national
Continuity (1)
Related Publication 20250056024A1 · Feb 13, 2025
References Cited (26)
US 11257254B2 · Minnen et al. · 2022 [cited by applicant]
US 11399198B1 · Pourreza et al. · 2022 [cited by applicant]
US 11412023B2 · Wang et al. · 2022 [cited by applicant]
US 20190373259A1 · Xu · 2019 [cited by examiner]
US 20200084444A1 · Egilmez · 2020 [cited by examiner]
US 20220021870A1 · Jiang et al. · 2022 [cited by applicant]
US 20220239944A1 · Zhang et al. · 2022 [cited by applicant]
US 20220272355A1 · Singh · 2022 [cited by examiner]
US 20220295095A1 · Pourreza · 2022 [cited by examiner]
US 20220417540A1 · Adzic · 2022 [cited by applicant]
US 20230051066A1 · Li et al. · 2023 [cited by applicant]
US 20230127006A1 · De Bock · 2023 [cited by applicant]
US 20240378752A1 · Westcott · 2024 [cited by examiner]
US 20250379978A1 · Li · 2025 [cited by examiner]
CN 114501013A · 2022 [cited by applicant]
CN 115529457A · 2022 [cited by applicant]
EP 3745305A1 · 2020 [cited by applicant]
EP 4115602A1 · 2023 [cited by applicant]
Communication dated Mar. 24, 2025 issued by the European Patent Office in European Patent Application No. 24193196.3. [cited by applicant]
Communication issued on Dec. 19, 2024 from the European Patent Office for European Patent Application No. 24193196.3. [cited by applicant]
Mu-Jung Chen et al., “B-CANF: Adaptive B-frame Coding with Conditional Augmented Normalizing Flows”, Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), arXiv:2209.01769v2 [eess.IV], Aug.… [cited by applicant]
Yeongwoong Kim et al., “Neural Video Compression with Temporal Layer-Adaptive Hierarchical B-frame Coding”, arXiv:2308.15791v1 [cs.CV], Aug. 30, 2023, 10 pages, XP091601211. [cited by applicant]
Lu et al., “DVC: An End-to-end Deep Video Compression Framework”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015, doi:10.1109/CVPR.2019.01126. [cited by applicant]
Li et al., “Deep Contextual Video Compression”, 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Dec. 14, 2021, 19 total pages, arXiv:2109.15047v2 [eess.IV], doi:10.48550/arXiv.2109.15047. [cited by applicant]
Sheng et al., “Temporal Context Mining for Learned Video Compression”, IEEE Transactions on Multimedia, Jan. 30, 2023, 13 total pages, arXiv:2111.13850v2 [cs.CV], doi:10.48550/arXiv.2111.13850. [cited by applicant]
Yang et al., “Learning for video compression with hierarchical quality and recurrent enhancement”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 10 total pages, doi:10.1109/CVPR42600.… [cited by applicant]