IP Library › Granted Patent US 12,581,135
Granted Patent B2
US 12,581,135 · App. 18/415,247 · Granted Mar 17, 2026

Apparatus and method with video processing using neural network

Inventors: Seungeon Kim (Suwon-si, KR); Wonhee Lee (Suwon-si, KR); Kyungboo Jung (Suwon-si, KR); Woosuk Choi (Suwon-si, KR); Young Hun Sung (Suwon-si, KR); Dokwan Oh (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H04N19/91G06V10/82H04N19/119H04N19/159H04N19/184H04N19/51H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,581,135
App. No.
18/415,247
Granted
Mar 17, 2026
Kind
B2
Abstract

An apparatus with video processing includes one or more processors configured to generate a syntax element processable by a target standard codec by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to a neural network and compressing the plurality of frames, and generate a bitstream by performing entropy encoding on the syntax element.

Claims (51)

1 . An apparatus with video processing, the apparatus comprising:

one or more processors configured to:

generate, based on decoder information on a decoder among a plurality of decoders, syntax elements processable by a target decoder of a target standard codec by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to neural networks and compressing the plurality of frames, and

generate a bitstream by performing entropy encoding on the syntax elements.

2 . The apparatus of claim 1 , wherein the syntax elements comprises a coding unit (CU) partition, a prediction unit (PU) partition, a PU prediction mode, and a transform unit (TU) partition.

3 . The apparatus of claim 1 , wherein the pre-decoded reference image is decoded at a time point before a time point when input frame is encoded.

4 . The apparatus of claim 1 , wherein, for the generating of the pre-decoded reference image, the one or more processors are further configured to:

generate decoded syntax elements by performing entropy decoding on the bitstream, and

generate the pre-decoded reference image by decompressing the decoded syntax elements.

5 . The apparatus of claim 4 , wherein

the decoder is a target decoder comprising a decoder of a standard codec, and

the decoded syntax elements are decodable by the target decoder.

6 . The apparatus of claim 1 , wherein the neural networks comprise a first neural network and a second neural network.

7 . The apparatus of claim 6 , wherein, for the generating of the bitstream, the one or more processors are further configured to:

select one from an output of the first neural network and an output of the second neural network, and

perform entropy encoding on the selected output.

8 . The apparatus of claim 6 , wherein, for the generating of the syntax elements, the one or more processors are further configured to perform either one or both of:

intra-prediction through the first neural network; and

inter-prediction through the second neural network.

9 . The apparatus of claim 7 , wherein, for the generating of the syntax elements, the one or more processors are further configured to:

partition the plurality of frames into a plurality of blocks through the first neural network, and

perform motion estimation and compensation by inputting the plurality of blocks to the second neural network.

10 . The apparatus of claim 1 , wherein, for the generating of the syntax elements, the one or more processors are further configured to adjust the quantization parameter based on a shape of adaptive instance normalization of a layer constituting the neural networks and a product or sum of features at an arbitrary level.

11 . The apparatus of claim 1 , wherein, for the generating of the syntax elements, the one or more processors are further configured to:

receive the decoder information on the decoder to perform decoding among a plurality of decoders, and

generate the syntax elements by considering the decoder information.

12 . A processor-implemented method with video processing, the method comprising:

generating, based on decoder information on a decoder among a plurality of decoders, syntax elements processable by a target decoder of a target standard codec by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to neural networks and compressing the plurality of frames; and

generating a bitstream by performing entropy encoding on the syntax elements.

13 . The method of claim 12 , wherein the syntax elements comprises a coding unit (CU) partition, a prediction unit (PU) partition, a PU prediction mode, and a transform unit (TU) partition.

14 . The method of claim 12 , wherein the pre-decoded reference image is decoded at a time point before a time point when input frame is encoded.

15 . The method of claim 12 , wherein the generating of the pre-decoded reference image comprises:

generating decoded syntax elements by performing entropy decoding on the bitstream; and

generating the pre-decoded reference image by decompressing the decoded syntax elements.

16 . The method of claim 15 , wherein

the decoder is a target decoder comprising a decoder of a standard codec, and

the decoded syntax elements are decodable by the target decoder.

17 . The method of claim 12 , wherein the neural networks comprise a first neural network and a second neural network.

18 . The method of claim 17 , wherein the generating of the bitstream comprises:

selecting one from an output of the first neural network and an output of the second neural network; and

performing entropy encoding on the selected output.

19 . The method of claim 17 , wherein the generating of the syntax elements comprises either one or both of:

performing intra-prediction through the first neural network; and

performing inter-prediction through the second neural network.

20 . The method of claim 18 , wherein the generating of the syntax elements comprises:

partitioning the plurality of frames into a plurality of blocks through the first neural network; and

performing motion estimation and compensation by inputting the plurality of blocks to the second neural network.

21 . The method of claim 12 , wherein the generating of the syntax elements comprises adjusting the quantization parameter based on a shape of adaptive instance normalization of a layer constituting the neural networks and a product or sum of features at an arbitrary level.

22 . The method of claim 12 , wherein the generating of the syntax elements comprises:

receiving the decoder information on the decoder to perform decoding among a plurality of decoders; and

generating the syntax element by considering the decoder information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2024
From: KIM, SEUNGEON; LEE, WONHEE; JUNG, KYUNGBOO; CHOI, WOOSUK; SUNG, YOUNG HUN; OH, DOKWAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066538/0702 →
Priority Claims (2)
KR 10-2023-0006675 · Jan 17, 2023 · national
KR 10-2023-0101225 · Aug 2, 2023 · national
Continuity (2)
Continuation In Part 18343916 · Jun 29, 2023
Related Publication 20240244273A1 · Jul 18, 2024
References Cited (13)
US 10104391B2 · Su et al. · 2018 [cited by applicant]
US 10123053B2 · Sze et al. · 2018 [cited by applicant]
US 11190782B2 · Park et al. · 2021 [cited by applicant]
US 11330264B2 · Zhou et al. · 2022 [cited by applicant]
US 20230007246A1 · Li · 2023 [cited by examiner]
KR 1020220124622A · 2022 [cited by applicant]
Laude, Thorsten, et al., “Deep Learning-Based Intra Prediction Mode Decision For HEVC,” Picture Coding Symposium, IEEE, 2016, (5 Pages in English). [cited by applicant]
Liu, Zhenyu, et al., “CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network,” IEEE Transactions on Image Processing, vol. 25, No. 11, Nov. 2016, (p. 5088-5103). [cited by applicant]
Li, Jiahao, et al., “Intra Prediction Using Fully Connected Network for Video Coding,” IEEE International Conference on Image Processing, 2017, (5 Pages in English). [cited by applicant]
Lee, Jung Kyung, et al., “Convolution Neural Network Based Video Coding Technique Using Reference Video Synthesis,” Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, Nov. 12-15, 20… [cited by applicant]
Zhao, Zhenghui, et al., “CNN-Based Bi-Directional Motion Compensation for High Efficiency Video Coding,” 2018 IEEE International Symposium on Circuits and Systems (ISCAS), 2018, (4 Pages in English). [cited by applicant]
Xu, Mai, et al., “Reducing Complexity of HEVC: A Deep Learning Approach,” IEEE Transactions on Image Processing, vol. 27, No. 10, Oct. 2018, (p. 5044-5059). [cited by applicant]
Lu, Guo, et al., “DVC: An End-To-End Deep Video Compression Framework,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, (p. 11006-11015). [cited by applicant]