IP Library › Granted Patent US 12,563,196
Granted Patent B2
US 12,563,196 · App. 18/343,916 · Granted Feb 24, 2026

Apparatus and method with video processing using neural network

Inventors: Seungeon Kim (Suwon-si, KR); Wonhee Lee (Suwon-si, KR); Kyungboo Jung (Suwon-si, KR); Woosuk Choi (Suwon-si, KR); Young Hun Sung (Suwon-si, KR); Dokwan Oh (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H04N19/13H04N19/51H04N19/70H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,563,196
App. No.
18/343,916
Granted
Feb 24, 2026
Kind
B2
Abstract

An apparatus with video processing includes: one or more processors configured to: generate a syntax element processable by a target standard codec by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to a neural network and compressing the plurality of frames, and generate a bitstream by performing entropy encoding on the syntax element.

Claims (41)

1 . An apparatus with video processing, the apparatus comprising:

one or more processors configured to:

generate syntax elements processable by a target decoder of a target standard codec by selecting one of outputs generated by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to neural networks, and

generate a bitstream by performing entropy encoding on the syntax elements.

2 . The apparatus of claim 1 , wherein the syntax elements comprise a coding unit (CU) partition, a prediction unit (PU) partition, a PU prediction mode, and a transform unit (TU) partition.

3 . The apparatus of claim 1 , wherein the pre-decoded reference image is decoded at a time point before a time point when input frame is encoded.

4 . The apparatus of claim 1 , wherein, for the generating of the pre-decoded reference image, the one or more processors are further configured to:

generate decoded syntax elements by performing entropy decoding on the bitstream, and

generate the pre-decoded reference image by decompressing the decoded syntax elements.

5 . The apparatus of claim 4 , wherein the decoded syntax elements are decodable by the target decoder comprising a decoder of a standard codec.

6 . The apparatus of claim 1 , wherein the neural networks comprise a first neural network and a second neural network.

7 . The apparatus of claim 6 , wherein, for the generating of the bitstream, the one or more processors are further configured to:

select one from an output of the first neural network and an output of the second neural network, and

perform entropy encoding on the selected output.

8 . The apparatus of claim 6 , wherein, for the generating of the syntax elements, the one or more processors are further configured to perform either one or both of:

intra-prediction through the first neural network; and

inter-prediction through the second neural network.

9 . The apparatus of claim 7 , wherein, for the generating of the syntax elements, the one or more processors are further configured to:

partition the plurality of frames into a plurality of blocks through the first neural network, and

perform motion estimation and compensation by inputting the plurality of blocks to the second neural network.

10 . The apparatus of claim 1 , wherein, for the generating of the syntax elements, the one or more processors are further configured to adjust the quantization parameter based on a shape of adaptive instance normalization of a layer constituting the neural networks and a product or sum of features at an arbitrary level.

11 . A processor-implemented method with video processing, the method comprising:

generating syntax elements processable by a target decoder of a target standard codec by selecting one of outputs generated by inputting a quantization parameter, a pre-decoded reference image, and a plurality of frames comprised in a video to neural networks; and

generating a bitstream by performing entropy encoding on the syntax elements.

12 . The method of claim 11 , wherein the syntax elements comprise a coding unit (CU) partition, a prediction unit (PU) partition, a PU prediction mode, and a transform unit (TU) partition.

13 . The method of claim 11 , wherein the pre-decoded reference image is decoded at a time point before a time point when input frame is encoded.

14 . The method of claim 11 , wherein the generating of the pre-decoded reference image comprises:

generating decoded syntax elements by performing entropy decoding on the bitstream; and

generating the pre-decoded reference image by decompressing the decoded syntax elements.

15 . The method of claim 14 , wherein the decoded syntax elements are decodable by the target decoder comprising a decoder of a standard codec.

16 . The method of claim 11 , wherein the neural networks comprise a first neural network and a second neural network.

17 . The method of claim 16 , wherein the generating of the bitstream comprises:

selecting one from an output of the first neural network and an output of the second neural network; and

performing entropy encoding on the selected output.

18 . The method of claim 16 , wherein the generating of the syntax elements comprises either one or both of:

performing intra-prediction through the first neural network; and

performing inter-prediction through the second neural network.

19 . The method of claim 17 , wherein the generating of the syntax elements comprise:

partitioning the plurality of frames into a plurality of blocks through the first neural network; and

performing motion estimation and compensation by inputting the plurality of blocks to the second neural network.

20 . The method of claim 11 , wherein the generating of the syntax elements comprises adjusting the quantization parameter based on a shape of adaptive instance normalization of a layer constituting the neural networks and a product or sum of features at an arbitrary level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2023
From: KIM, SEUNGEON; LEE, WONHEE; JUNG, KYUNGBOO; CHOI, WOOSUK; SUNG, YOUNG HUN; OH, DOKWAN
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 064111/0348 →
Priority Claims (1)
KR 10-2023-0006675 · Jan 17, 2023 · national
Continuity (1)
Related Publication 20240244208A1 · Jul 18, 2024
References Cited (14)
US 10104391B2 · Su et al. · 2018 [cited by applicant]
US 10123053B2 · Sze et al. · 2018 [cited by applicant]
US 11190782B2 · Park et al. · 2021 [cited by applicant]
US 11330264B2 · Zhou et al. · 2022 [cited by applicant]
US 20230007246A1 · Li · 2023 [cited by examiner]
KR 1020220124622A · 2022 [cited by applicant]
Extended European search report issued on Apr. 15, 2024, in counterpart European Patent Application No. 24152384.4 (9 pages). [cited by applicant]
Laude, Thorsten, et al., “Deep Learning-Based Intra Prediction Mode Decision For HEVC,” Picture Coding Symposium, IEEE, 2016, (5 Pages in English). [cited by applicant]
Liu, Zhenyu, et al., “CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network,” IEEE Transactions on Image Processing, vol. 25, No. 11, Nov. 2016, (p. 5088-5103). [cited by applicant]
Li, Jiahao, et al., “Intra Prediction Using Fully Connected Network for Video Coding,” IEEE International Conference on Image Processing, 2017, (5 Pages in English). [cited by applicant]
Lee, Jung Kyung, et al., “Convolution Neural Network Based Video Coding Technique Using Reference Video Synthesis,” Asia-Pacific Signal and Information Processing Association Annual Summit and Conference, Nov. 12-15, 20… [cited by applicant]
Zhao, Zhenghui, et al., “CNN-Based Bi-Directional Motion Compensation for High Efficiency Video Coding,” 2018 IEEE International Symposium on Circuits and Systems (ISCAS), 2018, (4 Pages in English). [cited by applicant]
Xu, Mai, et al., “Reducing Complexity of HEVC: A Deep Learning Approach,” IEEE Transactions on Image Processing, vol. 27, No. 10, Oct. 2018, (p. 5044-5059). [cited by applicant]
Lu, Guo, et al., “DVC: An End-To-End Deep Video Compression Framework,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, (p. 11006-11015). [cited by applicant]