IP Library Granted Patent US 12,375,697
Granted Patent B2
US 12,375,697 · App. 18/186,006 · Granted Jul 29, 2025

Computer-implemented method and apparatus for video coding using super-resolution restoration with residual frame coding

Inventors: Byeongdoo Choi (Irvine, CA); Christopher Andrew Segall (Camas, WA); Kiran Mukesh Misra (Camas, WA)
Assignee: Amazon Technologies, Inc.
H04N19/42G06T3/4053H04N19/117H04N19/12H04N19/136H04N19/139H04N19/172H04N19/176H04N19/59H04N19/60H04N19/82H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,375,697
App. No.
18/186,006
Granted
Jul 29, 2025
Kind
B2
Abstract

The present disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media for video coding using super-resolution restoration with residual frame coding. According to some examples, a computer-implemented method includes receiving a coded frame of a video; performing a video coding on the coded frame of the video to generate a resultant for the coded frame at a second lower resolution than a first resolution; upsampling the resultant in at least a vertical direction to a higher resolution than the second lower resolution to generate an upsampled resultant; generating a decoded frame based on at least the upsampled resultant; and transmitting the decoded frame to a frame buffer or to a display device.

Claims (45)

1. A computer-implemented method comprising:

receiving a video at a content delivery service;

downsampling a frame of the video from a first resolution to a second lower resolution in a vertical direction and a horizontal direction;

performing a video coding on the frame of the video by the content delivery service that coverts the frame from a pixel domain to a transform domain and back to the pixel domain to generate a resultant for the frame at the second lower resolution;

upsampling the resultant in the vertical direction and the horizontal direction to a higher resolution than the second lower resolution to generate an upsampled resultant;

performing an entropy encode of the frame based on at least the upsampled resultant to generate a coded frame;

transmitting the coded frame from the content delivery service to a decoder;

generating a reference frame at the higher resolution for the frame based on the upsampled resultant;

storing the reference frame at the higher resolution in a frame buffer; and

storing a corresponding decoded motion vector at the second lower resolution for the frame in the frame buffer.

2. The computer-implemented method of claim 1 , wherein

the storing the reference frame at the higher resolution and

the storing the corresponding decoded motion vector at the second lower resolution for the frame are in a single frame unit in the frame buffer.

3. The computer-implemented method of claim 1 , wherein the upsampling comprises generating the upsampled resultant by a machine learning model.

4. A computer-implemented method comprising:

receiving a coded frame of a video;

performing a video coding on the coded frame of the video to generate a resultant for the coded frame at a second lower resolution than a first resolution;

upsampling the resultant in at least a vertical direction to a higher resolution than the second lower resolution to generate an upsampled resultant;

generating a decoded frame at the higher resolution based on at least the upsampled resultant;

causing a store of the decoded frame at the higher resolution in a frame buffer; and

causing a store of a corresponding decoded motion vector at the second lower resolution for the coded frame in the frame buffer.

5. The computer-implemented method of claim 4 , wherein the causing the store of the decoded frame at the higher resolution and the causing the store of the corresponding decoded motion vector at the second lower resolution for the coded frame are in a single frame unit in the frame buffer.

6. The computer-implemented method of claim 4 , wherein the frame buffer does not include a corresponding decoded motion vector at the higher resolution for the coded frame.

7. The computer-implemented method of claim 4 , wherein the upsampling comprises generating the upsampled resultant by a machine learning model.

8. The computer-implemented method of claim 7 , wherein the generating the decoded frame comprises performing a residual operation based on the upsampled resultant at the higher resolution to generate the decoded frame.

9. The computer-implemented method of claim 4 , wherein the generating the decoded frame comprises performing a residual operation based on the upsampled resultant at the higher resolution to generate the decoded frame.

10. The computer-implemented method of claim 4 , wherein a viewer device comprises a decoder including the frame buffer, and the upsampling the resultant occurs within a postprocessing operation of the decoder.

11. The computer-implemented method of claim 10 , wherein the postprocessing operation comprises the upsampling in at least the vertical direction and a loop restoration operation.

12. The computer-implemented method of claim 11 , wherein the postprocessing operation further comprises performing a residual operation based on the upsampled resultant at the higher resolution to generate the decoded frame.

13. The computer-implemented method of claim 4 , wherein the upsampling the resultant in at least the vertical direction to the higher resolution occurs within a decoder in a super-resolution mode.

14. The computer-implemented method of claim 4 , wherein a header for the video comprises a super-resolution scaling parameter for the vertical direction.

15. An apparatus comprising:

a coupling to a display; and

a video decoder to:

receive a coded frame of a video,

perform a video coding on the coded frame of the video to generate a resultant for the coded frame at a second lower resolution than a first resolution,

upsample the resultant in at least a vertical direction to a higher resolution than the second lower resolution to generate an upsampled resultant,

generate a decoded frame at the higher resolution based on the upsampled resultant,

store the decoded frame at the higher resolution in a frame buffer, and

store a corresponding decoded motion vector at the second lower resolution for the coded frame in the frame buffer.

16. The apparatus of claim 15 , wherein the store of the decoded frame at the higher resolution and the store of the corresponding decoded motion vector at the second lower resolution for the coded frame are in a single frame unit in the frame buffer.

17. The apparatus of claim 16 , wherein the frame buffer does not include a corresponding decoded motion vector at the higher resolution for the coded frame.

18. The apparatus of claim 15 , wherein the video decoder is to generate the upsampled resultant by a machine learning model.

19. The apparatus of claim 15 , wherein the video decoder is to upsample the resultant in at least the vertical direction to the higher resolution when the video decoder is in a super-resolution mode.

20. The apparatus of claim 15 , wherein the video decoder is further to read a header for the video that comprises a super-resolution scaling parameter for the vertical direction.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2023
From: CHOI, BYEONGDOO; SEGALL, CHRISTOPHER ANDREW; MISRA, KIRAN MUKESH
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 063022/0931 →
Continuity (2)
Provisional Application 63437957 · Jan 9, 2023
Related Publication 20240236366A1 · Jul 11, 2024
References Cited (40)
US 8989508B2 · Deshpande · 2015 [cited by applicant]
US 9049427B2 · Hattori · 2015 [cited by applicant]
US 9319703B2 · Wang · 2016 [cited by applicant]
US 9969299B2 · Murase et al. · 2018 [cited by applicant]
US 10313698B2 · Sullivan et al. · 2019 [cited by applicant]
US 10623753B2 · Skupin et al. · 2020 [cited by applicant]
US 11700390B2 · Wang · 2023 [cited by applicant]
US 11743505B2 · Wang · 2023 [cited by applicant]
US 11765394B2 · Wang et al. · 2023 [cited by applicant]
US 11812062B2 · Wang · 2023 [cited by applicant]
US 12022122B2 · Deshpande · 2024 [cited by applicant]
US 12034927B2 · Okawa et al. · 2024 [cited by applicant]
US 20060126952A1 · Suzuki · 2006 [cited by examiner]
US 20100220939A1 · Tourapis · 2010 [cited by examiner]
US 20140086336A1 · Wang · 2014 [cited by applicant]
US 20190068969A1 · Rusanovskyy et al. · 2019 [cited by applicant]
US 20200374524A1 · Gao · 2020 [cited by examiner]
US 20220321919A1 · Deshpande · 2022 [cited by examiner]
US 20240137577A1 · Lin · 2024 [cited by examiner]
US 20240205439A1 · Sjberg et al. · 2024 [cited by applicant]
US 20240214558A1 · Dumas et al. · 2024 [cited by applicant]
US 20240236366A1 · Choi et al. · 2024 [cited by applicant]
US 20240267548A1 · Du et al. · 2024 [cited by applicant]
US 20240292003A1 · Damghanian et al. · 2024 [cited by applicant]
US 20240422360A1 · Kang et al. · 2024 [cited by applicant]
Ding, Dandan et al., “Advances In Video Compression System Using Deep Neural Network: A Review And Case Studies”, arXiv:2101.06341v1, Jan. 16, 2021, 27 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2024/010748, May 3, 2024, 09 pages. [cited by applicant]
Kang, Jihong et al., “Multi-modal/multi-scale Convolutional Neural Network Based In-loop Filter Design for Next Generation Video Codec”, IEEE International Conference on Image Processing (ICIP), Sep. 2017, pp. 26-30. [cited by applicant]
Wang, Zhao et al., “AHG11: Separate Density Attention Network for Loop Filtering”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, 22nd Meeting, Apr. 2021, 3 pages. [cited by applicant]
Zhao, Yanchen et al., “Joint Luma and Chroma Multi-Scale CNN In-loop Filter for Versatile Video Coding”, IEEE International Symposium on Circuits and Systems (ISCAS), May 2022, pp. 3205-3209. [cited by applicant]
De Rivaz, P., et al., “AV1 Bitstream & Decoding Process Specification”, Version 1.0.0 with Errata 1, Jan. 8, 2019, available online at https://aomediacodec.github.io/av1-spec/av1-spec.pdf, 681 pages. [cited by applicant]
The Linux Foundation, “CONV2D”, PyTorch open source code, Dec. 2022, retrieved from https://pytorch.org/docs/stable/generated/torch.nn.Conv2d.html, 2 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/342,406, Dec. 4, 2024, 14 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/186,084, Aug. 14, 2024, 20 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 18/186,084, Mar. 7, 2025, 25 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/342,406, Mar. 19, 2025, 5 pages. [cited by applicant]
Kang, Jihong et al., “Multi-modal/Multi-scale Convolutional Neural Network Based In-loop Filter Design for Next Generation Video Codec”, IEEE International Conference on Image Processing, Sep. 2017, 5 pages. [cited by applicant]
Wang, Zhao et al., “AHG11: Separate Density Attention Network for Loop Filtering”, 22nd Meeting of the Joint Video Experts Team (JVET), Document No. JVET-V0074-v3, Apr. 2021, 3 pages. [cited by applicant]
Wang, Zhao et al., “Multi-Density Attention Network for Loop Filtering in Video Compression”, arXiv:2104.12865v1, Apr. 8, 2021, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/186,084, Jun. 24, 2025, 10 pages. [cited by applicant]