IP Library Granted Patent US 12,634,433
Granted Patent B2
US 12,634,433 · App. 18/250,949 · Granted May 19, 2026

Systems and methods for hybrid machine learning and DCT-based video compression

Inventors: Feng Liu (Beaverton, OR); Wu-Chi Feng (Tigard, OR)
Assignee: PORTLAND STATE UNIVERSITY
H04N19/105H04N19/154H04N19/172H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,433
App. No.
18/250,949
Granted
May 19, 2026
Kind
B2
Abstract

Systems and methods for hybrid video compression. A method includes receiving an encoded video including at least two compressed frames corresponding to at least two anchor frames and a compressed subset of at least one intermediate frame between the at least two anchor frames, generating, by inputting the at least two anchor frames into a deep machine learning model, a synthesized image frame corresponding to the at least one intermediate frame between the at least two anchor frames, reconstructing at least one hybrid image frame by combining the compressed subset of the at least one intermediate frame with the synthesized image frame, and outputting a video including the at least two anchor frames and the at least one hybrid image frame. Thus, DCT-based video compression techniques can leverage machine learning-based video interpolation techniques to provide encoded video streams with a reduced bitrate and thus reduced bandwidth while maintaining image quality.

Claims (33)

1 . A system, comprising:

a hybrid encoding system configured to selectively encode a portion of a video; and

a decoding system communicatively coupled to the hybrid encoding system and configured to, synthesize, with a deep machine learning model, a remainder of the video upon receiving the selectively encoded portion of the video from the hybrid encoding system, and selectively combine the synthesized remainder of the video with the portion of the video,

wherein the hybrid encoding system comprises a first processor and a first non-transitory memory, the first non-transitory memory configured with executable instructions that when executed by the first processor cause the first processor to:

compress two anchor frames in the video, wherein the two anchor frames include MPEG-encoded data from ground truth frames; and

generate, by inputting the two anchor frames into a second deep machine learning model identical to the deep machine learning model, at least one intermediate frame between the two anchor frames, wherein the at least one intermediate frame includes receiver-generated data.

2 . The system of claim 1 , wherein the first non-transitory memory is further configured with executable instructions that when executed by the first processor cause the first processor to:

compress a subset of the at least one intermediate frame; and

combine the two compressed anchor frames and the compressed subset of the at least one intermediate frame into a compressed video, the compressed video comprising the selectively encoded portion of the video.

3 . A system, comprising:

a hybrid encoding system configured to selectively encode a portion of a video; and

a decoding system communicatively coupled to the hybrid encoding system and configured to, synthesize, with a deep machine learning model, a remainder of the video upon receiving the selectively encoded portion of the video from the hybrid encoding system, and selectively combine the synthesized remainder of the video with the portion of the video;

wherein the hybrid encoding system comprises a first processor and a first non-transitory memory, the first non-transitory memory configured with executable instructions that when executed by the first processor cause the first processor to:

compress two anchor frames in the video;

generate, by inputting the two anchor frames into a second deep machine learning model identical to the deep machine learning model, at least one intermediate frame between the two anchor frames;

compress a subset of the at least one intermediate frame;

select the subset of the at least one intermediate frame based on an image quality metric of at least one macroblock of the at least one intermediate image frame below an image quality metric threshold; and

combine the two compressed anchor frames and the compressed subset of the at least one intermediate frame into a compressed video, the compressed video comprising the selectively encoded portion of the video.

4 . The system of claim 3 , wherein the image quality metric comprises a peak signal-to-noise ratio (PSNR) or a video multi-method assessment fusion (VMAF).

5 . The system of claim 3 , wherein the decoding system comprises a second processor and a second non-transitory memory, the second non-transitory memory configured with executable instructions that when executed by the second processor cause the second processor to:

decompress the at least two compressed anchor frames to obtain at least two decompressed anchor frames;

decompress the compressed subset of the at least one intermediate frame to obtain a decompressed subset of the at least one intermediate frame;

generate, by inputting the at least two decompressed anchor frames into the deep machine learning model, a synthesized image frame corresponding to the at least one intermediate frame between the at least two anchor frames;

reconstruct a hybrid image frame by combining the decompressed subset of the at least one intermediate frame with the synthesized image frame; and

output, to a display device, a video comprising the at least two decompressed anchor frames and the hybrid image frame.

6 . The system of claim 3 , wherein an anchor-frame distance between the two anchor frames is greater than two, and wherein the at least one intermediate frame between the two anchor frames comprises at least two intermediate frames between the two anchor frames.

7 . The system of claim 3 , wherein the deep machine learning model comprises a video frame synthesis neural network.

8 . The system of claim 1 , wherein the hybrid encoding system performs the MPEG-based encoding of the two anchor frames by using H.264/MPEG-4 or high efficiency video coding (HEVC).

9 . The system of claim 3 , wherein the first non-transitory memory is further configured with executable instructions that when executed by the first processor cause the first processor to transmit the compressed video to a decoding system.

10 . The system of claim 9 , wherein the decoding system is configured to:

decompress the two compressed anchor frames into two decompressed anchor frames and the compressed subset into a decompressed subset;

generate, by inputting the two decompressed anchor frames of the compressed video into a second deep machine learning model, at least one intermediate frame between the two decompressed anchor frames; and

combine the at least one intermediate frame between the two decompressed anchor frames with the decompressed subset of the at least one intermediate frame to generate a hybrid intermediate frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: LIU, FENG; FENG, WU-CHI
To: PORTLAND STATE UNIVERSITY
Reel/Frame 063468/0781 →
Continuity (2)
Provisional Application 63111514 · Nov 9, 2020
Related Publication 20230412796A1 · Dec 21, 2023
References Cited (21)
US 10445921B1 · Li · 2019 [cited by examiner]
US 10701394B1 · Caballero et al. · 2020 [cited by applicant]
US 10848734B1 · Waggoner · 2020 [cited by examiner]
US 20190289257A1 · Schroers et al. · 2019 [cited by applicant]
US 20200280730A1 · Wang et al. · 2020 [cited by applicant]
US 20200327702A1 · Wang et al. · 2020 [cited by applicant]
US 20200342234A1 · Gan · 2020 [cited by examiner]
US 20210280191A1 · Geng · 2021 [cited by examiner]
US 20210342686A1 · Kothari · 2021 [cited by examiner]
US 20220273139A1 · Mahapatra · 2022 [cited by examiner]
US 20220301563A1 · Chang · 2022 [cited by examiner]
US 20220303555A1 · Lu et al. · 2022 [cited by applicant]
US 20230276070A1 · Yang · 2023 [cited by examiner]
JP 2004048627A · 2004 [cited by applicant]
WO 2018199051A1 · 2018 [cited by applicant]
WO 2019168765A1 · 2019 [cited by applicant]
WO 2020150264A1 · 2020 [cited by applicant]
ISA United States Patent and Trademark Office, International Search Report and Written Opinion Issued in Application No. PCT/US2021/072288, Jan. 31, 2022, WIPO, 7 pages. [cited by applicant]
Japanese Patent Office, Office Action Issued in Application No. 2025-528092, Aug. 14, 2025, 7 pages. [cited by applicant]
Wu, C. et al., “Video Compression through Image Interpolation,” ArXiv Cornell University Website, Available Online at https://arxiv.org/abs/1804.06919, Apr. 18, 2018, 18 pages. [cited by applicant]
Japan Patent Office, Office Action Issued in Application No. 2023-528092, Mar. 24, 2026, 10 pages. (Submitted with English Translation). [cited by applicant]