IP Library › Granted Patent US 12,744,944
Granted Patent B2
US 12,744,944 · App. 18/680,485 · Granted Sep 22, 2026

Machine learning networks for hybrid video compression and corresponding decompression

Inventors: Matthew Lawrence Bronder (Bellevue, WA); Saswata Mandal (Bellevue, WA); Sameer Avinash Nene (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
H04N19/86H04N19/132H04N19/423H04N19/60H04N19/172H04N19/174
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,944
App. No.
18/680,485
Granted
Sep 22, 2026
Kind
B2
Abstract

Innovations in machine learning (“ML”) networks used in video processing scenarios are described. For example, an ML refinement network can be used to refine video after a video decoder has reconstructed the video. Using the ML refinement network for post-processing can mitigate compression artifacts introduced during encoding and otherwise improve the quality of the reconstructed video. Or, as another example, an ML encoder network and ML decoder network can be used, in combination with a core video encoder and core video decoder, for hybrid compression and corresponding decompression. In the hybrid compression, the ML encoder network can transform video before encoding in order to boost rate-distortion performance of the core video encoder. In corresponding decompression, the ML decoder network can enhance reconstructed video after decoding, thereby compensating for transformations applied by the ML encoder network, mitigating compression artifacts, and otherwise improving the quality of the reconstructed video.

Claims (50)

1 . A server computer system comprising a processor system and memory, wherein the server computer system is configured to perform operations comprising:

receiving a current unit of input video;

retrieving a given previous unit;

warping the given previous unit to spatially align sample values of the given previous unit with expected locations in a version of the current unit, thereby producing a given warped previous unit;

providing the given warped previous unit to a machine learning (“ML”) encoder network;

with the ML encoder network, transforming the current unit to facilitate preservation of image quality, thereby producing a transformed current unit, wherein, as part of a temporal feedback loop, the transforming the current unit is based at least in part on the given warped previous unit;

encoding the transformed current unit, thereby producing encoded data for the transformed current unit; and

outputting the encoded data as part of a bitstream.

2 . The server computer system of claim 1 , wherein the transforming the current unit also partially compresses the current unit by downsampling the current unit.

3 . The server computer system of claim 1 , wherein the ML encoder network is a convolutional neural network having a U-Net architecture.

4 . The server computer system of claim 1 , wherein the operations further comprise:

decoding the encoded data, thereby producing a decoded current unit; and

with an ML decoder network, enhancing the decoded current unit to compensate for transformations applied by the ML encoder network and mitigate compression artifacts, thereby producing an enhanced current unit.

5 . The server computer system of claim 4 , wherein the enhancing the decoded current unit also partially decompresses the decoded current unit by upsampling the decoded current unit.

6 . The server computer system of claim 4 , wherein the ML decoder network is a convolutional neural network having a U-Net architecture.

7 . The server computer system of claim 4 , wherein the operations further comprise:

storing, in a decoded video buffer, the decoded current unit for use in providing temporal feedback to the ML decoder network.

8 . The server computer system of claim 4 , wherein the given previous unit is a given decoded previous unit retrieved from a decoded video buffer, wherein the given warped previous unit is a given warped, decoded previous unit, and wherein the operations further comprise:

providing the given warped, decoded previous unit to the ML decoder network, wherein the enhancing the decoded current unit is based at least in part on the given warped, decoded previous unit.

9 . The server computer system of claim 4 , wherein the operations further comprise:

storing, in an enhanced video buffer, the enhanced current unit for use in providing temporal feedback to the ML encoder network and the ML decoder network.

10 . The server computer system of claim 4 , wherein the given previous unit is a given enhanced previous unit retrieved from an enhanced video buffer, wherein the given warped previous unit is a given warped, enhanced previous unit, and wherein the operations further comprise:

providing the given warped, enhanced previous unit to the ML decoder network, wherein the enhancing the decoded current unit is based at least in part on the given warped, enhanced previous unit.

11 . A computer system comprising a processor system and memory, wherein the computer system is configured to perform operations comprising:

receiving encoded data for a current unit;

decoding the encoded data, thereby producing a decoded current unit;

retrieving a given previous unit;

warping the given previous unit to spatially align sample values of the given previous unit with expected locations in a version of the current unit, thereby producing a given warped previous unit;

providing the given warped previous unit to a machine learning (“ML”) decoder network; and

with the ML decoder network, enhancing the decoded current unit to compensate for transformations applied by an ML encoder network and mitigate compression artifacts, thereby producing an enhanced current unit, wherein, as part of a temporal feedback loop, the enhancing the decoded current unit is based at least in part on the given warped previous unit.

12 . The computer system of claim 11 , wherein the enhancing the decoded current unit also partially decompresses the decoded current unit by upsampling the decoded current unit.

13 . The computer system of claim 11 , wherein the ML decoder network is a convolutional neural network having a U-Net architecture.

14 . The computer system of claim 11 , wherein the current unit is a frame, a slice, or a tile.

15 . The computer system of claim 11 , wherein the operations further comprise:

storing, in a decoded video buffer, the decoded current unit for use in providing temporal feedback to the ML decoder network.

16 . The computer system of claim 11 , wherein the given previous unit is a given decoded previous unit retrieved from a decoded video buffer, and wherein the given warped previous unit is a given warped, decoded previous unit.

17 . The computer system of claim 11 , wherein the operations further comprise:

storing, in an enhanced video buffer, the enhanced current unit for use in providing temporal feedback to the ML decoder network.

18 . The computer system of claim 11 , wherein the given previous unit is a given enhanced previous unit retrieved from an enhanced video buffer, and wherein the given warped previous unit is a given warped, enhanced previous unit.

19 . The computer system of claim 11 , wherein the operations further comprise:

processing the enhanced current unit for display; and

outputting results of the processing the enhanced current unit for display.

20 . In a computer system, a method of training a machine learning (“ML”) encoder network and an ML decoder network for hybrid compression of video and corresponding decompression, the method comprising:

receiving a current unit of input video;

with an ML encoder network, transforming the current unit to facilitate preservation of image quality, thereby producing a transformed current unit;

encoding the transformed current unit, thereby producing encoded data for the transformed current unit;

decoding the encoded data, thereby producing a decoded current unit;

with an ML decoder network, enhancing the decoded current unit to compensate for transformations applied by the ML encoder network and mitigate compression artifacts, thereby producing an enhanced current unit;

determining feedback based at least in part on differences between the current unit of input video and the enhanced current unit; and

adjusting at least one of the ML encoder network and the ML decoder network based at least in part on the feedback.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2025
From: BRONDER, MATTHEW LAWRENCE; NENE, SAMEER AVINASH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070992/0128 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2025
From: MANDAL, SASWATA
To: MICROSOFT CORPORATION
Reel/Frame 070992/0168 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2025
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 070992/0179 →
Continuity (1)
Related Publication 20250373859A1 · Dec 4, 2025
References Cited (37)
US 9298453B2 · Vangala et al. · 2016 [cited by applicant]
US 10979718B2 · Chou et al. · 2021 [cited by applicant]
US 11809998B2 · Mao et al. · 2023 [cited by applicant]
US 11895330B2 · Zhang et al. · 2024 [cited by applicant]
US 12034916B2 · Zhang et al. · 2024 [cited by applicant]
US 12501050B2 · Van Rozendaal et al. · 2025 [cited by applicant]
US 20190266436A1 · Prakash et al. · 2019 [cited by applicant]
US 20210136370A1 · Tourapis et al. · 2021 [cited by applicant]
US 20220405979A1 · Ding et al. · 2022 [cited by applicant]
US 20230065183A1 · Zirr et al. · 2023 [cited by applicant]
US 20230128106A1 · Otsuka · 2023 [cited by applicant]
US 20230274401A1 · Price et al. · 2023 [cited by applicant]
US 20230300341A1 · Wu et al. · 2023 [cited by applicant]
US 20230377096A1 · Dolgin et al. · 2023 [cited by applicant]
US 20240073438A1 · Wu et al. · 2024 [cited by applicant]
US 20240386704A1 · Zhang et al. · 2024 [cited by applicant]
US 20250373861A1 · Bronder et al. · 2025 [cited by applicant]
WO WO2025075787A1 · 2025 [cited by examiner]
Chan et al., “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment,” arXiv:2104.13371v1, 12 pp. (Apr. 27, 2021). [cited by applicant]
Liang et al., “Details or Artifacts: A Locally Discriminative Learning Approach to Realistic Image Super-Resolution,” IEEE Image and Video Processing, pp. 1-10 (Mar. 2022). [cited by applicant]
Mittag et al., “LSTM-Based Video Quality Prediction Accounting for Temporal Distortions in Videoconferencing Calls,” arXiv:2303.12761v1, 5 pp. (Mar. 2023). [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” arXiv:1505.04593v1, 8 pp. (May 18, 2015). [cited by applicant]
Thomas et al., “A Reduced-Precision Network for Image Reconstruction,” ACM Trans. Graph., vol. 39, No. 6, Article 231, 12 pp. (Dec. 2020). [cited by applicant]
Viola et al., “Rapid Object Detection Using a Boosted Cascade of Simple Features,” IEEE Conf. On Computer Vision and Pattern Recognition, pp. I-511-I-518 (Dec. 2001). [cited by applicant]
Wang et al., “ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks,” arXiv:1809.00219v2, 23 pp. (Sep. 2018). [cited by applicant]
Zhang et al., “Residual Dense Network for Image Super-Resolution,” arXiv:1802.08797v2, 10 pp. (Mar. 2018). [cited by applicant]
Benjak et al., “Learning-Based Scalable Video Coding with Spatial and Temporal Prediction,” Int'l Conf. on Visual Communications and Image Processing, 5 pp. (Dec. 2023). [cited by applicant]
International Search Report and Written Opinion dated Jun. 24, 2025, from International Patent Application No. PCT/US2025/018039, 12 pp. [cited by applicant]
International Search Report and Written Opinion dated Aug. 4, 2025, from International Patent Application No. PCT/US2025/029453, 17 pp. [cited by applicant]
Khani et al., “Efficient Video Compression via Content-Adaptive Super-Resolution,” IEEE/CVF Int'l Conf. on Computer Vision, pp. 4501-4510 (Oct. 2021). [cited by applicant]
Liu et al., “Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model,” IEEE Trans on Circuits and Systems for Video Technology, vol. 31, No. 8, pp. 3182-3196 (Aug. 2021). [cited by applicant]
Office Action dated Jun. 4, 2025, from U.S. Appl. No. 18/680,438, 54 pp. [cited by applicant]
Sivaraman et al., “Gemino: Practical and Robust Neural Compression for Video Conferencing,” arXiv:2209.10507v1, 18 pp. [cited by applicant]
Wei et al., “Video Compression based on Jointly Learned Down-Sampling and Super-Resolution Networks,” Int'l Conf. on Visual Communications and Image Processing, 5 pp. (Dec. 2021). [cited by applicant]
Xiang et al., “Learning Spatio-Temporal Downsampling for Effective Video Upscaling,” Springer International Publishing, pp. 162-181 (Nov. 2022). [cited by applicant]
Final Office Action dated Jan. 9, 2026, from U.S. Appl. No. 18/680,438, 50 pp. [cited by applicant]
Notice of Allowance dated Jun. 12, 2026, from U.S. Appl. No. 18/680,438, 9 pp. [cited by applicant]