IP Library Granted Patent US 12,652,404
Granted Patent B2
US 12,652,404 · App. 18/806,804 · Granted Jun 9, 2026

Systems and methods for video coding for machines using an autoencoder

Inventors: Hari Kalva (Boca Raton, FL); Borivoje Furht (Boca Raton, FL); Velibor Adzic (Canton, GA)
Assignee: OP Solutions LLC
H04N19/42H04N19/169
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,652,404
App. No.
18/806,804
Granted
Jun 9, 2026
Kind
B2
Abstract

Systems and methods for encoding and decoding video for machine consumption (video coding for machines) are provided in which an autoencoder is employed. The autoencoder has an encoder portion, a bottleneck portion, and a decoder portion. The autoencoder being distributed between a VCM encoder and VCM decoder such that the VCM encoder includes the encoder portion and bottleneck portion and the VCM decoder includes the bottleneck portion and the decoder portion.

Claims (34)

1 . An encoder for video for video coding for machine applications, the encoder comprising:

an autoencoder encoder portion;

a latent feature encoder coupled to the autoencoder encoder portion; and

an autoencoder description encoder, the autoencoder description encoder being coupled to the autoencoder encoder portion;

a multiplexor, the multiplexor coupled to the latent feature encoder and the autoencoder description encoder and providing a coded bitstream.

2 . The encoder of claim 1 , wherein the autoencoder encoder portion is a neural network comprising a plurality of successive encoder layers and a bottle neck layer after the final encoder layer.

3 . The encoder of claim 2 , wherein with each successive encoder layer having fewer neurons than the prior layer, and a bottle neck layer having fewer neurons than the final encoder layer.

4 . The encoder of claim 3 , wherein the bottle neck layer has fewer neurons than the final encoder layer.

5 . The encoder of claim 2 , wherein the number of encoder layers is the same as a number of layers in a decoder portion of a compatible decoder.

6 . An decoder for encoded video for video coding for machine applications, the decoder comprising:

a demultiplexor receiving a bitstream encoded with an autoencoder;

a latent feature decoder coupled tot the demultiplexor;

an autoencoder description decoder coupled to the demultiplexor;

an autoencoder decoder portion, the autoencoder decoder portion coupled to the autoencoder description decoder and the latent feature decoder, the autoencoder decoder portion providing decoded video for machine use.

7 . The decoder of claim 5 , wherein the autoencoder decoder portion is a neural network comprising a bottle neck layer as an input layer and a plurality of successive decoder layers.

8 . The decoder of claim 6 , wherein with each successive decoder layer has a higher number of neurons than the prior layer.

9 . The decoder of claim 7 , wherein the bottle neck layer has fewer neurons than the first decoder layer.

10 . The decoder of claim 6 , wherein the number of decoder layers is the same as a number of layers in an encoder portion of the encoder which generated the coded bitstream.

11 . A system for encoding image data in a video coding for machine application, the system comprising:

a video coding for machine (VCM) encoder, the VCM encoder comprising an autoencoder encoder portion and an autoencoder bottleneck portion, the VCM encoder receiving image information and generating a coded bitstream including feature data of the image information; and

a VCM decoder, the VCM decoder comprising the autoencoder bottleneck portion and an autoencoder decoder portion, the VCM decoder receiving the coded bitstream and reconstructing the image information from the feature data.

12 . The system of claim 11 , wherein the autoencoder encoder portion is a neural network comprising a plurality of successive encoder layers with each successive encoder layer having fewer neurons than the prior layer, and the bottle neck portion having fewer neurons than a final encoder layer.

13 . The system of claim 12 , wherein the autoencoder decoder portion is a neural network comprising a plurality of successive encoder layers with each successive decoder layer having a higher number of neurons than the prior layer, and the bottle neck portion having fewer neurons than the first decoder layer.

14 . The system of claim 13 , wherein the number of encoder layers in the encoder portion is the same as a number of layers in a decoder portion.

15 . The system of claim 11 , wherein the autoencoder encoder portion is in communication with the autoencoder decoder portion and updates to the autoencoder are communicated between the VCM encoder and VCM decoder.

16 . The system of claim 11 , wherein the VCM encoder further comprises:

an autoencoder description encoder, the autoencoder description encoder being coupled to the autoencoder encoder portion;

a latent feature encoder coupled to the autoencoder encoder portion; and

a multiplexor, the multiplexor coupled to the latent feature encoder and the autoencoder description encoder and providing a coded bitstream.

17 . The system of claim 11 , wherein VCM decoder further comprises:

a demultiplexor receiving the coded bitstream;

a latent feature decoder coupled to the demultiplexor;

an autoencoder description decoder coupled to the demultiplexor; and

the autoencoder decoder portion being coupled to the autoencoder description decoder and the latent feature decoder.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2025
From: FURHT, BORIVOJE; KALVA, HARI
To: FLORIDA ATLANTIC UNIVERSITY RESEARCH CORPORATION
Reel/Frame 073482/0048 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2025
From: FLORIDA ATLANTIC UNIVERSITY RESEARCH CORPORATION
To: OP SOLUTIONS, LLC
Reel/Frame 073482/0504 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2025
From: ADZIC, VELIBOR
To: OP SOLUTIONS, LLC
Reel/Frame 073482/0944 →
Continuity (3)
Continuation PCTUS2023013073 · Feb 15, 2023
Provisional Application 63311333 · Feb 17, 2022
Related Publication 20240406424A1 · Dec 5, 2024
References Cited (27)
US 5275553A · Frish · 1994 [cited by examiner]
US 6000612A · Xu · 1999 [cited by examiner]
US 11720994B2 · Luo · 2023 [cited by examiner]
US 20030151810A1 · Haisch · 2003 [cited by examiner]
US 20030152145A1 · Kawakita · 2003 [cited by examiner]
US 20050243171A1 · Ross · 2005 [cited by examiner]
US 20080148877A1 · Sim · 2008 [cited by examiner]
US 20090131811A1 · Morris · 2009 [cited by examiner]
US 20120062615A1 · Van Lier · 2012 [cited by examiner]
US 20120090757A1 · Buchan · 2012 [cited by examiner]
US 20130141558A1 · Jeon · 2013 [cited by examiner]
US 20170234709A1 · Mackie · 2017 [cited by examiner]
US 20230229918A1 · Park · 2023 [cited by examiner]
US 20230230228A1 · Liu · 2023 [cited by examiner]
US 20230237345A1 · Zahn · 2023 [cited by examiner]
US 20230273573A1 · Hildebrandt · 2023 [cited by examiner]
US 20230290438A1 · Alvarez · 2023 [cited by examiner]
US 20230307908A1 · Sun · 2023 [cited by examiner]
US 20240185043A1 · Yoon · 2024 [cited by examiner]
US 20240186018A1 · Chen · 2024 [cited by examiner]
US 20240187640A1 · Racape · 2024 [cited by examiner]
US 20240193412A1 · Bai · 2024 [cited by examiner]
US 20240193827A1 · Persson · 2024 [cited by examiner]
US 20240194303A1 · Holderrieth · 2024 [cited by examiner]
Wang et al. “Sparse Tensorbased Multiscale Representation for Point Cloud Geometry Compression” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, No. 7, pp. 9055-9071, Jul. 1, 2023, doi: 10.110… [cited by examiner]
Pessoa et al.entitled “End-to-End Learning of Video Compression Using Spatio-Temporal Autoencoders” 2020 IEEE Workshop on Signal Processing Systems (SiPS), Coimbra, Portugal, 2020, pp. 1-6, doi: 10.1109/SiPS50750.2020.9… [cited by examiner]
Ma et al. “Image and Video Compression with Neural Networks: A Review.” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 6, pp. 1683-1698, Jun. 2020, doi: 10.1109/TCSVT.2019.2910119. (Year… [cited by examiner]