IP Library › Granted Patent US 12,394,103
Granted Patent B2
US 12,394,103 · App. 17/695,684 · Granted Aug 19, 2025

Class-specific neural network for video compressed sensing

Inventors: Yifei Pei (Santa Clara, CA); Ying Liu (Santa Clara, CA); Nam Ling (Santa Clara, CA); Lingzhi Liu (San Jose, CA); Yongxiong Ren (San Jose, CA); Ming Kai Hsu (Fremont, CA)
Assignees: KWAI INC.; SANTA CLARA UNIVERSITY
G06T9/002H04N19/176H04N19/625
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,103
App. No.
17/695,684
Granted
Aug 19, 2025
Kind
B2
Abstract

A class-specific neural network for video compressed sensing and methods for training and testing the class-specific neural network are provided. The class-specific neural network includes a Gaussian-mixture model (GMM) and a plurality of encoders, where the GMM classifies video frame blocks with a plurality of clusters and assigns the video frame blocks to the plurality of clusters. Further, the plurality of encoders receive the video frame blocks and generate a plurality of compressed-sensed frame block vectors, where the plurality of encoders correspond to the plurality of clusters.

Claims (48)

1. A method for video compressed sensing by a class-specific neural network, comprising:

classifying, by a Gaussian-mixture model (GMM), video frame blocks with a plurality of clusters and assigning the video frame blocks to the plurality of clusters;

receiving, by a plurality of encoders, the video frame blocks;

generating, by the plurality of encoders, a plurality of compressed-sensed frame block vectors, wherein the plurality of encoders respectively correspond to the plurality of clusters,

wherein each encoder comprises a flatten layer, a discrete cosine transform (DCT) transform layer, and a trainable compressed sensing layer;

predicting, by a logistic regression classifier, class labels for the plurality of compressed-sensed frame block vectors without recording clustering information, wherein the logistic regression classifier predicts the class labels by respectively maximizing probabilities of the plurality of compressed-sensed frame block vectors;

sending, by the logistic regression classifier, the compressed-sensed frame block vectors to a plurality of decoders, and

wherein the plurality of compressed-sensed frame block vectors are assigned to the plurality of decoders based on the class labels for the plurality of compressed-sensed frame block vectors, wherein the class labels are saved in a hashmap data structure.

2. The method of claim 1 , wherein the GMM predicts labels for the video frame blocks.

3. The method of claim 1 , wherein each decoder comprises an expansion layer, a trainable reconstruction layer, an inverse DCT transform layer, and a reshape layer that converts the plurality of compressed-sensed frame block vectors to a plurality of predicted block matrices.

4. A method for training a class-specific neural network for video compressed sensing, comprising:

training a Gaussian-mixture model (GMM) in the class-specific neural network using a plurality of video frame blocks, wherein the GMM comprises a plurality of clusters;

assigning the plurality of video frame blocks to the plurality of clusters;

training a plurality of encoders that respectively correspond to the plurality of clusters using the plurality of video frame blocks,

wherein each encoder comprises a flatten layer, a discrete cosine transform (DCT) transform layer, and a trainable compressed sensing layer,

generating class labels for a plurality of compressed-sensed frame block vectors by labeling the plurality of compressed-sensed frame block vectors according to the plurality of clusters;

training a logistic regression classifier for the plurality of compressed-sensed frame block vectors based on the class labels without recording clustering information, wherein the logistic regression classifier predicts the labels by respectively maximizing probabilities of the plurality of compressed-sensed frame block vectors; and

sending, by the logistic regression classifier, the compressed-sensed frame block vectors to a plurality of decoders,

wherein the plurality of compressed-sensed frame block vectors are assigned to the plurality of decoders based on the class labels for the plurality of compressed-sensed frame block vectors, wherein the class labels are saved in a hashmap data structure.

5. A method for testing a class-specific neural network for video compressed sensing, comprising:

assigning and sending, by a trained Gaussian-mixture model (GMM) in the class-specific neural network, a plurality of video frame blocks to a plurality of clusters;

generating, by a plurality of encoders in the class-specific neural network, a plurality of compressed-sensed frame block vectors,

wherein each encoder comprises a flatten layer, a discrete cosine transform (DCT) transform layer, and a trainable compressed sensing layer,

classifying, by a trained logistic regression classifier, class labels of the plurality of compressed-sensed frame block vectors without recording clustering information, wherein the trained logistic regression classifier predicts the class labels by respectively maximizing probabilities of the plurality of compressed-sensed frame block vectors; and

assigning, by the trained logistic regression classifier, the plurality of compressed-sensed frame block vectors to corresponding decoders based on the class labels to reconstruct a whole video frame,

wherein the class labels are saved in a hashmap data structure.

6. An apparatus for training a class-specific neural network for video compressed sensing, comprising:

one or more processors; and

a memory configured to store instructions executable by the one or more processors,

wherein the one or more processors, upon execution of the instructions, are configured to:

train a Gaussian-mixture model (GMM) in the class-specific neural network using a plurality of video frame blocks, wherein the GMM comprises a plurality of clusters;

assign the plurality of video frame blocks to the plurality of clusters;

train a plurality of encoders that respectively correspond to the plurality of clusters using the plurality of video frame blocks,

wherein each encoder comprises a flatten layer, a discrete cosine transform (DCT) transform layer, and a trainable compressed sensing layer,

generate class labels for a plurality of compressed-sensed frame block vectors by labeling the plurality of compressed-sensed frame block vectors according to the plurality of clusters;

train a logistic regression classifier for the plurality of compressed-sensed frame block vectors based on the class labels without recording clustering information, wherein the logistic regression classifier predicts the labels by respectively maximizing probabilities of the plurality of compressed-sensed frame block vectors; and

send, by the logistic regression classifier, the compressed-sensed frame block vectors to a plurality of decoders,

wherein the plurality of compressed-sensed frame block vectors are assigned to the plurality of decoders based on the class labels for the plurality of compressed-sensed frame block vectors, wherein the class labels are saved in a hashmap data structure.

7. An apparatus for testing a class-specific neural network for video compressed sensing, comprising:

one or more processors; and

a memory configured to store instructions executable by the one or more processors,

wherein the one or more processors, upon execution of the instructions, are configured to:

assign and send, by a trained Gaussian-mixture model (GMM) in the class-specific neural network, a plurality of vectorized video frame blocks to corresponding clusters; and

generate, by a plurality of encoders in the class-specific neural network, a plurality of compressed-sensed frame block vectors,

wherein each encoder comprises a flatten layer, a discrete cosine transform (DCT) transform layer, and a trainable compressed sensing layer,

classify, by a trained logistic regression classifier, class labels of the plurality of compressed-sensed frame block vectors without recording clustering information, wherein the trained logistic regression classifier predicts the class labels by respectively maximizing probabilities of the plurality of compressed-sensed frame block vectors; and

assign, by the trained logistic regression classifier, the plurality of compressed-sensed frame block vectors to corresponding decoders based on the class labels to reconstruct a whole video frame,

wherein the class labels are saved in a hashmap data structure.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2026
From: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
To: BEIJING TRANSTREAMS TECHNOLOGY CO., LTD.
Reel/Frame 074656/0721 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2025
From: KWAI INC.
To: BEIJING DAJIA INTERNET INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 073219/0688 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2022
From: LIU, LINGZHI; REN, YONGXIONG; HSU, MING KAI
To: KWAI INC.
Reel/Frame 059330/0063 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2022
From: PEI, YIFEI; LIU, YING; LING, NAM
To: SANTA CLARA UNIVERSITY
Reel/Frame 059330/0124 →
Continuity (2)
Provisional Application 63161431 · Mar 15, 2021
Related Publication 20220292727A1 · Sep 15, 2022
References Cited (4)
US 20150222859A1 · Schweid · 2015 [cited by examiner]
US 20210185276A1 · Peters · 2021 [cited by examiner]
CA 3138340A1 · 2020 [cited by examiner]
Cheng “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules”, arXiv:2001.01568v3, Mar. 30, 2020. (Year: 2020). [cited by examiner]