IP Library › Patent Application 19390297
Patent Application
App. No. 19/390,297

PROCESSING MULTI-DIMENSIONAL DATA USING NEURAL STATE-SPACE MODELS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/390,297
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for processing an input sequence of multi-dimensional data using a state-space model (SSM) system. The SSM system includes a neural state-space model. The neural state-space model includes a stack of one or more SSM layers.

Claims (49)

1 . A method performed by one or more computers, the method comprising:

receiving an input sequence of multidimensional data;

partitioning the input sequence of multidimensional data into a sequence of data segments;

for each data segment that is after a first data segment in the sequence of data segments:

processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and

generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and

for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.

2 . The method of claim 1 , further comprising, for the first data segment in the sequence of data segments:

processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;

generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and

generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.

3 . The method of claim 1 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.

4 . The method of claim 3 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.

5 . The method of claim 1 , wherein the input sequence of multi-dimensional data comprises an input sequence of audio data.

6 . The method of claim 1 , wherein the state matrix is a diagonal matrix.

7 . The method of claim 3 , wherein the encoder neural network comprises a vision Transformer neural network, and wherein the encoded representation of the data segment comprises a plurality of visual tokens.

8 . The method of claim 3 , wherein the decoder neural network comprises a text decoder neural network, and wherein the output that characterizes the data segment comprises a text caption of a video segment that comprises one or more video frames.

9 . A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving an input sequence of multidimensional data;

partitioning the input sequence of multidimensional data into a sequence of data segments;

for each data segment that is after a first data segment in the sequence of data segments:

processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and

generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and

for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.

10 . The system of claim 9 , wherein the operations further comprise, for the first data segment in the sequence of data segments:

processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;

generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and

generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.

11 . The system of claim 9 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.

12 . The system of claim 11 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.

13 . The system of claim 9 , wherein the input sequence of multi-dimensional data comprises an input sequence of audio data.

14 . The system of claim 9 , wherein the state matrix is a diagonal matrix.

15 . The system of claim 11 , wherein the encoder neural network comprises a vision Transformer neural network, and wherein the encoded representation of the data segment comprises a plurality of visual tokens.

16 . The system of claim 11 , wherein the decoder neural network comprises a text decoder neural network, and wherein the output that characterizes the data segment comprises a text caption of a video segment that comprises one or more video frames.

17 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving an input sequence of multidimensional data;

partitioning the input sequence of multidimensional data into a sequence of data segments;

for each data segment that is after a first data segment in the sequence of data segments:

processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and

generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and

for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.

18 . The computer-readable storage media of claim 17 , wherein the operations further comprise, for the first data segment in the sequence of data segments:

processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;

generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and

generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.

19 . The computer-readable storage media of claim 17 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.

20 . The computer-readable storage media of claim 19 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2025
From: PIERGIOVANNI, ANTHONY JACOB; SATISH MALLYA, GANESH; KIM, DAHUN; ANGELOVA, ANELIA
To: GDM HOLDING LLC
Reel/Frame 073136/0974 →