PROCESSING MULTI-DIMENSIONAL DATA USING NEURAL STATE-SPACE MODELS
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for processing an input sequence of multi-dimensional data using a state-space model (SSM) system. The SSM system includes a neural state-space model. The neural state-space model includes a stack of one or more SSM layers.
1 . A method performed by one or more computers, the method comprising:
receiving an input sequence of multidimensional data;
partitioning the input sequence of multidimensional data into a sequence of data segments;
for each data segment that is after a first data segment in the sequence of data segments:
processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and
generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and
for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.
2 . The method of claim 1 , further comprising, for the first data segment in the sequence of data segments:
processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;
generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and
generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.
3 . The method of claim 1 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.
4 . The method of claim 3 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.
5 . The method of claim 1 , wherein the input sequence of multi-dimensional data comprises an input sequence of audio data.
6 . The method of claim 1 , wherein the state matrix is a diagonal matrix.
7 . The method of claim 3 , wherein the encoder neural network comprises a vision Transformer neural network, and wherein the encoded representation of the data segment comprises a plurality of visual tokens.
8 . The method of claim 3 , wherein the decoder neural network comprises a text decoder neural network, and wherein the output that characterizes the data segment comprises a text caption of a video segment that comprises one or more video frames.
9 . A system comprising:
one or more computers; and
one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:
receiving an input sequence of multidimensional data;
partitioning the input sequence of multidimensional data into a sequence of data segments;
for each data segment that is after a first data segment in the sequence of data segments:
processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and
generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and
for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.
10 . The system of claim 9 , wherein the operations further comprise, for the first data segment in the sequence of data segments:
processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;
generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and
generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.
11 . The system of claim 9 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.
12 . The system of claim 11 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.
13 . The system of claim 9 , wherein the input sequence of multi-dimensional data comprises an input sequence of audio data.
14 . The system of claim 9 , wherein the state matrix is a diagonal matrix.
15 . The system of claim 11 , wherein the encoder neural network comprises a vision Transformer neural network, and wherein the encoded representation of the data segment comprises a plurality of visual tokens.
16 . The system of claim 11 , wherein the decoder neural network comprises a text decoder neural network, and wherein the output that characterizes the data segment comprises a text caption of a video segment that comprises one or more video frames.
17 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
receiving an input sequence of multidimensional data;
partitioning the input sequence of multidimensional data into a sequence of data segments;
for each data segment that is after a first data segment in the sequence of data segments:
processing the data segment using an encoder neural network to generate an encoded representation of the data segment; and
generating a hidden state that corresponds to the data segment by using a neural state-space model to process the encoded representation of the data segment and a preceding hidden state that corresponds to a preceding data segment that precedes the data segment in the sequence of data segments, wherein the neural state-space model comprises one or more layers, each layer comprising a convolution kernel, the convolution kernel comprising parameters that represent a state matrix, an input matrix, and an output matrix of a state-space model; and
for each of one or more data segments in the sequence of data segments, generating, using a decoder neural network and based on a previous hidden state that corresponds to a preceding data segment, an output that characterizes the data segment.
18 . The computer-readable storage media of claim 17 , wherein the operations further comprise, for the first data segment in the sequence of data segments:
processing the first data segment using the encoder neural network to generate an encoded representation of the first data segment;
generating an output of the neural state-space model that corresponds to the first data segment by using the neural state-space model to process the encoded representation of the first data segment; and
generating, using the decoder neural network and from the output of the neural state-space model that corresponds to the first data segment, an output that characterizes the first data segment.
19 . The computer-readable storage media of claim 17 , wherein the input sequence of multi-dimensional data comprises an input sequence of video data.
20 . The computer-readable storage media of claim 19 , wherein the input sequence of video data comprises a live stream of video data that is being streamed.