Modeling disjoint manifolds
A computer model is trained to account for data samples in a high-dimensional space as lying on different manifolds, rather than a single manifold to represent the data set, accounting for the data set as a whole as a union of manifolds. Different data samples that may be expected to belong to the same underlying manifold are determined by grouping the data. For generative models, a generative model may be trained that includes a sub-model for each group trained on that group's data samples, such that each sub-model can account for the manifold of that group. The overall generative model includes information describing the frequency to sample from each sub-model to correctly represent the data set as a whole in sampling. Multi-class classification models may also use the grouping to improve classification accuracy by weighing group data samples according to the estimated latent dimensionality of the group.
1. A system for a training a generative model of data on disjoint manifolds, comprising:
one or more processors;
one or more non-transitory computer-readable media containing instructions for execution by the one or more processors for:
identifying a plurality of training samples for which to train a generative model;
grouping the plurality of training samples to a plurality of groups;
generating a plurality of generative sub-models corresponding to a number of the plurality of groups by, for each group of the plurality of groups:
identifying a sampling frequency for sampling the sub-model based on a number of training samples associated with the group relative to the plurality of training samples; and
training a generative sub-model for the group based on the training samples of the group; and
storing the generative model as the plurality of generative sub-models and the associated sampling frequency for each sub-model.
2. The system of claim 1 , wherein each sub-model models a different continuous manifold of a high-dimensional space of the training samples.
3. The system of claim 1 , wherein at least one of the generative sub-models is a pushforward model from a latent space having lower dimensionality than a dimensionality of a high-dimensional space of the training data samples.
4. The system of claim 1 , wherein training the generative sub-model for at least one group comprises:
determining a latent dimensionality of the group based on the data samples of the group;
setting one or more parameters for the generative sub-model based on the latent dimensionality of the group; and
training the generative sub-model for the group based on the one or more parameters.
5. The system of claim 1 , wherein the plurality of generative sub-models include modeling with respect to latent spaces that do not have the same latent dimensionality.
6. The system of claim 1 , the instructions further being for:
receiving a sampling request to generate a total number of samples from the generative model;
determining, based on the associated sampling frequency of each sub-model, a sub-model sample quantity for each sub-model;
generating a set of model samples by generating samples from each sub-model according to the sample quantity; and
providing the set of model samples as a response to the sampling request.
7. The system of claim 6 , wherein the associated sampling frequency for each sub-model is represented as a probability distribution; and determining the sub-model sample quantity for the sub-model comprises sampling from the probability distribution a number of times according to the total number of samples for the generative model.
8. The system of claim 6 , wherein generating samples from each sub-model according to the sample quantity comprises:
loading a first sub-model to a memory;
sampling the first sub-model at the associated sub-model sample quantity;
after generating all samples for the first sub-model, loading a second sub-model to the memory; and
sampling the second sub-model at the associated sub-model sample quantity.
9. The system of claim 1 , wherein grouping the plurality of training samples comprises an agglomerative clustering algorithm.
10. The system of claim 1 , wherein the plurality of training samples are images.
11. A method for a training a generative model of data on disjoint manifolds, comprising:
identifying a plurality of training samples for which to train a generative model;
grouping the plurality of training samples to a plurality of groups;
generating a plurality of generative sub-models corresponding to a number of the plurality of groups by, for each group of the plurality of groups:
identifying a sampling frequency for sampling the sub-model based on a number of training samples associated with the group relative to the plurality of training samples; and
training a generative sub-model for the group based on the training samples of the group; and
storing the generative model as the plurality of generative sub-models and the associated sampling frequency for each sub-model.
12. The method of claim 11 , wherein each sub-model models a different continuous manifold of a high-dimensional space of the training samples.
13. The method of claim 11 , wherein at least one of the generative sub-models is a pushforward model from a latent space having lower dimensionality than a dimensionality of a high-dimensional space of the training data samples.
14. The method of claim 11 , wherein training the generative sub-model for at least one group comprises:
determining a latent dimensionality of the group based on the data samples of the group;
setting one or more parameters for the generative sub-model based on the latent dimensionality of the group; and
training the generative sub-model for the group based on the one or more parameters.
15. The method of claim 11 , wherein the plurality of generative sub-models include modeling with respect to latent spaces that do not have the same latent dimensionality.
16. The method of claim 11 , the method further comprising:
receiving a sampling request to generate a total number of samples from the generative model;
determining, based on the associated sampling frequency of each sub-model, a sub-model sample quantity for each sub-model;
generating a set of model samples by generating samples from each sub-model according to the sample quantity; and
providing the set of model samples as a response to the sampling request.
17. The method of claim 16 , wherein the associated sampling frequency for each sub-model is represented as a probability distribution; and determining the sub-model sample quantity for the sub-model comprises sampling from the probability distribution a number of times according to the total number of samples for the generative model.
18. The method of claim 16 , wherein generating samples from each sub-model according to the sample quantity comprises:
loading a first sub-model to a memory;
sampling the first sub-model at the associated sub-model sample quantity;
after generating all samples for the first sub-model, loading a second sub-model to the memory; and
sampling the second sub-model at the associated sub-model sample quantity.
19. The method of claim 11 , wherein grouping the plurality of training samples comprises an agglomerative clustering algorithm.
20. The method of claim 11 , wherein the plurality of training samples are images.