IP Library Granted Patent US 10,885,111
Granted Patent B2
US 10,885,111 · App. 15/953,986 · Granted Jan 5, 2021

Generating cross-domain data using variational mapping between embedding spaces

Inventors: Subhajit Chaudhury (Kanagawa, JP); Sakyasingha Dasgupta (Tokyo, JP); Asim Munawar (Ichikawa, JP); Ryuki Tachibana (Kanagawa, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/86G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,885,111
App. No.
15/953,986
Granted
Jan 5, 2021
Kind
B2
Abstract

A computer-implemented method, computer program product, and system are provided for learning mapping information between different modalities of data. The method includes mapping, by a processor, high-dimensional modalities of data into a low-dimensional manifold to obtain therefor respective low-dimensional embeddings through at least a part of a first network. The method further includes projecting, by the processor, each of the respective low-dimensional embeddings to a common latent space to obtain therefor a respective one of separate latent space distributions in the common latent space through at least a part of a second network. The method also includes optimizing, by the processor, parameters of each of the networks by minimizing a distance between the separate latent space distributions in the common latent space using a variational lower bound. The method additionally includes outputting, by the processor, the parameters as the mapping information.

Claims (33)

1. A computer-implemented method for learning mapping information between different modalities of data, comprising:

mapping, by a processor, high-dimensional modalities of data into a low-dimensional manifold to obtain therefor respective low-dimensional embeddings through at least a part of a first network;

projecting, by the processor, each of the respective low-dimensional embeddings to a common latent space to obtain therefor a respective one of separate latent space distributions in the common latent space through at least a part of a second network;

optimizing, by the processor, parameters of each of the networks by minimizing a distance between the separate latent space distributions in the common latent space using a variational lower bound; and

outputting, by the processor, the parameters as the mapping information.

2. The computer-implemented method of claim 1 , wherein said optimizing step is performed to minimize a respective reconstruction error for reconstructed versions of each of the different modalities of data.

3. The computer-implemented method of claim 1 , wherein the variational lower bound comprises a set of terms for making the separate latent space distributions have a same semantic meaning by minimizing a distance between the separate latent space distributions in the common latent space.

4. The computer-implemented method of claim 1 , wherein the variational lower bound comprises a set of terms for forcing the separate latent space distributions to be identical to a prior latent distribution designed by a user to enable Bayesian inference between the different modalities.

5. The computer-implemented method of claim 1 , wherein the variational lower bound comprises a set of terms for minimizing a reconstruction error to equalize input and output data instances, the input data instances comprising the high dimensional modalities of data, the output data instances comprising reconstructed versions of the high dimensional modalities of data.

6. The computer-implemented method of claim 1 , wherein separate generative models are used for the first network and the second network.

7. The computer-implemented method of claim 1 , wherein the variational lower bound comprises a set of terms that comprise an embedding reconstruction loss, a prior latent distribution divergence, and a distance measure between the separate latent space distributions.

8. The computer-implemented method of claim 1 , wherein the second network comprises a set of stochastic neural networks.

9. The computer-implemented method of claim 1 , wherein said mapping step uses variational mapping to encode the high-dimensional modalities into the low-dimensional embeddings.

10. The computer-implemented method of claim 1 , wherein the different modalities of data comprise a first modality of data and a second modality of data, and wherein the method further comprises generating cross-domain data of one of the first modality or the second modality based on data of another one of the first modality or the second modality and the mapping information.

11. The computer-implemented method of claim 10 , further comprising populating a database with the cross-domain data responsive to the database being incompatible with the data of the other one of the first modality or the second modality.

12. The computer-implemented method of claim 10 , further comprising captioning an image represented by one of the first modality or the second modality with text represented by the other one of the first modality or the second modality.

13. The computer-implemented method of claim 1 , further comprising performing cross modal data generation based on the mapping information.

14. The computer-implemented method of claim 1 , wherein the common latent space is configured for decoupled encoding between the different modalities of data.

15. A computer program product for learning mapping information between different modalities of data, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

mapping, by a processor of the computer, high-dimensional modalities of data into a low-dimensional manifold to obtain therefor respective low-dimensional embeddings through at least a part of a first network;

projecting, by the processor, each of the respective low-dimensional embeddings to a common latent space to obtain therefor a respective one of separate latent space distributions in the common latent space through at least a part of a second network;

optimizing, by the processor, parameters of each of the networks by minimizing a distance between the separate latent space distributions in the common latent space using a variational lower bound; and

outputting, by the processor, the parameters as the mapping information.

16. The computer program product of claim 15 , wherein said optimizing step is performed to minimize a respective reconstruction error for reconstructed versions of each of the different modalities of data.

17. The computer program product of claim 15 , wherein the variational lower bound comprises a set of terms for making the separate latent space distributions have a same semantic meaning by minimizing a distance between the separate latent space distributions in the common latent space.

18. The computer program product of claim 15 , wherein the variational lower bound comprises a set of terms for forcing the separate latent space distributions to be identical to a prior latent distribution designed by a user to enable B ayesian inference between the different modalities.

19. The computer program product of claim 15 , wherein the variational lower bound comprises a set of terms for minimizing a reconstruction error to equalize input and output data instances, the input data instances comprising the high dimensional modalities of data, the output data instances comprising reconstructed versions of the high dimensional modalities of data.

20. A system for learning mapping information between different modalities of data, comprising:

a processor, configured to

map high-dimensional modalities of data into a low-dimensional manifold to obtain therefor respective low-dimensional embeddings through at least a part of a first network;

project each of the respective low-dimensional embeddings to a common latent space to obtain therefor a respective one of separate latent space distributions in the common latent space through at least a part of a second network;

optimize parameters of each of the networks by minimizing a distance between the separate latent space distributions in the common latent space using a variational lower bound; and

output the parameters as the mapping information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2018
From: CHAUDHURY, SUBHAJIT; DASGUPTA, SAKYASINGHA; MUNAWAR, ASIM; TACHIBANA, RYUKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 045553/0285 →
Continuity (1)
Related Publication 20190318040A1 · Oct 17, 2019
Cited By (2)
US 12,217,191 US 12,608,621