IP Library Granted Patent US 12,205,028
Granted Patent B2
US 12,205,028 · App. 17/958,597 · Granted Jan 21, 2025

Co-disentagled series/text multi-modal representation learning for controllable generation

Inventors: Yuncong Chen (Plainsboro, NJ); Zhengzhang Chen (Princeton Junction, NJ); Xuchao Zhang (Elkridge, MD); Wenchao Yu (Plainsboro, NJ); Haifeng Chen (West Windsor, NJ); LuAn Tang (Pennington, NJ); Zexue He (La Jolla, CA)
Assignee: NEC Corporation
G06N3/08G06F40/30G06F40/47
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,028
App. No.
17/958,597
Granted
Jan 21, 2025
Kind
B2
Abstract

A computer-implemented method for multi-model representation learning is provided. The method includes encoding, by a trained time series (TS) encoder, an input TS segment into a TS-shared latent representation and a TS-private latent representation. The method further includes generating, by a trained text generator, a natural language text that explains the input TS segment, responsive to the TS-shared latent representation, the TS-private latent representation, and a text-private latent representation.

Claims (32)

1. A computer-implemented method for multi-model representation learning, comprising:

encoding, by a trained time series (TS) encoder, an input TS segment into a TS-shared latent representation and a TS-private latent representation; and

generating, by a trained text generator, a natural language text that explains the input TS segment, responsive to the TS-shared latent representation, the TS-private latent representation, and a text-private latent representation.

2. The computer-implemented method of claim 1 , wherein the text-private latent representation is obtaining by inputting a list of example texts with a desired style into a trained text encoder which outputs the text-private latent representation in response thereto.

3. The computer-implemented method of claim 2 , further comprising pre-processing the example texts by translating each of the example texts into a foreign language and translating back using a different translation tool to obtain resultant text with a same meaning but different words for use as the list of example texts.

4. The computer-implemented method of claim 1 , wherein the text-private latent representation is selected from a list of style latent prototypes.

5. The computer-implemented method of claim 1 , wherein the natural language text that explains the input TS segment is in a user-selected style.

6. The computer-implemented method of claim 1 , further comprising pre-processing the input TS segment by normalizing the input TS segment to have values from 0 to 1, and augmenting the normalized TS segment by adding Gaussian noise to the normalized TS segment.

7. The computer-implemented method of claim 1 , further comprising:

adding a perturbation to dimension i of the TS-shared latent representation; and

visualizing generated pairs of TS and text as a function of changing perturbation for each dimension i of the TS-shared latent variable.

8. The computer-implemented method of claim 1 , wherein a hidden state vector of a last timestep is used as the TS-shared latent representation and the TS-private latent representation of the TS segment.

9. The computer-implemented method of claim 1 , further comprising controlling an operating parameter of a motorized device to prevent an impending device failure responsive to the natural language text that explains the input TS segment.

10. A computer program product for multi-model representation learning, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

encoding, by a trained time series (TS) encoder, an input TS segment into a TS-shared latent representation and a TS-private latent representation; and

generating, by a trained text generator, a natural language text that explains the input TS segment, responsive to the TS-shared latent representation, the TS-private latent representation, and a text-private latent representation.

11. The computer program product of claim 10 , wherein the text-private latent representation is obtaining by inputting a list of example texts with a desired style into a trained text encoder which outputs the text-private latent representation in response thereto.

12. The computer program product of claim 11 , further comprising pre-processing the example texts by translating each of the example texts into a foreign language and translating back using a different translation tool to obtain resultant text with a same meaning but different words for use as the list of example texts.

13. The computer program product of claim 10 , wherein the text-private latent representation is selected from a list of style latent prototypes.

14. The computer program product of claim 10 , wherein the natural language text that explains the input TS segment is in a user-selected style.

15. The computer program product of claim 10 , further comprising pre-processing the input TS segment by normalizing the input TS segment to have values from 0 to 1, and augmenting the normalized TS segment by adding Gaussian noise to the normalized TS segment.

16. The computer program product of claim 10 , further comprising:

adding a perturbation to dimension i of the TS-shared latent representation; and

visualizing generated pairs of TS and text as a function of changing perturbation for each dimension i of the TS-shared latent variable.

17. The computer program product of claim 10 , wherein a hidden state vector of a last timestep is used as the TS-shared latent representation and the TS-private latent representation of the TS segment.

18. The computer program product of claim 10 , further comprising controlling an operating parameter of a motorized device to prevent an impending device failure responsive to the natural language text that explains the input TS segment.

19. A computer processing system for multi-model representation learning, comprising:

a memory device for storing program code; and

a hardware processor operatively coupled to the memory device for running the program code to:

encode, using a trained time series (TS) encoder, an input TS segment into a TS-shared latent representation and a TS-private latent representation; and

generate, using a trained text generator, a natural language text that explains the input TS segment, responsive to the TS-shared latent representation, the TS-private latent representation, and a text-private latent representation.

20. The computer processing system of claim 19 , wherein the text-private latent representation is obtaining by inputting a list of example texts with a desired style into a trained text encoder which outputs the text-private latent representation in response thereto.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2022
From: CHEN, YUNCONG; CHEN, ZHENGZHANG; ZHANG, XUCHAO; YU, WENCHAO; CHEN, HAIFENG; TANG, LUAN; HE, ZEXUE
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 061286/0572 →
Continuity (3)
Provisional Application 63253169 · Oct 7, 2021
Provisional Application 63308081 · Feb 9, 2022
Related Publication 20230109729A1 · Apr 13, 2023
References Cited (19)
US 10789755B2 · Amer · 2020 [cited by examiner]
US 11294756B1 · Sadrieh · 2022 [cited by examiner]
US 20070078564A1 · Hoshino et al. · 2007 [cited by applicant]
US 20080208565A1 · Bisegna · 2008 [cited by applicant]
US 20180165554A1 · Zhang · 2018 [cited by examiner]
US 20190278489A1 · Saito et al. · 2019 [cited by applicant]
US 20200364596A1 · Zang et al. · 2020 [cited by applicant]
US 20210133590A1 · Amroabadi · 2021 [cited by examiner]
US 20220108082A1 · Morris · 2022 [cited by examiner]
KR 1020190109108A · 2019 [cited by applicant]
M. Lee and V. Pavlovic, “Private-Shared Disentangled Multimodal VAE for Learning of Latent Representations,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, TN, USA, 202… [cited by examiner]
O. T.. -C. Chen, Y. H. Tsai, C. W. Su, P. C. Kuo and W. C. Lai, “Voice-activity home care system,” 2016 IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI), Las VegaUSA, 2016, pp. 110-113, doi:… [cited by examiner]
Cheng, P., Hao, W., Dai, S., Liu, J., Gan, Z., & Carin, L. (Nov. 21, 2020). Club: A contrastive log-ratio upper bound of mutual information. In International conference on machine learning (pp. 1779-1788). PMLR. [cited by applicant]
Edunov, S., Ott, M., Auli, M., & Grangier, D. (Aug. 28, 2018). Understanding back-translation at scale. arXiv preprint arXiv:1808.09381. [cited by applicant]
Kim, H., & Mnih, A. (Jul. 3, 2018). Disentangling by factorising. In International Conference on Machine Learning (pp. 2649-2658). PMLR. [cited by applicant]
Kudo, T., & Richardson, J. (Aug. 19, 2018). Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226. [cited by applicant]
Nguyen, X., Wainwright, M. J., & Jordan, M. (Dec. 3, 2007). Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization. Advances in neural information processing systems, 20. [cited by applicant]
Qin, Y., Song, D., Chen, H., Cheng, W., Jiang, G., & Cottrell, G. (Apr. 7, 2017). A dual-stage attention-based recurrent neural network for time series prediction. arXiv preprint arXiv:1704.02971. [cited by applicant]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., . . . & Polosukhin, I. (Dec. 4, 2017). Attention is all you need. Advances in neural information processing systems, 30. [cited by applicant]