IP Library Granted Patent US 12,198,051
Granted Patent B2
US 12,198,051 · App. 18/096,198 · Granted Jan 14, 2025

Systems and methods for collaborative filtering with variational autoencoders

Inventors: William G. Macready (West Vancouver, CA); Jason T. Rolfe (Vancouver, CA)
Assignee: D-WAVE SYSTEMS INC.
G06N3/08G06F18/2148G06N3/045G06N10/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,051
App. No.
18/096,198
Granted
Jan 14, 2025
Kind
B2
Abstract

Collaborative filtering systems based on variational autoencoders (VAEs) are provided. VAEs may be trained on row-wise data without necessarily training a paired VAE on column-wise data (or vice-versa), and may optionally be trained via minibatches. The row-wise VAE models the output of the corresponding column-based VAE as a set of parameters and uses these parameters in decoding. In some implementations, a paired VAE is provided which receives column-wise data and models row-wise parameters; each of the paired VAEs may bind their learned column- or row-wise parameters to the output of the corresponding VAE. The paired VAEs may optionally be trained via minibatches. Unobserved data may be explicitly modelled. Methods for performing inference with such VAE-based collaborative filtering systems are also disclosed, as are example applications to search and anomaly detection.

Claims (35)

1. A method of performing inference with a collaborative filtering system defined over an input space of values, each value associated with a row dimension and a column dimension, the method executed by circuitry including at least one processor, the method comprising:

receiving an input row vector of values in the input space, the input row vector comprising one or more observed values associated with a first row element in the row dimension;

encoding, by a row-wise encoder of a trained variational autoencoder, the input row vector to an encoded row vector;

decoding, by a decoder of the trained variational autoencoder, a first model distribution over the input space for a first row-column pair based on the encoded row vector and a learned column vector of a set of learned column vectors, the learned column vector being a parameter of the encoder and comprising one or more learned values associated with a first column element of the first row-column pair; and

determining a predicted value based on the first model distribution.

2. The method according to claim 1 wherein the first model distribution is a joint probability distribution modelling, for the first row-column pair, at least:

one or more probabilities associated with one or more values in the input space; and

a probability associated with an absence of an observed value in the first row-column pair.

3. The method according to claim 2 wherein receiving an input row vector of values in the input space, the input row vector comprising one or more observed values associated with a first row element in the row dimension comprises receiving one or more categorical values, at least one category of the one or more categorical values corresponding to an unobserved designation.

4. The method according to claim 3 wherein the first row element corresponds to a user of a plurality of users, the first column element corresponds to an item of a plurality of items, values correspond to ratings of items of the plurality of items by users of the plurality of users, and wherein the first model distribution modelling, for the row-column pair, at least a probability associated with an absence of an observed value in the first row-column pair comprises the first model distribution modelling, for the row-column pair, at least a probability that the row-column pair is unrated.

5. The method according to claim 2 wherein determining a predicted value comprises determining a truncated distribution based on the first model distribution conditioned on an associated row-column pair having an observed value.

6. The method according to claim 5 wherein determining a predicted value further comprises determining a mean of the truncated distribution to yield an expectation value and determining the predicted value based on the expectation value.

7. The method according to claim 2 wherein the first model distribution comprises a probability distribution over a characteristic of the first row-column pair and determining a predicted value further comprises determining a probability that the first row-column pair has the characteristic.

8. The method according to claim 7 wherein the first row element corresponds to a user of a plurality of users, the first column element corresponds to an item of a plurality of items, values correspond to ratings of items of the plurality of items by users of the plurality of users, and the characteristic corresponds to an interaction between the user and the item that is independent of a rating, and wherein determining a predicted value further comprises determining a probability of the interaction between the user and the item of the first row-column pair.

9. The method according to claim 1 wherein encoding the input row vector to an encoded row vector comprises:

determining a latent distribution of the first row element in a latent space of the row-wise encoder; and

deterministically extracting an extracted value from the latent space based on the latent distribution, the extracted value associated with the first row element.

10. The method according to claim 9 wherein encoding the input row vector to an encoded row vector further comprises transforming the extracted value into the encoded row vector.

11. The method according to claim 9 wherein deterministically extracting an extracted value comprises determining an expected value or a mode of the latent distribution for the first row element.

12. A system for collaborative filtering over an input space comprising values, each value associated with a row dimension and a column dimension, the system comprising at least one processor and at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data which, when executed by the at least one processor cause the at least one processor to:

receive an input row vector of values in the input space, the input row vector comprising one or more observed values associated with a first row element in the row dimension;

encode, by a row-wise encoder of a trained variational autoencoder, the input row vector to an encoded row vector;

decode, by a decoder of the trained variational autoencoder, a first model distribution over the input space for a first row-column pair based on the encoded row vector and a learned column vector of a set of learned column vectors, the learned column vector being a parameter of the encoder and comprising one or more learned values associated with a first column element of the first row-column pair; and

determine a predicted value based on the first model distribution.

13. The system according to claim 12 , wherein the first model distribution is a joint probability distribution that models, for the first row-column pair, at least:

one or more probabilities associated with one or more values in the input space; and

a probability associated with an absence of an observed value in the first row-column pair.

14. The system according to claim 13 wherein the input row vector of values in the input space comprises one or more categorical values, wherein at least one category of the one or more categorical values corresponds to an unobserved designation.

15. The system according to claim 14 wherein the first row element corresponds to a user of a plurality of users, the first column element corresponds to an item of a plurality of items, values correspond to ratings of items of the plurality of items by users of the plurality of users, and wherein the first model distribution that models, for the row-column pair, at least a probability associated with an absence of an observed value in the first row-column pair models, for the row-column pair, at least a probability that the row-column pair is unrated.

16. The system according to claim 13 wherein the predicted value based on the first model distribution is determined based on a truncated distribution based on the first model distribution conditioned on an associated row-column pair having an observed value.

17. The system according to claim 16 wherein the predicted value based on the first model distribution is further determined based on an expectation value corresponding to a mean of the truncated distribution.

18. The system according to claim 13 wherein the first model distribution comprises a probability distribution over a characteristic of the first row-column pair and the predicted value is determined based on a probability that the first row-column pair has the characteristic.

19. The method according to claim 18 wherein the first row element corresponds to a user of a plurality of users, the first column element corresponds to an item of a plurality of items, values correspond to ratings of items of the plurality of items by users of the plurality of users, the characteristic corresponds to an interaction between the user and the item that is independent of a rating, and the predicted value is further determined based on a probability of the interaction between the user and the item of the first row-column pair.

20. The system according to claim 13 wherein the encoded row vector is encoded from the input row vector through transformation of a deterministically extracted value associated with the first row element from a latent spaced based on a latent distribution, wherein the latent distribution of the first row element is determined in the latent space of the row-wise encoder.

21. The system according to claim 20 wherein the extracted value is a determined expected value or a mode of the latent distribution for the first row element.

Assignments (6)
RELEASE OF SECURITY INTEREST Recorded Mar 11, 2025
From: PSPIB UNITAS INVESTMENTS II INC.
To: D-WAVE SYSTEMS INC.; 1372934 B.C. LTD.
Reel/Frame 070470/0098 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2024
From: ROLFE, JASON T.; MACREADY, WILLIAM G.
To: D-WAVE SYSTEMS INC.
Reel/Frame 069492/0889 →
CHANGE OF NAME Recorded Dec 5, 2024
From: DWSI HOLDINGS INC.
To: D-WAVE SYSTEMS INC.
Reel/Frame 069493/0142 →
CONTINUATION Recorded Dec 5, 2024
From: D-WAVE SYSTEMS INC.
To: D-WAVE SYSTEMS INC.
Reel/Frame 069513/0889 →
MERGER Recorded Dec 5, 2024
From: D-WAVE SYSTEMS INC.; DWSI HOLDINGS INC.
To: DWSI HOLDINGS INC.
Reel/Frame 069515/0937 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 14, 2023
From: D-WAVE SYSTEMS INC.; 1372934 B.C. LTD.
To: PSPIB UNITAS INVESTMENTS II INC., AS COLLATERAL AGENT
Reel/Frame 063340/0888 →
Continuity (4)
Continuation 16772094
Provisional Application 62598880 · Dec 14, 2017
Provisional Application 62658461 · Apr 16, 2018
Related Publication 20230222337A1 · Jul 13, 2023
References Cited (400)
US 5249122A · Stritzke · 1993 [cited by applicant]
US 6424933B1 · Agrawala et al. · 2002 [cited by applicant]
US 6671661B1 · Bishop · 2003 [cited by applicant]
US 7135701B2 · Amin et al. · 2006 [cited by applicant]
US 7418283B2 · Amin · 2008 [cited by applicant]
US 7493252B1 · Nagano et al. · 2009 [cited by applicant]
US 7533068B2 · Maassen et al. · 2009 [cited by applicant]
US 7876248B2 · Berkley et al. · 2011 [cited by applicant]
US 8008942B2 · Van et al. · 2011 [cited by applicant]
US 8073808B2 · Rose · 2011 [cited by applicant]
US 8175995B2 · Amin · 2012 [cited by applicant]
US 8190548B2 · Choi · 2012 [cited by applicant]
US 8195596B2 · Rose et al. · 2012 [cited by applicant]
US 8244650B2 · Rose · 2012 [cited by applicant]
US 8340439B2 · Mitarai et al. · 2012 [cited by applicant]
US 8421053B2 · Bunyk et al. · 2013 [cited by applicant]
US 8548828B1 · Longmire · 2013 [cited by applicant]
US 8560282B2 · Love et al. · 2013 [cited by applicant]
US 8700689B2 · Macready et al. · 2014 [cited by applicant]
US 8863044B1 · Casati et al. · 2014 [cited by applicant]
US 8977576B2 · Macready · 2015 [cited by applicant]
US 9015215B2 · Berkley et al. · 2015 [cited by applicant]
US 9378733B1 · Vanhoucke et al. · 2016 [cited by applicant]
US 9495644B2 · Chudak et al. · 2016 [cited by applicant]
US 9588940B2 · Hamze et al. · 2017 [cited by applicant]
US 9881256B2 · Hamze et al. · 2018 [cited by applicant]
US 10275422B2 · Israel et al. · 2019 [cited by applicant]
US 10296846B2 · Csurka et al. · 2019 [cited by applicant]
US 10318881B2 · Rose et al. · 2019 [cited by applicant]
US 10339466B1 · Ding et al. · 2019 [cited by applicant]
US 10486611B1 · Klindt · 2019 [cited by applicant]
US 10725422B2 · Amann et al. · 2020 [cited by applicant]
US 10789540B2 · King et al. · 2020 [cited by applicant]
US 10817796B2 · Macready et al. · 2020 [cited by applicant]
US 10846611B2 · Wabnig et al. · 2020 [cited by applicant]
US 11042811B2 · Rolfe et al. · 2021 [cited by applicant]
US 11062227B2 · Amin et al. · 2021 [cited by applicant]
US 11157817B2 · Rolfe · 2021 [cited by applicant]
US 11176484B1 · Dorner · 2021 [cited by examiner]
US 11386346B2 · Xue et al. · 2022 [cited by applicant]
US 11410067B2 · Rolfe et al. · 2022 [cited by applicant]
US 11461644B2 · Vahdat et al. · 2022 [cited by applicant]
US 11468293B2 · Chudak · 2022 [cited by applicant]
US 11481669B2 · Rolfe et al. · 2022 [cited by applicant]
US 11531852B2 · Vahdat · 2022 [cited by applicant]
US 11537926B2 · King et al. · 2022 [cited by applicant]
US 11586915B2 · Macready · 2023 [cited by examiner]
US 11645611B1 · Puthiyapurayil · 2023 [cited by examiner]
US 20020010691A1 · Chen · 2002 [cited by applicant]
US 20020077756A1 · Arouh et al. · 2002 [cited by applicant]
US 20030030575A1 · Frachtenberg et al. · 2003 [cited by applicant]
US 20050119829A1 · Bishop et al. · 2005 [cited by applicant]
US 20050171923A1 · Kiiveri et al. · 2005 [cited by applicant]
US 20050171932A1 · Nandhra · 2005 [cited by applicant]
US 20060041421A1 · Ta et al. · 2006 [cited by applicant]
US 20060047477A1 · Bachrach · 2006 [cited by applicant]
US 20060074870A1 · Brill et al. · 2006 [cited by applicant]
US 20060115145A1 · Bishop et al. · 2006 [cited by applicant]
US 20060117077A1 · Kiiveri et al. · 2006 [cited by applicant]
US 20070011629A1 · Shacham et al. · 2007 [cited by applicant]
US 20070162406A1 · Lanckriet · 2007 [cited by applicant]
US 20080069438A1 · Winn et al. · 2008 [cited by applicant]
US 20080103996A1 · Forman et al. · 2008 [cited by applicant]
US 20080132281A1 · Kim et al. · 2008 [cited by applicant]
US 20080312663A1 · Haimerl et al. · 2008 [cited by applicant]
US 20080313430A1 · Bunyk · 2008 [cited by applicant]
US 20090077001A1 · Macready et al. · 2009 [cited by applicant]
US 20090171956A1 · Gupta et al. · 2009 [cited by applicant]
US 20090254505A1 · Davis et al. · 2009 [cited by applicant]
US 20090278981A1 · Bruna et al. · 2009 [cited by applicant]
US 20090322871A1 · Ji et al. · 2009 [cited by applicant]
US 20100010657A1 · Do et al. · 2010 [cited by applicant]
US 20100185422A1 · Hoversten · 2010 [cited by applicant]
US 20100228694A1 · Le et al. · 2010 [cited by applicant]
US 20100332423A1 · Kapoor et al. · 2010 [cited by applicant]
US 20110022369A1 · Carroll et al. · 2011 [cited by applicant]
US 20110022820A1 · Bunyk et al. · 2011 [cited by applicant]
US 20110044524A1 · Wang et al. · 2011 [cited by applicant]
US 20110142335A1 · Ghanem et al. · 2011 [cited by applicant]
US 20110238378A1 · Allen et al. · 2011 [cited by applicant]
US 20110295845A1 · Gao et al. · 2011 [cited by applicant]
US 20120084235A1 · Suzuki et al. · 2012 [cited by applicant]
US 20120124432A1 · Pesetski et al. · 2012 [cited by applicant]
US 20120149581A1 · Fang · 2012 [cited by applicant]
US 20120209880A1 · Callan et al. · 2012 [cited by applicant]
US 20120215821A1 · Macready et al. · 2012 [cited by applicant]
US 20120254586A1 · Amin et al. · 2012 [cited by applicant]
US 20130071837A1 · Winters-Hilt et al. · 2013 [cited by applicant]
US 20130097103A1 · Chari et al. · 2013 [cited by applicant]
US 20130236090A1 · Porikli et al. · 2013 [cited by applicant]
US 20130245429A1 · Zhang et al. · 2013 [cited by applicant]
US 20140040176A1 · Balakrishnan et al. · 2014 [cited by applicant]
US 20140152849A1 · Bala et al. · 2014 [cited by applicant]
US 20140187427A1 · Macready et al. · 2014 [cited by applicant]
US 20140200824A1 · Pancoska · 2014 [cited by applicant]
US 20140201208A1 · Satish et al. · 2014 [cited by applicant]
US 20140214835A1 · Oehrle et al. · 2014 [cited by applicant]
US 20140214836A1 · Stivoric et al. · 2014 [cited by applicant]
US 20140278239A1 · Macaro et al. · 2014 [cited by applicant]
US 20140279727A1 · Baraniuk et al. · 2014 [cited by applicant]
US 20140297235A1 · Arora et al. · 2014 [cited by applicant]
US 20150006443A1 · Rose et al. · 2015 [cited by applicant]
US 20150161524A1 · Hamze · 2015 [cited by applicant]
US 20150205759A1 · Israel et al. · 2015 [cited by applicant]
US 20150242463A1 · Lin et al. · 2015 [cited by applicant]
US 20150248586A1 · Gaidon et al. · 2015 [cited by applicant]
US 20150317558A1 · Adachi et al. · 2015 [cited by applicant]
US 20160019459A1 · Audhkhasi et al. · 2016 [cited by applicant]
US 20160078359A1 · Csurka et al. · 2016 [cited by applicant]
US 20160078600A1 · Perez Pellitero et al. · 2016 [cited by applicant]
US 20160110657A1 · Gibiansky et al. · 2016 [cited by applicant]
US 20160174902A1 · Georgescu et al. · 2016 [cited by applicant]
US 20160180746A1 · Coombes et al. · 2016 [cited by applicant]
US 20160191627A1 · Huang et al. · 2016 [cited by applicant]
US 20160253597A1 · Bhatt et al. · 2016 [cited by applicant]
US 20160307305A1 · Madabhushi et al. · 2016 [cited by applicant]
US 20160328253A1 · Majumdar · 2016 [cited by applicant]
US 20170102984A1 · Jiang et al. · 2017 [cited by applicant]
US 20170132509A1 · Li et al. · 2017 [cited by applicant]
US 20170161633A1 · Clinchant et al. · 2017 [cited by applicant]
US 20180018584A1 · Nock et al. · 2018 [cited by applicant]
US 20180025291A1 · Dey et al. · 2018 [cited by applicant]
US 20180082172A1 · Patel et al. · 2018 [cited by applicant]
US 20180137422A1 · Wiebe et al. · 2018 [cited by applicant]
US 20180157923A1 · El et al. · 2018 [cited by applicant]
US 20180165554A1 · Zhang et al. · 2018 [cited by applicant]
US 20180165601A1 · Wiebe et al. · 2018 [cited by applicant]
US 20180203836A1 · Singh · 2018 [cited by examiner]
US 20180232649A1 · Wiebe et al. · 2018 [cited by applicant]
US 20180277246A1 · Zhong et al. · 2018 [cited by applicant]
US 20180365594A1 · Macready et al. · 2018 [cited by applicant]
US 20190005402A1 · Mohseni et al. · 2019 [cited by applicant]
US 20190018933A1 · Oono et al. · 2019 [cited by applicant]
US 20190030078A1 · Aliper et al. · 2019 [cited by applicant]
US 20190050534A1 · Apte et al. · 2019 [cited by applicant]
US 20190108912A1 · Spurlock et al. · 2019 [cited by applicant]
US 20190122404A1 · Freeman et al. · 2019 [cited by applicant]
US 20190180147A1 · Zhang et al. · 2019 [cited by applicant]
US 20190244680A1 · Rolfe et al. · 2019 [cited by applicant]
US 20190258907A1 · Rezende et al. · 2019 [cited by applicant]
US 20190258952A1 · Denchev · 2019 [cited by applicant]
US 20200090050A1 · Rolfe et al. · 2020 [cited by applicant]
US 20200167691A1 · Golovin et al. · 2020 [cited by applicant]
US 20200226197A1 · Woerner et al. · 2020 [cited by applicant]
US 20200234172A1 · King et al. · 2020 [cited by applicant]
US 20200257984A1 · Vahdat et al. · 2020 [cited by applicant]
US 20200311589A1 · Ollitrault et al. · 2020 [cited by applicant]
US 20200401916A1 · Rolfe et al. · 2020 [cited by applicant]
US 20200410384A1 · Aspuru-Guzik et al. · 2020 [cited by applicant]
US 20210089884A1 · Macready et al. · 2021 [cited by applicant]
US 20210279631A1 · Pichler et al. · 2021 [cited by applicant]
US 20220101170A1 · Denchev · 2022 [cited by applicant]
CN 101088102A · 2007 [cited by applicant]
CN 101329731A · 2008 [cited by applicant]
CN 101473346A · 2009 [cited by applicant]
CN 102651073A · 2012 [cited by applicant]
CN 102831402A · 2012 [cited by applicant]
CN 1023240478 · 2013 [cited by applicant]
CN 102364497B · 2013 [cited by applicant]
CN 104050509A · 2014 [cited by applicant]
CN 104426822A · 2015 [cited by applicant]
CN 104766167A · 2015 [cited by applicant]
CN 104919476A · 2015 [cited by applicant]
CN 106569601A · 2017 [cited by applicant]
CN 112771549A · 2021 [cited by applicant]
JP 2011008631A · 2011 [cited by applicant]
KR 20130010181A · 2013 [cited by applicant]
WO 2005009364A2 · 2005 [cited by applicant]
WO 2005093649A1 · 2005 [cited by applicant]
WO 2007008507A2 · 2007 [cited by applicant]
WO 2007085074A1 · 2007 [cited by applicant]
WO 2010071997A1 · 2010 [cited by applicant]
WO 2016037300A1 · 2016 [cited by applicant]
WO 2016089711A1 · 2016 [cited by applicant]
WO 2016210018A1 · 2016 [cited by applicant]
WO 2017124299A1 · 2017 [cited by applicant]
WO 2020163455A1 · 2020 [cited by applicant]
Mart Van Baalen (Deep Matrix Factorization for recommendation, Master's Thesis, University of Amsterdam, Sep. 30, 2016) (Year: 2016). [cited by examiner]
Achille et Soatto, “Information Dropout: Learning Optimal Representations Through Noise” Nov. 4, 2016, ICLR, arXiv: 1611.01353v1, pp. 1-12. (Year: 2016). [cited by applicant]
Adachi, S.H. et al., “Application of Quantum Annealing to Training of Deep Neural Networks,” URL:https://arxiv.org/ftp/arxiv/papers/151 0/1510.06356.pdf, Oct. 21, 2015, 18 pages. [cited by applicant]
Amin et al., “Quatum Boltzmann Machine”. arXiv:1601.02036v1, Jan. 8, 2016. [cited by applicant]
Amin, “Effect of Local Minima on Adiabatic Quantum Optimization,” Physical Review Letters 100(130503), 2008, 4 pages. [cited by applicant]
B. Sallans and G.E. Hitton , “Reinforcement Learning with Factored States and Actions”. JMLR, 5:1063-1088, 2004. [cited by applicant]
Bearman , et al., “What's the Point: Semantic Segmentation with Point Supervision”. ECCV, Jul. 23, 2016. https://arxiv.org/abs/1506.02106. [cited by applicant]
Bell , et al., “The “Independent Components” of Natural Scenes are Edge Filters”, Vision Res. 37(23) 1997, pp. 3327-3338. [cited by applicant]
Bian , et al., “The Ising Model: teaching an old problem new tricks”, D-wave systems. 2 (year 2010), 32 pages. [cited by applicant]
Bolton , et al., “Statistical fraud detection: A review”, Statistical Science 17(3) Aug. 1, 2002. https://projecteuclid.org/journals/statistical-science/volume-17/issue-3/Statistical-Fraud-Detection-A-Review/10.1214/ss/… [cited by applicant]
Chen , et al., “Stochastic Gradient Hamiltonian Monte Carlo”, arXiv:1402.4102 May 12, 2014. https://arxiv.org/abs/1402.4102. [cited by applicant]
Cho, K-H., Raiko, T, & Ilin, A. , “Parallel tempering is efficient for learning restricted Boltzmann machines”, 2010. [cited by applicant]
Doersch , “Tutorial on variational autoencoders”, arXiv:1606.05908 Jan. 3, 2021. https://arxiv.org/abs/1606.05908. [cited by applicant]
Fabius, Otto , et al., “Variational Recurrent Auto-Encoders”, Accepted as workshop contributions at ICLR 2015, 5 pages. [cited by applicant]
Hamze , “Sampling From a Set Spins With Clamping”. U.S. Appl. No. 61/912,385, filed Dec. 5, 2013, 35 pages. [cited by applicant]
Hees , “Setting up a Linked Data mirror from RDF dumps”. Jörn's Blog, Aug. 26, 2015. SciPy Hierarchical Clustering and Dendrogram Tutorial | Jörn's Blog (joernhees.de). [cited by applicant]
Heidrich-Meisner , et al., “Reinforcement Learning in a Nutshell”. http://image.diku.dk/igel/paper/RLiaN.pdf. [cited by applicant]
Hidasi , et al., “Session-based recommendations with recurrent neural networks”, ICRL Mar. 29, 2016. https://arxiv.org/abs/1511.06939. [cited by applicant]
Hinton, Geoffrey E, et al., “Reducing the Dimensionality of Data with Neural Networks”, Science, wwwsciencemag.org, vol. 313, Jul. 28, 2006, pp. 504-507. [cited by applicant]
Hurley, Barry , et al., “Proteus: A hierarchical Portfolio of Solvers and Transformations”, arXiv:1306.5606v2 [cs.AI], Feb. 17, 2014, 17 pages. [cited by applicant]
Kingma, Diederik P, et al., “Semi-Supervised Learning with Deep Generative Models”, arXiv:1406.5298v2 [cs.LG], Oct. 31, 2014, 9 pages. [cited by applicant]
Krause , et al., “The Unreasonable Effectiveness of Noisy Data for Fine-Grained Recognition”, 2016, Springer International Publishing AG, ECCV 2016, Part III, LNCS 9907, pp. 301-320 (Year:2016). [cited by applicant]
L.Wan, M. Zieler, et. al. , “Regularization of Neural Networks using DropConnect”. ICML, 2013. [cited by applicant]
Le Roux, Nicolas , et al., “Representational Power of Restricted Boltzmann Machines and Deep Belief Networks”, Dept. IRO, University of Montréal Canada, Technical Report 1294, Apr. 18, 2007, 14 pages. [cited by applicant]
Lee, H. , et al., “Sparse deep belief net model for visual area v2”. Advances in Neural Information Processing Systems, 20 . MIT Press, 2008. [cited by applicant]
Li, et al., “R/'enyi Divergence Variational Inference”, arXiv:1602.02311 Oct. 28, 2016. https://arxiv.org/abs/1602.02311. [cited by applicant]
Lovasz, et al., “Orthogonal Representations and Connectivity of Graphs”, Linear Algebra and its applications 114/115; 1989, pp. 439-454. [cited by applicant]
Maddison , et al., “The concrete distribution: A continuous relaxation of discrete random variables”, arXiv:1611.00712 Mar. 5, 2017. https://arxiv.org/abs/1611.00712. [cited by applicant]
Misra , et al., “Seeing through the Human Reporting Bias: Visual Classifiers from Noisy Human-Centric Labels”, 2016 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2016, pp. 2930-2939. [cited by applicant]
Misra , et al., “Visual classifiers from noisy humancentric labels”. In the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. [cited by applicant]
Mnih , et al., “Variational inference for Monte Carlo objectives”. arXiv:1602.06725, Jun. 1, 2016. https://arxiv.org/abs/1602.06725. [cited by applicant]
Mnih, Andriy , et al., “Variational Inference for Mote Carlo Objectives”, Proceedings of the 33rd International Conference on Machine Learning, New York, NY USA, 2016, JMLR: W&CP vol. 48, 9 pages. [cited by applicant]
Neven , et al., “Training a binary classifier with the quantum adiabatic algorithm”, arXiv preprint arXivc:0811.0416, 2008, 11 pages. [cited by applicant]
Rasmus, Antti , et al., “Semi-Supervised Learning with Ladder Networks”, arXiv:1507.02672v2 [cs.NE] Nov. 24, 2015, 19 pages. [cited by applicant]
Raymond , et al., “Systems and Methods for Comparing Entropy and KL Divergence of Post-Processed Samplers,” U.S. Appl. No. 62/322,116, filed Apr. 13, 2016, 47 pages. [cited by applicant]
Rezende , et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” arXiv:1401.4082v3 [stat.ML] May 30, 2014, 14 pages. https://arxiv.org/abs/1401.4082. [cited by applicant]
Rezende, Danilo J, et al., “Variational Inference with Normalizing Flows”, Proceedings of the 32nd International Conference on Machine Learning, Lille, France 2015, JMLR: W&CP vol. 37, 9 pages. [cited by applicant]
Rose , et al., “Systems and Methods for Quantum Processing of Data, for Example Functional Magnetic Resonance Image Data”. U.S. Appl. No. 61/841,129, filed Jun. 28, 2013, 129 pages. [cited by applicant]
Rose , et al., “Systems and Methods for Quantum Processing of Data, for Example Imaging Data”. U.S. Appl. No. 61/873,303, filed Sep. 3, 2013, 38 pages. [cited by applicant]
Shahriari , et al., “Taking the human out of the loop: A review of bayesian optimization”, Proceedings of the IEEE 104 Jan. 1, 2016. [cited by applicant]
Sutton, R., et al., “Policy gradient methods for reinforcement learning with function approximation”. Advances in Neural Information Processing Sytems, 12, pp. 1057-1063, MIT Press, 2000. [cited by applicant]
Szegedy , et al., “Rethinking the Inception Architecture for Computer Vision”, 2016, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2818-2826 (Year: 2016). [cited by applicant]
Tieleman, T. & Hinton, G. , “Using fast weights to improve persistent contrastive divergence”, 2009. [cited by applicant]
Tripathi , et al., “Survey on credit card fraud detection methods”, Internation Journal of Emerging Technology and Advanced Engineering Nov. 12, 2012. [cited by applicant]
Tucker , et al., “Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models”. arXiv:1703.07370, Nov. 6, 2017. https://arxiv.org/abs/1703.07370. [cited by applicant]
Vahdat , “Toward Robustness against Label Noise in Training Deep Disciminative Neural Networks”. arXiv:1706.00038v2, Nov. 3, 2017. https://arxiv.org/abs/1706.00038. [cited by applicant]
Wan, L. , et al., “Regularization of Neural Networks using DropConnec”. ICML 2013. [cited by applicant]
Wiebe, Nathan , et al., “Quantum Inspired Training for Boltzmann Machines”, arXiv:1507.02642v1 [cs.LG] Jul. 9, 2015, 18 pages. [cited by applicant]
Williams , “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” Springer, College of Computer Science, Northeastern University, Boston, MA, 1992, 27 pages. https://link.springer.c… [cited by applicant]
Wittek, Peter , “What Can We Expect from Quantum Machine Learning”. Yandex 1-32 School of Data Analysis Conference Machine Learning: Prospects and Applications, Oct. 5, 2015. pp. 1-16. [cited by applicant]
“Neuro-computing for Parallel and Learning Information Systems”, 2019-516164, www.jstage.jst.go.jp/article/sicej/1962/27/3/27_3_255/_article/-char/ja,Nov. 14, 2021, 17 pages. [cited by applicant]
“On the Challenges of Physical Implementations of RBMs”, arXiv:1312.5258V1 [stat.ML] Dec. 18, 2013, XP-002743443, 9 pages. [cited by applicant]
Bach , et al., “Optimization with Sparsity-Inducing Penalties”. arXiv:1108.0775v2, Nov. 22, 2011. [cited by applicant]
Bahnsen , et al., “Feature Engineering Strategies for Credit Card Fraud Detection”, Expert systems with applications Elsevier Jun. 1, 2016. https://www.sciencedirect.com/science/article/abs/pil/S0957417415008386?via%3Di… [cited by applicant]
Bellman, R. E., “Dynamic Programming”. Princeton University Press, Princeton, NJ. Republished 2003: Dover, ISBN 0-486-42809-5. [cited by applicant]
Burda , et al., “Importance Weighted Autoencoders”, arXiv:1509.00519 Nov. 7, 2016. https://arxiv.org/abs/1509.00519. [cited by applicant]
Buss , “Introduction to Inverse Kinematics with Jacobian Transpose, Pseudoinverse and Damped Least Squares methods”, Mathematics UCS 2004. https://www.math.ucsd.edu/˜sbuss/ResearchWeb/ikmethods/iksurvey.pdf. [cited by applicant]
Chen , et al., “Domain Adaptive Faster R-CNN for Object Detection in the Wild”. IEEE Xplore, 2018. https://arxiv.org/abs/1803.03243. [cited by applicant]
Cho, Kyunghyun , et al., “On the Properties of Neural Machine Translation: Encoder-Decoder Approaches”, arXiv:1409.1259v2. [cs.CL] Oct. 7, 2014, 9 pages. [cited by applicant]
Dai , et al., “Generative Modeling of Convolutional Neural Networks”. ICLR 2015. [cited by applicant]
Friedman , et al., “Learning Bayesan Networks from Data”, Stanford Robotics. http://robotics.stanford.edu/people/nir/tutorial/index.html. [cited by applicant]
G. Hinton, N. Srivastava, et. al., “Improving neural networks by preventing co-adaptation of feature detectors”. CoRR , abs/1207.0580, 2012. [cited by applicant]
G.A. Rummery and M. Niranjan , “Online Q-Learning using Connectionist Systems”. CUED/FINFENG/TR 166, Cambridge, UK, 1994. [cited by applicant]
Glynn , “Likelihood ratio gradient estimation for stochastic systems”. Communications of the ACM, 1990. https://dl.acm.org/doi/10.1145/84537.84552. [cited by applicant]
Gregor, Karol , et al., “DRAW: A Recurrent Neural Network For Image Generation”, Proceedings of the 32nd International Conference on Machine Leaning, Lille, France, 2015, JMLR: W&CP vol. 37. Copyright 2015, 10 pages. [cited by applicant]
Gu , et al., “Muprop: Unbiased backpropagation for stochastic neural networks”. arXiv:1511.05176, Feb. 25, 2016 https://arxiv.org/abs/1511.05176. [cited by applicant]
Husmeier , “Introduction to Learning Bayesian Networks from Data”, Probabilistic Modeling in Bioinformatics and Medical Informatics 2005. https://link.springer.com/chapter/10.1007/1-84628-119-9 2. [cited by applicant]
Jiang , et al., “Learning a discriminative dictionary for sparse coding via label consistent K-SVD”, In CVPR 2011 (pp. 1697-1704) IEEE. June, Year 2011). [cited by applicant]
Kuzelka, Ondrej , et al., “Fast Estimation of First-Order Clause Coverage through Randomization and Maximum Likelihood”, In proceeding of the 25th International Conference on Machine Learning (pp. 504-5112). Association… [cited by applicant]
Lee, et al., “Efficient sparse coding algorithm”, NIPS, 2007, pp. 801-808. [cited by applicant]
Zheng , et al., “Graph regularized sparse coding for image representation”, IEEE transaction on image processing, 20 (5), (Year: 2010) 1327-1336. [cited by applicant]
Lin , et al., “Efficient Piecewise Training of Deep Structured Models for Semantic Segmentation”. arXiv:1504.01013v4, 2016. [cited by applicant]
Mnih , et al., “Neural variational inference and learning in bellef networks”. arXiv:1402.0030 Jun. 4, 2016. https://arxiv.org/abs/1402.0030. [cited by applicant]
Murphy , “Machine Learning: a probalistic perspective”, MIT Press, 2012. http://noiselab.ucsd.edu/ECE228/Murphy_Machine_Learning.pdf. [cited by applicant]
N. Srivastava, G. Hinton, et. al., “Dropout: A Simple Way to Prevent Neural Networks from Overtting”. ICML 15 (Jur):19291958, 2014. [cited by applicant]
Neal , et al., “Mcmc Using Hamiltonian Dynamics”, Handbook of Markov Chain Monte Carlo 2011. [cited by applicant]
Olshausen, Bruno A, et al., “Emergence of simple cell receptive field properties by learning a sparse code for natural images”, Nature, vol. 381, Jun. 13, 1996, pp. 607-609. [cited by applicant]
Pozzolo, et al., “Learned Lessons in credit card fraud detection from a practitioner perspective”, Feb. 18, 2014. https://www.semanticscholar.org/paper/Learned-lessons-in-credit-card-fraud-detection-from-Pozzolo-Caelen/… [cited by applicant]
Rolfe , “Discrete variational autoencoders” arXiv:1609.02200 Apr. 22, 2017. https://arxiv.org/abs/1609.02200. [cited by applicant]
Salakhutdinov, R. , “Learning deep Boltzmann machines using adaptive MCMC”, 2010. [cited by applicant]
Salakhutdinov, R. , “Learning in Markov random transitions.elds using tempered”, 2009. [cited by applicant]
Salakhutdinov. R. & Murray, I. , “On the quantitative analysis of deep belief networks”, 2008. [cited by applicant]
Salimans, Tim , et al., “Markov Chain Monte Carlo and Variational Inference: Bridging the Gap”, arXiv: 1410.6460v4 [stat.CO] May 19, 2015, 9 pages. [cited by applicant]
Schulman , et al., “Gradient estimation using stochastic computing graphs”. arXiv:1506.05254, Jan. 5, 2016. https://arxiv.org/abs/1506.05254. [cited by applicant]
Schwartz-Ziv , et al., “Opening the black box of Deep Neural Networks via Information”, arXiv:1703.00810 Apr. 29, 2017. https://arxiv.org/abs/1703.00810. [cited by applicant]
Silver, et al., “Mastering the game of Go with deep neural networks and tree search”. Nature, 529, 484489, 2016. [cited by applicant]
Sonderby , et al., “Ladder Variational Autoencoders”, arXiv:1602.02282v3 [stat.ML] May 27, 2016, 12 pages. [cited by applicant]
Sprechmann , et al., “Dictionary learning and sparse coding for unsupervised clustering”, in 2010 IEEE international conference on acoustics, speech and signal processing (pp. 2042-2045) IEEE (year:2010). [cited by applicant]
Sutton , “Learning to Predict by the Methods of Temporal Differences”. https://webdocs.cs.ualberta.ca/sutton/papers/sutton-88-with-erratum.pdf. [cited by applicant]
Suzuki, et al., “Joint Multimodal Learning With Deep Generative Models”, Nov. 7, 2016, arXiv:1611.0189v1 (Year: 2016). [cited by applicant]
Tokui , et al., “Evaluating the variance of likelihood-ratio gradient estimators”, Proceedings of the 34th International Conference on Machine Learning, 2017. http://proceedings.mlr.press/v70/tokui17a.html. [cited by applicant]
Vahdat , “Machine Learning Systems and Methods for Training With Noisy Labels,” U.S. Appl. No. 62/427,020, filed Nov. 28, 2016, 30 pages. [cited by applicant]
Vahdat , “Machine Learning Systems and Methods for Training With Noisy Labels,” U.S. Appl. No. 62/508,343, filed May 18, 2017, 46 pages. [cited by applicant]
Vahdat , et al., “Dvae++: Discrete variational autoencoders with overlapping transformations”, arXiv:1802.04920 May 25, 2018. https://arxiv.org/abs/1802.04920. [cited by applicant]
Van Det Maaten , et al., “Hidden unit conditional random Fields”. 14th International Conference on Artificial Intelligence and Statistics, 2011. [cited by applicant]
Veit , et al., “Learning From Noisy Large-Scale Datasets With Minimal Supervision”. arXiv:1701.01619v2, Apr. 10, 2017. https://arxiv.org/abs/1701.01619. [cited by applicant]
Xiao , et al., “Learning from massive noisy labeled data for image classification”. The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2015. [cited by applicant]
Lee, et al., “Local Low-Rank Matrix Approximation”, 2013. [cited by applicant]
Li, et al., “Low-Rank Matrix Approximation with Stability”, International Conference on Machine Learning, New York, NY, 2016. [cited by applicant]
Miao, Y., “Neural Variational Inference For Text Processing”, 2016. [cited by applicant]
Mikolov, T., “Distributed Representations of Words and Phrases and Their Compositionality”, arXiv:1310.4546v1, 2013. [cited by applicant]
Mittal, et al., “Symbolic Music Generation With Diffusion Models”, arXiv:2103.16091v2, 2021. [cited by applicant]
Muphy,“A Brief Introduction to Graphical Models and Bayesian Networks”, 1998. [cited by applicant]
Neal, et al., “A View of the EM Algorithm That Justifies Incremental, Sparse, And Other Variants”, 1998. [cited by applicant]
Nichol, “Improved Denoising Diffusion Probabilistic Models”, 2021. [cited by applicant]
Nichol,“Glide: Towards Photorealistic Image Generation”, 2022. [cited by applicant]
Nickel, et al., “Reducing the Rank of Relational Factorization Models by Including Observable Patterns”, 2014. [cited by applicant]
Nickel, et al. “A Three-Way Model For Collective Learning”, 2011. [cited by applicant]
Nickel, et al., “A Review of Relational Machine Learning for Knowledge Graphs”, arXiv:1503.00759v3 [stat.ML] Sep. 28, 2015. [cited by applicant]
Non-Final Office Action Issued in U.S. Appl. No. 16/968,465, mailed Aug. 24, 2023, 23 pages. [cited by applicant]
Oger, Julie, Emmanuel Lesigne, and Philippe Leduc. “A Random Field Model and its Application in industrial Production.”arXiv:1312.1653 (2013) 29 pages. [cited by applicant]
Pan, et al., “One Class Collaborative Filtering”, 2008. [cited by applicant]
Paterek, “Improving Regularized Singular Value Decomposition For Collaborative Filtering”, 2007. [cited by applicant]
Ramesh, et al. “Hierarchical Text-Conditional Image Generation With CLIP Latents”, 2022. [cited by applicant]
Rudolph, M. et al., “Generation of High-Resolution Handwritten Digits With an Ion-Trap Quantum Computer”, 2020. [cited by applicant]
Salakhutdinov, et al., “Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo”, 2008. [cited by applicant]
Salakhutdinov, et al., “Deep Boltzman Machines”, 2009. [cited by applicant]
Salakhutdinov, et al., “Probabilistic Matrix Factorization”, 2007. [cited by applicant]
Salimans, et al., “PixelCNN++: Improving PixelCNN With Discretized Logistic Mixture Likelihood and other Modifications”, arXiv:1701.05517v1 Jan. 19, 2017. [cited by applicant]
Vincent, et al., “Extracting And Composing Robust Features With Denoising Autoencoders”, 2008, 8 pages.. [cited by applicant]
Wang, “An Overview of SPSA: Recent Development And Applications”, 2020. [cited by applicant]
Wilson, et al., “Quantum-assisted associative adversarial network: applying quantum annealing in deep learning”, Quantum Machin Intelligence (2021), 14 pages. [cited by applicant]
Zhang., “Fake Detector: Effective Fake News Detection With Deep Diffusive Neural Network”, 2019. [cited by applicant]
Zheng, “A Neural Autoregressive Approach to Collaborative Filtering”, 2016. [cited by applicant]
Zheng, “Neural Autoregressive Collaborative Filtering For Implicit Feedback”, 2016. [cited by applicant]
Graves, et al., “Stochastic Backpropagation Through Mixture Density Distributions”, arXiv:1607.05690v1 [cs.NE] Jul. 19, 2016, 6 pages. [cited by applicant]
Sanchez-Vega, et al., “Learning Multivariate Distributions by Competitive Assembly of Marginals”, IEEE Transactions on Pattern Analysis and Machine Intelligence (vol. 35, Issue 2, pp. 398-410). [cited by applicant]
Van den Oord, et al., “Conditional Image Generation with PixelCNN Decorders”, ArXiv:1606.05328v2, Jun. 18, 2016. [cited by applicant]
Ren, et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, Microsoft Research, University of Science and Technology of China, 2015—9 pages. [cited by applicant]
Rezende et al., “Stochastic Backpropagation and Approximate Inference in Deep Generative Models,” arXiv:1401.4082v3 [stat.ML] May 30, 2014, 14 pages. [cited by applicant]
Rolfe et al., “Discrete Variational Auto-Encoder Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” United States U.S. Appl. No. 62/462,821, filed Feb. 23, 2017, 113 pages. [cited by applicant]
Rolfe et al., “Discrete Variational Auto-Encoder Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” United States U.S. Appl. No. 62/404,591, filed Oct. 5, 2016, 87 pages. [cited by applicant]
Rolfe et al., “Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” U.S. Appl. No. 62/207,057, filed Aug. 19, 2015, 39 pages. [cited by applicant]
Rolfe, “Discrete Variational Auto-Encoder Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” U.S. Appl. No. 62/206,974, filed Aug. 19, 2015, 43 pages. [cited by applicant]
Rolfe, “Discrete Variational Auto-Encoder Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” U.S. Appl. No. 62/268,321, filed Dec. 16, 2015, 52 pages. [cited by applicant]
Rolfe, “Discrete Variational Auto-Encoder Systems and Methods for Machine Learning Using Adiabatic Quantum Computers,” U.S. Appl. No. 62/307,929, filed Mar. 14, 2016, 67 pages. [cited by applicant]
Rolfe, Jason Tyler “Discrete Variational Autoencoders” Sep. 7, 2016, arXiv: 1609.02200v1, pp. 1-29. (Year: 2016). [cited by applicant]
Rose et al., “First ever DBM trained using a quantum computer”, Hack the Multiverse, Programming quantum computers for fun and profit, XP-002743440, Jan. 6, 2014, 8 pages. [cited by applicant]
Ross, S. et al., “Learning Message-Passing Inference Machines for Structured Prediction,” CVPR 2011, 2011,8 pages. [cited by applicant]
Sakkaris, et al., “QuDot Nets: Quantum Computers and Bayesian Networks”, arXiv:1607.07887v1 [quant-ph] Jul. 26, 2016, 22 page. [cited by applicant]
Salakhutdinov, et al., “Restricted Boltzmann Machines for Collaborative Filtering”, International Conference on Machine Learning, Corvallis, OR, 2007, 8 pages. [cited by applicant]
Salimans, Tim, and David A. Knowles. “Fixed-form variational posterior approximation through stochastic linear regression.” Bayesian Analysis 8.4 (2013): 837-882. (Year: 2013). [cited by applicant]
Salimans, Tim. “A structured variational auto-encoder for learning deep hierarchies of sparse features.” arXiv preprint arXiv: 1602.08734 (2016). (Year: 2016). [cited by applicant]
Scarselli, F. et al., “The Graph Neural Network Model,” IEEE Transactions on Neural Networks, vol. 20, No. 1,2009, 22 pages. [cited by applicant]
Sedhain, et al., “AutoRec: Autoencoders Meet Collaborative Filtering”, WWW 2015 Companion, May 18-22, 2015, Florence, Italy, 2 pages. [cited by applicant]
Serban et al., “Multi-Modal Variational Encoder-Decoders” Dec. 1, 2016, arXiv: 1612.00377v1, pp. 1-18. (Year: 2016). [cited by applicant]
Shah et al., “Feeling the Bern: Adaptive Estimators for Bernoulli Probabilities of Pairwise Comparisons” Mar. 22, 2016, pp. 1-33. (Year: 2016). [cited by applicant]
Somma, R., S Boixo, and H Barnum. Quantum simulated annealing. arXiv preprint arXiv:0712.1008, 2007. [cited by applicant]
Somma, Rd, S Boixo, H Barnum, and E Knill. Quantum simulations of classical annealing processes. Physical review letters, 101(13):130504, 2008. [cited by applicant]
Spall, “Multivariate Stochastic Approximation Using a Simultaneous Perturbation Gradient Approximation,” IEEE Transactions on Automatic Control 37(3):332-341, 1992. [cited by applicant]
Strub, F., et al. “Hybrid Collaborative Filtering with Autoencoders,” arXiv:1603.00806v3 [cs.IR], Jul. 19, 2016, 10 pages. [cited by applicant]
Sukhbaatar et al., “Training Convolutional Networks with Noisy Labels,” arXiv:1406.2080v4 [cs.CV] Apr. 10, 2015, 11 pages. [cited by applicant]
Suzuki, “Natural quantum reservoir computing for temporal information processing”, Scientific Reports, Nature Portfolio, Jan. 25, 2022. [cited by applicant]
Tieleman, T., “Training Restricted Boltzmann Machines using Approximation to the Likelihood Gradient,” ICML '08: Proceedings of the 25th international conference on Machine learning, 2008, 8 pages. [cited by applicant]
Tosh, Christopher, “Mixing Rates for the Alternating Gibbs Sampler over Restricted Boltzmann Machines and Friends” Jun. 2016. (Year: 2016). [cited by applicant]
Tucci, “Use of a Quantum Computer to do Importance and Metropolis-Hastings Sampling of a Classical Bayesian Network”, arXiv:0811.1792v1 [quant-ph] Nov. 12, 2008, 41 pages. [cited by applicant]
Van Baalen, M. “Deep Matrix Factorization for Recommendation,” Master's Thesis, Univ.of Amsterdam, Sep. 30, 2016, URL: https://scholar.google.co.kr/scholar?q=Deep+Matrix+Factorization+for+Recommendation&hl=ko&as_sdt=O&a… [cited by applicant]
Van de Meent, J-W., Paige, B., & Wood, “Tempering by subsampling”, 2014. [cited by applicant]
Van der Maaten, L. et al., “Hidden-Unit Conditional Random Fields,” Journal of Machine Learning Research 15, 2011, 10 Pages. [cited by applicant]
Van Rooyen, et al., “Learning with Symmetric Label Noise: The Importance of Being Unhinged” May 28, 2015, arXiv: 1505.07634v1, pp. 1-30. (Year: 2016). [cited by applicant]
Venkatesh, et al., “Quantum Fluctuation Theorems and Power Measurements,” New J. Phys., 17, 2015, pp. 1-19. [cited by applicant]
Wang et al., “Paired Restricted Boltzmann Machine for Linked Data” Oct. 2016. (Year: 2016). [cited by applicant]
Wang, Discovering phase transitions with unsupervised learning, Physical Review B 94, 195105 (2016), 5 pages. [cited by applicant]
Wang, W., Machta, J., & Katzgraber, H. G. “Population annealing: Theory and applications in spin glasses”, 2015. [cited by applicant]
Williams, “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” College of Computer Science, Northeastern University, Boston, MA, 1992, 27 pages. [cited by applicant]
Written Opinion of the International Searching Authority, mailed Nov. 18, 2016, for International Application No. PCT/US2016/047627, 9 pages. [cited by applicant]
Xu et Ou “Joint Stochastic Approximation Learning of Helmholtz Machines” Mar. 20, 2016, ICLR arXiv: 1603.06170v1, pp. 1-8. (Year: 2016). [cited by applicant]
Yoshihara et al., “Estimating the Trend of Economic Indicators by Deep Learning”, 2019-516164, Graduate School of System Informatics, Kobe University, 28 Annual Conferences of Japanese Society for Artificial Intelligenc… [cited by applicant]
Zhang et al., “Understanding Deep Learning Requires Re-Thinking Generalization”, arXiv:1611.03530 Feb. 26, 2017. https://arxiv.org/abs/1611.03530. [cited by applicant]
Zhao et al., “Towards a Deeper Understanding of Variational Autoencoding Models”, arXiv:1702.08658 Feb. 28, 2017. https://arxiv.org/abs/1702.08658. [cited by applicant]
Zhu, X. et al., “Combining Active Learning and Semi-Supervised Learning Using Gaussian Fields and Harmonic Functions,” ICML 2003 workshop on The Continuum from Labeled to Unlabeled Data in Machine Learning and Data Mini… [cited by applicant]
Zojaji et al., “A Survey of Credit Card Fraud Detection Techniques: Data and Technique Oriented Perspective”, arXiv:1611.06439 Nov. 19, 2016. https://arxiv.org/abs/1611.06439. [cited by applicant]
Amari, S., “Neural Computation—Toward Parallel Learning Information Processing”, https://www.jstage.jst.go.jp/article/sicejl1962/27/3/27_3_255/_article/-char/ja/. [cited by applicant]
Arici, T. et al., “Associate Adversarial Networks”, arXiv:1611.06953v1, Nov. 18, 2016, 8 pages. [cited by applicant]
Ba, J.L. et al., “Layer Normalization”, arXiv:1607.06450v1, Jul. 21, 2016, 14 pages. [cited by applicant]
Bell, et al., “Lessons From The Netflix Prize Challenge” 2007, 5 pages. [cited by applicant]
Benedetti, et al., “Quantum-assisted learning” Sep. 8, 2016, arXiv: 1609.02542v4, 13 pages. [cited by applicant]
Bordes, et al., “Translating Embeddings for Modeling Multi-relational Data”, 9 pages. [cited by applicant]
Bordes, et al. “Learning Structured Embeddings of Knowledge Bases”, 2011, 6 pages. [cited by applicant]
Bowman, et al., “Generating Sentences from a Continuous Space”, 2016, 12 pages. [cited by applicant]
Chen, et al., WEMAREC: Accurate and Scalable Recommendation Through Weighted and Ensemble Matrix Approximation_2015. [cited by applicant]
Cohen, et al., “Image Restoration via Ising Theory and Automatic Noise Estimation”, 2022. [cited by applicant]
Dhariwal, P. et al., “Diffusion Models Beat GANS on Image Synthesis”, arXiv:2105.052334v4, 2021, 44 pages. [cited by applicant]
Georgiev, et al., “A non-IID Framework for Collaborative Filtering with Restricted Boltzmann Machines”, 2013, 9 pags. [cited by applicant]
Gomez-Uribe, “The Netflix Recommender System: Algorithms, Business Value, and Innovation”, 2015. [cited by applicant]
Gopalan, et al., “Bayesian Nonparametric Poisson Factorization for Recommendation Systems”, 2014. [cited by applicant]
Gopalan, et al., “Scalable Recommendation with Hierarchical Poisson Factorization”, 2015. [cited by applicant]
Gopalan, et al., Content-based Recommendations with Poisson Factorization, 2014, 9 pages. [cited by applicant]
Gulrajani, et al., “PixeIVAE: A Latent Variable Model For Natural Images”, 2016, 9 pages. [cited by applicant]
Hees, “SciPy Hierarchical Clustering and Dendrogram Tutorial” | Jörn's Blog (joernhees.de). 2015, 40 pages. [cited by applicant]
Hernandez-Lobato, “Probabilistic Matrix Factorization with Non-random Missing Data”, 2014, 9 pages. [cited by applicant]
Ho, et al., “Denoising Diffusion Probabilistic Models”, 2020, 9 pages. [cited by applicant]
Ho, et al., Video Diffusion Models, 2022, 11 pages. [cited by applicant]
Jain, et al., “Estimating the class prior and posterior from noisy positives and unlabeled data” , arXiv: 1606.08561v2, 2017 19 pages. [cited by applicant]
Jenatton, R. et al., A Latent Factor Model For Highly Multi-Relational Data, 2012. [cited by applicant]
Johnson, C., “Logistic Matrix Factorization for Implicit Feedback Data”, 2014, 10 pages. [cited by applicant]
Khoshaman, et al., “Quantum Variational Autoencoder”, 2019. [cited by applicant]
King, et al., “Quantum-Assisted Genetic Algorithm”, 2019. [cited by applicant]
Koren, Y., “Factorization Meets the Neighborhood: A Multifaceted Collaborative Filtering Model”, 2008. [cited by applicant]
Larochelle, et al., “The Neural Autoregressive Distribution Estimator”, 2011. [cited by applicant]
Awasthi et al., “Efficient Learning of Linear Seperators under Bounded Noise” Mar. 12, 2015, arXiv: 1503.03594v1, pp. 1-23. (Year: 2015). [cited by applicant]
Awasthi et al., “Learning and 1-bit Compressed Sensing under Asymmetric Noise” Jun. 6, 2016, JMLR, pp. 1-41. (Year: 2016). [cited by applicant]
Bach et al., “On the Equivalence between Herding and Conditional Gradient Algorithms,” Proceedings of the 29th International Conference on Machine Learning, 2012, 8 pages. [cited by applicant]
Bach, F. et al., “Optimization with Sparsity-Inducing Penalties,” arXiv:1108.0775v2 [cs.LG], Nov. 22, 2011, 116 pages. [cited by applicant]
Benedetti et al., “Quantum-assisted learning of graphical models with arbitrary pairwise connectivity” Sep. 8, 2016, arXiv: 1609.02542v1, pp. 1-13. (Year: 2016). [cited by applicant]
Berkley, A.J. et al., “Tunneling Spectroscopy Using a Probe Qubit,” arXiv:1210.6310v2 [cond-mat.supr-con], Jan. 3, 2013, 5 pages. [cited by applicant]
Blanchard et al., “Classification with Asymmetric Label Noise: Consistency and Maximal Denoising” Aug. 5, 2016, arXiv: 1303.1208v3, pp. 1-47. (Year: 2016). [cited by applicant]
Blume-Kohout et al., “Streaming Universal Distortion-Free Entanglement Concentration”; IEEE Transactions on Information Theory Year: 2014; vol. 60, Issue 1; Journal Article' Publisher: IEEE; 17 pages. [cited by applicant]
Bornschein et al., “Bidirectional Helmholtz Machines” May 25, 2016, arXiv: 1506.03877v5. (Year: 2016). [cited by applicant]
Brakel, P., Dieleman, S., & Schrauwen. “Training restricted Boltzmann machines with multi-tempering: Harnessing parallelization”, 2012. [cited by applicant]
Chen et al., “Variational Lossy Autoencoder” Nov. 8, 2016, arXiv: 1611.02731v1, pp. 1-13. (Year: 2016). [cited by applicant]
Chen et al., “Herding as a Learning System with Edge-of-Chaos Dynamics,” arXiv:1602.030142V2 [stat.ML], Mar. 1, 2016, 48 pages. [cited by applicant]
Chen et al., “Parametric Herding,” Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS), 2010, pp. 97-104. [cited by applicant]
Chinese Office Action for Application No. CN 2016800606343, dated May 8, 2021, 21 pages (with English translation). [cited by applicant]
Courville, A. et al., “A Spike and Slab Restricted Boltzmann Machine,” Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS), 2011, 9 pages. [cited by applicant]
Covington, et al., “Deep Neural Networks for YouTube Recommendations”, RecSys '16, Sep. 15-19, 2016, Boston MA,8 pages. [cited by applicant]
Deng, J. et al., “ImageNet: A Large-Scale Hierarchical Image Database,” Proceedings / CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2009, 8 pages. [cited by applicant]
Desjardins, G., Courville, A., Bengio, Y., Vincent, P., & Delalleau, O. “Parallel tempering for training of restricted Boltzmann machines”, 2010. [cited by applicant]
Dziugaite, et al., “Neural Network Matrix Factorization”, arXiv:1511.06443v2 [cs.LG] Dec. 15, 2015, 7 pages. [cited by applicant]
Elkan, C., “Learning Classifiers from Only Positive and Unlabeled Data,” KDD08: The 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining Las Vegas Nevada USA Aug. 24-27, 2008, 8 pages. [cited by applicant]
Extended European Search Report for EP Application No. 16837862.8,dated Apr. 3, 2019, 12 pages. [cited by applicant]
Fergus, R. et al., “Semi-Supervised Learning in Gigantic Image Collections,” Advances in Neural Information Processing Systems, vol. 22, 2009, 8 pages. [cited by applicant]
First Office Action issued Nov. 29, 2021 in CN App No. 2016800731803. (English Translation). [cited by applicant]