IP Library Granted Patent US 12,614,113
Granted Patent B2
US 12,614,113 · App. 18/146,075 · Granted Apr 28, 2026

Using consistency metadata for filtering of machine learning data across jobs

Inventors: Leo Parker Dirac (Seattle, WA); Jin Li (Bellevue, WA); Tianming Zheng (Seattle, WA); Donghui Zhuo (Kirkland, WA)
Assignee: Amazon Technologies, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,113
App. No.
18/146,075
Granted
Apr 28, 2026
Kind
B2
Abstract

Consistency metadata, including a parameter for a pseudo-random number source, are determined for training-and-evaluation iterations of a machine learning model. Using the metadata, a first training set comprising records of at least a first chunk is identified from a plurality of chunks of a data set. The first training set is used to train a machine learning model during a first training-and-evaluation iteration. A first test set comprising records of at least a second chunk is identified using the metadata, and is used to evaluate the model during the first training-and-evaluation iteration.

Claims (46)

1 . A computer-implemented method, comprising:

scheduling, at a first set of servers by a machine learning service of a cloud computing environment in response to one or more programmatic requests from a client, a training job for a particular machine learning model, wherein a parameter of the one or more programmatic requests comprises consistency metadata to be used to ensure that a test set used for the particular machine learning model does not overlap with a training set of the machine learning model, wherein the first set of servers has access to a first pseudo-random number source, and wherein one or more pseudo-random numbers obtained from the first pseudo-random number source are used at the first set of servers to select, as part of the training job, a training set for the particular machine learning model from a data set;

scheduling, at a second set of servers by the machine learning service in response to the one or more programmatic requests, a test job for the particular machine learning model, wherein the second set of servers has access to a second pseudo-random number source; and

selecting, at the second set of servers as part of the test job, a test set for the particular machine learning model from the data set, wherein the test set is selected using one or more pseudo-random numbers obtained from the second pseudo-random number source after synchronizing, using at least the consistency metadata, a state of the second pseudo-random number source with a state of the first pseudo-random number source.

2 . The computer-implemented method as recited in claim 1 , wherein the consistency metadata comprises a seed value.

3 . The computer-implemented method as recited in claim 1 , wherein synchronizing the state of the second pseudo-random number source comprises:

obtaining a plurality of pseudo-random numbers from the second pseudo-random number source.

4 . The computer-implemented method as recited in claim 1 , wherein another parameter of the one or more parameters indicates a ratio of a size of the training set to a size of the test set, and wherein the test set is selected in accordance with the ratio.

5 . The computer-implemented method as recited in claim 1 , wherein another parameter of the one or more parameters indicates that a plurality of training-and-evaluation iterations is to be performed for the machine learning model, wherein the training set is utilized for a particular training-and-evaluation iteration of the plurality of training-and-evaluation iterations, and wherein the test set is utilized for the particular training-and-evaluation iteration.

6 . The computer-implemented method as recited in claim 1 , wherein the one or more parameters indicate that the data set is distributed among a plurality of data sources, wherein the one or more programmatic requests are received via a data-source-agnostic interface of the machine learning service, and wherein selecting the training set comprises:

extracting at least a first record of a plurality of records of the training set from a first data source of the plurality of data sources; and

extracting at least a second record of the plurality of records from a second data source of the plurality of data sources.

7 . The computer-implemented method as recited in claim 1 , further comprising:

receiving, at the machine learning service, a first request to shuffle another data set;

saving, at the machine learning service, additional consistency metadata which includes state information associated with one or more other pseudo-random numbers used to shuffle the other data set; and

re-obtaining, at the machine learning service in response to a second request to shuffle the other data set, results of a first shuffle operation performed in response to the first shuffle request, wherein said re-obtaining comprises using the additional consistency metadata.

8 . A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices:

schedule, at a first set of servers by a machine learning service of a cloud computing environment in response to one or more programmatic requests from a client, a training job for a particular machine learning model, wherein a parameter of the one or more programmatic requests comprises consistency metadata to be used to ensure that a test set used for the particular machine learning model does not overlap with a training set of the machine learning model, wherein the first set of servers has access to a first pseudo-random number source, and wherein one or more pseudo-random numbers obtained from the first pseudo-random number source are used at the first set of servers to select, as part of the training job, a training set for the particular machine learning model from a data set;

schedule, at a second set of servers by the machine learning service in response to the one or more programmatic requests, a test job for the particular machine learning model, wherein the second set of servers has access to a second pseudo-random number source; and

select, at the second set of servers as part of the test job, a test set for the particular machine learning model from the data set, wherein the test set is selected using one or more pseudo-random numbers obtained from the second pseudo-random number source after synchronizing, using at least the consistency metadata, a state of the second pseudo-random number source with a state of the first pseudo-random number source.

9 . The system as recited in claim 8 , wherein the consistency metadata comprises a seed value.

10 . The system as recited in claim 8 , wherein to synchronize the state of the second pseudo-random number source, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices:

obtain a plurality of pseudo-random numbers from the second pseudo-random number source.

11 . The system as recited in claim 8 , wherein another parameter of the one or more parameters indicates a ratio of a size of the training set to a size of the test set, and wherein the test set is selected in accordance with the ratio.

12 . The system as recited in claim 8 , wherein another parameter of the one or more parameters indicates that a plurality of training-and-evaluation iterations is to be performed for the machine learning model, wherein the training set is utilized for a particular training-and-evaluation iteration of the plurality of training-and-evaluation iterations, and wherein the test set is utilized for the particular training-and-evaluation iteration.

13 . The system as recited in claim 8 , wherein the one or more parameters indicate that the data set is distributed among a plurality of data sources, wherein the one or more programmatic requests are received via a data-source-agnostic interface of the machine learning service, and wherein to select the training set, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices:

extract at least a first record of a plurality of records of the training set from a first data source of the plurality of data sources; and

extract at least a second record of the plurality of records from a second data source of the plurality of data sources.

14 . The system as recited in claim 8 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices:

receive, at the machine learning service, a first request to shuffle another data set;

save, at the machine learning service, additional consistency metadata which includes state information associated with one or more other pseudo-random numbers used to shuffle the other data set; and

re-obtain, at the machine learning service in response to a second request to shuffle the other data set, results of a first shuffle operation performed in response to the first shuffle request, wherein said re-obtaining comprises using the additional consistency metadata.

15 . One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors:

schedule, at a first set of servers by a machine learning service of a cloud computing environment in response to one or more programmatic requests from a client, a training job for a particular machine learning model, wherein a parameter of the one or more programmatic requests comprises consistency metadata to be used to ensure that a test set used for the particular machine learning model does not overlap with a training set of the machine learning model, wherein the first set of servers has access to a first pseudo-random number source, and wherein one or more pseudo-random numbers obtained from the first pseudo-random number source are used at the first set of servers to select, as part of the training job, a training set for the particular machine learning model from a data set;

schedule, at a second set of servers by the machine learning service in response to the one or more programmatic requests, a test job for the particular machine learning model, wherein the second set of servers has access to a second pseudo-random number source; and

select, at the second set of servers as part of the test job, a test set for the particular machine learning model from the data set, wherein the test set is selected using one or more pseudo-random numbers obtained from the second pseudo-random number source after synchronizing, using at least the consistency metadata, a state of the second pseudo-random number source with a state of the first pseudo-random number source.

16 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein the consistency metadata comprises a seed value.

17 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein to synchronize the state of the second pseudo-random number source, the one or more non-transitory computer-accessible storage media store further program instructions that when executed on or across the one or more processors:

obtain a plurality of pseudo-random numbers from the second pseudo-random number source.

18 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein another parameter of the one or more parameters indicates a ratio of a size of the training set to a size of the test set, and wherein the test set is selected in accordance with the ratio.

19 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein another parameter of the one or more parameters indicates that a plurality of training-and-evaluation iterations is to be performed for the machine learning model, wherein the training set is utilized for a particular training-and-evaluation iteration of the plurality of training-and-evaluation iterations, and wherein the test set is utilized for the particular training-and-evaluation iteration.

20 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein the one or more parameters indicate that the data set is distributed among a plurality of data sources, wherein the one or more programmatic requests are received via a data-source-agnostic interface of the machine learning service, and wherein to select the training set, the one or more non-transitory computer-accessible storage media store further program instructions that when executed on or across the one or more processors:

extract at least a first record of a plurality of records of the training set from a first data source of the plurality of data sources; and

extract at least a second record of the plurality of records from a second data source of the plurality of data sources.

Continuity (4)
Continuation 16591521 · Oct 2, 2019
Continuation 14460314 · Aug 14, 2014
Continuation In Part 14319902 · Jun 30, 2014
Related Publication 20230126005A1 · Apr 27, 2023
References Cited (211)
US 4821333A · Gillies · 1989 [cited by applicant]
US 6230131B1 · Kuhn et al. · 2001 [cited by applicant]
US 6408290B1 · Thiesson et al. · 2002 [cited by applicant]
US 6615209B1 · Gomes et al. · 2003 [cited by applicant]
US 6658423B1 · Pugh et al. · 2003 [cited by applicant]
US 6681383B1 · Pastor et al. · 2004 [cited by applicant]
US 6804691B2 · Coha et al. · 2004 [cited by applicant]
US 7305373B1 · Cunningham et al. · 2007 [cited by applicant]
US 7328218B2 · Steinberg et al. · 2008 [cited by applicant]
US 7366718B1 · Pugh et al. · 2008 [cited by applicant]
US 7392262B1 · Alspector et al. · 2008 [cited by applicant]
US 7624274B1 · Alspector et al. · 2009 [cited by applicant]
US 7725475B1 · Alspector et al. · 2010 [cited by applicant]
US 7743003B1 · Tong et al. · 2010 [cited by applicant]
US 7809695B2 · Conrad et al. · 2010 [cited by applicant]
US 7827123B1 · Yagnik · 2010 [cited by applicant]
US 7930322B2 · MacLennan · 2011 [cited by applicant]
US 8046372B1 · Thirumalai et al. · 2011 [cited by applicant]
US 8078556B2 · Adi et al. · 2011 [cited by applicant]
US 8229864B1 · Lin et al. · 2012 [cited by applicant]
US 8283576B2 · Lin · 2012 [cited by applicant]
US 8370280B1 · Lin et al. · 2013 [cited by applicant]
US 8428915B1 · Nipko · 2013 [cited by applicant]
US 8429103B1 · Aradhye et al. · 2013 [cited by applicant]
US 8438122B1 · Mann · 2013 [cited by applicant]
US 8463071B2 · Snavely et al. · 2013 [cited by applicant]
US 8499010B2 · Gracie et al. · 2013 [cited by applicant]
US 8510238B1 · Aradhye et al. · 2013 [cited by applicant]
US 8583576B1 · Lin et al. · 2013 [cited by applicant]
US 8606730B1 · Tong et al. · 2013 [cited by applicant]
US 8682814B2 · DiCorpo et al. · 2014 [cited by applicant]
US 8886576B1 · Sanketi et al. · 2014 [cited by applicant]
US 9020861B2 · Lin et al. · 2015 [cited by applicant]
US 9069737B1 · Kimotho et al. · 2015 [cited by applicant]
US 9081817B2 · Arasu et al. · 2015 [cited by applicant]
US 9380032B2 · Resch et al. · 2016 [cited by applicant]
US 9697248B1 · Ahire · 2017 [cited by applicant]
US 9935318B1 · Surdoval et al. · 2018 [cited by applicant]
US 10169715B2 · Dirac et al. · 2019 [cited by applicant]
US 10467547B1 · Range et al. · 2019 [cited by applicant]
US 10496927B2 · Achin · 2019 [cited by applicant]
US 11295229B1 · Kumar · 2022 [cited by applicant]
US 11379755B2 · Dirac et al. · 2022 [cited by applicant]
US 11386351B2 · Dirac et al. · 2022 [cited by applicant]
US 11544623B2 · Dirac et al. · 2023 [cited by applicant]
US 11915104B2 · Range et al. · 2024 [cited by applicant]
US 20010027408A1 · Nakisa · 2001 [cited by applicant]
US 20010034580A1 · Skolnick et al. · 2001 [cited by applicant]
US 20030033194A1 · Ferguson · 2003 [cited by applicant]
US 20030176931A1 · Pednault et al. · 2003 [cited by applicant]
US 20030191795A1 · Bernadin · 2003 [cited by applicant]
US 20040059966A1 · Chan · 2004 [cited by applicant]
US 20050097068A1 · Graepel et al. · 2005 [cited by applicant]
US 20050105712A1 · Williams et al. · 2005 [cited by applicant]
US 20050119999A1 · Zait et al. · 2005 [cited by applicant]
US 20060050953A1 · Farmer · 2006 [cited by applicant]
US 20060179016A1 · Forman et al. · 2006 [cited by applicant]
US 20060195508A1 · Bernardin et al. · 2006 [cited by applicant]
US 20060212142A1 · Madani et al. · 2006 [cited by applicant]
US 20060248054A1 · Kirshenbaum · 2006 [cited by applicant]
US 20070005556A1 · Ganti et al. · 2007 [cited by applicant]
US 20070185896A1 · Jagannath et al. · 2007 [cited by applicant]
US 20080008116A1 · Buga · 2008 [cited by applicant]
US 20080010642A1 · MacLellan · 2008 [cited by applicant]
US 20080027916A1 · Asai et al. · 2008 [cited by applicant]
US 20080033900A1 · Zhang et al. · 2008 [cited by applicant]
US 20080082316A1 · Tsui et al. · 2008 [cited by applicant]
US 20080233576A1 · Weston · 2008 [cited by applicant]
US 20080275861A1 · Baluja et al. · 2008 [cited by applicant]
US 20090024586A1 · Zhou · 2009 [cited by applicant]
US 20100076913A1 · Yang · 2010 [cited by applicant]
US 20100100416A1 · Herbrich et al. · 2010 [cited by applicant]
US 20100115519A1 · Chechik · 2010 [cited by applicant]
US 20100179930A1 · Teller et al. · 2010 [cited by applicant]
US 20100223211A1 · Johnson et al. · 2010 [cited by applicant]
US 20100262568A1 · Schwaighofer et al. · 2010 [cited by applicant]
US 20100306249A1 · Hill et al. · 2010 [cited by applicant]
US 20110044533A1 · Cobb et al. · 2011 [cited by applicant]
US 20110145920A1 · Mahaffey et al. · 2011 [cited by applicant]
US 20110185230A1 · Agrawal et al. · 2011 [cited by applicant]
US 20110225594A1 · Iyengar et al. · 2011 [cited by applicant]
US 20110282932A1 · Ramjee et al. · 2011 [cited by applicant]
US 20110313953A1 · Lane et al. · 2011 [cited by applicant]
US 20110320767A1 · Eren et al. · 2011 [cited by applicant]
US 20120054658A1 · Chuat et al. · 2012 [cited by applicant]
US 20120089446A1 · Gupta et al. · 2012 [cited by applicant]
US 20120131088A1 · Liu et al. · 2012 [cited by applicant]
US 20120158791A1 · Kasneci et al. · 2012 [cited by applicant]
US 20120191630A1 · Breckenridge · 2012 [cited by applicant]
US 20120191631A1 · Breckenridge et al. · 2012 [cited by applicant]
US 20120253927A1 · Qin et al. · 2012 [cited by applicant]
US 20120284212A1 · Lin · 2012 [cited by applicant]
US 20130024170A1 · Dannecker · 2013 [cited by applicant]
US 20130097706A1 · Titonis et al. · 2013 [cited by applicant]
US 20130132963A1 · Lukyanov · 2013 [cited by applicant]
US 20130159376A1 · Moore · 2013 [cited by applicant]
US 20130185729A1 · Vasic · 2013 [cited by applicant]
US 20130191513A1 · Kamen et al. · 2013 [cited by applicant]
US 20130268457A1 · Wang · 2013 [cited by applicant]
US 20130297330A1 · Kame · 2013 [cited by applicant]
US 20130316421A1 · Scott et al. · 2013 [cited by applicant]
US 20130318240A1 · Hebert et al. · 2013 [cited by applicant]
US 20130323720A1 · Watelet et al. · 2013 [cited by applicant]
US 20130346347A1 · Patterson et al. · 2013 [cited by applicant]
US 20130346594A1 · Banerjee · 2013 [cited by applicant]
US 20140019542A1 · Rao et al. · 2014 [cited by applicant]
US 20140046879A1 · Maclennan et al. · 2014 [cited by applicant]
US 20140095521A1 · Blount · 2014 [cited by applicant]
US 20140121564A1 · Raskin · 2014 [cited by applicant]
US 20140122381A1 · Nowozin · 2014 [cited by applicant]
US 20140188919A1 · Huffman et al. · 2014 [cited by applicant]
US 20140304238A1 · Halla-Aho et al. · 2014 [cited by applicant]
US 20140344193A1 · Bilenko · 2014 [cited by applicant]
US 20140358825A1 · Phillipps et al. · 2014 [cited by applicant]
US 20140358831A1 · Adams · 2014 [cited by applicant]
US 20140365450A1 · Trimble et al. · 2014 [cited by applicant]
US 20150016461A1 · Qiang · 2015 [cited by applicant]
US 20150178811A1 · Chen · 2015 [cited by applicant]
US 20150280959A1 · Vincent · 2015 [cited by applicant]
US 20150379072A1 · Dirac et al. · 2015 [cited by applicant]
US 20150379423A1 · Dirac et al. · 2015 [cited by applicant]
US 20150379424A1 · Dirac et al. · 2015 [cited by applicant]
US 20150379425A1 · Dirac · 2015 [cited by applicant]
US 20150379426A1 · Steele et al. · 2015 [cited by applicant]
US 20150379427A1 · Dirac et al. · 2015 [cited by applicant]
US 20150379428A1 · Dirac et al. · 2015 [cited by applicant]
US 20150379429A1 · Lee et al. · 2015 [cited by applicant]
US 20150379430A1 · Dirac et al. · 2015 [cited by applicant]
US 20160026720A1 · Lehrer et al. · 2016 [cited by applicant]
US 20160078361A1 · Brueckner · 2016 [cited by applicant]
US 20180046926A1 · Achin · 2018 [cited by applicant]
US 20200050968A1 · Lee et al. · 2020 [cited by applicant]
US 20210374610A1 · Dirac et al. · 2021 [cited by applicant]
US 20220335338A1 · Dirac et al. · 2022 [cited by applicant]
US 20220391763A1 · Dirac et al. · 2022 [cited by applicant]
US 20230126005A1 · Dirac et al. · 2023 [cited by applicant]
US 20240185130A1 · Range et al. · 2024 [cited by applicant]
CN 102622441 · 2012 [cited by applicant]
CN 102770847A · 2012 [cited by applicant]
CN 103154936A · 2013 [cited by applicant]
CN 103218263A · 2013 [cited by applicant]
CN 103336869A · 2013 [cited by applicant]
CN 104123192 · 2014 [cited by applicant]
CN 104536902 · 2015 [cited by applicant]
EP 2393043 · 2011 [cited by applicant]
EP 2629247A1 · 2013 [cited by applicant]
JP 2009282577 · 2009 [cited by applicant]
WO 2012151198 · 2012 [cited by applicant]
U.S. Appl. No. 14/489,449, filed Sep. 17, 2014, Leo Parker Dirac. [cited by applicant]
International Search Report and Written Opinion from PCT/US2015/038610, Date of mailing Sep. 25, 2015, Amazon Technologies, Inc., pp. 1-12. [cited by applicant]
Kolo, B., “Binary and Multiclass Classification, Passage”, Binary and Multiclass Classification, XP002744526, Aug. 12, 2010, pp. 78-80. [cited by applicant]
Gamma, E., et al., “Design Patterns, Passage”, XP002286644, Jan. 1, 1995, pp. 293-294; 297, 300-301. [cited by applicant]
International Search Report and Written Opinion from PCT/US2015/038589, Date of mailing Sep. 23, 2015, Amazon Technologies, Inc., pp. 1-12. [cited by applicant]
“API Reference”, Google Prediction API, Jun. 12, 2013, 1 page. [cited by applicant]
“Google Prediction API”, Google developers, Jun. 9, 2014, 1 page. [cited by applicant]
U.S. Appl. No. 14/319,902, filed Jun. 30, 2014, Leo Parker Dirac. [cited by applicant]
U.S. Appl. No. 14/319,880, filed Jun. 30, 2014, Leo Parker Dirac. [cited by applicant]
U.S. Appl. No. 14/460,312, filed Aug. 14, 2014, Leo Parker Dirac. [cited by applicant]
U.S. Appl. No. 14/460,163, Filed Aug. 14, 2014, Zuohua Zhang. [cited by applicant]
U.S. Appl. No. 14/489,448, filed Sep. 17, 2014, Leo Parker Dirac, et al. [cited by applicant]
U.S. Appl. No. 14/463,434, filed Aug. 19, 2014, Robert Matthias Steele, et al. [cited by applicant]
U.S. Appl. No. 14/569,458, filed Dec. 12, 2014, Leo Parker Dirac, et al. [cited by applicant]
U.S. Appl. No. 14/484,201, filed Sep. 11, 2014, Michael Brueckner, et al. [cited by applicant]
U.S. Appl. No. 14/538,723, filed Nov. 11, 2014, Polly Po Yee Lee, et al. [cited by applicant]
U.S. Appl. No. 14/923,237, filed Oct. 26, 2015, Leo Parker Dirac, et al. [cited by applicant]
U.S. Appl. No. 15/132,959, filed Apr. 19, 2016, Pooja Ashok Kumar, et al. [cited by applicant]
Soren Sonnenburg, et al., “The SHOGUN Maching Learning Toolbox,” Journal of Machine Learning Research, Jan. 1, 2010, pp. 1799-1802, XP055216366, retrieved from http://www.jmlr.org/papers/volume11/sonnenburg10a/sonnenbur… [cited by applicant]
Anonymous: “GSoC 2014 Ideas,” Internet Citation, Jun. 28, 2014, XP882745876, Retrieved from http://web.archive.org/web/20140628051115/http://shogun-toolbox.org/page/Events/gsoc2014 ideas [retrieved on Sep. 25, 2015] Sec… [cited by applicant]
Anonymous: “Blog Aug. 21, 2013”, The Shogun Machine Learning Toolbox, Aug. 21, 2013, XP002745300, Retrieved from http://shogun-toolbox.org/page/contact/irclog/2013-08-21/, [retrieved on 2815-18-81], pp. 1-5. [cited by applicant]
Pyrathon D.: “Shogun as a SaaS”, Shogun Machine Learning Toolbox Mailing List Archive, Mar. 4, 2014, XP882745382, Retrieved from http://comments.gmane.org/gmane.comp.ai.machine-leaming.shogun/4359, [retrieved on 2815-18… [cited by applicant]
Yu, H-F, et al., “Large linear classification when data cannot fit in memory”, ACM Trans. on Knowledge Discovery from Data (TKDD), vol. 5, No. 4, 2012, 23 Pages. [cited by applicant]
Soroush, E. et al., “ArrayStore: a store manager for complex parallel array processing”, Proc. of the 2011 ACM Sigmod Intrl. Conf. on Management of Data, ACM, pp. 253-264. [cited by applicant]
Esposito, et al., “A Comparative Analysis of Methods for Pruning Decision Trees”, 1997, IEEE, 0162-8828/97, pp. 476-491. [cited by applicant]
Golbandi, et al., “Adaptive Bootstrapping of Recommender Systems Using Decision Trees”, 2011, WSDM, pp. 595-604. [cited by applicant]
Goldstein, et al., Penalized Split Criteria for Interpretable Trees, The Wharton School, University of Pennsylvania, 2013, pp. 1-25. [cited by applicant]
Zhan, et al., The State Problem for Test Generation in Si mu link, GECC0'06, Jul. 8-12, 2006, pp. 1941-1948. [cited by applicant]
The MathWorks, Inc., Simulink Projects Source Control Adapter Software Development Kit, SOK Version 1.2 for R2013b, Mar. 2013, pp. 1-9. [cited by applicant]
Eric Brochu, et al., “A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning”, arX1v:1012.2599v1, Dec. 14, 2010, pp. 1-49. [cited by applicant]
Jasper Snoek, et al., “Practical Bayesian Optimization of Machine Learning Algorithms”, arXiv:1206.2944v2 [stat.ML] Aug. 29, 2012, pp. 1-12. [cited by applicant]
AWS, “Amazon Machine Learning Developer Guide”, 2015, pp. 1-133. [cited by applicant]
Michael A. Osborne, et al., “Gaussian Processes for Global Optimization”, Published in the 3rd International Conference on Learning and Intelligent Optimization, 2009, pp. 1-1. [cited by applicant]
Wikipedia, “Multilayer perception”, Retrieved from URL: https://en.wikipedia.org/wiki/Multilayer_perceptron on Jan. 21, 2016, pp. 1-5. [cited by applicant]
Spark, “Spark Programming Guide—Spark 1.2.0 Documentation”, Retrieved from URL: https://spark.apache.org/docs/1.2.0/programmingguide.html on Jan. 15, 2016, pp. 1-18. [cited by applicant]
Wikipedia, “Stochastic gradient descent”, Retrieved from URL: https://en.wikipedia.org/wiki/Stochastic_gradient descent on Jan. 21, 2016, pp. 1-9. [cited by applicant]
Adomavicius et al., “Context-Aware Recommender Systems”, AI Magazine, Fall 2011, pp. 67-80. [cited by applicant]
Beach, et al., “Fusing Mobile, Sensor, and Social Data to Fully Enable Context-Aware Computing”, HOTMOBILE 2010, ACM, pp. 1-6. [cited by applicant]
Baltrunas, et al., “Context Relevance Assessment for Recommender Systems”, IUI ″11, Feb. 13-16, 2011, ACM, pp. 1-4. [cited by applicant]
U.S. Appl. No. 15/060,439, filed Mar. 3, 2016, Saman Zarandioon. [cited by applicant]
U.S. Appl. No. 14/460,314, filed Aug. 14, 2014, Leo Parker Dirac. [cited by applicant]
Borut Sluban, “Ensemble-Based Noise and Outlier Detection,” Doctoral Dissertation, Jozef Stefan International Postgraduate School, 2014, pp. 1-135. [cited by applicant]
Office Action in Canadian Patent Application No. 2,953,969 mailed Feb. 11, 2021, Amazon Technologies, Inc., pp. 1-9. [cited by applicant]
Office Action in European Patent Application No. 15739125.1 mailed Mar. 18, 2021, Amazon Technologies, Inc., pp. 1-9. [cited by applicant]
Office Action in European Patent Application No. 15739127.7 mailed Mar. 18, 2021, Amazon Technologies, Inc., pp. 1-9. [cited by applicant]
Kranjc, Janez, et al.,“Real-time data analysis in ClowdFlows”, 2013 IEEE International Conference on Big Data, Oct. 6, 2013, pp. 15-22; IEEE. [cited by applicant]
Kranjc, Janez, et al., “ClowdFlows: A Cloud Based Scientific Workflow Platform”, Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Part II, Lecture Notes in Computer Science 7524, Sep. … [cited by applicant]
Kranjc, Janez, et al., “Active learning for sentiment analysis on data streams: Methodology and workflow Implementation in the ClowdFlows platform”, Information Processing & Management, Mar. 2015, vol. 51, No. 2, pp. 18… [cited by applicant]
Thorton, et al., “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algoirthms,” 2013, KDD, pp. 847-855, 2013. [cited by applicant]
Duan, et al., “A Method of Determine the Hyper-parameter Range for Tuning RBF Support Vector Machines,” 2010, IEEE pp. 1-4, 2010. [cited by applicant]
Office Action mailed Jul. 20, 2023 in Chinese Patent application No. 202110397530.4, Amazon Technologies, Inc., pp. 1-14 (including translation). [cited by applicant]
Mattias Varewyck, et al., “A Practical Approach to Model Selection for Support Vector Machines With a Gaussian Kernel”, IEEE Transactions on Systems, Man. and Cybernetics, Part B: Cybernetics, Apr. 2011, pp. 330-340. Vo… [cited by applicant]
Chun-Xiang Li, et al., “Study on algorithm with optimized parameter of least squares supporting vector”, Journal of Hangzhou University of Electronics Science and Technology Aug. 15, 2010. [cited by applicant]
“[Machine Leaning] Linear Regression with one variable”, Jul. 18, 2013, Retrieved from https://blog.51cto.com/u_15127657/4081678 on Aug. 11, 2023. [cited by applicant]
Jain, et al., Using Bloom Filters to Refine Web Search Results:, Jun. 2005, WebDB, 2005, pp. 1-6. [cited by applicant]
Bardenet, et al., “Collaborative Hyperparameter Tuning,” 2013, Proceedings of the 30th International Conference on Machine Learning, vol. 28, pp. 1-9, 2013. [cited by applicant]
U.S. Appl. No. 18/775,912, filed Jul. 19, 2024, Leo Parker Dirac, et al. [cited by applicant]
Gopfert, et al., Measurement Extraction with Natural Language Processing: A Review, Finding of the Association for Computational Linguistics: EMNLP 2022, Dec. 11, 2022, pp. 2191-2215. [cited by applicant]
Di Pierro, et al, LPG-Based Knowledge Graphs: A Survey, a Proposal and Current Trends, Information 2023, 14, 154, Mar. 2023, pp. 1-32. [cited by applicant]
Huser, Forecasting Intracranial Hypertension Using Time Series and Waveform Features, Masters Thesis, ETH Zurich, 2015, pp. 1-78 (Year: 2015). [cited by applicant]
Chapelle, et al., “Choosing Multiple Parameters for Support Vector Machines,” Machine Learning, 2002, vol. 46, pp. 131-159. [cited by applicant]
Karatzoglou, et al., “Support Vector Machines in R,” Journal of Statistical Software, Apr. 2006, vol. 15, Issue 9, pp. 1-28. [cited by applicant]
First Examination Report mailed Dec. 5, 2025 in Indian Patent Application No. 202318087522, Amazon Technologies, Inc., 9 pages. [cited by applicant]