IP Library › Granted Patent US 12,530,564
Granted Patent B1
US 12,530,564 · App. 17/456,752 · Granted Jan 20, 2026

Combined neural network

Inventors: Jacob Vincent Bouvrie (Arlington, MA); Tianbai Cui (Seabrook, NH)
Assignee: Kayak Software Corporation
G06N3/045G06F18/2193
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,564
App. No.
17/456,752
Granted
Jan 20, 2026
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for creating a combined neural network. One of the methods includes creating a combined neural network by combining a) a first neural network that includes a first plurality of weights and was trained to predict single output values that have a type with b) a second neural network that includes a second plurality of weights, the combined neural network created to predict a sequence of values, each value of which has the type; training the combined neural network by: determining a loss function for the combined neural network using a result of a comparison of training output data and expected output data; and updating one or more weights in the second plurality of weights for the second neural network using the loss function; and storing the combined neural network in memory.

Claims (57)

1 . A system comprising one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

creating a combined neural network by combining a) a first neural network that includes a first plurality of weights and was trained to predict single output values that have a type with b) a second neural network that includes a second plurality of weights, wherein combining the first neural network and the second neural network comprises removing an output layer from the first neural network and combining an output from an earlier layer to an input layer for the second neural network, the combined neural network created to predict a sequence of values, each value of which has the type;

training the combined neural network by:

determining a loss function for the combined neural network using a result of a comparison of training output data and expected output data, the training output data generated by the combined neural network using input data; and

updating one or more weights in the second plurality of weights for the second neural network using the loss function; and

storing the combined neural network in memory.

2 . The system of claim 1 , wherein determining the loss function comprises:

computing one or more loss-based metrics; and

defining the loss function using the one or more loss-based metrics.

3 . The system of claim 1 , wherein training the combined neural network comprises maintaining, unchanged, one or more of the weights in the first plurality of weights.

4 . The system of claim 3 , wherein training the combined neural network comprises:

updating one or more initial states for the second neural network while maintaining, unchanged, the weights in the first plurality of weights and the second plurality of weights.

5 . The system of claim 1 , the operations comprising:

after training the combined neural network:

receiving a runtime output data from the combined neural network that the combined neural network generated using runtime input data; and

providing, to a device, the runtime output data to cause the device to present the runtime output data.

6 . The system of claim 1 , the operations comprising initializing the second neural network with random weights prior to creating the combined neural network.

7 . The system of claim 6 , wherein the second neural network is not trained prior to creating the combined neural network.

8 . The system of claim 1 , wherein creating the combined neural network comprises creating the combined neural network by connecting an output of a layer from the first neural network with an input of a first layer in the second neural network.

9 . The system of claim 8 , wherein connecting the output of the layer from the first neural network with the input of the first layer in the second neural network comprises connecting the output of a second to last layer from the first neural network with the input of the first layer in the second neural network.

10 . The system of claim 1 , wherein:

the first neural network was trained to generate a single output value; and

training the combined neural network comprises training the combined neural network to generate a time-series output.

11 . The system of claim 1 , wherein:

the first neural network was trained to predict a value for a time period; and

training the combined neural network comprises training the combined neural network to predict a sequence of values for the time period.

12 . The system of claim 1 , wherein:

the first neural network was trained to predict a value for a time period; and

training the combined neural network comprises training the combined neural network to predict a sequence of values for a time window that includes the time period.

13 . The system of claim 1 , wherein the first neural network comprises a multilayer perceptron neural network model.

14 . The system of claim 1 , wherein the second neural network comprises an autoregressive model.

15 . The system of claim 1 , wherein determining the loss function for the combined neural network using the result of the comparison of the training output data and the expected output data comprises:

providing input data to the combined neural network;

in response to providing input data to the combined neural network, receiving the training output data; and

determining a difference between the training output data and the expected output data that is associated with the input data.

16 . The system of claim 1 , the operations comprising:

training the first neural network that includes the first plurality of weights until a first convergence criteria is satisfied, wherein:

training the combined neural network comprises training the combined neural network until a second convergence criteria is satisfied.

17 . A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

creating a combined neural network by combining a) a first neural network that includes a first plurality of weights and was trained to predict single output values that have a type with b) a second neural network that includes a second plurality of weights, wherein combining the first neural network and the second neural network comprises removing an output layer from the first neural network and combining an output from an earlier layer to an input layer for the second neural network, the combined neural network created to predict a sequence of values, each value of which has the type;

training the combined neural network by:

determining a loss function for the combined neural network using a result of a comparison of training output data and expected output data, the training output data generated by the combined neural network using input data; and

updating one or more weights in the second plurality of weights for the second neural network using the loss function;

and

storing the combined neural network in memory.

18 . The non-transitory computer storage medium of claim 17 , wherein:

the first neural network was trained to generate a single output value; and

training the combined neural network comprises training the combined neural network to generate a time-series output.

19 . The non-transitory computer storage medium of claim 17 , wherein:

the first neural network was trained to predict a value for a time period; and

training the combined neural network comprises training the combined neural network to predict a sequence of values for the time period.

20 . A computer-implemented method comprising:

creating a combined neural network by combining a) a first neural network that includes a first plurality of weights and was trained to predict single output values that have a type with b) a second neural network that includes a second plurality of weights, wherein combining the first neural network and the second neural network comprises removing an output layer from the first neural network and combining an output from an earlier layer to an input layer for the second neural network, the combined neural network created to predict a sequence of values, each value of which has the type;

training the combined neural network by:

determining a loss function for the combined neural network using a result of a comparison of training output data and expected output data, the training output data generated by the combined neural network using input data; and

updating one or more weights in the second plurality of weights for the second neural network using the loss function; and

storing the combined neural network in memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2022
From: BOUVRIE, JACOB VINCENT; CUI, TIANBAI
To: KAYAK SOFTWARE CORPORATION
Reel/Frame 060055/0570 →
References Cited (37)
US 10690806B2 · Dail · 2020 [cited by examiner]
US 10802488B1 · Abeloe · 2020 [cited by examiner]
US 11164093B1 · Zappella · 2021 [cited by examiner]
US 11443623B1 · Ratrout · 2022 [cited by examiner]
US 11733427B1 · Thielke · 2023 [cited by examiner]
US 11889112B2 · Ding · 2024 [cited by examiner]
US 20030209893A1 · Breed · 2003 [cited by examiner]
US 20040129478A1 · Breed · 2004 [cited by examiner]
US 20060208169A1 · Breed · 2006 [cited by examiner]
US 20140195466A1 · Phillipps · 2014 [cited by examiner]
US 20170299772A1 · Yuzhakov · 2017 [cited by examiner]
US 20180101147A1 · Khabibrakhmanov · 2018 [cited by examiner]
US 20190286970A1 · Karaletsos · 2019 [cited by examiner]
US 20200175691A1 · Zhang · 2020 [cited by examiner]
US 20200210809A1 · Kaizerman · 2020 [cited by examiner]
US 20200309993A1 · Ganshin · 2020 [cited by examiner]
US 20200375549A1 · Wexler · 2020 [cited by examiner]
US 20210146531A1 · Tremblay · 2021 [cited by examiner]
US 20210334644A1 · Yu · 2021 [cited by examiner]
US 20210342760A1 · Patel · 2021 [cited by examiner]
US 20210374502A1 · Roth · 2021 [cited by examiner]
US 20220041299A1 · Wankewycz · 2022 [cited by examiner]
US 20220076133A1 · Yang · 2022 [cited by examiner]
US 20220319018A1 · Gervais · 2022 [cited by examiner]
US 20220343221A1 · Cook · 2022 [cited by examiner]
US 20230011970A1 · Serra Lleti · 2023 [cited by examiner]
US 20230409572A1 · Bouvrie · 2023 [cited by examiner]
US 20240062515A1 · Oh · 2024 [cited by examiner]
US 20240232729A1 · Bouvrie · 2024 [cited by examiner]
US 20250137689A1 · Rigney · 2025 [cited by examiner]
Cheng et al., “Wide & Deep Learning for Recommender Systems,” DLRS., Sep. 15, 2016, pp. 7-10. [cited by applicant]
Cho et al., “On the Properties of Neural Machine Translation: Encoder-Decoder Approaches,” Semantics and Structure in Statistical Translation, Oct. 25, 2014, 103-111. [cited by applicant]
Gal et al., “A Theoretically Grounded Application of Dropout in Recurrent Neural Networks,” Advances in Neural Information Processing Systems (NIPS 2016), Oct. 5, 2016, vol. 29, 14 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, 1997, 9(8):1735-1780. [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research, 2014, 15:1929-1958. [cited by applicant]
Wikipedia.com, “LightGBM,” last updated May 14, 2021, Retrieved from URL <https://en.wikipedia.org/w/index.php?title=LightGBM&oldid=1023071208>, 3 pages. [cited by applicant]