IP Library › Granted Patent US 12,524,650
Granted Patent B2
US 12,524,650 · App. 17/839,252 · Granted Jan 13, 2026

Efficient cross-platform serving of deep neural networks for low latency applications

Inventor: Aaron Andalman (Berkeley, CA)
Assignee: Cognitiv Corp.
G06N3/044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,650
App. No.
17/839,252
Granted
Jan 13, 2026
Kind
B2
Abstract

Systems, apparatuses, and methods for implementation of an inference or prediction process using a recurrent neural network (RNN) that is particularly advantageous for low-latency applications. Embodiments introduce an implementation of a recurrent neural network-based system which results in a fixed inference time (i.e., a constant computation time to perform an inference stage) that is independent of input data sequence length. Embodiments may be used to implement real-time data mapping and management and perform an inference strategy that enables the system to be used for serving different types of models, including sequential deep neural networks for low latency (i.e., real-time, or close to real-time) applications.

Claims (34)

1 . A method of implementing an inference process for a low latency application, wherein for each of a plurality of time segments, the method comprises:

obtaining data corresponding to a current time segment for a low latency application;

retrieving a previously generated hidden layer of a Recurrent Neural Network (RNN) used to implement an inference process for the low latency application, wherein the previously generated hidden layer of the RNN is a result of executing the RNN for an earlier time segment and the hidden layer is retrieved from a low latency key value store;

executing the RNN using the obtained data corresponding to the current time segment and the retrieved previously generated hidden layer;

providing an output of the executed RNN to a user as an output of the low latency application; and

storing at least part of the hidden layer of the executed RNN in the low latency key value store for retrieval during a subsequent time segment.

2 . The method of claim 1 , wherein a representation for the data corresponding to a current time segment is obtained by generating the representation using a pre-trained model.

3 . The method of claim 1 , wherein the low latency application is responding to a bid-request for serving content to a viewer of a web page.

4 . The method of claim 1 , wherein the data corresponding to the current time segment comprises one or more of viewer data and contextual data, the contextual data comprising data regarding a web page viewed by the viewer.

5 . The method of claim 1 , wherein the data corresponding to a current time segment comprises a key used to retrieve the previously generated state of the RNN.

6 . The method of claim 1 , further comprising aggregating the data corresponding to the current time segment with other data prior to executing the RNN.

7 . The method of claim 1 , further comprising using the output of the RNN as an input to a decision process, the decision process comprising one or more of a resource allocation, a selection of content to present to a viewer, or a decision whether to initiate an event or action.

8 . A system for implementing an inference process for a low latency application, comprising:

one or more electronic processors configured to execute a set of computer-executable instructions; and

one or more non-transitory electronic data storage media containing the set of computer-executable instructions, wherein when executed, the instructions cause the one or more electronic processors to

obtaining data corresponding to a current time segment for a low latency application;

retrieving a previously generated hidden layer a Recurrent Neural Network (RNN) used to implement an inference process for the low latency application, wherein the previously generated hidden layer of the RNN is a result of executing the RNN for an earlier time segment and the hidden layer is retrieved from a low latency key value store;

executing the RNN using the obtained data corresponding to the current time segment and the retrieved previously generated hidden layer;

providing an output of the executed RNN to a user as an output of the low latency application; and

storing at least part of the hidden layer of the executed RNN in the low latency key value store for retrieval during a subsequent time segment.

9 . The system of claim 8 , wherein the low latency application is responding to a bid-request for serving content to a viewer of a web page.

10 . The system of claim 8 , wherein the data corresponding to the current time segment comprises one or more of viewer data and contextual data, the contextual data comprising data regarding a web page viewed by the viewer.

11 . The system of claim 8 , wherein the data corresponding to a current time segment comprises a key used to retrieve the previously generated layer of the RNN.

12 . The system of claim 8 , wherein the instructions further cause the one or more electronic processors to use the output of the RNN as an input to a decision process, the decision process comprising one or more of a resource allocation, a selection of content to present to a viewer, or a decision whether to initiate an event or action.

13 . One or more non-transitory computer-readable media comprising a set of computer-executable instructions for implementing an inference process for a low latency application, that when executed by one or more programmed electronic processors, cause the processors to

obtaining data corresponding to a current time segment for a low latency application;

retrieving a previously generated hidden layer a Recurrent Neural Network (RNN) used to implement an inference process for the low latency application, wherein the previously generated hidden layer of the RNN is a result of executing the RNN for an earlier time segment and the hidden layer is retrieved from a low latency key value store,

executing the RNN using the obtained data corresponding to the current time segment and the retrieved previously generated hidden layer;

providing an output of the executed RNN to a user as an output of the low latency application; and

storing at least part of the hidden layer of the executed RNN in the low latency key value store for retrieval during a subsequent time segment.

14 . The one or more non-transitory computer-readable media of claim 13 , wherein the inference model is used as part of an application responding to a bid-request for serving content to a viewer of a web page.

15 . The one or more non-transitory computer-readable media of claim 13 , wherein the data corresponding to the current time segment comprises one or more of viewer data and contextual data, the contextual data comprising data regarding a web page viewed by the viewer.

16 . The one or more non-transitory computer-readable media of claim 13 , wherein the data corresponding to a current time segment comprises a key used to retrieve the previously generated layer of the RNN.

17 . The one or more non-transitory computer-readable media of claim 13 , wherein the instructions further cause the one or more electronic processors to use the output of the executed RNN as an input to a decision process, the decision process comprising one or more of a resource allocation, a selection of content to present to a viewer, or a decision whether to initiate an event or action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2022
From: ANDALMAN, AARON
To: COGNITIV CORP.
Reel/Frame 061529/0596 →
Continuity (2)
Provisional Application 63210635 · Jun 15, 2021
Related Publication 20220398433A1 · Dec 15, 2022
References Cited (22)
US 10671908B2 · Simard · 2020 [cited by applicant]
US 11704299B1 · Bansal · 2023 [cited by examiner]
US 20050193414A1 · Horvitz · 2005 [cited by examiner]
US 20180168515A1 · Farahmand et al. · 2018 [cited by applicant]
US 20190279100A1 · Cline · 2019 [cited by applicant]
US 20190332952A1 · Nonaka et al. · 2019 [cited by applicant]
US 20200226675A1 · Mitra · 2020 [cited by examiner]
US 20200310540A1 · Hussami et al. · 2020 [cited by applicant]
US 20210073995A1 · Yang et al. · 2021 [cited by applicant]
US 20220039756A1 · Mikhno · 2022 [cited by examiner]
US 20220292339A1 · Byrne · 2022 [cited by examiner]
US 20220391459A1 · Lim · 2022 [cited by examiner]
US 20220398433A1 · Andalman · 2022 [cited by examiner]
EP 3627399A1 · 2019 [cited by applicant]
WO 2020245936A1 · 2020 [cited by applicant]
European Patent Office; “Extended European Search Report” dated Mar. 13, 2025; EP Patent Application No. 22825633.5; pp. 1-10 (2025). [cited by applicant]
Gharibshah Zhabiz et al.: “Deep Learning for User Interest and Response Prediction in Online Display Advertising”, Data Science and Engineering, [Online] vol. 5, No. 1, Jan. 17, 2020 (Jan. 17, 2020), pp. 12-26, XP055893… [cited by applicant]
International Searching Authority of the PCT (US); “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration” dated Oct. 14, 202… [cited by applicant]
Hazelwood et al., “Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective”; Published in: 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA); Date of Conference: Fe… [cited by applicant]
Shea et al., “Heterogeneous Scheduling of Deep Neural Networks for Low-power Real-time Designs”; ACM Journal on Emerging Technologies in Computing Systems; vol. 15; Issue 4; Oct. 2019; Article No. 36; pp. 1-31; Online:D… [cited by applicant]
Neil; Thesis submitted to attain the degree of Doctor of Sciences of Eth Zurich; PhD diss. 24392; University of Zurich; “Deep Neural Networks and Hardware Systems for Event-driven Data”; pp. 1-154 (2017). [cited by applicant]
Graves et al.; “Improving neural language models with a continuous cache”; Facebook AI Research; arXiv preprint arXiv:1612.04426; Dec. 13, 2016; pp. 1-9. [cited by applicant]