IP Library Granted Patent US 11,449,796
Granted Patent B2
US 11,449,796 · App. 16/578,060 · Granted Sep 20, 2022

Machine learning inference calls for database query processing

Inventors: Sangil Song (Bellevue, WA); Yongsik Yoon (Sammamish, WA); Kamal Kant Gupta (Belmont, CA); Saileshwar Krishnamurthy (Palo Alto, CA); Stefano Stefani (Issaquah, WA); Sudipta Sengupta (Sammamish, WA); Jaeyun Noh (Sunnyvale, CA)
Assignee: Amazon Technologies, Inc.
G06N20/00G06F16/2433G06F16/24542G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,449,796
App. No.
16/578,060
Granted
Sep 20, 2022
Kind
B2
Abstract

Techniques for making machine learning inference calls for database query processing are described. In some embodiments, a method of making machine learning inference calls for database query processing may include generating a first batch of machine learning requests based at least on a query to be performed on data stored in a database service, wherein the query identifies a machine learning service, sending the first batch of machine learning requests to an input buffer of an asynchronous request handler, the asynchronous request handler to generate a second batch of machine learning requests based on the first batch of machine learning requests, and obtaining a plurality of machine learning responses from an output buffer of the asynchronous request handler, the machine learning responses generated by the machine learning service using a machine learning model in response to receiving the second batch of machine learning requests.

Claims (37)

1. A computer-implemented method comprising:

receiving a request at a database service, wherein the request includes a structured query language (SQL) query to be performed on at least a portion of a dataset in the database service and wherein the request identifies a machine learning service to be used in processing the SQL query;

creating a virtual operator to perform at least a portion of the SQL query;

generating a first batch of machine learning requests based at least on the portion of the SQL query performed by the virtual operator;

sending the first batch of machine learning requests to an input buffer of an asynchronous request handler, the asynchronous request handler to generate a second batch of machine learning requests based on the first batch of machine learning requests;

obtaining a plurality of machine learning responses from an output buffer of the asynchronous request handler, the machine learning responses generated by the machine learning service using a machine learning model in response to receiving the second batch of machine learning requests, wherein the machine learning service publishes an application programming interface (API) to perform inference using the machine learning model in response to requests received from a plurality of users; and

generating a query response based on the machine learning responses.

2. The computer-implemented method of claim 1 , wherein generating a first batch of machine learning requests based at least on the SQL query, further comprises:

determining a query execution plan that minimizes a number of records associated with machine learning request.

3. The computer-implemented method of claim 1 , wherein the machine learning service adds a flag to the output buffer when the second batch of machine learning requests has been processed.

4. A computer-implemented method comprising:

executing at least a portion of a query on data stored in a database service using a temporary data structure to generate a first batch of machine learning requests, wherein the query is a structured query language (SQL) query, wherein the query identifies a machine learning service using an application programming interface (API) call to the machine learning service, and wherein the machine learning service publishes the API to perform inference using the machine learning model in response to requests received from a plurality of users;

generating a second batch of machine learning requests based on the first batch of machine learning requests and based on the machine learning service; and

obtaining a plurality of machine learning responses, the machine learning responses generated by the machine learning service using a machine learning model in response to receiving the second batch of machine learning requests.

5. The computer-implemented method of claim 4 , wherein the query identifies the machine learning service using an endpoint associated with the machine learning model hosted by the machine learning service.

6. The computer-implemented method of claim 4 , wherein the second batch of machine learning requests is sent to the machine learning service over at least one network.

7. The computer-implemented method of claim 4 , further comprising:

sending a request to the machine learning service for the machine learning model;

receiving the machine learning model from the machine learning service, the machine learning model compiled for the database service by the machine learning service; and

wherein the second batch of machine learning requests is sent to the machine learning model hosted by the database service.

8. The computer-implemented method of claim 7 , further comprising:

storing a copy of the machine learning model in a plurality of nodes of the database service, wherein machine learning requests generated during the query processing by a particular node of the database service are sent to the copy of the machine learning model stored on the particular node.

9. The computer-implemented method of claim 4 , wherein the second batch size is different from the first batch size, and wherein the second batch size is associated with the machine learning service.

10. The computer-implemented method of claim 4 , wherein the first batch of machine learning requests includes machine learning requests generated in response to multiple queries received from a plurality of different users.

11. A system comprising:

a machine learning service implemented by a first one or more electronic devices; and

a database service implemented by a second one or more electronic devices, the database service including instructions that upon execution cause the database service to:

execute at least a portion of a query on data stored in a database service using a temporary data structure to generate a first batch of machine learning requests, wherein the query is a structured query language (SQL) query, wherein the query identifies the machine learning service using an application programming interface (API) call to the machine learning service, and wherein the machine learning service publishes the API to perform inference using the machine learning model in response to requests received from a plurality of users;

generating a second batch of machine learning requests based on the first batch of machine learning requests and based on the machine learning service; and

obtain a plurality of machine learning responses, the machine learning responses generated by the machine learning service using a machine learning model in response to receiving the second batch of machine learning requests.

12. The system of claim 11 , wherein the query identifies the machine learning service using an endpoint associated with the machine learning model hosted by the machine learning service.

13. The system of claim 11 , wherein the second batch of machine learning requests is sent to the machine learning service over at least one network.

14. The system of claim 11 , further comprising:

sending a request to the machine learning service for the machine learning model;

receiving the machine learning model from the machine learning service, the machine learning model compiled for the database service by the machine learning service;

wherein the second batch of machine learning requests is sent to the machine learning model hosted by the database service; and

storing a copy of the machine learning model in a plurality of nodes of the database service, wherein machine learning requests generated during the query processing by a particular node of the database service are sent to the copy of the machine learning model stored on the particular node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2020
From: SONG, SANGIL; YOON, YONGSIK; GUPTA, KAMAL KANT; KRISHNAMURTHY, SAILESHWAR; STEFANI, STEFANO; SENGUPTA, SUDIPTA; NOH, JAEYUN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 052800/0066 →
Cited By (3)
US 12,475,365 US 12,511,282 US 12,632,775