IP Library › Granted Patent US 11,422,853
Granted Patent B2
US 11,422,853 · App. 16/554,769 · Granted Aug 23, 2022

Dynamic tree determination for data processing

Inventors: Govindaswamy Bacthavachalu (Sammamish, WA); Peter Grant Gavares (San Francisco, CA); Ahmed A. Badran (Issaquah, WA); James E. Scharf, Jr. (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06F9/4881G06F9/5066G06F16/22G06F16/2246G06F16/2455G06F16/2471G06F16/24532G06F16/24556G06F16/275G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,422,853
App. No.
16/554,769
Granted
Aug 23, 2022
Kind
B2
Abstract

Data can be processed in parallel across a cluster of nodes using a parallel processing framework. Using Web services calls between components allows the number of nodes to be scaled as necessary, and allows developers to build applications on the framework using a Web services interface. A job scheduler works together with a queuing service to distribute jobs to nodes as the nodes have capacity, such that jobs can be performed in parallel as quickly as the nodes are able to process the jobs. Data can be loaded efficiently across the cluster, and levels of nodes can be determined dynamically to process queries and other requests on the system.

Claims (37)

1. A system, comprising:

a plurality of computing devices, respectively comprising at least one processor and a memory, to implement a Web service;

the Web service, configured to:

receive, via an interface for the Web service, a request that causes a job received at the Web service to execute, wherein the job obtains data from one or more input sources and stores a transformed version of the data to one or more output locations;

automatically scale a number of instances to perform at least a portion of the job, based at least in part, on a size of the data to be obtained from the one or more input sources; and

execute the portion of the job at the automatically scaled number of instances as part of executing the job at the Web service.

2. The system of claim 1 , wherein to automatically scale the number of instances to perform the portion of the job, the Web service is configured to scale the number of nodes up or down according to a load caused by the job.

3. The system of claim 2 , wherein the number of nodes are scaled up or down from an initially selected number of nodes via the interface for the Web service.

4. The system of claim 1 , wherein the Web service is further configured to:

receive, via the interface for the Web service, a query directed to the transformed version of the data; and

perform, by the Web service, the query with respect to the transformed version of the data to provide a query result.

5. The system of claim 1 , wherein the transformed version of the data is in a user-specified schema for the data.

6. The system of claim 1 , wherein the automatic scaling of the number of instances to perform the portion of the job is performed as part of monitoring the performance of the job.

7. The system of claim 1 , wherein the interface for the Web service is a graphical user interface (GUI).

8. A method, comprising:

receiving, via an interface for a Web service, a request that causes a job received at the Web service to execute, wherein the job obtains data from one or more input sources and stores a transformed version of the data to one or more output locations;

automatically scaling, by the Web service, a number of instances to perform at least a portion of the job, based at least in part, on a size of the data to be obtained from the one or more input sources; and

executing, by the Web service, the portion of the job at the scaled number of instances as part of executing the job at the Web service.

9. The method of claim 8 , wherein automatically scaling the number of instances to perform the portion of the job comprises scaling the number of nodes up or down according to a load caused by the job.

10. The method of claim 9 , wherein the number of nodes are scaled up or down from an initially selected number of nodes via the interface for the Web service.

11. The method of claim 8 , further comprising:

receiving, via the interface for the Web service, a query directed to the transformed version of the data; and

performing, by the Web service, the query with respect to the transformed version of the data to provide a query result.

12. The method of claim 8 , wherein the transformed version of the data is in a user-specified schema for the data.

13. The method of claim 8 , wherein the automatically scaling the number of instances to perform the portion of the job is performed as part of monitoring the performance of the job by the Web service.

14. The method of claim 8 , wherein the interface for the Web service is a graphical user interface (GUI).

15. One or more non-transitory, computer-readable storage media, storing program instructions that when executed by the one or more computing devices, cause the one or more computing devices to implement:

receiving, via an interface for a Web service, a request that causes a job received at the Web service to execute, wherein the job obtains data from one or more input sources and stores a transformed version of the data to one or more output locations;

automatically scaling, by the Web service, a number of instances to perform at least a portion of the job, based at least in part, on a size of the data to be obtained from the one or more input sources; and

executing, by the Web service, the portion of the job at the scaled number of instances as part of executing the job at the Web service.

16. The or more non-transitory, computer-readable storage media of claim 15 , wherein, in automatically determining the number of instances to perform the portion of the job, the program instructions when executed on or across the one or more computing devices cause the one or more computing devices to implement scaling the number of nodes up or down according to a load caused by the job.

17. The or more non-transitory, computer-readable storage media of claim 16 , wherein the number of nodes are scaled up or down from an initially selected number of nodes via the interface for the Web service.

18. The or more non-transitory, computer-readable storage media of claim 15 , wherein the one or more non-transitory, computer-readable storage media store further program instructions that when executed on or across the one or more computing devices cause the one or more computing devices to further implement:

receiving, via the interface for the Web service, a query directed to the transformed version of the data; and

performing, by the Web service, the query with respect to the transformed version of the data to provide a query result.

19. The or more non-transitory, computer-readable storage media of claim 15 , wherein the transformed version of the data is in a user-specified schema for the data.

20. The or more non-transitory, computer-readable storage media of claim 15 , wherein the automatically scaling the number of instances to perform the portion of the job is performed as part of monitoring the performance of the job by the Web service.

Continuity (6)
Continuation 15344180 · Nov 4, 2016
Continuation 14745178 · Jun 19, 2015
Continuation 14107570 · Dec 16, 2013
Continuation 13620240 · Sep 14, 2012
Continuation 12200821 · Aug 28, 2008
Related Publication 20200057672A1 · Feb 20, 2020
Cited By (1)
US 12,664,024