IP Library Granted Patent US 9,389,913
Granted Patent B2
US 9,389,913 · App. 13/383,594 · Granted Jul 12, 2016

Resource assignment for jobs in a system having a processing pipeline that satisfies a data freshness query constraint

Inventors: Kimberly Keeton (San Francisco, CA); Charles B. Morrey, III (Palo Alto, CA); Craig A. Souies (San Francisco, CA); Alistair Veitch (Mountain View, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F9/50G06F9/5011G06F17/30286
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,389,913
App. No.
13/383,594
Granted
Jul 12, 2016
Kind
B2
Abstract

A set of jobs to be scheduled is identified ( 402 ) in a system including a processing pipeline having plural processing stages that apply corresponding different processing to a data update to allow the data update to be stored. The set of jobs is based on one or both of the data update and a query that is to access data in the system. The set of jobs is scheduled ( 404 ) by assigning resources to perform the set of jobs, where assigning the resources is subject to at least one constraint selected from at least one constraint associated with the data update and at least one constraint associated with the query.

Claims (38)

1. A method comprising:

receiving, by at least one processor, a query having at least one query constraint comprising a freshness constraint specifying how up-to-date data in a response to the query should be;

identifying, by the at least one processor, a set of jobs to be scheduled in a system including a processing pipeline having plural processing stages that are to apply corresponding different processing to a data update to allow the data update to be stored, wherein the set of jobs is based on the data update and the query that requests access of data in the system, wherein the freshness constraint causes the set of lobs to access data of selected processing stages of the plural processing stages of the processing pipeline, the selected processing stages based on the freshness constraint; and

scheduling, by the at least one processor, the set of jobs by assigning resources to perform the set of jobs, wherein assigning the resources is subject to at least one constraint associated with the data update and the at least one query constraint of the query, the resources assigned being based on the selected processing stages.

2. The method of claim 1 , wherein assigning the resources comprises:

assigning resources to at least one of the processing stages of the processing pipeline; and assigning resources to a query processing engine, wherein the query processing engine performs query processing in response to the query.

3. The method of claim 1 ,

wherein assigning the resources to perform the set of jobs is based on the freshness constraint included in the query, the set of jobs comprising processing of the query using a portion of the assigned resources.

4. The method of claim 1 , wherein the at least one constraint associated with the data update is selected from among an input data constraint relating to reading input data from one or more different computer nodes, a precedence constraint, an execution time constraint, and a resource constraint.

5. The method of claim 1 , wherein assigning the resources comprises making decisions selected from the group consisting of:

determining a degree of parallelism used for a given job;

determining specific ones of the resources to allocate to the given job;

determining a fraction of each of the resources to allocate to the given job;

determining the given job's start time; and

determining the given job's end time.

6. The method of claim 1 , further comprising performing the data update in the processing pipeline that has a stage to transform the data update to allow content of the data update to be stored into a database.

7. The method of claim 6 , wherein transforming the data update comprises at least one selected from among remapping identifiers of the data update, sorting the data update, and merging the data update.

8. The method of claim 1 , wherein the system comprises the at least one processor, and wherein the identifying and the scheduling are performed by a resource allocation and scheduling mechanism in the system.

9. A computer system comprising:

at least one central processing unit (CPU); and

a scheduling module executable on the at least one CPU to:

receive a query having at least one query constraint comprising a freshness constraint specifying how up-to-date data in a response to the query should be, wherein the query is for execution in a system having a processing pipeline including plural processing stages to process a data update, the plural processing stages selected from among: an ingest stage, an identifier remapping stage, a sorting stage, and a merging stage;

identify a set of jobs to be scheduled based on the received query and the data update to be processed by the processing pipeline, wherein the freshness constraint causes the set of lobs to access data of selected processing stages of the plural processing stages of the processing pipeline, the selected processing stages based on the freshness constraint; and

assign resources of the system having the processing pipeline to the set of jobs according to the at least one query constraint and at least one constraint associated with the data update, the resources assigned being based on the selected processing stages.

10. The method of claim 3 , wherein the at least one query constraint of the query further comprises at least one selected from among a query performance goal, an input data constraint relating to reading input data from one or more different computer nodes, a precedence constraint, an execution time constraint, and a resource constraint.

11. The method of claim 3 , wherein different levels of the freshness constraint included in the query cause the processing of the query to access different respective combinations of the plural processing stages.

12. The method of claim 3 , wherein the query further includes a response time constraint specifying a target response time of the response to the query,

wherein assigning the resources to perform the set of jobs is further based on the response time constraint.

13. The computer system of claim 9 , wherein different levels of the freshness constraint included in the query cause the processing of the query to access different respective combinations of the plural processing stages.

14. The computer system of claim 9 , wherein the query further includes a response time constraint specifying a target response time of the response to the query,

wherein the assigning of the resources to the set of jobs is further based on the response time constraint.

15. An article comprising at least one non-transitory computer-readable storage medium storing instructions that upon execution cause a computer system to:

receive a query having at least one query constraint comprising a freshness constraint specifying how up-to-date data in a response to the query should be;

identify a set of jobs to be scheduled in a system that has a processing pipeline to process a data update received at the processing pipeline, wherein the processing pipeline has plural processing stages that apply corresponding different processing to the data update to allow the data update to be stored, wherein the set of jobs identified is based on the data update and the query, wherein the freshness constraint causes the set of lobs to access data of selected processing stages of the plural processing stages of the processing pipeline, the selected processing stages based on the freshness constraint; and

allocate resources to perform the set of jobs that is based on a constraint associated with the data update and the at least one query constraint of the query, the resources assigned being based on the selected processing stages.

16. The article of claim 15 , wherein the constraint associated with the data update comprises at least one selected from among an input data constraint relating to reading input data from one or more different computer nodes, a precedence constraint, an execution time constraint, and a resource constraint.

17. The article of claim 15 , wherein allocating the resources comprises employing fitness tests that measure performance of different jobs, and wherein employing the fitness tests determines a level of parallelism to use for query processing and for each of the plural processing stages.

18. The article of claim 15 , wherein different levels of the freshness constraint included in the query cause the processing of the query to access different respective combinations of the plural processing stages.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2012
From: KEETON, KIMBERLY; MORREY III, CHARLES B; SOULES, CRAIG A; VEITCH, ALISTAIR
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 027824/0301 →
Continuity (1)
Related Publication 20140223444A1 · Aug 7, 2014