IP Library Granted Patent US 10,922,133
Granted Patent B2
US 10,922,133 · App. 16/072,701 · Granted Feb 16, 2021

Method and apparatus for task scheduling

Inventors: Le He (Hangzhou, CN); Yan Huang (Hangzhou, CN); Yingjie Shi (Hangzhou, CN); Jie Zhang (Hangzhou, CN); Chen Zhang (Hangzhou, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G06F9/4881G06F9/50G06F9/5033G06F9/5072H04L67/10G06F2209/486
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,922,133
App. No.
16/072,701
Granted
Feb 16, 2021
Kind
B2
Abstract

The disclosed embodiments provide a task scheduling method and apparatus. Network resources needed for a task to perform cross-cluster reading and writing are analyzed to obtain usage information of the occupied network resources for reading and writing; and the task is scheduled according to the usage information of the network resources needed for reading and writing. Because the usage information of the network resources occupied for reading and writing respectively represent network resources that can be saved by the cluster where access data is located when the task is scheduled for reading and writing, it can be determined that the cluster to which the task is scheduled can enable the task to occupy less network resources, thereby solving the problem of high bandwidth usage across clusters in current systems.

Claims (48)

1. A method comprising:

obtaining, by a task management device of a distributed system, network resources needed to perform cross-cluster reading and writing, the network resources associated with a task and comprising an amount of input data and an amount of output data for the task based on a history record of past executions of the task;

analyzing, by the task management device of the distributed system, the network resources by calculating an input-output ratio representing a proportion of the network resources needed for reading and writing for the task, the input-output ratio equal to a ratio of the amount of input data to the amount of output data; and

scheduling, by the task management device, the task according to the network resources.

2. The method of claim 1 , the scheduling the task comprising scheduling the task to a target cluster where read dependent data is located if the network resources needed for reading are more than network resources needed for writing.

3. The method of claim 1 , the scheduling the task further comprising:

determining whether the task meets a preset selection criteria, the preset selection criteria comprising that the input-output ratio is greater than a preset first threshold, the preset first threshold greater than one; and

scheduling the task to a target cluster where the dependent data read by the task locates.

4. The method of claim 3 , further comprising:

obtaining, by the task management device, a task identifier for a task meeting the selection criteria; and

generating, by the task management device, scheduling information that includes the task identifier.

5. The method of claim 4 , the scheduling the task to a target cluster where the dependent data read by the task locates comprising:

obtaining the task identifier for the task when the task is received; and

scheduling the task to a target cluster where dependent data for the task locates if the task identifier of the task matches the task identifier in the scheduling information.

6. The method of claim 4 , the obtaining a task identifier comprising:

determining whether a type of the task is Structured Query Language (SQL);

performing hash processing on the task to obtain a hash digest if the task type is SQL, and using the hash digest as the task identifier; and

using a number of the task as the task identifier if the task type is not SQL.

7. The method of claim 3 , the selection criteria comprising one or more of:

an amount of output data being smaller than a second threshold; and

occupied cluster resources being less than a preset quota, the occupied cluster resources comprising at least one of operating overhead, operating frequency, and cluster load.

8. The method of claim 1 , the network resources comprising at least one of a network bandwidth and a network bandwidth-delay product.

9. The method of claim 1 , further comprising determining, by the task management device, that the task is a cross-cluster task by determining, based on the history record, if the input or output data of the task is located in a cluster different from a current cluster of the task.

10. An apparatus comprising:

a processor; and

a storage medium for tangibly storing thereon program logic for execution by the processor, the stored program logic comprising:

logic, executed by the processor, for obtaining network resources needed to perform cross-cluster reading and writing, the network resources associated with a task and comprising an amount of input data and an amount of output data for the task based on a history record of past executions of the task;

logic, executed by the processor, for analyzing the network resources calculating an input-output ratio representing a proportion of the network resources needed for reading and writing for the task, the input-output ratio equal to a ratio of the amount of input data to the amount of output data; and

logic, executed by the processor, for scheduling the task according to the network resources.

11. The apparatus of claim 10 , the logic for scheduling the task comprising logic, executed by the processor, for scheduling the task to a target cluster where read dependent data is located if the network resources needed for reading are more than network resources needed for writing.

12. The apparatus of claim 10 , the logic for scheduling the task further comprising:

logic, executed by the processor, for determining whether the task meets a preset selection criteria, the preset selection criteria comprising that the input- output ratio is greater than a preset first threshold, the preset first threshold greater than one; and

logic, executed by the processor, for scheduling the task to a target cluster where the dependent data read by the task locates.

13. The apparatus of claim 12 , the stored program logic further comprising:

logic, executed by the processor, for obtaining a task identifier for a task meeting the selection criteria; and

logic, executed by the processor, for generating scheduling information that includes the task identifier.

14. The apparatus of claim 13 , the logic for scheduling the task to a target cluster where the dependent data read by the task locates comprising:

logic, executed by the processor, for obtaining the task identifier for the task when the task is received; and

logic, executed by the processor, for scheduling the task to a target cluster where dependent data for the task locates if the task identifier of the task matches the task identifier in the scheduling information.

15. The apparatus of claim 13 , the logic for obtaining a task identifier comprising:

logic, executed by the processor, for determining whether a type of the task is Structured Query Language (SQL);

logic, executed by the processor, for performing hash processing on the task to obtain a hash digest if the task type is SQL, and using the hash digest as the task identifier; and

logic, executed by the processor, for using a number of the task as the task identifier if the task type is not SQL.

16. The apparatus of claim 12 , the selection criteria comprising one or more of:

an amount of output data being smaller than a second threshold; and

occupied cluster resources being less than a preset quota, the occupied cluster resources comprising at least one of operating overhead, operating frequency, and cluster load.

17. The apparatus of claim 10 , the network resources comprising at least one of a network bandwidth and a network bandwidth-delay product.

18. The apparatus of claim 10 , the stored program logic further comprising logic, executed by the processor, for determining that the task is a cross-cluster task by determining, based on the history record, if the input or output data of the task is located in a cluster different from a current cluster of the task.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2026
From: ALIBABA GROUP HOLDING LIMITED
To: CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PRIVATE LIMITED
Reel/Frame 075478/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2019
From: HE, LE; HUANG, YAN; SHI, YINGJIE; ZHANG, JIE; ZHANG, CHEN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 048182/0448 →
Priority Claims (1)
CN 201610179807.5 · Mar 25, 2016 · national
Continuity (1)
Related Publication 20190034228A1 · Jan 31, 2019
Cited By (1)
US 12,268,872