IP Library Granted Patent US 9,485,197
Granted Patent B2
US 9,485,197 · App. 14/156,169 · Granted Nov 1, 2016

Task scheduling using virtual clusters

Inventors: Debojyoti Dutta (Santa Clara, CA); Madhav Marathe (Cupertino, CA); Senhua Huang (Fremont, CA); Raghunath Nambiar (San Ramon, CA)
Assignee: Cisco Technology, Inc.
H04L49/3045H04L41/0896
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,485,197
App. No.
14/156,169
Granted
Nov 1, 2016
Kind
B2
Abstract

In one embodiment, a device receives information regarding a data set to be processed by a map-reduce process. The device generates a set of virtual clusters for the map-reduce process based on network bandwidths between nodes of the virtual clusters, each node of the virtual cluster corresponding to a resource device, and associates the data set with a map-reduce process task. The device then schedules the execution of the task by a node of the virtual clusters based on the network bandwidth between the node and a source node on which the data set resides.

Claims (46)

1. A method comprising:

receiving, at a device, information regarding a data set to be processed by a map-reduce process, wherein the map-reduce process comprises a rack-aware scheduler;

generating a set of virtual clusters for the map-reduce process based on network bandwidths between nodes of the virtual clusters, each node of a virtual cluster corresponding to a resource device, wherein the set of virtual clusters are generated such that intra-cluster bandwidths in the set are greater than inter-cluster bandwidths in the set;

associating the data set with a map-reduce process task; and

scheduling, by the rack-aware scheduler, the execution of the task by a node of the virtual clusters based on the network bandwidth between the node and a source node on which the data set resides, wherein the virtual clusters are used by the rack-aware scheduler in lieu of a physical rack to make scheduling decisions.

2. The method as in claim 1 , wherein the bandwidth between two nodes of the virtual clusters corresponds to a maximum possible bandwidth between the two nodes.

3. The method as in claim 1 , wherein the bandwidth between two nodes of the virtual clusters corresponds to an available bandwidth calculated by subtracting the bandwidth used by other processes on the two nodes from the maximum possible bandwidth between the two nodes.

4. The method as in claim 1 , further comprising:

receiving a network topology change notification;

determining bandwidths between resource devices in the changed network topology; and

generating a new set of virtual clusters based on the bandwidths in the changed network topology.

5. The method as in claim 1 , further comprising:

compressing data transferred between nodes of different virtual clusters.

6. The method as in claim 1 , further comprising:

receiving network data that comprises a network topology;

determining bandwidths between computing resource nodes in the network topology;

forming a weighted graph using the computing resource nodes as vertices of the graph, wherein edges in the graph between vertices represent network connections between the resource nodes that are weighted using the bandwidth between the resource nodes;

assigning the graph vertices to the virtual clusters.

7. The method as in claim 1 , wherein the virtual clusters are treated as virtual mainframe racks by the rack-aware scheduler.

8. An apparatus comprising:

one or more network interfaces configured to communicate in a computer network;

a processor configured to execute one or more processes; and

a memory configured to store a process executable by the processor, the process when executed operable to:

receive information regarding a data set to be processed by a map-reduce process, wherein the map-reduce process comprises a rack-aware scheduler;

generate a set of virtual clusters for the map-reduce process based on network bandwidths between nodes of the virtual clusters, each node of a virtual cluster corresponding to a resource device, wherein the set of virtual clusters are generated such that intra-cluster bandwidths in the set are greater than inter-cluster bandwidths in the set;

associating the data set with a map-reduce process task; and

scheduling, by the rack-aware scheduler, the execution of the task by a node of the virtual clusters based on the network bandwidth between the node and a source node on which the data set resides, wherein the virtual clusters are used by the rack-aware scheduler in lieu of a physical rack to make scheduling decisions.

9. The apparatus as in claim 8 , wherein the bandwidth between two nodes of the virtual clusters corresponds to a maximum possible bandwidth between the two nodes.

10. The apparatus as in claim 8 , wherein the bandwidth between two nodes of the virtual clusters corresponds to an available bandwidth calculated by subtracting the bandwidth used by other processes on the two nodes from the maximum possible bandwidth between the two nodes.

11. The apparatus as in claim 8 , wherein the process is operable to:

receive a network topology change notification;

determine bandwidths between resource devices in the changed network topology; and

generate a new set of virtual clusters based on the bandwidths in the changed network topology.

12. The apparatus as in claim 8 , wherein the process is operable to:

compress data transferred between nodes of different virtual clusters.

13. The apparatus as in claim 8 , wherein the task comprises a reducer task.

14. The apparatus as in claim 8 , wherein the virtual clusters are treated as virtual mainframe racks by the rack-aware scheduler.

15. A tangible, non-transitory, computer-readable media having software encoded thereon, the software, when executed by a processor, operable to:

receive information regarding a data set to be processed by a map-reduce process, wherein the map-reduce process comprises a rack-aware scheduler;

generate a set of virtual clusters for the map-reduce process based on network bandwidths between nodes of the virtual clusters, each node of a virtual cluster corresponding to a resource device, wherein the set of virtual clusters are generated such that intra-cluster bandwidths in the set are greater than inter-cluster bandwidths in the set;

associate the data set with a map-reduce process task; and

schedule, via the rack-aware scheduler, the execution of the task by a node of the virtual clusters based on the network bandwidth between the node and a source node on which the data set resides, wherein the virtual clusters are used by the rack-aware scheduler in lieu of a physical rack to make scheduling decisions.

16. The computer-readable media as in claim 15 , wherein the software, when executed by the processor, is operable to:

receive a network topology change notification;

determine bandwidths between resource devices in the changed network topology; and

generate a new set of virtual clusters based on the bandwidths in the changed network topology.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2014
From: DUTTA, DEBOJYOTI; MARATHE, MADHAV; HUANG, SENHUA; NAMBIAR, RAGHUNATH
To: CISCO TECHNOLOGY, INC.
Reel/Frame 032288/0844 →
Continuity (1)
Related Publication 20150200867A1 · Jul 16, 2015