IP Library Granted Patent US 10,120,904
Granted Patent B2
US 10,120,904 · App. 14/588,250 · Granted Nov 6, 2018

Resource management in a distributed computing environment

Inventor: Jairam Ranganathan (Palo Alto, CA)
Assignee: Cloudera, Inc.
G06F17/3048G06F9/5066G06F9/5083G06F17/30194G06F2209/505
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,120,904
App. No.
14/588,250
Granted
Nov 6, 2018
Kind
B2
Abstract

Systems and methods are disclosed for resource management in a distributed computing environment. In some embodiments, a resource manager for a large distributed cluster needs to be able to provide resource responses very quickly. But each query may also not be accurate in initial resource request and will often have to come back to the resource manager multiple times. An artifact may provide low latency query responses by using resource request caching that can handle re-requests of resources. According to some embodiments, a queuing mechanism may take into account resources currently expended and any resource requirement estimates available in order to make queuing decisions that meet policies set by an administrator. In some embodiments, scheduling decisions are distribute across a cluster of computing systems while still maintaining approximate compliance with resource management policies set by an administrator.

Claims (21)

1. A computer-implemented method for managing computing resources in a distributed computing cluster operable to execute queries on data stored in the cluster, the computing cluster having multiple data nodes storing the data, the method comprising:

instantiating the plurality of data nodes in the cluster, a respective data node operating an instance of a query engine that is capable of planning a query and executing at least a query portion that corresponds to data stored on the respective data node;

receiving, by a master node in the distributed computing cluster, an administrator usage policy, wherein the administrator usage policy includes limitations on usage of computing resources by a user and/or group of users;

receiving, by the master node, a query and an initial request for computing resources to process the query; and

determining, by the master node, if sufficient computing resources are available to process the query based on the initial request for computing resources and a limitation on usage of computing resources applicable to the query and resource request based on the administrator usage policy, wherein,

if sufficient resources are not available:

the master node is configured to place the query into a queue until sufficient resources are available.

2. The method of claim 1 , wherein computing resources include one or more of, memory, processing power, data storage, and network bandwidth.

3. The method of claim 1 , wherein the distributed computing cluster includes a distributed file system or a data store.

4. The method of claim 3 , wherein the distributed computing cluster is an Apache Hadoop cluster, the distributed file system is a Hadoop Distributed File System (HDFS) and the data store is a NoSQL (No Structured Query Language) data store, wherein the NoSQL data store includes Apache HBase.

5. The method of claim 1 , wherein,

if sufficient resources are available, the master node is configured to:

reserve a cache of computing resources at one or more slave nodes in the distributed computing cluster based on the initial request for resources and a heuristic associated with a temporal clustering of received queries;

allocate a first set computing resources from the reserved cache of computing resources based on the initial request for computing resources to process the query;

receive, during processing of the query, an updated computing resource request; and

if additional computing resources are required:

allocate an additional second set of computing resources from the reserved cache based on the updated resource request.

6. The method of claim 5 , wherein,

if fewer resources are required:

the master node is configured to remove a third set of computing resources from the reserved cache based on the updated resource request.

7. The method of claim 5 , wherein the heuristic includes a prediction of a plurality of queries to follow the received query.

Assignments (5)
RELEASE OF SECURITY INTERESTS IN PATENTS Recorded Oct 14, 2021
From: CITIBANK, N.A.
To: CLOUDERA, INC.; HORTONWORKS, INC.
Reel/Frame 057804/0355 →
FIRST LIEN NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Oct 12, 2021
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 057776/0185 →
SECOND LIEN NOTICE AND CONFIRMATION OF GRANT OF SECURITY INTEREST IN PATENTS Recorded Oct 12, 2021
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 057776/0284 →
SECURITY INTEREST Recorded Dec 22, 2020
From: CLOUDERA, INC.; HORTONWORKS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 054832/0559 →
EMPLOYMENT AGREEMENT Recorded Jun 7, 2018
From: RANGANATHAN, JAIRAM
To: CLOUDERA, INC.
Reel/Frame 047062/0212 →
Continuity (1)
Related Publication 20160188594A1 · Jun 30, 2016
Cited By (2)
US 12,360,961 US 12,393,608