IP Library Granted Patent US 10,678,671
Granted Patent B2
US 10,678,671 · App. 16/434,157 · Granted Jun 9, 2020

Triggering the increased collection and distribution of monitoring information in a distributed processing system

Inventors: Thomas Rothschilds (Seattle, WA); Remi Bernotavicius (Seattle, WA); Edward Brow (Boulder, CO); William Ehlhardt (Seattle, WA)
Assignee: Qumulo, Inc.
G06F11/3495G06F11/3006G06F11/3065G06F11/3409G06F21/554G06F21/6218G06F11/3419G06F2221/2111G06F2221/2135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,678,671
App. No.
16/434,157
Granted
Jun 9, 2020
Kind
B2
Abstract

A facility comprising systems and method for automatically triggering the collection of comprehensive monitoring information in a distributed processing system. The facility compares the overall performance of distributed processing system to one or more performance metrics and, in response to determining that one or more performance metrics is not satisfied, triggers one or more of the nodes within the distributed processing system to increase one or more of its monitoring rate or its distribution rate. The facility collects and analyzes the collected information to provide resources that can be used to assess and diagnose failures within the distributed processing system. In this manner, the facility reacts to performance anomalies by triggering nodes within in the system to provide comprehensive performance information over a trigger period for diagnostic purposes.

Claims (66)

1. A method for managing data in a file system over a network using one or more processors that execute instructions to perform actions, comprising:

monitoring one or more metrics to collect data that is associated with one or more nodes that are part of the file system, wherein the data associated with the one or more nodes is truncated to include data that corresponds to an overlapping time period and to omit data that corresponds to one or more non-overlapping time periods;

determining the one or more nodes having computer resources that are associated with the one or more metrics that exceed one or more trigger levels;

employing a trigger time period to modify an original monitor rate associated with the one or more determined nodes to be further associated with the trigger time period;

in response to an expiration of the trigger time period, restoring the modified monitor rate to the original monitor rate; and

providing one or more reports that improve identification of each node having computer resources that exceed the one or more trigger levels during the trigger period.

2. The method of claim 1 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to increase a rate of collection of the one or more metrics for the one or more nodes; and

decreasing the rate of collection of the one or more metrics for the one or more nodes from after expiration of the trigger time period.

3. The method of claim 1 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to provide separate increases in a rate of collection for each metric of the one or more nodes; and

decreasing the separate rate of collection for each metric of the one or more nodes after expiration of the trigger time period.

4. The method of claim 1 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises: employing initiation of the trigger time period to provide current values for a lock graph, performance stack information, stack traces, or performance counters.

5. The method of claim 1 , further comprising: selecting a duration of the trigger time period based on a longest time period that is associated with the one or more metrics that exceed the one or more trigger levels.

6. The method of claim 1 , further comprising: invoking a trigger component of an identified node by sending a remote procedure call to the identified node.

7. A system for managing data in a file system over a network, comprising:

a network computer, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that perform actions, including:

monitoring one or more metrics to collect data that is associated with one or more nodes that are part of the file system, wherein the data associated with the one or more nodes is truncated to include data that corresponds to an overlapping time period and to omit data that corresponds to one or more non-overlapping time periods;

determining the one or more nodes having computer resources that are associated with the one or more metrics that exceed one or more trigger levels;

employing a trigger time period to modify an original monitor rate associated with the one or more determined nodes to be further associated with the trigger time period;

in response to an expiration of the trigger time period, restoring the modified monitor rate to the original monitor rate; and

providing one or more reports that improve identification of each node having computer resources that exceed the one or more trigger levels during the trigger period; and

a client computer, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that perform actions, including:

receiving, the one or more reports.

8. The system of claim 7 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to increase a rate of collection of the one or more metrics for the one or more nodes; and

decreasing the rate of collection of the one or more metrics for the one or more nodes from after expiration of the trigger time period.

9. The system of claim 7 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to provide separate increases in a rate of collection for each metric of the one or more nodes; and

decreasing the separate rate of collection for each metric of the one or more nodes after expiration of the trigger time period.

10. The system of claim 7 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises: employing initiation of the trigger time period to provide current values for a lock graph, performance stack information, stack traces, or performance counters.

11. The system of claim 7 , further comprising: selecting a duration of the trigger time period based on a longest time period that is associated with the one or more metrics that exceed the one or more trigger levels.

12. The system of claim 7 , further comprising: invoking a trigger component of an identified node by sending a remote procedure call to the identified node.

13. A processor readable non-transitory storage media that includes instructions for managing data in a file system over a network, wherein execution of the instructions by one or more processors on one or more network computers performs actions, comprising:

monitoring one or more metrics to collect data that is associated with one or more nodes that are part of the file system, wherein the data associated with the one or more nodes is truncated to include data that corresponds to an overlapping time period and to omit data that corresponds to one or more non-overlapping time periods;

determining the one or more nodes having computer resources that are associated with the one or more metrics that exceed one or more trigger levels;

employing a trigger time period to modify an original monitor rate associated with the one or more determined nodes to be further associated with the trigger time period;

in response to an expiration of the trigger time period, restoring the modified monitor rate to the original monitor rate; and

providing one or more reports that improve identification of each node having computer resources that exceed the one or more trigger levels during the trigger period.

14. The processor readable non-transitory storage media of claim 13 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to increase a rate of collection of the one or more metrics for the one or more nodes; and

decreasing the rate of collection of the one or more metrics for the one or more nodes from after expiration of the trigger time period.

15. The processor readable non-transitory storage media of claim 13 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to provide separate increases in a rate of collection for each metric of the one or more nodes; and

decreasing the separate rate of collection for each metric of the one or more nodes after expiration of the trigger time period.

16. The processor readable non-transitory storage media of claim 13 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to provide current values for a lock graph, performance stack information, stack traces, or performance counters.

17. A network computer for managing data in a file system over a network, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that perform actions, including:

monitoring one or more metrics to collect data that is associated with one or more nodes that are part of the file system, wherein the data associated with the one or more nodes is truncated to include data that corresponds to an overlapping time period and to omit data that corresponds to one or more non-overlapping time periods;

determining the one or more nodes having computer resources that are associated with the one or more metrics that exceed one or more trigger levels;

employing a trigger time period to modify an original monitor rate associated with the one or more determined nodes to be further associated with the trigger time period;

in response to an expiration of the trigger time period, restoring the modified monitor rate to the original monitor rate; and

providing one or more reports that improve identification of each node having computer resources that exceed the one or more trigger levels during the trigger period.

18. The network computer of claim 17 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to increase a rate of collection of the one or more metrics for the one or more nodes; and

decreasing the rate of collection of the one or more metrics for the one or more nodes from after expiration of the trigger time period.

19. The network computer of claim 17 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises:

employing initiation of the trigger time period to provide separate increases in a rate of collection for each metric of the one or more nodes; and

decreasing the separate rate of collection for each metric of the one or more nodes after expiration of the trigger time period.

20. The network computer of claim 17 , wherein monitoring the one or more metrics to collect data for the one or more nodes, further comprises: employing initiation of the trigger time period to provide current values for a lock graph, performance stack information, stack traces, or performance counters.

Assignments (3)
SECURITY INTEREST Recorded Jun 24, 2022
From: QUMULO, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 060439/0967 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2019
From: ROTHSCHILDS, THOMAS; BERNOTAVICIUS, REMI; BROW, EDWARD
To: QUMULO, INC.
Reel/Frame 049399/0521 →
EMPLOYEE INVENTION ASSIGNMENT, CONFIDENTIALITY AND NON-COMPETITION AGREEMENT Recorded Jun 6, 2019
From: EHLHARDT, WILLIAM
To: QUMULO, INC.
Reel/Frame 049408/0980 →
Continuity (3)
Continuation 15957809 · Apr 19, 2018
Provisional Application 62488028 · Apr 20, 2017
Related Publication 20190286543A1 · Sep 19, 2019
Cited By (9)
US 12,222,903 US 12,292,853 US 12,346,290 US 12,443,559 US 12,443,568 US 12,481,625 US 12,585,563 US 12,619,582 US 12,670,081