IP Library › Granted Patent US 12,181,999
Granted Patent B1
US 12,181,999 · App. 17/589,534 · Granted Dec 31, 2024

Machine learning modeling of candidate models for software installation usage

Inventors: Yanpei Chen (Sunnyvale, CA); Archana Ganapathi (Palo Alto, CA)
Assignee: Cisco Technology, Inc.
G06F11/3438G06N20/00G06Q30/0201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,181,999
App. No.
17/589,534
Filed
Jan 31, 2022
Granted
Dec 31, 2024
Kind
B1
Art Unit
2442
USPC
709/224
Abstract

This document discloses methods and systems for modeling product usage. In one practical application, the systems and methods may be utilized to model product usage based on large volume, machine generated product usage data to optimize product pricing and operations. Specifically, the systems and methods described herein may utilize methods with key components to select the maximum number of dimensions that can be modeled based on the number of data points, use a logarithm kernel function to normalize machine data with long-tailed statistical distributions on different numerical scales, compare a large number of candidate models with different candidate dimensions and different structures, and quantify the amount of change and drift in models over time.

Claims (88)

1. A computing device, comprising:

one or more hardware processors; and

a non-transitory computer-readable medium having stored thereon instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations including:

generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;

determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;

selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;

for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;

evaluating each candidate model and generating respective model testing results for each candidate model;

selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models.

2. The computing device of claim 1 , wherein the operations further comprise:

receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and

based on the additional data point and the generated model, responding to the request with the predicted value.

3. The computing device of claim 1 , wherein the operations further comprise:

determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and

comparing B to the predetermined threshold.

4. The computing device of claim 3 , wherein the determining of B uses

B

=

1

D

⁢

log

⁢

2

⁢

(

N

)

.

5. The computing device of claim 3 , wherein the operations further comprise:

based on a result of the comparing, rejecting the candidate number of dimensions.

6. The computing device of claim 1 , wherein the generating of the candidate model for each dimension set comprises generating a polynomial model.

7. The computing device of claim 1 , wherein the operations further comprise:

generating the N data points by linking customer relational information from a customer relationship management (CRM) system and the product usage data using shared identifiers.

8. The computing device of claim 7 , wherein the operations further comprise:

accessing CRM data that indicates a parent-subsidiary relationship between a parent account and a subsidiary account; and

accessing first product usage data that is linked to both the parent account and the subsidiary account; wherein

the generating of the N data points comprises generating a data point that links the first product usage data to the subsidiary account.

9. The computing device of claim 8 , wherein the generating of the N data points excludes generating a second data point that links the first product usage data to the parent account.

10. A computer-implemented method, comprising:

generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;

determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;

selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;

for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;

evaluating each candidate model and generating respective model testing results for each candidate model;

selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models, wherein the selected model models product usage.

11. The computer-implemented method of claim 10 , further comprising:

receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and

based on the additional data point and the generated model, responding to the request with the predicted value.

12. The computer-implemented method of claim 10 , further comprising:

determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and

comparing B to the predetermined threshold.

13. The computer-implemented method of claim 12 , wherein the determining of B uses

B

=

1

D

⁢

log

⁢

2

⁢

(

N

)

.

14. The computer-implemented method of claim 12 , further comprising:

based on a result of the comparing, rejecting the candidate number of dimensions.

15. The computer-implemented method of claim 10 , wherein the generating of the candidate model for each dimension set comprises generating a polynomial model.

16. The computer-implemented method of claim 10 , further comprising:

generating the N data points by linking customer relational information from a customer relationship management (CRM) system and the product usage data using shared identifiers.

17. The computer-implemented method of claim 16 , further comprising:

accessing CRM data that indicates a parent-subsidiary relationship between a parent account and a subsidiary account; and

accessing first product usage data that is linked to both the parent account and the subsidiary account; wherein

the generating of the N data points comprises generating a data point that links the first product usage data to the subsidiary account.

18. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:

generating a first number (N) of normalized data points from N data points, each data point of the N data points having data for a second number (C) of dimensions, by taking a logarithm of the data for each dimension of the C dimensions for each data point of the N data points, and wherein the N data points relate to product usage data;

determining, based on a predetermined threshold and the N normalized data points, a third number (D) of dimensions to use for modeling the N normalized data points;

selecting a plurality of dimension sets, wherein each dimension set includes a different combination of a plurality of dimensions and includes no greater than D dimensions;

for each dimension set in the plurality of dimension sets, generating, via a machine learning model, a candidate model based on the different combination of the plurality of dimensions in the dimension set, wherein each candidate model models product usage;

evaluating each candidate model and generating respective model testing results for each candidate model;

selecting a model from the candidate models based on a plurality of model quality measures from model testing results associated with the candidate models, wherein the selected model models product usage.

19. The non-transitory computer-readable medium of claim 18 , wherein the operations further comprise:

receiving a request to identify a predicted value for an additional data point using the selected model, the additional data point including data on dimensions corresponding to the dimension set corresponding to the selected model; and

based on the additional data point and the generated model, responding to the request with the predicted value.

20. The non-transitory computer-readable medium of claim 18 , wherein the operations further comprise:

determining, for a candidate number of dimensions representing the D dimensions, a number of bits of information (B) per dimension based on the candidate number of dimensions and the first number; and

comparing B to the predetermined threshold.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: CHEN, YANPEI; GANAPATHI, ARCHANA
To: SPLUNK INC.
Reel/Frame 059297/0045 →
Continuity (1)
Provisional Application 63143545 · Jan 29, 2021
Cited By (1)
US 12,730,725