IP Library Granted Patent US 11,423,059
Granted Patent B2
US 11,423,059 · App. 16/525,987 · Granted Aug 23, 2022

System and method for restrictive clustering of datapoints

Inventor: Shubhojit Mallick (New Delhi, IN)
Assignee: Innoplexus AG
G06F16/287G06F16/2264G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,059
App. No.
16/525,987
Granted
Aug 23, 2022
Kind
B2
Abstract

Disclosed is a system for restrictive clustering of datapoints. The system comprises server arrangement that acquires data record for performing clustering operation, determines datapoints for the data record, plots the datapoints in a multi-dimensional space, determines a cluster threshold, and performs a first iteration of clustering on the datapoints plotted in the multi-dimensional space, determines a segment threshold for the datapoints plotted in the multi-dimensional space, derives boundary conditions for determining segments based on the segment threshold and superimposes the boundary conditions corresponding to each of the segments based on the segment threshold onto the first iteration of clustering. Moreover, the server arrangement re-iterates the first iteration of clustering to obtain a second iteration of clustering, wherein the second iteration of clustering has an error value lower than an error value for the first iteration of clustering.

Claims (32)

1. A system for restrictive clustering of datapoints, wherein the system comprises:

a database arrangement; and

a server arrangement communicably coupled via a data communication network to the database arrangement, wherein the server arrangement comprises:

an extraction module, that when operated, acquires at least one data record for performing clustering operation thereupon, from the database arrangement;

a mapping module, that when operated, determines a set of datapoints for the at least one data record, and plots the set of datapoints in a multi-dimensional space, wherein the multi-dimensional space is plotted to depict a variation of at least one dependent variable against a variation of an independent variable, and wherein the set of datapoints provides values relating to the variation of the at least one dependent variable against the variation of the independent variable;

a clustering module employing a machine learning algorithm which is implemented as an unsupervised learning algorithm, that when operated:

determines a cluster threshold based on iterative operation of the clustering module; and

performs a first iteration of clustering on the set of datapoints plotted in the multi-dimensional space;

a regression module employing a machine learning algorithm which is implemented as a supervised learning algorithm, that when operated:

determines a segment threshold for the set of datapoints plotted in the multi-dimensional space, based on iterative operation of the regression module; and

derives a set of boundary conditions for determining segments based on the segment threshold; and

a restrictive clustering module, that when operated:

superimposes the derived set of boundary conditions corresponding to each of the segments based on the segment threshold onto the first iteration of clustering, on the set of datapoints; and

re-iterates the first iteration of clustering by utilizing the superimposed set of boundary conditions to obtain a second iteration of clustering, wherein the second iteration of clustering has an error value lower than an error value for the first iteration of clustering.

2. The system of claim 1 , wherein machine learning algorithms are employed by at least one of the: extraction module, mapping module, clustering module, regression module, restrictive clustering module.

3. The system of claim 2 , wherein the clustering module, employing the machine learning algorithms, employs k-means algorithm for operation thereof.

4. The system of claim 2 , wherein the regression module, employing the machine learning algorithms, employs splice regression algorithm for operation thereof.

5. The system of claim 1 , wherein the restrictive clustering module further re-iterates the second iteration of clustering based on an input provided by a user.

6. The system of claim 5 , wherein the input provided by the user is based on at least one of: the optimal number for cluster thresholds, the optimal number for segment thresholds, the error value.

7. A method for supervised clustering of datapoints, the method is implemented using a system comprising a server arrangement communicably coupled via a data communication network to a database arrangement, wherein the method comprises:

acquiring at least one data record for performing clustering operation thereupon from the database arrangement;

determining a set of datapoints for the at least one data record;

plotting the set of datapoints in a multi-dimensional space, wherein the multi-dimensional space is plotted to depict a variation of at least one dependent variable against a variation of an independent variable, and wherein the set of datapoints provides values relating to the variation of the at least one dependent variable against the variation of the independent variable;

determining, a cluster threshold;

performing a first iteration of clustering on the set of datapoints plotted in the multi-dimensional space by employing a machine learning algorithm which is implemented as an unsupervised learning algorithm;

determining a segment threshold for the set of datapoints plotted in the multi-dimensional space;

deriving a set of boundary conditions for determining segments based on the segment threshold by employing a machine learning algorithm which is implemented as a supervised learning algorithm;

superimposing the derived set of boundary conditions corresponding to each of the segments based on the segment threshold onto the first iteration of clustering on the set of datapoints; and

re-iterating the first iteration of clustering by utilizing the superimposed set of boundary conditions to obtain a second iteration of clustering, wherein the second iteration of clustering has an error value lower than an error value for the first iteration of clustering.

8. The method of claim 7 , wherein the method further comprises re-iterating the second iteration of clustering based on an input provided by a user.

9. The method of claim 8 , wherein the input provided by the user is based on at least one of: the optimal number for cluster thresholds, the optimal number for segment thresholds, the error value.

10. A computer program product comprising non-transitory computer-readable storage media having computer-readable instructions stored thereon, the computer-readable instructions being executable by a computerized device comprising processing hardware to execute a method of claim 7 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2020
From: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
To: INNOPLEXUS AG
Reel/Frame 052881/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2019
From: MALLICK, SHUBHOJIT
To: INNOPLEXUS CONSULTING SERVICES PVT. LTD.
Reel/Frame 049902/0038 →
Continuity (1)
Related Publication 20210034648A1 · Feb 4, 2021