IP Library Granted Patent US 11,481,679
Granted Patent B2
US 11,481,679 · App. 16/805,919 · Granted Oct 25, 2022

Adaptive data ingestion rates

Inventor: Tinniam Venkataraman Ganesh (Bangalore, IN)
Assignee: KYNDRYL, INC.
G06N20/00G06F8/60H04L69/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,679
App. No.
16/805,919
Granted
Oct 25, 2022
Kind
B2
Abstract

Described are techniques for data ingestion including determining a respective moving average streaming rate for each of a plurality of incoming data streams to a cluster-computing framework. The techniques further include determining a respective ingestion frequency for each of the plurality of incoming data streams by dividing a platform-preferred ingestion rate of the cluster-computing framework by a respective moving average streaming rate of each of the plurality of incoming data streams. The techniques further include ingesting each of the plurality of incoming data streams to the cluster-computing framework at the respective ingestion frequency.

Claims (40)

1. A computer-implemented method comprising:

determining a respective moving average streaming rate for each of a plurality of incoming data streams to a cluster-computing framework, wherein the cluster-computing framework includes a plurality of worker nodes;

determining a respective ingestion frequency for each of the plurality of incoming data streams by dividing a platform-preferred ingestion rate of the cluster-computing framework by a respective moving average streaming rate of each of the plurality of incoming data streams; and

adjusting ingestion frequencies of the plurality of worker nodes to the determined ingestion frequencies.

2. The method of claim 1 , wherein the cluster-computing framework is configured to transform the plurality of incoming data streams into a first format.

3. The method of claim 1 , comprising:

outputting, in response to the worker nodes ingesting the incoming data streams at the adjusted ingestion frequencies, transformed data corresponding to the plurality of incoming data streams, wherein the transformed data is in a first format.

4. The method of claim 3 , comprising:

performing machine learning on the transformed data; and

outputting a result based on performing the machine learning on the transformed data.

5. The method of claim 3 , comprising: training a machine learning model using the transformed data.

6. The method of claim 1 , wherein the platform-preferred ingestion rate of the cluster-computing framework is based on characteristics of the worker nodes.

7. The method of claim 6 , wherein the characteristics include an amount of read-access memory (RAM) on the worker nodes.

8. The method of claim 1 , wherein the moving average streaming rates of the incoming data streams are weighted moving averages.

9. The method of claim 1 , wherein the moving average streaming rates of the incoming data streams are exponentially weighted moving averages.

10. The method of claim 1 , wherein the method is performed according to software that is downloaded to the cluster-computing framework from a remote data processing system, and comprising:

metering a usage of the software; and

generating an invoice based on metering the usage.

11. A system comprising:

a processor; and

logic integrated with the processor, executable by the processor, or integrated with and executable by the processor, the logic being configured to:

determine a respective moving average streaming rate for each of a plurality of incoming data streams to a cluster-computing framework,

wherein the cluster-computing framework includes a plurality of worker nodes;

determine a respective ingestion frequency for each of the plurality of incoming data streams by dividing a platform-preferred ingestion rate of the cluster-computing framework by a moving average streaming rate of each of the plurality of incoming data streams; and

adjust ingestion frequencies of the plurality of worker nodes to the determined ingestion frequencies.

12. The system of claim 11 , the logic being configured to:

output, in response to the worker nodes ingesting the incoming data streams at the adjusted ingestion frequencies, transformed data corresponding to the plurality of incoming data streams, wherein the transformed data is in a first format.

13. The system of claim 11 , wherein the platform-preferred ingestion rate of the cluster-computing framework is based on a number of worker nodes associated with the cluster-computing framework, a number of processing cores associated with the cluster-computing framework, and an amount of read-access memory (RAM) associated with the cluster-computing framework.

14. The system of claim 11 , wherein the moving average streaming rates of the incoming data streams are weighted moving averages.

15. The system of claim 11 , wherein the moving average streaming rates of the incoming data streams are exponentially weighted moving averages.

16. A computer program product, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable and/or executable by a processor to cause the processor to:

determine, by the processor, a respective moving average streaming rate for each of a plurality of incoming data streams to a cluster-computing framework,

wherein the cluster-computing framework includes a plurality of worker nodes;

determine, by the processor, a respective ingestion frequency incoming data streams by dividing a platform-preferred ingestion rate of the cluster-computing framework by a respective moving average streaming rate of each of the plurality of incoming data streams; and

adjust, by the processor, ingestion frequencies of the plurality of worker nodes to the determined ingestion frequencies.

17. The computer program product of claim 16 , the program instructions readable and/or executable by the processor to cause the processor to:

output, by the processor, in response to the worker nodes ingesting the incoming data streams at the adjusted ingestion frequencies, transformed data corresponding to the plurality of incoming data streams, wherein the transformed data is in a first format.

18. The computer program product of claim 16 , wherein the platform-preferred ingestion rate of the cluster-computing framework is based on a number of worker nodes associated with the cluster-computing framework, a number of processing cores associated with the cluster-computing framework, and an amount of read-access memory (RAM) associated with the cluster-computing framework.

19. The computer program product of claim 16 , wherein the moving average streaming rates of the incoming data streams are weighted moving averages.

20. The computer program product of claim 16 , wherein the moving average streaming rates of the incoming data streams are exponentially weighted moving averages.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: KYNDRYL, INC.
Reel/Frame 058213/0912 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2020
From: VENKATARAMAN GANESH, TINNIAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051974/0469 →
Cited By (2)
US 12,260,659 US 12,438,823