IP Library Granted Patent US 12,443,437
Granted Patent B2
US 12,443,437 · App. 17/709,519 · Granted Oct 14, 2025

Informed optimization of thread group scheduling

Inventors: Jaeyoo Jung (Shrewsbury, MA); Peter Linden (Boston, MA); Robert Lucey (Whitechurch, IE); Wednesday Wolf (Brookline, MA)
Assignee: Dell Products L.P.
G06F9/4881G06F9/5038
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,437
App. No.
17/709,519
Granted
Oct 14, 2025
Kind
B2
Abstract

Individual processors of a storage system are analyzed to determine which thread types are most important for servicing a run workload, where importance is measured by number of CPU cycles used. Permutations of differentiated access to CPU cycles are calculated, where the most important thread types are provided with greater access to CPU cycles than thread types of lesser importance. The permutations are tested with the same run workload to determine which permutation yields the greatest average IOPS. The thread scheduler for the processor is configured with that permutation.

Claims (24)

1. A method comprising:

in a computer-based data storage system comprising processors that run a plurality of workload-supporting thread types comprising read-supporting threads and write-supporting threads for accessing disk drives to service processor-specific run workloads characterized by different ratios of reads to writes,

identifying, for each processor based on characteristics of a current run workload of that processor, a permutation of a plurality of predefined permutations of thread type differentiated access to central processing unit (CPU) cycles for the read-supporting threads and write-supporting threads that yields greatest average input-output operations per second (IOPS) for that processor for servicing the current run workload, each permutation comprising a record of quantified differentiated CPU cycle allocations to each of a plurality of thread type groups, including:

grouping thread types based on tasks performed by those thread types;

counting CPU cycles used by different types of threads running on the storage system processor to service the run workload and identifying one or more most important thread type groups in terms of CPU cycles used; and

characterizing T thread type groups that used a greatest number of CPU cycles as the most important thread type groups, where T is a predetermined integer value; and

configuring a thread scheduler for each processor with the respective identified permutation of differentiated access to CPU cycles, thereby improving performance of the data storage system in terms of IOPS based on current run workload characteristics.

2. The method of claim 1 further comprising assigning greater CPU cycle access to threads of the most important thread type groups than to threads of other thread types.

3. The method of claim 2 further comprising calculating permutations of differentiated CPU cycle access to the threads of the most important thread type groups.

4. The method of claim 3 further comprising measuring average IOPS yielded by each permutation responsive to servicing the run workload to identifying the permutation that yields the greatest average IOPS.

5. A non-transitory computer-readable storage medium that stores instructions that when executed by a computer perform a method comprising:

in a computer-based data storage system comprising processors that run a plurality of workload-supporting thread types comprising read-supporting threads and write-supporting threads for accessing disk drives to service processor-specific run workloads characterized by different ratios of reads to writes,

identifying, for each processor based on characteristics of a current run workload of that processor, a permutation of a plurality of predefined permutations of thread type differentiated access to central processing unit (CPU) cycles for the read-supporting threads and write-supporting threads that yields greatest average input-output operations per second (IOPS) for that processor for servicing the current run workload, each permutation comprising a record of quantified differentiated CPU cycle allocations to each of a plurality of thread type groups, including:

grouping thread types based on tasks performed by those thread types;

counting CPU cycles used by different types of threads running on the storage system processor to service the run workload and identifying one or more most important thread type groups in terms of CPU cycles used; and

characterizing T thread type groups that used a greatest number of CPU cycles as the most important thread type groups, where T is a predetermined integer value; and

configuring a thread scheduler for each processor with the respective identified permutation of differentiated access to CPU cycles, thereby improving performance of the data storage system in terms of IOPS based on current run workload characteristics.

6. The non-transitory computer-readable storage medium of claim 5 in which the method further comprises assigning greater CPU cycle access to threads of the most important thread type groups than to threads of other thread types.

7. The non-transitory computer-readable storage medium of claim 6 in which the method further comprises calculating permutations of differentiated CPU cycle access to the threads of the most important thread type groups.

8. The non-transitory computer-readable storage medium of claim 7 in which the method further comprises measuring average IOPS yielded by each permutation responsive to servicing the run workload to identifying the permutation that yields the greatest average IOPS.

9. An apparatus comprising:

a computer-based storage system comprising processors that run a plurality of workload-supporting thread types comprising read-supporting threads and write-supporting threads for accessing disk drives to service processor-specific run workloads characterized by different ratios of reads to writes, and a thread group scheduling optimizer configured to identify, for each processor based on characteristics of a current run workload of that processor, a permutation of a plurality of predefined permutations of thread type differentiated access to central processing unit (CPU) cycles for the read-supporting threads and write-supporting threads that yields greatest average input-output operations per second (IOPS) for that processor for servicing the current run workload, each permutation comprising a record of quantified differentiated CPU cycle allocations to each of a plurality of thread type groups, including the thread group scheduling optimizer being configured to group thread types based on tasks performed by those thread types, count CPU cycles used by different types of threads of the storage system processor to service the run workload and identify most important thread type groups in terms of CPU cycles used, and characterize T thread type groups that used a greatest number of CPU cycles as the most important thread type groups, where T is a predetermined integer value, and configure a thread scheduler for each processor with the respective identified permutation of differentiated access to CPU cycles, thereby improving performance of the data storage system in terms of IOPS based on current run workload characteristics.

10. The apparatus of claim 9 in which the thread group scheduling optimizer is further configured to assign greater CPU cycle access to threads of the most important thread type groups than to threads of other thread types.

11. The apparatus of claim 10 in which the thread group scheduling optimizer is further configured to calculate permutations of differentiated CPU cycle access to the threads of the most important thread type groups and measure average IOPS yielded by each permutation for servicing the run workload to identify the permutation that yields the greatest average IOPS.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2022
From: JUNG, JAEYOO; LINDEN, PETER; LUCEY, ROBERT; WOLF, WEDNESDAY
To: DELL PRODUCTS L.P.
Reel/Frame 059453/0488 →
Continuity (1)
Related Publication 20230315518A1 · Oct 5, 2023
References Cited (24)
US 4510565A · Dummermuth · 1985 [cited by examiner]
US 9654442B2 · Tung · 2017 [cited by examiner]
US 10880204B1 · Shalev · 2020 [cited by examiner]
US 10963189B1 · Neelakantam · 2021 [cited by examiner]
US 20030204552A1 · Zuberi · 2003 [cited by examiner]
US 20060212840A1 · Kumamoto · 2006 [cited by examiner]
US 20110283286A1 · Wu · 2011 [cited by examiner]
US 20120284718A1 · Dahlstedt · 2012 [cited by examiner]
US 20160077948A1 · Goel · 2016 [cited by examiner]
US 20160092108A1 · Karaje · 2016 [cited by examiner]
US 20160132214A1 · Koushik · 2016 [cited by examiner]
US 20170109251A1 · Das · 2017 [cited by examiner]
US 20190004710A1 · Ebsen · 2019 [cited by examiner]
US 20210034419A1 · Krasner · 2021 [cited by examiner]
US 20210064430A1 · Srivastava · 2021 [cited by examiner]
US 20210311852A1 · Doddaiah · 2021 [cited by examiner]
US 20220100573A1 · Allen · 2022 [cited by examiner]
US 20230168934A1 · Vaka · 2023 [cited by examiner]
US 20230315518A1 · Jung · 2023 [cited by examiner]
Uiseok Song, Optimizing communication performance in scale-out storage system. (Year: 2019). [cited by examiner]
R. Andersen, Harvesting Idle Windows CPU Cycles for Grid Computing. (Year: 2006). [cited by examiner]
Anastasios Papagiannis, Memory-Mapped I/O on Steroids. (Year: 2021). [cited by examiner]
Junjie Qian, Energy-efficient I/O Thread Schedulers for NVMe SSDs on NUMA. (Year: 2017). [cited by examiner]
Jeffery A. Brown, The Shared-Thread Multiprocessor. (Year: 2008). [cited by examiner]