IP Library › Granted Patent US 12,386,656
Granted Patent B2
US 12,386,656 · App. 17/465,021 · Granted Aug 12, 2025

Dynamic decomposition and thread allocation

Inventors: Skyler Arron Windh (McKinney, TX); Tony M. Brewer (Plano, TX); Patrick Estep (Rowlett, TX)
Assignee: Micron Technology, Inc.
G06F9/4881G06F9/3851G06F9/3888G06F9/5038G06F11/3409G06F2209/501G06F2209/5022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,656
App. No.
17/465,021
Granted
Aug 12, 2025
Kind
B2
Abstract

Devices and techniques for thread scheduling control and memory splitting in a processor are described herein. An apparatus includes a hardware interface configured to receive a first request to execute a first thread, the first request including an indication of a workload; and processing circuitry configured to: determine the workload to produce a metric based at least in part on the indication; compare the metric with a threshold to determine that the metric is beyond the threshold; divide, based at least in part on the comparison, the workload into a set of sub-workloads consisting of predefined number of equal parts from the workload; create a second request to execute a second thread, the second request including a first member of the set of sub-workloads; and process a second member of the set of sub-workloads in the first thread.

Claims (48)

1. An apparatus comprising:

a hardware interface configured to receive a first request to execute a first thread, the first request including an indication of a workload and a busy-fail field that is set; and

processing circuitry configured to:

determine the workload to produce a metric based at least in part on the indication;

compare the metric with a threshold to determine that the metric is beyond the threshold;

divide, based at least in part on the comparison, the workload into a set of sub-workloads consisting of a predefined number of equal parts from the workload;

create a second request to execute a second thread, the second request including a first member of the set of sub-workloads; and

process a second member of the set of sub-workloads in the first thread,

wherein the second request to execute the second thread fails, and wherein, in response to the second request to execute the second thread failing and the busy-fail field being set, the processing circuitry is configured to:

continue processing the second member of the set of sub-workloads; and

create a third request to execute the second thread, the third request including the first member of the set of sub-workloads.

2. The apparatus of claim 1 , wherein the predefined number is two.

3. The apparatus of claim 1 , wherein the first thread is a master thread.

4. The apparatus of claim 1 , wherein the second thread is a fiber thread.

5. The apparatus of claim 1 , wherein the busy-fail field is bit in a chip-to-chip protocol interface (CTCPI) packet.

6. The apparatus of claim 1 , wherein the second request includes a no-return field that is set.

7. The apparatus of claim 6 , wherein the no-return field is bit in a chip-to-chip protocol interface (CTCPI) packet.

8. The apparatus of claim 6 , wherein the no-return field is used to signal that the second thread does not return a value to a stack position.

9. The apparatus of claim 6 , wherein the no-return field releases the first thread from having to wait for the second thread to return.

10. The apparatus of claim 1 , wherein, to process the second member of the set of sub-workloads, the processing circuitry is configured to:

determine the second member to produce a second metric;

compare the second metric with the threshold to determine that the second metric is beyond the threshold;

divide, based at least in part on the comparison, the second member into a further set of sub-workloads consisting of the predefined number of equal parts from the second member;

create a third request to execute a third thread, the second request including a first member of the further set of sub-workloads; and

process a second member of the further set of sub-workloads.

11. The apparatus of claim 1 , wherein to create the second request to execute the second thread fails, and in response, the first thread is to process the second member of the set of sub-workloads in the first thread up to the threshold.

12. The apparatus of claim 11 , wherein the processing circuitry is to repeat the operation to create the second request to execute the second thread after processing the second member of the set of sub-workloads up to the threshold.

13. A method comprising:

receiving a first request to execute a first thread, the first request including an indication of a workload and a busy-fail field that is set;

determining the workload to produce a metric based at least in part on the indication;

comparing the metric with a threshold to determine that the metric is beyond the threshold;

dividing, based at least in part on the comparison, the workload into a set of sub-workloads consisting of predefined number of equal parts from the workload;

creating a second request to execute a second thread, the second request including a first member of the set of sub-workloads; and

processing a second member of the set of sub-workloads in the first thread,

wherein the second request to execute the second thread fails, and wherein, in response to the second request to execute the second thread failing and the busy-fail field being set, the method comprises:

continuing processing the second member of the set of sub-workloads; and

creating a third request to execute the second thread, the third request including the first member of the set of sub-workloads.

14. The method of claim 13 , wherein, processing the second member of the set of sub-workloads, comprises:

determining the second member to produce a second metric;

comparing the second metric with the threshold to determine that the second metric is beyond the threshold;

dividing, based at least in part on the comparison, the second member into a further set of sub-workloads consisting of the predefined number of equal parts from the second member;

creating a third request to execute a third thread, the second request including a first member of the further set of sub-workloads; and

processing a second member of the further set of sub-workloads.

15. The method of claim 13 , wherein the busy-fail field is bit in a chip-to-chip protocol interface (CTCPI) packet.

16. The method of claim 13 , wherein the second request includes a no-return field that is set.

17. The method of claim 16 , wherein the no-return field is bit in a chip-to-chip protocol interface (CTCPI) packet.

18. The method of claim 16 , wherein the no-return field is used to signal that the second thread does not return a value to a stack position.

19. The method of claim 16 , wherein the no-return field releases the first thread from having to wait for the second thread to return.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: WINDH, SKYLER ARRON; BREWER, TONY M; ESTEP, PATRICK
To: MICRON TECHNOLOGY, INC.
Reel/Frame 058491/0564 →
Continuity (2)
Provisional Application 63132754 · Dec 31, 2020
Related Publication 20220206846A1 · Jun 30, 2022
References Cited (29)
US 5434972A · Hamlin · 1995 [cited by examiner]
US 8169437B1 · Legakis · 2012 [cited by examiner]
US 8229946B1 · Crane · 2012 [cited by examiner]
US 8813091B2 · Maessen · 2014 [cited by examiner]
US 9043401B2 · Wong · 2015 [cited by examiner]
US 20070061519A1 · Barrett et al. · 2007 [cited by applicant]
US 20070067606A1 · Lin · 2007 [cited by examiner]
US 20100251257A1 · Kim · 2010 [cited by examiner]
US 20120246158A1 · Ke · 2012 [cited by examiner]
US 20140244643A1 · Basak · 2014 [cited by examiner]
US 20150026698A1 · Malakhov · 2015 [cited by examiner]
US 20150103677A1 · Lee · 2015 [cited by examiner]
US 20150193959A1 · Shah · 2015 [cited by examiner]
US 20160011901A1 · Hurwitz · 2016 [cited by examiner]
US 20160179574A1 · Merrill, III · 2016 [cited by examiner]
US 20160335127A1 · Artmeier · 2016 [cited by examiner]
US 20180103088A1 · Blainey · 2018 [cited by examiner]
US 20190303387A1 · Smarda · 2019 [cited by examiner]
US 20200184366A1 · Mandal · 2020 [cited by examiner]
US 20200233706A1 · Smith · 2020 [cited by examiner]
US 20200280511A1 · Gapin et al. · 2020 [cited by applicant]
US 20210019219A1 · Patel · 2021 [cited by examiner]
US 20210124611A1 · Saillet · 2021 [cited by examiner]
US 20210183005A1 · Du · 2021 [cited by examiner]
US 20210342200A1 · Gupta · 2021 [cited by examiner]
US 20230031226A1 · Lee · 2023 [cited by examiner]
“Cilk”, [Online]. Retrieved from the Internet: URL: https: web.archive.org web 20200412001219 https: en.wikipedia.org wiki Cilk, (Accessed on Apr. 12, 2020), 10 pgs. [cited by applicant]
“Chinese Application Serial No. 202111626756.3, Office Action mailed Nov. 26, 2024”, w/ English translation, 18 pgs. [cited by applicant]
“Chinese Application Serial No. 202111626756.3, Response filed Mar. 20, 2025 to Office Action mailed Nov. 26, 2024”, with English claims, 16 pages, English claims only. [cited by applicant]