IP Library › Granted Patent US 12,743,396
Granted Patent B1
US 12,743,396 · App. 18/916,648 · Granted Sep 22, 2026

Recursive generation and distribution of command bundles at core array

Inventors: Micah Villmow (Santa Clara, CA); Jiasheng Chen (Folsom, CA); Mark Leather (San Rafael, CA); Rajabali M. Koduri (San Francisco, CA)
Assignee: OXMIQ Labs Inc.
G06F15/8023G06F8/443G06F8/4436G06F8/45G06F8/451G06F8/452G06F8/453G06F8/456G06F8/457G06F8/458G06F9/3017G06F9/3836G06F9/3885G06F9/3887G06F9/45516G06F9/5066G06F9/52G06F15/80G06F15/8053
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,396
App. No.
18/916,648
Granted
Sep 22, 2026
Kind
B1
Abstract

Some embodiments provide a method for a first processor core of a multi-core arrangement having multiple processor cores. From a queue of command bundles for execution by the processor cores, the method retrieves a first instance of a first command bundle to execute at the first processor core. Multiple cores of the multi-core array also retrieve other instances of the first command bundle to execute. The method executes a first command of the first command bundle instance at the first processor core. Based on a second command of the first command bundle instance, the method adds multiple instances of a second command bundle to the queue for execution by the processor cores.

Claims (28)

1 . A method comprising:

at a first worker processor core of a multi-core arrangement, the multi-core arrangement comprising a set of processor cores, wherein one of the processor cores other than the first worker processor core is designated as a leader processor core that executes a scheduler program for executing programs dispatched to the multi-core arrangement:

from a queue of command bundles for execution by the plurality of processor cores, retrieving a first instance of a first command bundle to execute at the first processor core, wherein (i) the leader processor core places a plurality of instances of the first command bundle in the queue, (ii) each instance of the first command bundle specifies a same first set of commands to be performed on different sets of data, and (iii) a plurality of worker processor cores of the multi-core arrangement also retrieve from the queue other instances of the first command bundle to execute;

executing a first command of the first instance of the first command bundle at the first worker processor core; and

based on a second command of the first instance of the first command bundle, adding a plurality of instances of a second command bundle to the queue for execution by the plurality of worker processor cores, each instance of the second command bundle specifying a same second set of commands to be performed on different sets of data,

wherein the worker processor cores that are not designated as the leader processor core (i) are capable of generating pluralities of instances of command bundles and adding the plurality of instances of command bundles to the queue for explicitly parallelizable function calls that specify a number of instances of a specific function to be called and (ii) are not capable of generating pluralities of instances of command bundles for tensor computation commands,

wherein the second command is an explicitly parallelizable function call.

2 . The method of claim 1 , wherein a second worker processor core of the multi-core arrangement that executes a second instance of the first command bundle adds a plurality of instances of a third command bundle to the queue based on a particular command of the second instance of the first command bundle, each instance of the third command bundle specifying a same third set of commands to be performed on different sets of data.

3 . The method of claim 2 , wherein the second worker processor core adds the plurality of instances of the third command bundle to the queue based on a command of the third command bundle that corresponds to the second command of the first command bundle.

4 . The method of claim 1 , wherein retrieving the first instance of the first command bundle comprises:

receiving a notification from the leader processor core that data specifying a plurality of instances of the first command bundle have been placed in the queue by the leader processor core; and

retrieving data specifying the first instance of the first command bundle.

5 . The method of claim 4 , wherein the queue is located in a storage shared with the set of processor cores of the multi-core arrangement.

6 . The method of claim 5 further comprising retrieving data for executing the first instance of the first command bundle from the storage shared with the set of processor cores, wherein the retrieved data is used for executing the first instance of the first command bundle.

7 . The method of claim 6 , wherein other processor cores of the multi-core arrangement that execute respective instances of the first command bundle retrieve other data from the storage for executing the respective instances of the first command bundle.

8 . The method of claim 4 , wherein receiving the notification comprises receiving an assignment of the first instance of the first command bundle to the first worker processor core.

9 . The method of claim 8 , wherein receiving the assignment of the first instance of the first command bundle comprises receiving assignment of a group of instances of the first command bundle, said group including the first instance of the first command bundle.

10 . The method of claim 8 , wherein receiving the assignment of the first instance of the first command bundle comprises receiving a global identifier specifying the first instance of the first command bundle, said global identifier used to retrieve the data used for executing the first instance of the first command bundle.

11 . The method of claim 10 further comprising using the global identifier to read from memory (i) a first memory state common to all of the processor cores of the plurality of worker processor cores and (ii) a second memory state specific to the first worker processor core.

12 . The method of claim 11 , wherein each respective worker processor core in the plurality of worker processor cores that executes a respective instance of the first command bundle reads the first memory state from memory while retrieving data for executing the instance of the particular command bundle assigned to the respective worker processor core.

13 . The method of claim 1 further comprising, at the first worker processor core, after executing the first command bundle instance, notifying the leader processor core of the multi-core arrangement of the completion of the first command bundle instance.

14 . The method of claim 1 , wherein the first command bundle comprises a set of commands, said set of commands comprising at least one scalar command and at least one of (i) a vector command and (ii) a matrix command.

15 . The method of claim 14 , wherein executing the first instance of the first command bundle comprises decoding the commands of the first instance of the first command bundle at a scalar processor of the first worker processor core.

16 . The method of claim 15 , wherein the scalar processor executes the at least one scalar command and distributes (i) any vector instructions to a vector processor of the first worker processor core and (ii) any matrix instructions to a matrix processor of the first worker processor core.

17 . The method of claim 1 further comprising, at the first worker processor core:

executing additional commands of the first instance of the first command bundle until a synchronization point for the second command is reached; and

upon reaching the synchronization point, retrieving an additional command bundle instance from the queue of command bundles and executing commands of the additional command bundle instance.

18 . The method of claim 17 further comprising, after receiving notification that all of the instances of the second command bundle added to the queue based on the second command have been executed, continuing to execute commands of the first instance of the first command bundle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2024
From: VILLMOW, MICAH; CHEN, JIASHENG; LEATHER, MARK; KODURI, RAJABALI M.
To: OXMIQ LABS INC.
Reel/Frame 069655/0228 →
Continuity (2)
Provisional Application 63574855 · Apr 4, 2024
Provisional Application 63574850 · Apr 4, 2024
References Cited (87)
US 4949250A · Bhandarkar et al. · 1990 [cited by applicant]
US 6192384B1 · Dally · 2001 [cited by examiner]
US 6389489B1 · Stone et al. · 2002 [cited by applicant]
US 6624818B1 · Mantor et al. · 2003 [cited by applicant]
US 6708331B1 · Schwartz · 2004 [cited by examiner]
US 8776030B2 · Grover et al. · 2014 [cited by applicant]
US 8842122B2 · Nordlund et al. · 2014 [cited by applicant]
US 8943119B2 · Hansen et al. · 2015 [cited by applicant]
US 9189828B2 · Gorchetchnikov et al. · 2015 [cited by applicant]
US 9483318B2 · Vajapeyam · 2016 [cited by applicant]
US 9710748B2 · Ross et al. · 2017 [cited by applicant]
US 10540156B2 · Mineda · 2020 [cited by examiner]
US 11409568B2 · Jiang et al. · 2022 [cited by applicant]
US 11487846B2 · Mills et al. · 2022 [cited by applicant]
US 11893392B2 · Kim et al. · 2024 [cited by applicant]
US 20050223199A1 · Grochowski · 2005 [cited by examiner]
US 20070283337A1 · Kasahara et al. · 2007 [cited by applicant]
US 20090083516A1 · Saleem et al. · 2009 [cited by applicant]
US 20090113404A1 · Takayama · 2009 [cited by examiner]
US 20100023731A1 · Ito · 2010 [cited by examiner]
US 20100146495A1 · Song · 2010 [cited by examiner]
US 20100158408A1 · El-Mahdy et al. · 2010 [cited by applicant]
US 20120246654A1 · Eichenberger · 2012 [cited by examiner]
US 20130155080A1 · Nordlund · 2013 [cited by examiner]
US 20130298133A1 · Jones · 2013 [cited by examiner]
US 20140130052A1 · Lin · 2014 [cited by examiner]
US 20140181476A1 · Srinivasan et al. · 2014 [cited by applicant]
US 20140310467A1 · Shalf · 2014 [cited by examiner]
US 20150106596A1 · Vorbach et al. · 2015 [cited by applicant]
US 20150149745A1 · Eble · 2015 [cited by examiner]
US 20150220369A1 · Vajapeyam · 2015 [cited by applicant]
US 20160283240A1 · Mishra et al. · 2016 [cited by applicant]
US 20160321048A1 · Matsuura · 2016 [cited by examiner]
US 20170083318A1 · Burger et al. · 2017 [cited by applicant]
US 20180181380A1 · Kasahara · 2018 [cited by examiner]
US 20180301119A1 · Koker · 2018 [cited by examiner]
US 20180322381A1 · Liu et al. · 2018 [cited by applicant]
US 20180336456A1 · Norrie · 2018 [cited by examiner]
US 20190065150A1 · Heddes et al. · 2019 [cited by applicant]
US 20190146857A1 · LeBeane · 2019 [cited by examiner]
US 20190347125A1 · Sankaran et al. · 2019 [cited by applicant]
US 20200026568A1 · Harris · 2020 [cited by examiner]
US 20200218540A1 · Kesiraju et al. · 2020 [cited by applicant]
US 20220012573A1 · Nagendran et al. · 2022 [cited by applicant]
US 20220114234A1 · George · 2022 [cited by examiner]
US 20220171631A1 · Kim · 2022 [cited by examiner]
US 20220197718A1 · Drepper · 2022 [cited by applicant]
US 20220405228A1 · Tan et al. · 2022 [cited by applicant]
US 20220413865A1 · Masanamuthu Chinnathurai et al. · 2022 [cited by applicant]
US 20220414052A1 · Li et al. · 2022 [cited by applicant]
US 20220414817A1 · Zad Tootaghaj et al. · 2022 [cited by applicant]
US 20230004393A1 · Madduri et al. · 2023 [cited by applicant]
US 20230029176A1 · Ray et al. · 2023 [cited by applicant]
US 20230090973A1 · Ray · 2023 [cited by examiner]
US 20230205692A1 · Nori · 2023 [cited by examiner]
US 20230229588A1 · Aggarwal et al. · 2023 [cited by applicant]
US 20230297643A1 · Shivam et al. · 2023 [cited by applicant]
US 20230342211A1 · Kim · 2023 [cited by applicant]
US 20240020119A1 · Tran · 2024 [cited by applicant]
US 20240037183A1 · Du et al. · 2024 [cited by applicant]
US 20240054081A1 · Blixt · 2024 [cited by applicant]
US 20240160448A1 · Nagata et al. · 2024 [cited by applicant]
US 20240220250A1 · Wu · 2024 [cited by applicant]
US 20240394117A1 · Hu et al. · 2024 [cited by applicant]
US 20240403013A1 · Gartmann · 2024 [cited by examiner]
US 20250004861A1 · Barik et al. · 2025 [cited by applicant]
US 20250181551A1 · Chakraborty et al. · 2025 [cited by applicant]
US 20250298759A1 · Guim Bernat et al. · 2025 [cited by applicant]
US 20250307343A1 · Garegrat · 2025 [cited by applicant]
Author Unknown, “What is PyTorch, and How Does It Work: All You Need to Know,” Simplilearn, Nov. 7, 2023, retrieved from https://semiwiki.com/semiconductor-manufacturers/tsmc/313540-inverse-lithography-technology-a-stat… [cited by applicant]
Diamant, Ron, et al., “Accelerate deep learning and innovate faster with AWS Trainium”, Amazon Web Services, AWS re:Invent, Nov. 28-Dec. 2, 2022, Las Vegas, NV. [cited by applicant]
Gupta, Pradeep, “CUDA Refresher: The CUDA Programming Model”, Nvidia, Jun. 26, 2020, retrieved from https://developer.nvidia.com/blog/cuda-refresher-cuda-programming-model/. [cited by applicant]
Hsu, Kuan-Chieh, et al., “Simultaneous and Heterogeneous Multithreading,” MICRO '23: Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, Toronto, ON, CA, Oct. 28-Nov. 1, 2023, pp. 137-1… [cited by applicant]
Jiang, Lijuan, et al., “Towards Highly Efficient DGEMM on the Emerging SW26010 Many-core Processor,” 46th International Conference on Parallel Processing, Bristol, UK, Aug. 14-17, 2017, pp. 422-431, IEEE. [cited by applicant]
Lam, Chester, “Hot Chips 2023: AMD's Phoenix SoC,” Chips and Cheese, Sep. 16, 2023, retrieved from https://chipsandcheese.com/p/hot-chips-2023-amds-phoenix-soc. [cited by applicant]
Mcafee, David, “On-Chip AI Integration is the Future of PC Computing”, AMD, Sep. 28, 2023, retrieved from https://www.amd.com/en/blogs/2023/on-chip-ai-integration-is-the-future-of-pc-computi.html. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,645, filed Oct. 15, 2024, 113 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,646, filed Oct. 15, 2024, 113 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,647, filed Oct. 15, 2024, 113 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,649, filed Oct. 15, 2024, 84 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,650, filed Oct. 15, 2024, 85 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,652, filed Oct. 15, 2024, 85 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,654, filed Oct. 15, 2024, 84 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,656, filed Oct. 15, 2024, 103 pages, OXMIQ Labs Inc. [cited by applicant]
Non-Published Commonly Owned U.S. Appl. No. 18/916,657, filed Oct. 15, 2024, 85 pages, OXMIQ Labs Inc. [cited by applicant]
Sathe, Tejas, “Tutorial: A quick overview of tensorflow2.0,” Analytics Vidhya, Aug. 1, 2020, retrieved from https://medium.com/analytics-vidhya/tutorial-a-quick-overview-of-tensorflow2-0-b28e5c6906fa. [cited by applicant]
Screen captures from YouTube video clip entitled “Meteor Lake: AI Acceleration and NPU Explained”, 4 pages, uploaded on Dec. 11, 2023 by Intel Technology. Retrieved from Internet: https://www.youtube.com/watch?v=QSzNoX0… [cited by applicant]