IP Library Granted Patent US 12,236,237
Granted Patent B2
US 12,236,237 · App. 18/129,808 · Granted Feb 25, 2025

Processor cores using content object identifiers for routing and computation

Inventors: Davor Capalija (Cupertino, CA); Ljubisa Bajic (Toronto, CA); Jasmina Vasiljevic (Cupertino, CA); Yongbum Kim (Los Altos Hills, CA)
Assignee: Tenstorrent Inc.
G06F9/30G06F8/427
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,237
App. No.
18/129,808
Granted
Feb 25, 2025
Kind
B2
Abstract

Processor cores using content object identifiers for routing and computation are disclosed. One method includes executing a complex computation using a set of processing cores. The method includes routing a set of content objects using a set of content object identifiers and executing a set of instructions. The set of instructions are defined using a set of operand identifiers. The operand identifiers represent content object identifiers in the set of content object identifiers. The content objects can be routed according to a named data networking (NDN) or content-centric networking (CCN) paradigm with the content object identifiers mentioned above serving as the names for the computation data being routed by the network.

Claims (75)

1. A method, wherein each step is conducted by a set of processing cores executing a complex computation, comprising:

indirectly transmitting, from a first processing core in the set of processing cores to a second processing core in the set of processing cores, a request message having a content object identifier;

receiving, at the second processing core in the set of processing cores, the request message;

directly transmitting, from the second processing core to the first processing core, a content object in response to receiving the request message;

buffering the content object on a memory on the first processing core using a memory address;

obtaining the content object from the memory using an operand identifier; and

executing an instruction using a processing pipeline on the first processing core, wherein the instruction includes the operand identifier.

2. The method of claim 1 , wherein:

the request message does not identify a processing core in the set of processing cores.

3. The method of claim 1 , wherein:

the content object identifier and the operand identifier are the same.

4. The method of claim 1 , wherein indirectly transmitting the request message comprises:

routing the request message through a subset of processing cores in the set of processing cores; and

adding a set of entries to a set of content object routing tables with the content object identifier while routing the request message through the subset of processing cores.

5. The method of claim 4 , wherein directly transmitting the content object comprises:

checking the set of content object routing tables for the content object identifier on the processing cores in the subset of processing cores; and

routing the content object through the subset of processing cores using the set of content object routing tables.

6. The method of claim 1 , further comprising:

directly transmitting, from the second processing core to the first processing core, a set of content objects in response to receiving the request message;

wherein the content object identifier identifies the set of content objects.

7. The method of claim 6 , wherein:

the content object identifier is a shared root identifier for the set of content objects.

8. The method of claim 1 , wherein:

the content object identifier is from a set of content object identifiers;

the operand identifier is from a set of operand identifiers;

the set of operand identifiers are unambiguously mapped to an underlying set of application datums of the complex computation throughout the executing of the complex computation; and

the set of content object identifiers are unambiguously mapped to the underlying set of application datums of the complex computation throughout the executing of the complex computation.

9. The method of claim 8 , wherein:

the content object is from a set of content objects;

the instruction is from a set of instructions;

the set of instructions define composite computations of the complex computation; and

the underlying set of application datums are a set of variables in the complex computation.

10. The method of claim 9 , wherein:

the set of content objects contain data values for the underlying set of application datums; and

the set of instructions are executed using the data values for the underlying set of application datums.

11. The method of claim 9 , further comprising:

parsing an application code definition of the complex computation;

defining, based on the parsing, the set of content objects, the set of content objects including the content object;

defining, based on the parsing, the set of operand identifiers; and

defining, based on the parsing, a set of processing core operational codes to execute the complex computation;

wherein the set of instructions are the set of processing core operational codes combined with the set of operand identifiers.

12. The method of claim 11 , wherein:

the transmitting of the request message is conducted using a set of routers distributed across the set of processing cores; and

the processing pipeline is in a set of processing pipelines distributed across the set of processing cores.

13. The method of claim 12 , wherein:

the memory is in a set of memories distributed across the set of processing cores;

the set of memories are blocks of static random access memory located on the set of processing cores; and

a processing core controller conducts the obtaining of the content object from the memory by providing the operand identifier to the memory.

14. The method of claim 13 , wherein executing the set of instructions further comprises:

unpacking content objects from the set of content objects, using the set of processing pipelines, after obtaining data from the set of memories for the executing of the set of instructions; and

packing content objects from the set of content objects, using the set of processing pipelines, prior to writing data from the set of processing pipelines to the set of memories.

15. The method of claim 1 , wherein:

the operand identifier and the content object identifier are commonly mapped to the memory address.

16. A method, wherein each step is conducted by a set of processing cores executing a complex computation, comprising:

indirectly transmitting, from a first processing core in the set of processing cores, a solicitation message having a content object identifier;

receiving, at a second processing core in the set of processing cores, the solicitation message;

directly transmitting, from the second processing core to the first processing core, a request message in response to receiving the solicitation message;

receiving, at the first processing core in the set of processing cores, the request message;

directly transmitting, from the first processing core to the second processing core, a content object in response to receiving the request message;

buffering the content object on a memory on the second processing core using a memory address;

obtaining the content object from the memory using an operand identifier; and

executing an instruction using a processing pipeline in the set of processing cores, wherein the instruction includes the operand identifier.

17. The method of claim 16 , wherein:

the solicitation message does not identify a processing core in the set of processing cores.

18. The method of claim 16 , wherein:

the operand identifier and the content object identifier are commonly mapped to the memory address.

19. The method of claim 16 , wherein:

the operand identifier and the content object identifier are the same.

20. A set of processing cores storing executable instructions in a set of non-transitory computer-readable media which, when executed by the set of processing cores, cause the set of processing cores to execute a method comprising:

indirectly transmitting, from a first processing core in the set of processing cores to a second processing core in the set of processing cores, a request message having a content object identifier;

receiving, at the second processing core in the set of processing cores, the request message;

directly transmitting, from the second processing core to the first processing core, a content object in response to receiving the request message;

buffering the content object on a memory on the first processing core using a memory address;

obtaining the content object from the memory using an operand identifier; and

executing an instruction using a processing pipeline on the first processing core, wherein the instruction includes the operand identifier.

Assignments (3)
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT INC.
To: TENSTORRENT AI INC.
Reel/Frame 070298/0922 →
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT AI INC.
To: TENSTORRENT AI ULC
Reel/Frame 070298/0944 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2023
From: CAPALIJA, DAVOR; BAJIC, LJUBISA; VASILJEVIC, JASMINA; KIM, YONGBUM
To: TENSTORRENT INC.
Reel/Frame 063201/0945 →
Continuity (4)
Continuation In Part 17686003 · Mar 3, 2022
Continuation 16902035 · Jun 15, 2020
Provisional Application 62863042 · Jun 18, 2019
Related Publication 20230236831A1 · Jul 27, 2023
References Cited (29)
US 6453360B1 · Muller et al. · 2002 [cited by applicant]
US 6650640B1 · Muller et al. · 2003 [cited by applicant]
US 8001266B1 · Gonzalez · 2011 [cited by examiner]
US 20030108053A1 · Inaba · 2003 [cited by applicant]
US 20040250046A1 · Gonzalez et al. · 2004 [cited by applicant]
US 20050108518A1 · Pandya · 2005 [cited by applicant]
US 20050226238A1 · Hoskote et al. · 2005 [cited by applicant]
US 20080215820A1 · Conway · 2008 [cited by examiner]
US 20090016355A1 · Moyes · 2009 [cited by applicant]
US 20120185633A1 · Sano · 2012 [cited by applicant]
US 20140029616A1 · Chang et al. · 2014 [cited by applicant]
US 20140052923A1 · Ikeda · 2014 [cited by applicant]
US 20150319086A1 · Tripathi et al. · 2015 [cited by applicant]
US 20160294710A1 · Sreeramoju · 2016 [cited by applicant]
US 20170242697A1 · Baghsorkhi · 2017 [cited by applicant]
US 20170302530A1 · Wolting · 2017 [cited by applicant]
US 20170315726A1 · White · 2017 [cited by examiner]
US 20180139153A1 · Moradi · 2018 [cited by examiner]
US 20190286972A1 · Husseini et al. · 2019 [cited by applicant]
Notice of Allowance dated Oct. 17, 2023 from U.S. Appl. No. 17/686,003, 7 pages. [cited by applicant]
Non-final Office Action from U.S. Appl. No. 17/686,003 dated May 3, 2023, 11 pages. [cited by applicant]
Extended European Search Report dated Dec. 23, 2020 from European Application No. 20179923.6, 8 pages. [cited by applicant]
First Office Action from Chinese Application No. 202010557993.8 dated Sep. 5, 2022, 9 pages. [cited by applicant]
Hayashi et al., “Parallelization in an HPF Language Processor”, NEC Research and Development, Nippon electric Ltd. Tokyo, JP, vol. 39, No. 4, Oct. 1, 1998, pp. 414-421. [cited by applicant]
Astovetsky, “Parallel Computing on Heterogeneous Networks”, XP055759409, Hoboken, NJ, Retrieved from the Internet: URL:https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.216.8140&rep=rep1&type=pdf, Jan. 1, 2003, … [cited by applicant]
Notice of Allowance dated Dec. 30, 2021 from U.S. Appl. No. 19/902,035, 49 pages. [cited by applicant]
Second Office Action from Chinese Application No. 202010557993.8 dated Jan. 19, 2023, 4 pages. [cited by applicant]
Tanguy E. Raynaud et al., “A Cache Only Memory Architecture for Big Data Applications”, Technical Report No. 8, Jul. 1, 2014, 36 pages. [cited by applicant]
Extended European Search Report dated Aug. 16, 2024 from European Application No. 24167397.9, 9 pages. [cited by applicant]