IP Library › Granted Patent US 12,443,416
Granted Patent B2
US 12,443,416 · App. 18/363,333 · Granted Oct 14, 2025

Distributed geometry

Inventors: Todd Martin (Orlando, FL); Tad Robert Litwiller (Orlando, FL); Nishank Pathak (Orlando, FL); Randy Wayne Ramsey (Orlando, FL)
Assignee: Advanced Micro Devices, Inc.
G06F9/4411G06F9/3009G06F9/544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,416
App. No.
18/363,333
Granted
Oct 14, 2025
Kind
B2
Abstract

Systems, apparatuses, and methods for performing geometry work in parallel on multiple chiplets are disclosed. A system includes a chiplet processor with multiple chiplets for performing graphics work in parallel. Instead of having a central distributor to distribute work to the individual chiplets, each chiplet determines on its own the work to be performed. For example, during a draw call, each chiplet calculates which portions to fetch and process of one or more index buffer(s) corresponding to one or more graphics object(s) of the draw call. Once the portions are calculated, each chiplet fetches the corresponding indices and processes the indices. The chiplets perform these tasks in parallel and independently of each other. When the index buffer(s) are processed, one or more subsequent step(s) in the graphics rendering process are performed in parallel by the chiplets.

Claims (35)

1. A processor comprising:

a plurality of geometry circuits;

wherein responsive to a draw call, each geometry circuit of the plurality of geometry circuits is configured to:

independently calculate a location of a portion of geometry data related to one or more graphics primitives of the draw call;

fetch the portion of the geometry data, based on the calculated location; and

process the fetched portion of the geometry data.

2. The processor as recited in claim 1 , wherein each geometry circuit is configured to circuits independently determine the location of the portion of geometry data to be fetched, without coordination or communication with other geometry circuits of the plurality.

3. The processor as recited in claim 1 , wherein each geometry circuit is configured to calculate a location of a portion of data to fetch based at least in part on one or more of a total number of geometry circuits processing the draw call, a logical geometry circuit identifier of each geometry circuit, or a number of indices to fetch for the draw call.

4. The processor as recited in claim 1 , wherein a location calculated by a given geometry circuit is indicative of a portion of an index buffer associated with the draw call, wherein each of a plurality of indices of the index buffer corresponds to a distinct portion of data to be fetched for the draw call.

5. The processor as recited in claim 4 , wherein the index buffer comprises a list of pointers that point to vertices corresponding to a plurality of graphics primitives that form one or more graphics objects corresponding to the draw call.

6. The processor as recited in claim 4 , wherein a number of indices corresponding to a portion of data to be fetched, for each portion of the index buffer, is determined based at least in part on a size of a primitive group.

7. The processor as recited in claim 1 , wherein a portion of data fetched by each geometry circuit corresponds to a non-contiguous portion of an index buffer.

8. A method comprising:

responsive to a draw call:

independently calculating, by each geometry circuit of a plurality of geometry circuits, a location of a portion of geometry data associated with one or more graphics primitives of the draw call;

fetching, by each geometry circuit, the portion of the geometry data based on the calculated location; and

processing, by each geometry circuit, the fetched portion of the geometry data.

9. The method as recited in claim 8 , further comprising independently determining, by each geometry circuit, the location of the portion of geometry data to be fetched, without coordination or communication with other geometry circuits of the plurality of geometry circuits.

10. The method as recited in claim 8 , further comprising calculating, by each geometry circuit, the location of the portion of data to fetch based at least in part on one or more parameters comprising a total number of geometry circuits processing the draw call, a logical geometry circuit identifier of each geometry circuit, a number of indices to fetch for the draw call, or a combination thereof.

11. The method as recited in claim 8 , wherein the location calculated by each geometry circuit is indicative of a portion of an index buffer associated with the draw call, wherein each of a plurality of indices of the index buffer corresponds to a portion of data to be fetched for the draw call.

12. The method as recited in claim 11 , wherein the index buffer comprises a list of pointers that point to vertices corresponding to a plurality of graphics primitives that form one or more graphics objects corresponding to the draw call.

13. The method as recited in claim 11 , wherein a number of indices corresponding to a portion of data to be fetched, for each portion of the index buffer, is determined based at least in part on a size of a primitive group.

14. The method as recited in claim 8 , wherein the portion of data fetched by a given geometry circuit corresponds to a non-contiguous portion of an index buffer.

15. A system comprising:

a memory configured to store data; and

a processor comprising a plurality of geometry circuits, wherein each geometry circuit is configured to:

calculate a memory address corresponding to a portion of data to be processed by the geometry circuit;

fetch the portion of data from the memory using the calculated memory address; and

process the portion of data;

wherein the geometry circuits are configured to perform the fetch and process without requiring coordination with other geometry circuits.

16. The system as recited in claim 15 , wherein the portion of data comprises geometry data associated with one or more graphics primitives specified by a draw call.

17. The system as recited in claim 15 , wherein each geometry circuit is configured to calculate a location of a portion of data based at least in part on one or more of a number of geometry circuits processing a draw call, a geometry circuit identifier of each geometry circuit, or a number of indices to fetch for the draw call.

18. The system as recited in claim 15 , wherein a location calculated by a given geometry circuit is indicative of a portion of a buffer in the memory.

19. The system as recited in claim 16 , wherein the buffer comprises a list of pointers that point to vertices corresponding to a plurality of graphics primitives that form one or more graphics objects corresponding to the draw call.

20. The system as recited in claim 15 , wherein a portion of data fetched by each geometry circuit corresponds to a non-contiguous portion of an index buffer.

Continuity (2)
Continuation 17489059 · Sep 29, 2021
Related Publication 20230376318A1 · Nov 23, 2023
References Cited (52)
US 4434461A · Puhl · 1984 [cited by applicant]
US 4965750A · Matsuo · 1990 [cited by examiner]
US 4967375A · Pelham · 1990 [cited by examiner]
US 5159320A · Matsuo · 1992 [cited by examiner]
US 5276780A · Sugiura · 1994 [cited by examiner]
US 6166724A · Paquette · 2000 [cited by examiner]
US 6233599B1 · Nation et al. · 2001 [cited by applicant]
US 6362828B1 · Morgan · 2002 [cited by examiner]
US 6636224B1 · Blythe · 2003 [cited by examiner]
US 7015914B1 · Bastos · 2006 [cited by examiner]
US 7477266B1 · Bastos · 2009 [cited by examiner]
US 8730248B2 · Sasaki et al. · 2014 [cited by applicant]
US 11016929B2 · Ray et al. · 2021 [cited by applicant]
US 11755336B2 · Martin · 2023 [cited by examiner]
US 20020015039A1 · Moore · 2002 [cited by examiner]
US 20040236874A1 · Largman · 2004 [cited by examiner]
US 20040257622A1 · Shibaki et al. · 2004 [cited by applicant]
US 20050041031A1 · Diard · 2005 [cited by examiner]
US 20050204015A1 · Steinhart · 2005 [cited by examiner]
US 20080001960A1 · Chen · 2008 [cited by applicant]
US 20080313589A1 · Maixner et al. · 2008 [cited by applicant]
US 20100049477A1 · Sivan · 2010 [cited by examiner]
US 20130021360A1 · Gruber · 2013 [cited by applicant]
US 20140176588A1 · Duluk, Jr. · 2014 [cited by examiner]
US 20140362102A1 · Cerny et al. · 2014 [cited by applicant]
US 20150039860A1 · Sundar et al. · 2015 [cited by applicant]
US 20150145880A1 · Smith et al. · 2015 [cited by applicant]
US 20160104264A1 · Arulesan et al. · 2016 [cited by applicant]
US 20170178401A1 · Agrawal et al. · 2017 [cited by applicant]
US 20180024938A1 · Paltashev et al. · 2018 [cited by applicant]
US 20180300933A1 · Burke et al. · 2018 [cited by applicant]
US 20190236749A1 · Gould et al. · 2019 [cited by applicant]
US 20200074726A1 · Gierach · 2020 [cited by examiner]
US 20200151926A1 · Uhrenholt · 2020 [cited by examiner]
US 20210142438A1 · Appu · 2021 [cited by examiner]
US 20210263853A1 · Waters et al. · 2021 [cited by applicant]
US 20220139021A1 · Frisinger et al. · 2022 [cited by applicant]
US 20230097097A1 · Martin · 2023 [cited by applicant]
JP 05266201A · 1993 [cited by applicant]
JP 2001167075A · 2001 [cited by applicant]
WO 2007132746A1 · 2007 [cited by applicant]
Islam, et al., “Improving Node-Level MapReduce Performance using Processing-in-Memory Technologies”, European Conference on Parallel Processing, 12 pages, https://csrl.cse.unt.edu/kavi/Research/UCHPC-2014.pdf. [Retrieve… [cited by applicant]
Nyasulu, Peter M., “System Design for a Computational-RAM Login-In-Memory Parallel Processing Machine”, PhD Thesis, May 1999, 196 pages, https://central.bac-lac.gc.ca/.item?id=NQ42803&op=pdf&app=Library&oclc_number=1006… [cited by applicant]
Pugsley, et al. “NDC: Analyzing the impact of 3D-stacked Memory+Logic Devices on MapReduce Workloads”, 2014 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Mar. 2014, pp. 190-200, … [cited by applicant]
Yang et al., “A Processing-in-Memory Architecture Programming Paradigm for Wireless Internet-of-Things Applications”, Sensors Journal, Jan. 2019, 23 pages, https://pdfs.semanticscholar.org/81cd/bda211fd479c23de2213fb610… [cited by applicant]
Alexander, Esther C., “MPC8240 and MPC8245: Comparison and Compatibility”, Freescale Semiconductor, Inc., Document No. AN2128, Oct. 2006, 28 pages, Revision 5, https://www.nxp.com/docs/en/application-note/AN2128.pdf. [R… [cited by applicant]
“Intel® Pentium® 4 Processor-Based System Integration Overview for Processors” Intel, Dec. 21, 2015, 7 pages, https://www.intel.co.uk/content/www/uk/en/support/processors/desktop-processors/000006865.html. [Retrieved Ju… [cited by applicant]
Kalamatianos et al., U.S. Appl. No. 17/139,496, entitled “Reusing Remote Registers in Processing in Memory”, filed Dec. 31, 2020, 28 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2022/077848, mailed Feb. 2, 2023, 11 pages. [cited by applicant]
Dong et al., U.S. Appl. No. 17/499,494, entitled “Duplicated Registers in Chiplet Processing Units”, filed Oct. 12, 2021, 33 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2022/076223, mailed Jan. 11, 2023, 10 pages. [cited by applicant]
Office Action in Japanese Application No. 2024-518308 dated May 13, 2025, 14 pages. [cited by applicant]