IP Library Granted Patent US 12,536,732
Granted Patent B2
US 12,536,732 · App. 18/675,746 · Granted Jan 27, 2026

Apparatus and method for efficient graphics processing including ray tracing

Inventors: Sven Woop (Voelklingen, DE); Michael J. Doyle (San Jose, CA); Sreenivas Kothandaraman (Sammamish, WA); Karthik Vaidyanathan (San Francisco, CA); Abhishek R. Appu (El Dorado Hills, CA); Carsten Benthin (Voelklingen, DE); Prasoonkumar Surti (Folsom, CA); Holger Gruen (Peissenberg, DE); Stephen Junkins (Bend, OR); Adam Lake (Portland, OR); Bret G. Alfieri (Forest Grove, OR); Gabor Liktor (San Francisco, CA); Joshua Barczak (Timonium, MD); Won-Jong Lee (Santa Clara, CA)
Assignee: Intel Corporation
G06T15/06G06T1/20G06T1/60G06T15/005G06T2210/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,732
App. No.
18/675,746
Granted
Jan 27, 2026
Kind
B2
Abstract

Apparatus and method for efficient graphics processing including ray tracing. For example, one embodiment of a graphics processor comprises: execution hardware logic to execute graphics commands and render images; an interface to couple functional units of the execution hardware logic to a tiled resource; and a tiled resource manager to manage access by the functional units to the tiled resource, a functional unit of the execution hardware logic to generate a request with a hash identifier (ID) to request access to a portion of the tiled resource, wherein the tiled resource manager is to determine whether a portion of the tiled resource identified by the hash ID exists, and if not, to allocate a new portion of the tiled resource and associate the new portion with the hash ID.

Claims (32)

1 . A graphics processor comprising:

circuitry to schedule shaders for execution; and

execution hardware logic coupled to the circuitry to execute the shaders using a plurality of fixed sized blocks of memory, a functional unit of the execution hardware logic to request a write access with a hash identifier (ID) that is mapped to a portion of the plurality of fixed sized blocks of memory in executing a first shader, wherein the first shader is to write content to the portion of the plurality of fixed sized blocks of memory, and wherein a second shader is to request a read access with the hash ID to read the content from the portion of the plurality of fixed sized blocks of memory.

2 . The graphics processor of claim 1 , wherein the hash ID is generated based on one or more of a ray tracing instance ID, a geometry ID, and a frame counter.

3 . The graphics processor of claim 1 , wherein the first shader is to lock the portion of the plurality of fixed sized blocks of memory during writing the content to the portion of the plurality of fixed sized blocks of memory.

4 . The graphics processor of claim 1 , wherein responsive to the portion of the plurality of fixed sized blocks of memory identified by the hash ID not existing, a new portion of the plurality of fixed sized blocks of memory is allocated and associated with the hash ID.

5 . The graphics processor of claim 4 , wherein a flag is set for the new portion of the plurality of fixed sized blocks of memory to indicate that the new portion of the plurality of fixed sized blocks of memory is available for allocating.

6 . The graphics processor of claim 4 , wherein upon no new portion of the plurality of fixed sized blocks of memory being found, an existing portion of the plurality of fixed sized blocks of memory is evicted and reallocated as the new portion of the plurality of fixed sized blocks of memory associated with the hash ID.

7 . The graphics processor of claim 6 , wherein evicting the existing portion of the plurality of fixed sized blocks of memory follows a least recently used (LRU) eviction policy and the existing portion of the plurality of fixed sized blocks of memory is used least recently.

8 . The graphics processor of claim 1 , wherein upon a new portion of the plurality of fixed sized blocks of memory being allocated, the execution hardware logic is to execute a user compute shader, the user compute shader to write triangles in the portion of the plurality of fixed sized blocks of memory.

9 . The graphics processor of claim 1 , wherein the read access with the hash ID from the second shader is blocked until the write access with the hash ID in executing the first shader is fulfilled.

10 . The graphics processor of claim 1 , wherein the plurality of fixed sized blocks of memory comprises a memory buffer subdivided into tiles.

11 . A method comprising:

scheduling shaders for execution;

executing the shaders by execution hardware logic using a plurality of fixed sized blocks of memory;

requesting a write access with a hash identifier (ID) that is mapped to a portion of the plurality of fixed sized blocks of memory in executing a first shader;

writing, by the first shader, content to the portion of the plurality of fixed sized blocks of memory; and

requesting, by a second shader, a read access with the hash ID to read the content from the portion of the plurality of fixed sized blocks of memory.

12 . The method of claim 11 , wherein the hash ID is generated based on one or more of a ray tracing instance ID, a geometry ID, and a frame counter.

13 . The method of claim 11 , wherein the first shader lacks-lock the portion of the plurality of fixed sized blocks of memory during writing the content to the portion of the plurality of fixed sized blocks of memory.

14 . The method of claim 11 , wherein responsive to the portion of the plurality of fixed sized blocks of memory identified by the hash ID not existing, a new portion of the plurality of fixed sized blocks of memory is allocated and associated with the hash ID.

15 . The method of claim 11 , wherein the read access with the hash ID from the second shader is blocked until the write access with the hash ID in executing the first shader is fulfilled.

16 . The method of claim 11 , wherein the plurality of fixed sized blocks of memory comprises a memory buffer subdivided into tiles.

17 . A non-transitory machine-readable medium having program code stored thereon which, when executed by a machine, causes the machine to perform:

scheduling shaders for execution;

executing the shaders by execution hardware logic using a plurality of fixed sized blocks of memory;

requesting a write access with a hash identifier (ID) that is mapped to a portion of the plurality of fixed sized blocks of memory in executing a first shader;

writing, by the first shader, content to the portion of the plurality of fixed sized blocks of memory; and

requesting, by a second shader, a read access with the hash ID to read the content from the portion of the plurality of fixed sized blocks of memory.

18 . The non-transitory machine-readable medium of claim 17 , wherein the hash ID is generated based on one or more of a ray tracing instance ID, a geometry ID, and a frame counter.

19 . The non-transitory machine-readable medium of claim 17 , wherein the first shader locks the portion of the plurality of fixed sized blocks of memory during writing the content to the portion of the plurality of fixed sized blocks of memory.

20 . The non-transitory machine-readable medium of claim 17 , wherein responsive to the portion of the plurality of fixed sized blocks of memory identified by the hash ID not existing, a new portion of the plurality of fixed sized blocks of memory is allocated and associated with the hash ID.

Continuity (3)
Continuation 17133573 · Dec 23, 2020
Provisional Application 63066799 · Aug 17, 2020
Related Publication 20240394956A1 · Nov 28, 2024
References Cited (132)
US 5263136A · Deaguiar et al. · 1993 [cited by applicant]
US 6047088A · Van Beek et al. · 2000 [cited by applicant]
US 6219058B1 · Trika · 2001 [cited by applicant]
US 6344852B1 · Zhu et al. · 2002 [cited by applicant]
US 6445390B1 · Aftosmis et al. · 2002 [cited by applicant]
US 7102646B1 · Rubinstein et al. · 2006 [cited by applicant]
US 7570271B1 · Brandt · 2009 [cited by applicant]
US 9223789B1 · Seigle et al. · 2015 [cited by applicant]
US 10839475B2 · Benthin et al. · 2020 [cited by applicant]
US 11321910B2 · Doyle et al. · 2022 [cited by applicant]
US 20040155885A1 · Munshi et al. · 2004 [cited by applicant]
US 20050179686A1 · Christensen et al. · 2005 [cited by applicant]
US 20060106836A1 · Masugi · 2006 [cited by examiner]
US 20080122841A1 · Brown et al. · 2008 [cited by applicant]
US 20080186310A1 · Kondo · 2008 [cited by applicant]
US 20090128562A1 · Mccombe et al. · 2009 [cited by applicant]
US 20090189898A1 · Dammertz et al. · 2009 [cited by applicant]
US 20090284523A1 · Peterson et al. · 2009 [cited by applicant]
US 20090322752A1 · Peterson · 2009 [cited by examiner]
US 20100079451A1 · Zhou et al. · 2010 [cited by applicant]
US 20100289799A1 · Hanika et al. · 2010 [cited by applicant]
US 20110080403A1 · Ernst et al. · 2011 [cited by applicant]
US 20110216068A1 · Sathe · 2011 [cited by applicant]
US 20120011144A1 · Transier · 2012 [cited by examiner]
US 20120293515A1 · Clarberg et al. · 2012 [cited by applicant]
US 20130179787A1 · Brockmann et al. · 2013 [cited by applicant]
US 20130249915A1 · Stich · 2013 [cited by applicant]
US 20140015835A1 · Akenine-Moller et al. · 2014 [cited by applicant]
US 20140168221A1 · Fishwick · 2014 [cited by applicant]
US 20140176589A1 · Duluk, Jr. et al. · 2014 [cited by applicant]
US 20140292782A1 · Fishwick et al. · 2014 [cited by applicant]
US 20150022521A1 · Loop · 2015 [cited by applicant]
US 20150332429A1 · Fernandez et al. · 2015 [cited by applicant]
US 20160027144A1 · Fernandez et al. · 2016 [cited by applicant]
US 20160063753A1 · Peterson · 2016 [cited by examiner]
US 20160071312A1 · Laine et al. · 2016 [cited by applicant]
US 20160086303A1 · Bae · 2016 [cited by examiner]
US 20160292908A1 · Obert · 2016 [cited by applicant]
US 20170178387A1 · Woop et al. · 2017 [cited by applicant]
US 20170287203A1 · Vaidyanathan et al. · 2017 [cited by applicant]
US 20180061124A1 · Prokopenko et al. · 2018 [cited by applicant]
US 20180293782A1 · Benthin et al. · 2018 [cited by applicant]
US 20190308099A1 · Lalonde · 2019 [cited by examiner]
US 20190311521A1 · Nevraev et al. · 2019 [cited by applicant]
US 20200051314A1 · Laine et al. · 2020 [cited by applicant]
US 20200175741A1 · Gierach · 2020 [cited by examiner]
US 20200211264A1 · Janus et al. · 2020 [cited by applicant]
US 20200301826A1 · Appu · 2020 [cited by examiner]
US 20210383591A1 · Gupta et al. · 2021 [cited by applicant]
US 20220051466A1 · Doyle et al. · 2022 [cited by applicant]
TW I546770B · 2016 [cited by applicant]
TW I564839B · 2017 [cited by applicant]
TW 201842478A · 2018 [cited by applicant]
TW 202013308A · 2020 [cited by applicant]
Pantaleoni et al., HLBVH: Hierarchical LBVH Construction for Real Time Ray Tracing of Dynamic Geometry, NVIDIA, High-Performance Graphics, 2010, 49 pages. [cited by applicant]
Popov et al., “Object Partitioning Considered Harmful: Space Subdivision for BVHs”, Proceedings of the 1st ACM conference on High Performance Graphics, 2009, pp. 1-8. [cited by applicant]
Rasmusson, Jim, “Lossy and Lossless Compression Techniques for Graphics Processors”, Department of Computer Science, Lund University, 2012, 120 pages. [cited by applicant]
Rivera, Kris Krishna, “Ray Collection Bounding Volume Hierarchy,” Electronic Theses and Dissertations, 2004-2019, University of Central Florida, STARS, 2011, 93 pages. [cited by applicant]
Sander et al., “Fast Hardware Construction and Refitting of Quantized Bounding Volume Hierarchies”, Eurographics Symposium on Rendering 2017, vol. 36, No. 4, paper1037, 2017, pp. 1-12. [cited by applicant]
Search Report and Written Opinion, NL App. No. 2028744, Apr. 13, 2022, 9 pages of Original Document Only. [cited by applicant]
Search Report and Written Opinion, NL App. No. 2028745, Mar. 3, 2022, 9 pages of Original Document Only. [cited by applicant]
Segovia et al., “Memory efficient ray tracing with hierarchical mesh quantization,” in Graphics Interface '10: Proceedings of Graphics Interface 2010, 2010, pp. 153-160. [cited by applicant]
Stich et al., “Spatial Splits in Bounding Volume Hierarchies”, Available Online at <https://www.nvidia.in/docs/IO/77714/sbvh.pdf>, 2009, 7 pages. [cited by applicant]
Stich et al., “Spatial splits in bounding volume hierarchies,” in HPG '09: Proceedings of the Conference on High Performance Graphics 2009, 2009, 23 pages. [cited by applicant]
Vaidyanathan et al., “Watertight Ray Traversal with Reduced Precision”, High Performance Graphics, Eurographics Proceedings, 2016, 8 pages. [cited by applicant]
Vaidyanathan et al., “Watertight ray traversal with reduced precision,” in HPG '16: Proceedings of High Performance Graphics, 2016, 8 pages. [cited by applicant]
Viitanen et al., “Merge Tree: A Fast Hardware HLBVH Constructor for Animated Ray Tracing”, ACM Transactions on Graphics, vol. 36, No. 5, Article 169, Oct. 2017, pp. 169:1-169:14. [cited by applicant]
Viitanen et al., “PLOCTree: A Fast, High-Quality Hardware BVH Builder”, Proc. ACM Comput. Graph. Interact. Tech., vol. 1, No. 2, Article 35, Aug. 2018, pp. 35:1-35:19. [cited by applicant]
Wald et al., “Embree: A Kernel Framework for Efficient CPU Ray Tracing”, Available Online at <https://www.embree.org/papers/2014-Siggraph-Embree.pdf>, ACM Transactions on Graphics, 2014, 8 pages. [cited by applicant]
Wald, Ingo, “Fast Construction of SAH BVHs on the Intel Many Integrated Core (MIC) Architecture”, IEEE Transactions on Visualization and Computer Graphics, vol. 18, No. 1, Jan. 2012, pp. 47-57. [cited by applicant]
Wald, Ingo, “Fast Construction of SAH BVHs on the Intel Many Integrated Core (MIC) Architecture,” IEEE Transactions on Visualization and Computer Graphics, pp. 47-57, 2012, 9 pages. [cited by applicant]
Woop, Sven, “A Ray Tracing Hardware Architecture for Dynamic Scenes”, A thesis submitted in partial fulfillment of the requirements for the Diploma in Computer Science, Universitat des Saarlandes, Mar. 15, 2004, 100 pag… [cited by applicant]
Woop, Sven, “Drpu: A Programmable Hardware Architecture for Real-time Ray Tracing of Coherent Dynamic Scenes”, PhD Thesis, Saarland University, Jun. 29, 2007, 207 pages. [cited by applicant]
Ylitie et al., “Efficient Incoherent Ray Traversal on GPUs Through Compressed Wide BVHs”, HPG '17, ACM, Jul. 28-30, 2017, 13 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/133,547, Jan. 8, 2025, 19 pages. [cited by applicant]
Office Action and Search Report, TW App. No. 110125338, Jan. 3, 2025, 21 pages (10 pages of English Translation and 11 pages of Original Document). [cited by applicant]
Office Action and Search Report, TW App. No. 110125339, Jan. 3, 2025, 19 pages (9 pages of English Translation and 10 pages of Original Document). [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21858783.0, Sep. 11, 2024, 9 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/133,547, Aug. 28, 2024, 18 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 21858782.2, Sep. 16, 2024, 11 pages. [cited by applicant]
Bauszat et al., “The Minimal Bounding vol. Hierarchy”, The Eurographics Association, 2010, 8 pages. [cited by applicant]
Benthin et al., “PLOC++ : Parallel Locally-Ordered Clustering for Bounding Volume Hierarchy Construction Revisited”, Proceedings of the ACM on Computer Graphics and Interactive Techniques, 2022, p. 1-13. [cited by applicant]
Bittner et al., “Incremental BVH Construction for Ray Tracing”, Computers & Graphics, Dec. 8, 2014, 12 pages. [cited by applicant]
Clark, James H., “Hierarchical Geometric Models for Visible Surface Algorithms”, Communications of the ACM, vol. 19, No. 10, Oct. 1976, pp. 547-554. [cited by applicant]
Cline et al., “Lightweight Bounding Volumes for Ray Tracing”, Journal of Graphics, GPU, and Game Tools, Feb. 13, 2006, 10 pages. [cited by applicant]
Dammertz et al., “The Edge Volume Heuristic—Robust Triangle Subdivision for Improved BVH Performance”, IEEE Symposium on Interactive Ray Tracing, 2008, 4 pages. [cited by applicant]
Doyle et al., “A Hardware Unit for Fast SAH-optimised BVH Construction”, ACM Transactions on Graphics, vol. 32, No. 4, Article 139, Jul. 2013, 10 pages. [cited by applicant]
Doyle et al., “Apparatus and Method for Reduced Precision Bounding Volume Hierarchy Construction”, U.S. Appl. No. 16/746,636, filed Jan. 17, 2020, 166 pages. [cited by applicant]
Eisemann et al., “Implicit Object Space Partitioning: The No-Memory BVH”, Computer Graphics Lab, Technical Report, Dec. 23, 2011, 25 pages. [cited by applicant]
Ernst et al., “Early Split Clipping for Bounding Volume Hierarchies”, In IEEE Symposium on Interactive Ray Tracing, Sep. 2007, 3 pages. [cited by applicant]
Ernst et al., “Early Split Clipping for Bounding Volume Hierarchies,” in Proceedings Eurographics/IEEE Symposium on Interactive Ray Tracing 2007, 2007, 6 pages. [cited by applicant]
Fabianowski et al., “Compact BVH Storage for Ray Tracing and Photon Mapping”, In Proceedings of Eurographics Ireland Workshop, 2009, pp. 1-8. [cited by applicant]
Final Office Action, U.S. Appl. No. 16/746,636, filed Sep. 24, 2021, 23 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/133,547, filed Mar. 6, 2024, 17 pages. [cited by applicant]
Ganestam et al., “Bonsai: Rapid Bounding Volume Hierarchy Generation using Mini Trees”, Journal of Computer Graphics Techniques (JCGT), vol. 4, No. 3, Sep. 2015, pp. 23-42. [cited by applicant]
Ganestam et al., “SAH Guided Spatial Split Partitioning for Fast BVH Construction”, The Eurographics Association and John Wiley & Sons Ltd., vol. 35, No. 2, 2016, 10 pages. [cited by applicant]
Gu et al., “Efficient BVH Construction via Approximate Agglomerative Clustering”, HPG, Jul. 19-21, 2013, pp. 81-88. [cited by applicant]
Havran, Vlastimil, “Cache Sensitive Representation for BSP Trees”, In Compugraphics, vol. 97, 1997, 8 pages. [cited by applicant]
Hendrich et al., “Parallel BVH Construction using Progressive Hierarchical Refinement”, Eurographics, vol. 39, No. 2, 2017, pp. 487-494. [cited by applicant]
IEEE Computer Society, “IEEE Standard for Floating-Point Arithmetic”, Microprocessor Standards Committee, IEEE Std 75(trademark), Aug. 29, 2008, 70 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2021/042374, Mar. 2, 2023, 8 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US21/42381, Mar. 2, 2023, 8 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US21/42374, Oct. 22, 2021, 9 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US21/42381, Oct. 29, 2021, 8 pages. [cited by applicant]
Karras et al., “Fast parallel construction of high-quality bounding volume hierarchies,” in HPG '13: Proceedings of the 5th High-Performance Graphics Conference, 2013, 11 pages. [cited by applicant]
Keely, Sean, “Reduced Precision Hardware for Ray Tracing”, in GPUs, in Eurographics/ACM SIGGRAPH Symposium on High Performance Graphics, 2014, 55 pages. [cited by applicant]
Keely, Sean, “Reduced Precision for Hardware Ray Tracing in GPUs”, In High-Performance Graphics, 2014, 12 pages. [cited by applicant]
Laine, “Restart Trail for Stackless BVH Traversal”, NVIDIA Research, High Performance Graphics, 2010, 5 pages. [cited by applicant]
Lauterbach et al., “Fast BVH Construction on GPUs,” Computer Graphics Forum, Eurographics, vol. 28, No. 2, 2009, pp. 375-384. [cited by applicant]
Lauterbach et al., “Fast BVH Construction on GPUs”, Euro Graphics, vol. 28, No. 2, 2009, 10 pages. [cited by applicant]
Liktor et al., “Bandwidth-Efficient BVH Layout for Incremental Hardware Traversal”, High Performance Graphics, 2016, 11 pages. [cited by applicant]
Liu et al., “FastTree: A Hardware KD-Tree Construction Acceleration Engine for Real-Time Ray Tracing”, EDAA, 2015, pp. 1595-1598. [cited by applicant]
Lloyd et al., “Implementing Stochastic Levels of Detail with Microsoft DirectX Raytracing”, NVIDIA Corporation, Technical Blog, Jun. 15, 2020, 6 pages. [cited by applicant]
Macdonald et al., “Heuristics for Ray Tracing Using Space Subdivision”, The Visual Computer, Springer, 1990, pp. 153-166. [cited by applicant]
Mahovsky et al., “Memory-Conserving Bounding Volume Hierarchies with Coherent Ray Tracing”, IEEE Transactions On Visualization And Computer Graphics, 2006, pp. 1-8. [cited by applicant]
Meister et al., “Parallel Locally-Ordered Clustering for Bounding Volume Hierarchy Construction,” IEEE Transactions on Visualization and Computer Graphics , 2018, pp. 1345-1353. [cited by applicant]
Meister et al., “Parallel Locally-Ordered Clustering for Bounding Volume Hierarchy Construction”, DCGI, Mar. 2018, 89 pages. [cited by applicant]
Nah et al., “HART: A Hybrid Architecture for Ray Tracing Animated Scenes”, IEEE Transactions On Visualization And Computer Graphics, 2014, pp. 1-14. [cited by applicant]
Nah et al., “RayCore: A Ray-Tracing Hardware Architecture for Mobile Devices”, ACM Transactions on Graphics, vol. 33, No. 5, Article 162, pp. 162:1-162:15. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/746,636, Apr. 14, 2021, 24 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/083,123, Sep. 14, 2023, 12 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/133,547, Oct. 6, 2023, 15 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/133,573, Oct. 23, 2023, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/746,636, Jan. 7, 2022, 10 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/083,123, Jan. 24, 2024, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/133,573, Feb. 1, 2024, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/735,902, Jan. 31, 2023, 12 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/306,821, Feb. 21, 2024, 9 pages. [cited by applicant]
Notice of Grant, NL App. No. 2028744, Aug. 11, 2022, 6 pages of Original Document Only. [cited by applicant]
Notice of Grant, NL App. No. 2028745, Jul. 4, 2022, 6 pages of Original Document Only. [cited by applicant]
Notice of Allowance, TW App. No. 110125338, Aug. 5, 2025, 3 pages (1 page of English Translation and 2 pages of Original Document). [cited by applicant]
Office Action, TW App. No. 110125339, Jun. 27, 2025, 10 pages (4 pages of English Translation and 6 pages of Original Document). [cited by applicant]