IP Library Granted Patent US 12,198,253
Granted Patent B2
US 12,198,253 · App. 18/239,876 · Granted Jan 14, 2025

Method for handling of out-of-order opaque and alpha ray/primitive intersections

Inventors: Samuli Laine (Uusimaa, FI); Tero Karras (Uusimaa, FI); Greg Muthler (Austin, TX); William Parsons Newhall, Jr. (Woodside, CA); Ronald Charles Babich, Jr. (Murrysville, PA); Ignacio Llamas (Palo Alto, CA); John Burgess (Austin, TX)
Assignee: NVIDIA Corporation
G06T15/06G06T1/20G06T15/005G06T2210/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,253
App. No.
18/239,876
Granted
Jan 14, 2025
Kind
B2
Abstract

A hardware-based traversal coprocessor provides acceleration of tree traversal operations searching for intersections between primitives represented in a tree data structure and a ray. The primitives may include opaque and alpha triangles used in generating a virtual scene. The hardware-based traversal coprocessor is configured to determine primitives intersected by the ray, and return intersection information to a streaming multiprocessor for further processing. The hardware-based traversal coprocessor is configured to provide a deterministic result of intersected triangles regardless of the order that the memory subsystem returns triangle range blocks for processing, while opportunistically eliminating alpha intersections that lie further along the length of the ray than closer opaque intersections.

Claims (31)

1. A system including:

memory configured to store data that defines vertices of a plurality of primitives; and

hardware circuitry comprising a result queue and operatively coupled to the memory, wherein the hardware circuitry is configured to:

determine, in an order based on the presentation of the primitives to the hardware circuitry, primitives which are intersected by a ray;

store intersection results for a plurality of primitives the ray is determined to intersect in the result queue while omitting, from the result queue, intersection results for one or more primitives which have been determined to be intersected by the ray; and

report the intersection results stored in the result queue.

2. The system of claim 1 , wherein the omitted intersection results from the result queue are omitted to provide a deterministic visualization result regardless of an order the plurality of primitives are presented to the hardware circuitry.

3. The system of claim 1 , wherein omitting the intersection results from the result queue includes replacing intersection result stored in the result queue for a primitive intersection with another intersection result for a primitive intersection.

4. The system of claim 1 , wherein the primitives which are intersected by the ray are determined in an order the primitives in a primitive range are received from the memory that is different from an order the primitives are stored in the memory.

5. The system of claim 1 , wherein the data that defines vertices of the plurality of primitives is stored in cache line sized blocks, each of the cache line sized blocks including a different set of primitives associated with at least one node of an acceleration data structure, and the cache line sized blocks are returned from the memory in an order that is different from the order the cache line sized blocks are stored in the memory.

6. The system of claim 1 , wherein each of the one or more of the primitives the ray is determined to intersect is an alpha primitive.

7. The system of claim 1 , wherein the omitted intersection results for the one or more primitives from the result queue include primitives which are provably capable of being omitted without a functional impact on visualizing a virtual scene.

8. The system of claim 7 , wherein the primitives intersected by the ray which are provably capable of being omitted without a functional impact on visualizing the virtual scene are parametrically farther away from an origin of the ray than another primitive determined to be intersected by the ray.

9. The system of claim 1 , wherein the omitted intersection results for the one or more primitives from the result queue include at least one primitive that is parametrically closer to an origin of the ray than another primitive determined to be intersected by the ray.

10. A method implemented by a hardware circuitry including a result queue, the method comprising:

receiving, from a memory external to the hardware circuitry, data that defines vertices of a plurality of primitives;

determining, in an order based on the presentation of the primitives to the hardware circuitry, primitives which are intersected by a ray;

storing intersection results for a plurality of primitives the ray is determined to intersect in the result queue while omitting, from the result queue, intersection results for one or more primitives which have been determined to be intersected by the ray; and

reporting, to a processor, the intersection results stored in the result queue.

11. The method of claim 10 , wherein the omitted intersection results from the result queue are omitted to provide a deterministic visualization result regardless of an order the plurality of primitives are returned from the memory.

12. The method of claim 10 , wherein omitting the intersection results from the result queue includes replacing intersection result stored in the result queue for a primitive intersection with another intersection result for a primitive intersection.

13. The method of claim 10 , wherein the primitives which are intersected by the ray are determined in an order the primitives are received from the memory that is different from an order the primitives are stored in the memory.

14. The method of claim 10 , wherein each of the one or more of the primitives the ray is determined to intersect is an alpha primitive.

15. The method of claim 10 , wherein the omitted intersection results for the one or more primitives from the result queue include primitives which are provably capable of being omitted without a functional impact on visualizing a virtual scene.

16. The method of claim 15 , wherein the primitives intersected by the ray which are provably capable of being omitted without a functional impact on visualizing the virtual scene are parametrically farther away from an origin of the ray than another primitive determined to be intersected by the ray.

17. The method of claim 10 , wherein the data that defines vertices of the plurality of primitives is stored in cache line sized blocks, each of the cache line sized blocks including a different set of primitives associated with at least one node of an acceleration data structure, and the cache line sized blocks are returned from the memory in an order that is different from the order the cache line sized blocks are stored in the memory.

18. The method of claim 10 , wherein the omitted intersection results for the one or more primitives from the result queue include at least one primitive that is parametrically closer to an origin of the ray than another primitive determined to be intersected by the ray and reported.

19. The method of claim 10 , wherein the plurality of primitives include opaque primitives and alpha primitives, and reporting the primitives the ray is determined to intersect includes:

identifying, to the processor, an opaque primitive intersected by the ray which is parametrically closest intersected opaque primitive to an origin of the ray; and

identifying a plurality of alpha primitives intersected by the ray and spatially located between the origin of the ray and the parametrically closest opaque primitive to the processor in an order the intersected alpha primitives are stored in the memory.

20. The method of claim 10 , wherein the plurality of primitives include opaque primitives and alpha primitives, and the result queue includes a single location for storing an opaque primitive determined to be intersected by the ray and a plurality of locations for storing a plurality of alpha primitive determined to be intersected by the ray.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2024
From: LAINE, SAMULI; KARRAS, TERO; MUTHLER, GREG; NEWHALL, WILLIAM PARSONS, JR.; BABICH, RONALD CHARLES, JR.; LLAMAS, IGNACIO; BURGESS, JOHN
To: NVIDIA CORPORATION
Reel/Frame 069508/0412 →
Continuity (4)
Continuation 17490024 · Sep 30, 2021
Continuation 16919700 · Jul 2, 2020
Continuation 16101196 · Aug 10, 2018
Related Publication 20230410410A1 · Dec 21, 2023
References Cited (46)
US 6489955B1 · Newhall, Jr. · 2002 [cited by applicant]
US 6760024B1 · Lokovic · 2004 [cited by applicant]
US 8773422B1 · Garland et al. · 2014 [cited by applicant]
US 9552664B2 · Laine et al. · 2017 [cited by applicant]
US 9569559B2 · Karras et al. · 2017 [cited by applicant]
US 9582607B2 · Laine et al. · 2017 [cited by applicant]
US 9607426B1 · Peterson · 2017 [cited by applicant]
US 10025879B2 · Karras et al. · 2018 [cited by applicant]
US 10235338B2 · Laine et al. · 2019 [cited by applicant]
US 20040125103A1 · Kaufman et al. · 2004 [cited by applicant]
US 20140078143A1 · Lee et al. · 2014 [cited by applicant]
US 20160070767A1 · Karras et al. · 2016 [cited by applicant]
US 20160070820A1 · Laine et al. · 2016 [cited by applicant]
US 20160071234A1 · Lehtinen et al. · 2016 [cited by applicant]
US 20160071310A1 · Karras et al. · 2016 [cited by applicant]
US 20160071313A1 · Laine et al. · 2016 [cited by applicant]
US 20180089885A1 · Wald · 2018 [cited by applicant]
US 20180225862A1 · Petkov · 2018 [cited by applicant]
US 20190019326A1 · Clark et al. · 2019 [cited by applicant]
US 20190057539A1 · Stanard et al. · 2019 [cited by applicant]
U.S. Appl. No. 14/563,872, filed Dec. 8, 2014. [cited by applicant]
U.S. Appl. No. 14/697,480, filed Apr. 27, 2015. [cited by applicant]
U.S. Appl. No. 14/737,343, filed Jun. 11, 2015. [cited by applicant]
U.S. Appl. No. 16/101,066, filed Aug. 10, 2018. [cited by applicant]
U.S. Appl. No. 16/101,109, filed Aug. 10, 2018. [cited by applicant]
U.S. Appl. No. 16/101,148, filed Aug. 10, 2018. [cited by applicant]
U.S. Appl. No. 16/101,180, filed Aug. 10, 2018. [cited by applicant]
U.S. Appl. No. 16/101,232, filed Aug. 10, 2018. [cited by applicant]
U.S. Appl. No. 16/101,247, filed Aug. 10, 2018. [cited by applicant]
IEEE 754-2008 Standard for Floating-Point Arithmetic, Aug. 29, 2008, 70 pages. [cited by applicant]
The Cg Tutorial, Chapter 7, “Environment Mapping Techniques,” NVIDIA Corporation, 2003, 32 pages. [cited by applicant]
Akenine-Möller, Tomas, et al., “Real-Time Rendering,” Section 9.8.2, Third Edition CRC Press, 2008, p. 412. [cited by applicant]
Appel, Arthur, “Some techniques for shading machine renderings of solids,” AFIPS Conference Proceedings: 1968 Spring Joint Computer Conference, 9 pages. [cited by applicant]
Foley, James D., et al., “Computer Graphics: Principles and Practice,” 2nd Edition Addison-Wesley 1996 and 3rd Edition Addison-Wesley 2014. [cited by applicant]
Glassner, Andrew, “An Introduction to Ray Tracing,” Morgan Kaufmann, 1989. [cited by applicant]
Hall, Daniel, “Advanced Rendering Technology,” Graphics Hardware 2001, 7 pages. [cited by applicant]
Hery, Christophe, et al., “Towards Bidirectional Path Tracing at Pixar,” 2016, 20 pages. [cited by applicant]
Kajiya, James T., “The Rendering Equation,” SIGGRAPH, vol. 20, No. 4, 1986, pp. 143-150. [cited by applicant]
Parker, Steven G., et al., “OptiX: A General Purpose Ray Tracing Engine,” ACM Transactions on Graphics, vol. 29, Issue 4, Article No. 66, Jul. 2010, 13 pages. [cited by applicant]
Quinnell, Eric Charles, “Floating-Point Fused Multiply-Add Architectures,” Dissertation, University of Texas at Austin, 2007, 163 pages. [cited by applicant]
Stich, Martin, “Introduction to NVIDIA RTX and DirectX Ray Tracing,” NVIDIA Developer Blog, Mar. 19, 2018, 13 pages. [cited by applicant]
Whitted, Turner, “An Improved Illumination Model for Shaded Display,” Communications of the ACM, vol. 23, No. 6, Jun. 1990, pp. 343-349. [cited by applicant]
Woop, Sven, “A Ray Tracing Hardware Architecture for Dynamic Scenes,” Thesis, Universität des Saarlandes, 2004, 100 pages. [cited by applicant]
Woop, Sven, et al., “RPU: A Programmable Ray Processing Unit for Realtime Ray Tracing,” ACM Transactions on Graphics, Jul. 2005, 11 pages. [cited by applicant]
Woop, Sven, et al., “Watertight Ray/Triangle Intersection,” Journal of Computer Graphics Techniques, vol. 2, No. 1, 2013, pp. 65-82. [cited by applicant]
Office Action dated Jul. 11, 2019, issued in U.S. Appl. No. 16/101,066. [cited by applicant]