IP Library › Granted Patent US 12,725,345
Granted Patent B2
US 12,725,345 · App. 18/668,599 · Granted Sep 1, 2026

Ray tracing hardware acceleration with alternative world space transforms

Inventors: Gregory Muthler (Chapel Hill, NC); John Burgess (Austin, TX); James Robertson (Austin, TX); Magnus Andersson (Skane, SE)
Assignee: NVIDIA Corporation
G06T15/06G06F9/5027G06T1/20G06T15/005G06T15/08G06T17/10G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,345
App. No.
18/668,599
Granted
Sep 1, 2026
Kind
B2
Abstract

Enhanced techniques applicable to a ray tracing hardware accelerator for traversing a hierarchical acceleration structure are disclosed. The traversal efficiency of such hardware accelerators are improved, for example, by transforming a ray, in hardware, from the ray's coordinate space to two or more coordinate spaces at respective points in traversing the hierarchical acceleration structure. In one example, the hardware accelerator is configured to transform a ray, received from a processor, from the world space to at least one alternate world space and then to an object space in hardware before a corresponding ray-primitive intersection results are returned to the processor. The techniques disclosed herein facilitate the use of additional coordinate spaces to orient acceleration structures in a manner that more efficiently approximate the space occupied by the underlying primitives being ray-traced.

Claims (56)

1 . A ray tracing acceleration hardware device configured to be connected to a processor, the ray tracing acceleration hardware device comprising:

at least one state storage configured to store a traversal state of an acceleration structure and a ray; and

circuitry configured to, using the ray, traverse a traversal path in the acceleration structure from a root node in a first world space to a leaf node in an object space, and during the traversing, in response to a first transform node in the traversal path, update the stored traversal state upon transforming the ray from the first world space to a second world space and, in response to a second transform node in the traversal path, from the second world space to the object space,

wherein the first transform node is an instance node that specifies a first transform from the first world space to the second world space and the second transform node is an instance node that specifies a second transform from the second world space to the object space, and wherein the traversing includes transforming the ray according to the first transform and the second transform at respective nodes in the acceleration structure.

2 . The ray tracing acceleration hardware device according to claim 1 , wherein the transforming the ray from the first world space to the second world space is in accordance with information stored in the acceleration structure.

3 . The ray tracing acceleration hardware device according to claim 1 , wherein the circuitry is further configured to test the ray, in the object space, for intersection with bounding volumes defined by the acceleration structure.

4 . The ray tracing acceleration hardware device according to claim 1 , wherein, before the update, the at least one state is initialized based on information received from the processor relating to the ray, defined in the first world space, and the acceleration structure.

5 . The ray tracing acceleration hardware device according to claim 1 , wherein the circuitry receives the ray and information of the acceleration structure from the processor, wherein the acceleration structure comprises hierarchically arranged nodes defining plural bounding volumes bounding objects in a scene, and wherein the first world space is a coordinate space defined for an application running on the processor, and the object space is a coordinate space in which one of more of the objects in the scene are defined.

6 . The ray tracing acceleration hardware device according to claim 1 , wherein the ray tracing acceleration hardware device is a coprocessor to the processor.

7 . The ray tracing acceleration hardware device according to claim 1 , wherein the circuitry is further configured to maintain the traversal state as a top level traversal state and top level ray state for traversing the acceleration structure in the first and second world spaces, and to maintain a bottom level traversal state and bottom level ray state for traversing the acceleration structure in the object space, and wherein the circuitry is further configured to store information about the transformed ray in the top level ray state or the bottom level ray state.

8 . The ray tracing acceleration hardware device according to claim 1 , configured to, in response to reaching a node of an instance node type in the acceleration data-structure during the traversing, send information about a state of a continued traversing to a processor.

9 . The ray tracing acceleration hardware device according to claim 1 , wherein the circuitry is part of a server or a data center employed in generating an image, and the image is streamed to a user device.

10 . The ray tracing acceleration hardware device according to claim 1 , wherein the circuitry is employed in generating an image, and the image is used for training, testing, or certifying a neural network employed in a machine, robot, or autonomous vehicle.

11 . A system comprising a processor, a parallel processing unit (PPU) and a traversal coprocessor coupled to a memory, wherein the traversal coprocessor comprises:

at least one state storage configured to store a traversal state of an acceleration structure and a ray; and

circuitry configured to, using the ray, traverse a traversal path in the acceleration structure from a root node in a first world space to a leaf node in an object space, and during the traversing, in response to a first transform node in the traversal path, update the stored state upon transforming the ray from a first world space to a second world space and, in response to a second transform node in the traversal path, from the second world space to the object space,

wherein the first transform node is an instance node that specifies a first transform from the first world space to the second world space and the second transform node is an instance node that specifies a second transform from the second world space to the object space, and wherein the traversing includes transforming the ray according to the first transform and the second transform at respective nodes in the acceleration structure.

12 . The system according to claim 11 , wherein the memory is configured to store the acceleration structure of hierarchically-arranged nodes, the hierarchically-arranged nodes defining a plurality of bounding volumes bounding portions of a scene and being axis-aligned in a first coordinate space, and including a plurality of leaf nodes corresponding to surfaces in the scene, the surfaces being in respective object coordinate spaces.

13 . The system according to claim 12 , wherein the processor is configured to:

access the acceleration structure in the memory; and

for each of a plurality of frames containing the scene:

select an alternate world space for the scene; and

when the selected world space is different from the first world space, rebuild the acceleration structure in the selected world space and including, for a top level node of the acceleration structure, a transform from the first coordinate space to the selected world space.

14 . The system according to claim 13 , wherein the PPU is configured to, for each of said plurality of frames containing the scene;

access the rebuilt acceleration structure; and

in response to detecting the transform for the top level node, transmit information about the ray in the first world space, the rebuilt acceleration structure, and information about the transform to the traversal coprocessor;

receive intersection information for the ray and the rebuilt acceleration structure from the traversal coprocessor; and

render the scene in the frame in accordance with the received intersection information; and

wherein the traversal coprocessor is further configured to:

receive said information about the ray, the rebuilt acceleration structure, and information about the transform, wherein the transforming the ray from the first world space to the second world space is in accordance with the received information about the transform; and

provide determined one or more intersections to the PPU.

15 . A method of testing, in a hardware device configured to accelerate traversal of an acceleration structure, whether a ray intersects a primitive, comprising:

initializing a traversal state stored in the hardware device based on information received from a processor connected to the hardware device, the information comprising an acceleration structure and a ray in a first world space; and

traversing, with the ray and by circuitry in the hardware device, a traversal path in the acceleration structure from a root node in the first world space to a leaf node in an object space, and during the traversing, in response to a first transform node in the traversal path, updating the stored state upon transforming the ray from a first world space to a second world space and, in response to a second transform node in the traversal path, from the second world space to the object space,

wherein the first transform node is an instance node that specifies a first transform from the first world space to the second world space and the second transform node is an instance node that specifies a second transform from the second world space to the object space, and wherein the traversing includes transforming the ray according to the first transform and the second transform at respective nodes in the acceleration structure.

16 . The method of claim 15 , wherein the transforming the ray from the first world space to the second world space is performed in accordance with transform information stored in the acceleration structure.

17 . The method of claim 15 , further comprising maintaining distinct top-level and bottom-level traversal states for world-space and object-space traversals, respectively, and storing information about transformed rays within said distinct states.

18 . The method of claim 15 , further comprising, in response to reaching a node of an instance node type in the acceleration structure during the traversing, sending information about a state of a continued traversing from the hardware device to the processor.

19 . The method of claim 15 , further comprising generating an image based on results of the testing, and using the image for training, testing, or certifying a neural network employed in a machine, robot, or autonomous vehicle.

20 . A ray tracing acceleration hardware device configured to be connected to a processor, the ray tracing acceleration hardware device comprising:

at least one state storage configured to store a traversal state of an acceleration structure and a ray; and

circuitry configured to, using the ray, traverse a traversal path in the acceleration structure from a root node in a first world space to a leaf node in an object space, and during the traversing, update the stored traversal state upon transforming the ray from a first world space to a second world space and from the second world space to the object space,

wherein the circuitry receives the ray and information of the acceleration structure from the processor, wherein the acceleration structure comprises hierarchically arranged nodes defining plural bounding volumes bounding objects in a scene, and wherein the first world space is a coordinate space defined for an application running on the processor, and the object space is a coordinate space in which one of more of the objects in the scene are defined.

21 . A ray tracing acceleration hardware device configured to be connected to a processor, the ray tracing acceleration hardware device comprising:

at least one state storage configured to store a traversal state of an acceleration structure and a ray; and

circuitry configured to, using the ray, traverse a traversal path in the acceleration structure from a root node in a first world space to a leaf node in an object space, and during the traversing, update the stored traversal state upon transforming the ray from a first world space to a second world space and from the second world space to the object space,

wherein the circuitry is further configured to maintain the traversal state as a top level traversal state and top level ray state for traversing the acceleration structure in the first and second world spaces, and to maintain a bottom level traversal state and bottom level ray state for traversing the acceleration structure in the object space, and wherein the circuitry is further configured to store information about the transformed ray in the top level ray state or the bottom level ray state.

22 . A system comprising a processor, a parallel processing unit (PPU) and a traversal coprocessor coupled to a memory, wherein the traversal coprocessor comprises:

at least one state storage configured to store a traversal state of an acceleration structure and a ray; and

circuitry configured to, using the ray, traverse a traversal path in the acceleration structure from a root node in a first world space to a leaf node in an object space, and during the traversing, update the stored state upon transforming the ray from a first world space to a second world space and from the second world space to the object space,

wherein the memory is configured to store the acceleration structure of hierarchically-arranged nodes, the hierarchically-arranged nodes defining a plurality of bounding volumes bounding portions of a scene and being axis-aligned in a first coordinate space, and including a plurality of leaf nodes corresponding to surfaces in the scene, the surfaces being in respective object coordinate spaces,

wherein the processor is configured to:

access the acceleration structure in the memory; and

for each of a plurality of frames containing the scene:

select an alternate world space for the scene; and

when the selected world space is different from the first world space, rebuild the acceleration structure in the selected world space and including, for a top level node of the acceleration structure, a transform from the first coordinate space to the selected world space.

Continuity (3)
Continuation 17669430 · Feb 11, 2022
Continuation 16897745 · Jun 10, 2020
Related Publication 20240303906A1 · Sep 12, 2024
References Cited (17)
US 11282261B2 · Muthler · 2022 [cited by applicant]
US 20160070767A1 · Karras · 2016 [cited by applicant]
US 20160070820A1 · Laine · 2016 [cited by examiner]
US 20160071313A1 · Laine · 2016 [cited by applicant]
US 20160292908A1 · Obert · 2016 [cited by applicant]
US 20170116760A1 · Laine · 2017 [cited by examiner]
US 20190088002A1 · Howson · 2019 [cited by applicant]
US 20190318530A1 · Hunt · 2019 [cited by applicant]
US 20200050451A1 · Babich · 2020 [cited by examiner]
US 20200051317A1 · Muthler · 2020 [cited by applicant]
CN 110827385A · 2020 [cited by applicant]
CN 110827388A · 2020 [cited by applicant]
CN 110827389A · 2020 [cited by applicant]
KR 20160029601A · 2016 [cited by applicant]
KR 20160038640A · 2016 [cited by applicant]
WO WO2019183664A1 · 2019 [cited by examiner]
Chinese Office Action for Application No. 202110631675 dated Jun. 6, 2023, 14 pages. [cited by applicant]