Techniques for training and inference using multiple processor resources
Apparatuses, systems, and techniques for neural network training and inference using multiple processor resources. In at least one embodiment, one or more neural networks are used to generate one or more second versions of one or more images based, at least in part, on a first version of the one or more images and a three-dimensional representation of the one or more first versions of the one or more images.
1 . A plurality of processors comprising: circuitry to:
cause a first processor of the plurality of processors to generate one or more first images using a three-dimensional (3D) representation of a scene depicted by the one or more first images, the one or more first images and the 3D representation of the scene being stored in a memory shared between the first processor and a second processor of the plurality of processors; and
cause the second processor of the plurality of processors to use the one or more first images and the 3D representation of the scene stored in the shared memory to train one or more neural networks used by the first processor.
2 . The plurality of processors of claim 1 , wherein the circuitry is to:
use the first processor of the plurality of processors to generate the one or more first images based at least in part on the 3D representation of the scene;
provide the one or more first images and the 3D representation of the scene to the second processor; and
cause the second processor of the plurality of processors to train the one or more neural networks to be used by the first processor to generate the one or more first images using the one or more first images and the 3D representation of the scene stored in the shared memory.
3 . The plurality of processors of claim 2 , wherein the circuitry is to:
provide the one or more first images and the 3D representation of the scene to the second processor of the plurality of processors via a buffer, wherein:
the first processor of the plurality of processors is to write the one or more first images and the 3D representation of the scene to the buffer; and
the second processor of the plurality of processors is to read the one or more first images and the 3D representation of the scene from the buffer; and
wherein the buffer is allocated in memory shared between the first processor and the second processor of the plurality of processors.
4 . The plurality of processors of claim 3 , wherein the second processor of the plurality of processors is to read the one or more first images and the 3D representation of the scene from the buffer with a four-frame delay after the first processor writes the one or more first images to the buffer.
5 . The plurality of processors of claim 3 , wherein the buffer is a ring buffer.
6 . The plurality of processors of claim 5 , wherein the one or more first images further comprise additional image data including depth data, normal data, albedo data, roughness data, or motion vector data.
7 . The plurality of processors of claim 2 , wherein the first processor of the plurality of processors comprises a first processor core and the second processor of the plurality of processors comprises a second processor core.
8 . A system comprising: a plurality of processors to:
cause a first processor of the plurality of processors to generate one or more first images using a three-dimensional (3D) representation of a scene depicted by the one or more first images, the one or more first images and the 3D representation of the scene being stored in a memory shared between the first processor and a second processor of the plurality of processors; and
cause the second processor of the plurality of processors to use the one or more first images and the 3D representation of the scene stored in the shared memory to train one or more neural networks used by the first processor.
9 . The system of claim 8 , wherein the plurality of processors comprises the first processor to execute a software application comprising a plugin that provides the 3D representation of the scene to a second processor that generates the one or more first images and controls training of the one or more neural networks on a third processor resource.
10 . The system of claim 9 , wherein the second processor of the plurality of processors is connected to a display device for presenting the one or more first images.
11 . The system of claim 10 , wherein the plugin controls whether a first image or a second image of the one or more images is to be presented on the display device.
12 . The system of claim 11 , wherein the plugin is to:
determine a set of parameters from training the one or more neural networks; and
update a different one or more neural networks used by the first processor of the plurality of processors to generate the one or more first images to use the set of parameters.
13 . The system of claim 10 , wherein:
the plugin is to receive training information from the third processor of the plurality of processors; and
the first processors of the plurality of processors provide the training information to the second processors of the plurality of processors to be presented using the display device.
14 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by a plurality of processors, cause the plurality of processors to at least:
cause a first processor of the plurality of processors to at least generate one or more first images using a three-dimensional (3D) representation of a scene depicted by the one or more first images, the one or more first images and the 3D representation of the scene being stored in a memory shared between the first processor and a second processor of the plurality of processors; and
cause the second processor of the plurality of processors to use the one or more first images and the 3D representation of the scene stored in the shared memory to train one or more neural networks used by the first processor.
15 . The non-transitory machine-readable medium of claim 14 , wherein the set of instructions include instructions to:
use the first processor of the plurality of processors to render the one or more first images based at least in part on the 3D representation of the scene and a second one or more neural networks;
provide the one or more first images and the 3D representation of the scene to the second processor of the plurality of processors; and
cause the second processor of the plurality of processors to train the one or more neural networks to be used by the first processor to generate the one or more first images using the one or more first images and the 3D representation of the scene stored in the shared memory.
16 . The machine-readable medium of claim 15 , wherein the first processor of the plurality of processors is to:
render a noisy version of the one or more first images using the 3D representation; and
use the second one or more neural networks to generate a denoised version of the one or more first images from the noisy version of the one or more image.
17 . The machine-readable medium of claim 16 , wherein the noisy version of the one or more first images is to be rendered using a non-deterministic algorithm.
18 . The machine-readable medium of claim 17 , wherein the non-deterministic algorithm is a Monte Carlo path tracing algorithm.
19 . The machine-readable medium of claim 15 , wherein the first processor of the plurality of processors is to render the one or more first images using a first number of samples, and the second processor of the plurality of processors is to render a higher sample version the one or more first images using a second number of samples that is greater than the first number.
20 . The machine-readable medium of claim 15 , wherein the one or more first images are used as ground truth data to train the one or more neural networks.
21 . The non-transitory machine-readable medium of claim 15 , wherein the first processor is a graphics processing unit (GPU) and the second processor comprises a plurality of GPUs to collectively train the one or more neural networks.
22 . A plurality of processors comprising: two or more processors to perform one or more inference operations, the two or more processors comprising circuitry to:
cause a first processor of the plurality of processors to generate one or more first images using a three-dimensional (3D) representation of a scene depicted by the one or more first images, the one or more first images and the 3D representation of the scene being stored in a memory shared between the first processor and a second processor of the plurality of processors; and
cause the second processor of the plurality of processors to use the one or more first images and the 3D representation of the scene stored in the shared memory to train one or more neural networks used by the first processor.
23 . The plurality of processors of claim 22 , wherein:
the first processor of the plurality of processors is to:
render the one or more first images based, at least in part, on a 3D representation of the scene;
generate the one or more first images from one or more rendered images using the one or more neural networks; and
provide the one or more first images and the 3D representation of the scene to the second processor; and the second processor of the plurality of processors is to use the one or more first images and 3D representation of the scene, stored in the shared memory between the first processor and the second processor, to train the one or more neural networks.
24 . The plurality of processors of claim 23 , wherein the first processor of the plurality of processors is to order a queue including at least the one or more first images.
25 . The plurality of processors of claim 24 , wherein the first processor of the plurality of processors orders a queue based, at least in part, on additional data associated with the one or more first images including depth data, normal data, albedo data, roughness data, or motion vector data.
26 . The plurality of processors of claim 23 , wherein the 3D representation is to be used by the second processor of the plurality of processors to generate a ground truth image that is to be compared against the one or more first images as part of training the second version of the one or more neural networks.
27 . The plurality of processors of claim 22 , wherein the one or more neural networks comprise a denoiser neural network, wherein the first version and the second version of the one or more neural networks have differing weights.
28 . The plurality of processors of claim 22 , wherein the 3D representation of the scene comprises is a spatial representation of the one or more first images.
29 . A system comprising: two or more processors to train one or more neural networks to perform one or more inference operations, the two or more processors comprising circuitry to;
cause a first processor of a plurality of processors to generate one or more first images using a three-dimensional (3D) representation of a scene depicted by the one or more first images, the one or more first images and the 3D representation of the scene being stored in a memory shared between the first processor and a second processor of the plurality of processors; and
cause the second processor of the plurality of processors to use the one or more first images and the 3D representation of the scene stored in the shared memory to train the one or more neural networks used by the first processor.
30 . The system of claim 29 , wherein the first processor of the plurality of processors comprises a first graphics processing unit (GPU) and the second processor of the plurality of processors comprises a second GPU.
31 . The system of claim 29 , wherein the one or more first images comprise a noised image generated from the 3D representation to train the one or more neural networks.
32 . The system of claim 29 , wherein a denoised version of the one or more images is generated by denoising the one or more first images.
33 . The system of claim 29 , wherein the 3D representation of the scene comprises a spatial representation of the one or more first images.
34 . The system of claim 29 , wherein:
the first processor of the plurality of processors is to perform the one or more inference operations;
the second processor of the plurality of processors is to perform training of the one or more neural networks;
and the first processor has a greater processing capability than the second processor.