Deep learning based causal image reprojection for temporal supersampling in AR/VR systems
Generating synthesized data includes capturing one or more frames of a scene at a first frame rate by one or more cameras of a wearable device, determining body position parameters for the frames, and obtaining geometry data for the scene in accordance with the one or more frames. The frames, body position parameters, and geometry data are applied to a trained network which predicts one or more additional frames. With respect to virtual data, generating a synthesized frame includes determining current body position parameters in accordance with the one or more frames, predicting a future gaze position based on the current body position parameters, and rendering, at a first resolution, a gaze region of a frame in accordance with the future gaze position. A peripheral region is predicted for the frame at a second resolution, and the combined regions form a frame that is used to drive a display.
1 . A method comprising:
capturing one or more frames of a scene at a first frame rate by one or more cameras of a wearable device;
determining body position parameters in accordance with the one or more frames;
obtaining a 3D geometric representation of at least a portion of the scene in accordance with the one or more frames; and
applying the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene to a trained network to generate one or more additional frames to follow the one or more frames, wherein the trained network is trained based on a rendered set of frames,
wherein an error is minimized between a predicted frame and one or more of the rendered set of frames during training based on the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene.
2 . The method of claim 1 , wherein the body position parameters include body pose data and gaze data.
3 . The method of claim 1 , wherein the one or more frames comprises a plurality of frames including a first frame and a second frame, and wherein the trained network further predicts one or more intermediate frames between the first frame and the second frame.
4 . The method of claim 3 , wherein a frame rate is determined for a set of frames comprising the first frame, the second frame, and the intermediate frames, and wherein the one or more additional frames are predicted in accordance with the determined frame rate.
5 . The method of claim 1 , wherein the one or more frames are captured from a first viewing frustum, and wherein the one or more additional frames are from one or more additional viewing frustums, wherein the one or more additional viewing frustums are predicted by the trained network.
6 . The method of claim 1 , wherein the one or more frames are captured at a first resolution, and wherein the trained network predicts the one or more additional frames and a second resolution higher than the first resolution.
7 . The method of claim 6 , further comprising:
determining virtual content to be rendered in the one or more additional frames; and
rendering the virtual content at the second resolution in accordance with the determination.
8 . The method of claim 1 , wherein the one or more frames, the body position parameters, and the 3D geometric representation are applied to the trained network in accordance with power consumption parameters for the one or more frames or the wearable device.
9 . A non-transitory computer readable medium comprising computer readable code executable by one or more processors to:
render one or more frames of a scene;
determine body position parameters in accordance with the one or more frames;
obtain a 3D geometric representation of at least a portion of the scene in accordance with the one or more frames; and
apply the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene to a trained network to generate one or more additional frames to follow the one or more frames, wherein the trained network is trained based on a rendered set of frames,
wherein an error is minimized between a predicted frame and one or more of the rendered set of frames during training based on the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene.
10 . The non-transitory computer readable medium of claim 9 , wherein the body position parameters include body pose data and gaze data.
11 . The non-transitory computer readable medium of claim 9 , wherein the one or more frames comprises a plurality of frames including a first frame and a second frame, and wherein the trained network further predicts one or more intermediate frames between the first frame and the second frame.
12 . The non-transitory computer readable medium of claim 11 , wherein a frame rate is determined for a set of frames comprising the first frame, the second frame, and the intermediate frames, and wherein the one or more additional frames are predicted in accordance with the determined frame rate.
13 . The non-transitory computer readable medium of claim 9 , wherein the one or more frames are captured from a first viewing frustum, and wherein the one or more additional frames are from one or more additional viewing frustums, wherein the one or more additional viewing frustums are predicted by the trained network.
14 . The non-transitory computer readable medium of claim 9 , wherein the one or more frames, the body position parameters, and the 3D geometric representation are applied to a trained network in accordance with power consumption parameters for the one or more frames.
15 . A system comprising:
one or more processors; and
one or more computer readable media comprising computer readable code executable by the one or more processors to:
capture one or more frames of a scene at a first frame rate by one or more cameras of a wearable device;
determine body position parameters in accordance with the one or more frames;
obtain a 3D geometric representation of at least a portion of the scene in accordance with the one or more frames; and
apply the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene to a trained network to generate one or more additional frames to follow the one or more frames, wherein the trained network is trained based on a rendered set of frames,
wherein an error is minimized between a predicted frame and one or more of the rendered set of frames during training based on the one or more frames, the body position parameters, and the 3D geometric representation of at least the portion of the scene.
16 . The system of claim 15 , wherein the one or more frames comprises a plurality of frames including a first frame and a second frame, and wherein the trained network further predicts one or more intermediate frames between the first frame and the second frame.
17 . The system of claim 16 , wherein a frame rate is determined for a set of frames comprising the first frame, the second frame, and the intermediate frames, and wherein the one or more additional frames are predicted in accordance with the determined frame rate.