Pose prediction of objects for extended reality systems
Systems and techniques are described herein for providing virtual content for a display. A method for providing virtual content for a display is provided. The method may include obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of an object in an environment; predicting, based on the plurality of images, a pose of the object in a reference coordinate system associated with the environment; determining, based on the predicted pose of the object in the reference coordinate system, a pose of the object relative to the device; and providing, to a display of the device, virtual content based on the pose of the object relative to the device.
1 . A method of providing virtual content for display, the method comprising:
obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;
processing the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;
obtaining a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;
applying the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and
providing, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.
2 . The method of claim 1 , wherein the predicted pose of the real-world object in the reference coordinate system is further based on previously-determined poses of the real-world object.
3 . The method of claim 1 , wherein predicting the pose of the real-world object comprises:
predicting a number of future poses of the real-world object at a number of respective future times; and
predicting the pose of the real-world object based on interpolating between the predicted number of future poses.
4 . The method of claim 1 , wherein the transformation is based on a head-pose prediction model.
5 . The method of claim 1 , wherein:
the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and
the method further comprises displaying the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.
6 . The method of claim 1 , wherein the device is an extended-reality device.
7 . The method of claim 1 , wherein the device is a see-through extended-reality device.
8 . An apparatus for providing virtual content for display, the apparatus comprising:
at least one memory; and
at least one processor coupled to the at least one memory and configured to:
obtain a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;
process the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;
obtain a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;
apply the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and
provide, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.
9 . The apparatus of claim 8 , wherein the at least one processor is configured to predict the pose of the real-world object in the reference coordinate system is-further based on previously-determined poses of the real-world object.
10 . The apparatus of claim 8 , wherein, to predict the pose of the real-world object, the at least one processor is configured to:
predict a number of future poses of the real-world object at a number of respective future times; and
predict the pose of the real-world object based on interpolating between the predicted number of future poses.
11 . The apparatus of claim 8 , wherein the transformation is based on a head-pose prediction model.
12 . The apparatus of claim 8 , wherein:
the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and
the at least one processor is further configured to display the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.
13 . The apparatus of claim 8 , wherein the device comprises a display and a camera of an extended-reality device and wherein the apparatus comprises a processor of the extended-reality device.
14 . The apparatus of claim 8 , wherein the device comprises a display of a see-through extended-reality device and wherein the apparatus comprises a processor of the see-through extended-reality device.
15 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:
obtain a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;
process the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;
obtain a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;
apply the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and
provide, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein the predicted pose of the real-world object in the reference coordinate system is further based on previously-determined poses of the real-world object.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions, when executed by at least one processor, cause the at least one processor to, in predicting the pose of the real-world object:
predict a number of future poses of the real-world object at a number of respective future times; and
predict the pose of the real-world object based on interpolating between the predicted number of future poses.
18 . The non-transitory computer-readable storage medium of claim 15 , wherein the transformation is based on a head-pose prediction model.
19 . The non-transitory computer-readable storage medium of claim 15 , wherein:
the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and
the instructions, when executed by at least one processor, cause the at least one processor to display the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.
20 . The non-transitory computer-readable storage medium of claim 15 , wherein the device comprises a display and a camera of an extended-reality device and wherein the at least one processor is a component of a computing unit of the extended-reality device.
21 . The non-transitory computer-readable storage medium of claim 15 , wherein the device comprises a display of a see-through extended-reality device and wherein the at least one processor is a component of a computing unit of the see-through extended-reality device.
22 . An apparatus for providing virtual content for display, the apparatus comprising:
one or more means for obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;
one or more means for processing the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;
one or more means for obtaining a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;
one or more means for applying the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and
one or more means for providing, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.