IP Library Granted Patent US 12705777
Granted Patent B2
US 12705777 · App. 18/174,532 · Granted Aug 11, 2026

Pose prediction of objects for extended reality systems

Inventor: Robert Peter Viehauser (Bad Hofgastein, AT)
Assignee: QUALCOMM Incorporated
G06T7/70G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705777
App. No.
18/174,532
Granted
Aug 11, 2026
Kind
B2
Abstract

Systems and techniques are described herein for providing virtual content for a display. A method for providing virtual content for a display is provided. The method may include obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of an object in an environment; predicting, based on the plurality of images, a pose of the object in a reference coordinate system associated with the environment; determining, based on the predicted pose of the object in the reference coordinate system, a pose of the object relative to the device; and providing, to a display of the device, virtual content based on the pose of the object relative to the device.

Claims (56)

1 . A method of providing virtual content for display, the method comprising:

obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;

processing the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;

obtaining a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;

applying the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and

providing, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.

2 . The method of claim 1 , wherein the predicted pose of the real-world object in the reference coordinate system is further based on previously-determined poses of the real-world object.

3 . The method of claim 1 , wherein predicting the pose of the real-world object comprises:

predicting a number of future poses of the real-world object at a number of respective future times; and

predicting the pose of the real-world object based on interpolating between the predicted number of future poses.

4 . The method of claim 1 , wherein the transformation is based on a head-pose prediction model.

5 . The method of claim 1 , wherein:

the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and

the method further comprises displaying the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.

6 . The method of claim 1 , wherein the device is an extended-reality device.

7 . The method of claim 1 , wherein the device is a see-through extended-reality device.

8 . An apparatus for providing virtual content for display, the apparatus comprising:

at least one memory; and

at least one processor coupled to the at least one memory and configured to:

obtain a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;

process the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;

obtain a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;

apply the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and

provide, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.

9 . The apparatus of claim 8 , wherein the at least one processor is configured to predict the pose of the real-world object in the reference coordinate system is-further based on previously-determined poses of the real-world object.

10 . The apparatus of claim 8 , wherein, to predict the pose of the real-world object, the at least one processor is configured to:

predict a number of future poses of the real-world object at a number of respective future times; and

predict the pose of the real-world object based on interpolating between the predicted number of future poses.

11 . The apparatus of claim 8 , wherein the transformation is based on a head-pose prediction model.

12 . The apparatus of claim 8 , wherein:

the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and

the at least one processor is further configured to display the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.

13 . The apparatus of claim 8 , wherein the device comprises a display and a camera of an extended-reality device and wherein the apparatus comprises a processor of the extended-reality device.

14 . The apparatus of claim 8 , wherein the device comprises a display of a see-through extended-reality device and wherein the apparatus comprises a processor of the see-through extended-reality device.

15 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:

obtain a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;

process the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;

obtain a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;

apply the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and

provide, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the predicted pose of the real-world object in the reference coordinate system is further based on previously-determined poses of the real-world object.

17 . The non-transitory computer-readable storage medium of claim 15 , wherein the instructions, when executed by at least one processor, cause the at least one processor to, in predicting the pose of the real-world object:

predict a number of future poses of the real-world object at a number of respective future times; and

predict the pose of the real-world object based on interpolating between the predicted number of future poses.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein the transformation is based on a head-pose prediction model.

19 . The non-transitory computer-readable storage medium of claim 15 , wherein:

the plurality of images captured by the camera include the real-world object and the environment from a perspective of the camera; and

the instructions, when executed by at least one processor, cause the at least one processor to display the virtual content at a location of the display that is related to a pose of the real-world object within a line of sight of a user of the device according to an orientation of the device and a position of the device.

20 . The non-transitory computer-readable storage medium of claim 15 , wherein the device comprises a display and a camera of an extended-reality device and wherein the at least one processor is a component of a computing unit of the extended-reality device.

21 . The non-transitory computer-readable storage medium of claim 15 , wherein the device comprises a display of a see-through extended-reality device and wherein the at least one processor is a component of a computing unit of the see-through extended-reality device.

22 . An apparatus for providing virtual content for display, the apparatus comprising:

one or more means for obtaining a plurality of images captured by a camera of a device, each image of the plurality of images including a respective representation of a real-world object in an environment;

one or more means for processing the plurality of images using a pose-prediction machine-learning model to predict a pose of the real-world object in a reference coordinate system associated with the environment, wherein the pose-prediction machine-learning model is trained to predict poses of real-world objects in reference coordinate systems based on images of the real-world objects;

one or more means for obtaining a transformation between the reference coordinate system and a device coordinate system associated with an orientation of the device;

one or more means for applying the transformation to the predicted pose of the real-world object in the reference coordinate system to obtain a pose of the real-world object relative to the device; and

one or more means for providing, to a display of the device, virtual content that is associated with the real-world object, wherein a pose of the virtual content is based on the pose of the real-world object relative to the device.