Stationary extended reality device
Methods and systems are disclosed for generating an extended reality (XR) experience using a statically positioned device. The system receives, from a camera of a stationary device, an image depicting a real-world object, the camera being directed in a stationary manner towards a specified field of view of a real-world environment. The system analyzes the image using a machine learning model to predict tracking information for the real-world object, the machine learning model trained based on a plurality of training images depicting real-world objects in the specified field of view of the real-world environment and corresponding ground-truth tracking information for the real-world objects. The system selects an extended reality (XR) experience from a plurality of XR experiences and overlays one or more XR elements associated with the XR experience on the image based on the predicted tracking information to generate a modified image.
1 . A method comprising:
receiving, from a camera of a stationary device comprising an extended reality (XR) shoe mirror, an image depicting a real-world object, the camera being directed in a stationary manner towards a specified field of view of a real-world environment;
setting a focal length and one or more camera parameters to a specified value based on a first distance between the camera and a physical area comprising the specified field of view and one or more external conditions of the XR shoe mirror, the XR shoe mirror including a display screen that is placed at a second distance above the specified field of view, the XR shoe mirror including the physical area attached to the display screen that defines the specified field of view, and the XR shoe mirror including the camera pointed towards the physical area;
analyzing the image using a machine learning model to generate predicted tracking information for the real-world object, the machine learning model trained on a plurality of images depicting real-world objects in the specified field of view of the real-world environment and corresponding ground-truth tracking information for the real-world objects;
selecting an XR experience from a plurality of XR experiences; and
overlaying one or more XR elements associated with the XR experience on the image based on the predicted tracking information to generate a modified image.
2 . The method of claim 1 , further comprising:
generating a mask based on an area outside of the physical area of the XR shoe mirror, the mask being generated when the XR shoe mirror is manufactured or placed within the real-world environment.
3 . The method of claim 2 , wherein the real-world object comprises a foot, and wherein the one or more XR elements comprise one or more virtual shoes.
4 . The method of claim 1 , further comprising:
accessing lighting conditions of the real-world environment;
analyzing the lighting conditions with the machine learning model to predict a modification to one or more physical light parameters of the stationary device and to predict one or more light parameters of the one or more XR elements associated with the XR experience; and
automatically adjusting physical light that surrounds a screen on which the modified image is displayed based on the predicted one or more physical light parameters.
5 . The method of claim 4 , wherein the display screen is mounted on a physical surface that is tilted at an angle relative to the physical area, and wherein the camera is integrated into the physical surface and is directed such that a surface normal of the camera is perpendicular to the physical area.
6 . The method of claim 4 , wherein the camera points down towards the physical area.
7 . The method of claim 4 , further comprising:
identifying an area outside of the physical area of the XR shoe mirror;
generating a mask based on the area outside of the physical area of the XR shoe mirror; and
occluding pixels of the image that fall within the identified area outside of the physical area of the XR shoe mirror while overlaying the one or more XR elements on the image.
8 . The method of claim 7 , wherein the mask is generated when the XR shoe mirror is manufactured or placed within the real-world environment.
9 . The method of claim 1 , wherein the stationary device comprises an XR vanity mirror.
10 . The method of claim 9 , wherein the real-world object comprises a face, and wherein the one or more XR elements comprise one or more virtual fashion items or virtual makeup.
11 . The method of claim 9 , further comprising:
accessing lighting conditions of the real-world environment;
analyzing the lighting conditions with the machine learning model to predict a modification to one or more physical light parameters of the stationary device and to predict one or more light parameters of the one or more XR elements associated with the XR experience; and
automatically adjusting physical light that surrounds a screen on which the modified image is displayed based on the predicted one or more physical light parameters.
12 . The method of claim 11 , wherein the one or more physical light parameters comprise a light temperature, color, or intensity.
13 . The method of claim 1 , wherein the stationary device comprises an XR window placed on an opening of a building or on an interior wall of the building.
14 . The method of claim 13 , wherein the real-world object comprises a sky or city landscape, and wherein the one or more XR elements comprise one or more labels or virtual sky elements.
15 . The method of claim 14 , further comprising setting the focal length and the one or more camera parameters based on a second distance between the camera and the sky or city landscape.
16 . The method of claim 13 , further comprising:
receiving, by the XR window, physical movement information of a mobile device that is external to the stationary device; and
adjusting one or more visual properties of the one or more XR elements displayed by the XR window based on the physical movement information of the mobile device.
17 . The method of claim 1 , further comprising training the machine learning model by performing training operations comprising:
selecting a first training image from the plurality of training images, the first training image depicting an individual real-world object having associated individual ground-truth tracking information, the first training image being captured by the stationary device and depicting the individual real-world object in the specified field of view of the real-world environment;
analyzing the first training image using the machine learning model to predict estimated tracking information for the individual real-world object;
generating loss based on a deviation between the estimated tracking information for the individual real-world object and the associated individual ground-truth tracking information of the individual real-world object; and
updating one or more parameters of the machine learning model based on the loss.
18 . A system comprising:
at least one processor; and at least one memory component having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:
receiving, from a camera of a stationary device comprising an extended reality (XR) shoe mirror, an image depicting a real-world object, the camera being directed in a stationary manner towards a specified field of view of a real-world environment;
setting a focal length and one or more camera parameters to a specified value based on a first distance between the camera and a physical area comprising the specified field of view and one or more external conditions of the XR shoe mirror, the XR shoe mirror including a display screen that is placed at a second distance above the specified field of view, the XR shoe mirror including the physical area attached to the display screen that defines the specified field of view, and the XR shoe mirror including the camera pointed towards the physical area;
analyzing the image using a machine learning model to generate predicted tracking information for the real-world object, the machine learning model trained based on a plurality of images depicting real-world objects in the specified field of view of the real-world environment and corresponding ground-truth tracking information for the real-world objects;
selecting an XR experience from a plurality of XR experiences; and
overlaying one or more XR elements associated with the XR experience on the image based on the predicted tracking information to generate a modified image.
19 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:
receiving, from a camera of a stationary device comprising an extended reality (XR) shoe mirror, an image depicting a real-world object, the camera being directed in a stationary manner towards a specified field of view of a real-world environment;
setting a focal length and one or more camera parameters to a specified value based on a first distance between the camera and a physical area comprising the specified field of view and one or more external conditions of the XR shoe mirror, the XR shoe mirror including a display screen that is placed at a second distance above the specified field of view, the XR shoe mirror including the physical area attached to the display screen that defines the specified field of view, and the XR shoe mirror including the camera pointed towards the physical area;
analyzing the image using a machine learning model to generate predicted tracking information for the real-world object, the machine learning model trained based on a plurality of images depicting real-world objects in the specified field of view of the real-world environment and corresponding ground-truth tracking information for the real-world objects;
selecting an XR experience from a plurality of XR experiences; and
overlaying one or more XR elements associated with the XR experience on the image based on the predicted tracking information to generate a modified image.
20 . The non-transitory computer-readable storage medium of claim 19 , wherein the real-world object comprises a sky or city landscape, and wherein the one or more XR elements comprise one or more labels or virtual sky elements.