IP Library Granted Patent US 12700187
Granted Patent B2
US 12700187 · App. 18/162,453 · Granted Aug 4, 2026

Footwear mirror experience

Inventors: Matan Zohar (Rishon LeZion, IL); Itamar Berger (Hod Hasharon, IL); Gal Sasson (Kibbutz Ayyelet Hashahar, IL); Omri Berg (Tel Aviv, IL)
Assignee: Snap Inc.
G06T19/006G06F3/017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700187
App. No.
18/162,453
Granted
Aug 4, 2026
Kind
B2
Abstract

Aspects of the present disclosure involve a system for providing a footwear mirror experience. The system accesses a video stream captured by a camera directed at feet of users. The system detects a depiction of one or more feet of a user in the video stream captured by the camera. The system, in response to detecting the one or more feet in the video stream captured by the camera, activates a display screen that is positioned at an eye-level of the user and presents the video stream including the depiction of the one or more feet of the user on the display screen.

Claims (51)

1 . A method comprising:

tracking, using one or more first machine learning models, feet of a user in a first video stream captured by a first camera coupled to a display screen, the one or more first machine learning models trained by comparing tracking information for training images with estimated tracking information generated by the one or more first machine learning models based on the training images;

presenting, on the display screen, the first video stream including the feet of the user captured by the first camera on a first portion of the display screen and a second video stream including a body of the user captured by a second camera coupled to the display screen, the second video stream presented on a second portion of the display screen;

estimating, using one or more second machine learning models, an estimated shoe size of the user based on images extracted from the first video stream;

presenting, on the second video stream, an augmented reality (AR) shoe depicted over the feet in the second video stream, wherein the AR shoe in the second video stream is sized based on the estimated shoe size of the user; and

changing between different AR shoes that are depicted over the feet in the first video stream and the second video stream in response to gestures performed by the feet of the user, the gestures detected based on the estimated tracking information generated by the one or more first machine learning models based on the first video stream.

2 . The method of claim 1 , wherein the first camera is placed at a first specified distance from a floor level, and wherein at least a portion of the display screen is placed at a second specified distance from the floor level, the second specified distance being greater than the first specified distance.

3 . The method of claim 1 , further comprising:

detecting a foot gesture performed by the feet in the first video stream; and

triggering activation of the display screen to present the first video stream in response to detecting the foot gesture.

4 . The method of claim 1 , further comprising:

detecting a fashion item being worn by the feet of the user in the first video stream; and

extracting, from the first video stream, a portion comprising the fashion item.

5 . The method of claim 1 , further comprising:

receiving input that activates an AR experience associated with trying on shoes; and

in response to the input, adding the AR shoe to the feet in the first video stream captured by the first camera.

6 . The method of claim 5 , further comprising:

modifying the first video stream presented to the user on the display screen to depict the AR shoe that has been added to the feet.

7 . The method of claim 5 , further comprising:

receiving a selection of the AR shoe from a list of AR shoes presented on the display screen.

8 . The method of claim 7 , wherein the AR shoe is a first AR shoe, further comprising:

detecting a gesture associated with the feet depicted in the first video stream; and

selecting a second AR shoe from the list of AR shoes in response to detecting the gesture.

9 . The method of claim 8 , further comprising:

adding the second AR shoe instead of the first AR shoe to the feet.

10 . The method of claim 8 , wherein the gesture comprises touch input of a portion of the display screen.

11 . The method of claim 8 , wherein the gesture comprises a foot gesture performed by the feet depicted in the first video stream.

12 . The method of claim 8 , wherein the second AR shoe corresponds to a different style or color of the first AR shoe.

13 . The method of claim 1 , wherein the first camera is coupled to the display screen via one or more external wires.

14 . The method of claim 1 , wherein the display screen comprises a full-body display screen that is placed on a floor of a real-world environment, and wherein the first camera is integrated with the display screen at a bottom of the display screen.

15 . The method of claim 1 ,

wherein the one or more second machine learning models are trained based on training data comprising images of feet and images of shoes with corresponding shoe sizes.

16 . The method of claim 15 , further comprising:

overlaying the AR shoe on the feet in the first video stream to present the first video stream including the feet of the user and the AR shoe on the display screen.

17 . A system comprising:

at least one processor of a device configured to perform operations comprising:

tracking, using one or more first machine learning models, feet of a user in a first video stream captured by a first camera coupled to a display screen, the one or more first machine learning models trained by comparing tracking information for training images with estimated tracking information generated by the one or more first machine learning models based on the training images:

presenting, on the display screen, the first video stream including the feet of the user captured by the first camera on a first portion of the display screen and a second video stream including a body of the user captured by a second camera coupled to the display screen, the second video stream presented on a second portion of the display screen;

estimating, using one or more second machine learning models, an estimated shoe size of the user based on images extracted from the first video stream;

presenting, on the second video stream, an augmented reality (AR) shoe depicted over the feet in the second video stream, wherein the AR shoe in the second video stream is sized based on the estimated shoe size of the user; and

changing between different AR shoes that are depicted over the feet in the first video stream and the second video stream in response to gestures performed by the feet of the user, the gestures detected based on the estimated tracking information generated by the one or more first machine learning models based on the first video stream.

18 . The system of claim 17 , wherein the gestures comprise tapping heels of the feet together.

19 . The system of claim 17 , the operations comprising:

detecting a foot gesture comprising tapping heels of the feet together performed by the feet in the first video stream; and

triggering activation of the display screen to present the first video stream in response to detecting the foot gesture comprising tapping heels of the feet together.

20 . A non-transitory machine-readable storage medium that includes instructions that, when executed by one or more processors of a device, cause the device to perform operations comprising:

tracking, using one or more first machine learning models, feet of a user in a first video stream captured by a first camera coupled to a display screen, the one or more first machine learning models trained by comparing tracking information for training images with estimated tracking information generated by the one or more first machine learning models based on the training images;

presenting, on the display screen, the first video stream including the feet of the user captured by the first camera on a first portion of the display screen and a second video stream including a body of the user captured by a second camera coupled to the display screen, the second video stream presented on a second portion of the display screen;

estimating, using one or more second machine learning models, an estimated shoe size of the user based on images extracted from the first video stream;

presenting, on the second video stream, an augmented reality (AR) shoe depicted over the feet in the second video stream, wherein the AR shoe in the second video stream is sized based on the estimated shoe size of the user; and

changing between different AR shoes that are depicted over the feet in the first video stream and the second video stream in response to gestures performed by the feet of the user, the gestures detected based on the estimated tracking information generated by the one or more first machine learning models based on the first video stream.