IP Library Granted Patent US 12682635
Granted Patent B2
US 12682635 · App. 17/558,185 · Granted Jul 14, 2026

Software-based user interface element analogues for physical device elements

Inventors: Andrew Lovitt (Redmond, WA); Sebastian Sztuk (Menlo Park, CA)
Assignee: Meta Platforms Technologies, LLC
G06V20/20G02B27/017G06F3/017G06T19/006G06V10/255G02B2027/0178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682635
App. No.
17/558,185
Granted
Jul 14, 2026
Kind
B2
Abstract

A client device of a user (e.g., a headset) provides a software-based user interface, not relying solely on physical buttons built into the client device itself. The client device's user interface system can render the user interface using graphical, virtual user interface elements that take the place of physical buttons, reducing device manufacturing costs, improving user interface reliability, and allowing greater user interface flexibility. Information on user movements can then be obtained using a separate wearable device, and the client device can use such information to determine whether the user's movements indicate interactions with the virtual user interface elements, taking appropriate actions if so. The client device's user interface system can also render the user interface using sounds made to appear to emanate from particular locations. Subsequent user movements can then be analyzed to determine whether they indicate the particular locations corresponding to the sound-based elements.

Claims (103)

1 . A computer-implemented method comprising:

transitioning from a first power state to a second power state by dimming or shutting off a display of a client device of a user, wherein the second power state is a lower power state than the first power state;

rendering two or more user interface elements at two or more user-relative locations by, in response to transitioning to the second power state, using spatial sound to cause audio content to appear to originate from the two or more user-relative locations;

receiving sensor data associated with a movement of at least one of the client device or a wearable device of the user;

determining, using the sensor data, a location corresponding to the movement;

identifying, using the determined location, a user interface element among the two or more user interface elements; and

calling a function registered in association with the user interface element;

wherein the two or more user interface elements represent options to proceed with a current action and to go back, the two or more user-relative locations are to left and right sides of the user, the two or more user interface elements are rendered using text-to-speech on respective labels thereof and application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations, and the sensor data indicates turning of a head of the user toward one of the two or more user-relative locations.

2 . The computer-implemented method of claim 1 , wherein the application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations includes application of head-related transfer functions.

3 . The computer-implemented method of claim 1 , further comprising determining user-relative locations for a plurality of user interface elements, the determining comprising:

performing visual analysis on an image captured by an imaging device of an augmented reality or virtual reality headset of a user;

identifying objects present within the image based on the visual analysis;

determining locations of the identified objects within the image;

ranking at least some of the user interface elements according to likelihoods of the user interface elements being selected by the user into high-likelihood user interface elements and low-likelihood user interface elements;

identifying lower-priority regions and higher-priority regions within the image displayed on the display based on the performed visual analysis of the image displayed on the display of the augmented reality or the virtual reality headset of the user, wherein lower-priority regions and higher-priority regions are distinguished from one another based on one or more of a presence of text, a uniformity of backgrounds, or an existence of movement; and

assigning locations of the user interface elements to be within the lower-priority regions.

4 . The computer-implemented method of claim 1 , further comprising:

receiving second sensor data associated with a second movement of a wearable device of the user;

determining that the second movement corresponds to a second location;

identifying, using the second determined location, a second user interface element among the rendered user interface elements;

determining that the second movement is part of the user removing a headset; and

determining, based on the determination that the second movement is part of the user removing the headset, that the user is not selecting the second user interface element.

5 . The computer-implemented method of claim 1 , further comprising:

determining, using the sensor data, a type of gesture corresponding to the movement; and identifying the called function at least in part based on the type of gesture.

6 . The computer-implemented method of claim 1 , further comprising:

determining second user-relative locations for a second plurality of user interface elements; and

rendering the second plurality of user interface elements at the second user-relative locations, the rendering comprising:

generating synthetic speech corresponding to textual labels of the second plurality of user interface elements;

applying transfer functions to the generated synthetic speech to produce transformed speech that appears to emanate from the second user-relative locations; and

aurally outputting the transformed speech.

7 . The computer-implemented method of claim 1 , further comprising:

receiving second sensor data associated with a second movement of the user;

determining, using the second sensor data, a second location associated with the second movement; and

identifying, using the second location, a selected one of a second plurality of user interface elements.

8 . The computer-implemented method of claim 1 , wherein

a headset includes a physical interaction site, the method further comprising:

receiving second sensor data associated with a second movement of the wearable device;

determining, using the second sensor data, a second location corresponding to the second movement;

determine that the determined second location corresponds to a location of the physical interaction site; and

calling a function registered in association with the physical interaction site.

9 . A non-transitory computer-readable storage medium storing instructions that when executed by a computer processor perform actions comprising:

transitioning from a first power state to a second power state by dimming or shutting off a display of a client device of a user, wherein the second power state is a lower power state than the first power state;

rendering two or more user interface elements at two or more user-relative locations by, in response to transitioning to the second power state, using spatial sound to cause audio content to appear to originate from the two or more user-relative locations;

receiving sensor data associated with a movement of at least one of the client device or a wearable device of the user;

determining, using the sensor data, a location corresponding to the movement;

identifying, using the determined location, a user interface element among the two or more user interface elements; and

calling a function registered in association with the user interface element,

wherein the two or more user interface elements represent options to proceed with a current action and to go back, the two or more user-relative locations are to left and right sides of the user, the two or more user interface elements are rendered using text-to-speech on respective labels thereof and application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations, and the sensor data indicates turning of a head of the user toward one of the two or more user-relative locations.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations includes application of head-related transfer functions.

11 . The non-transitory computer-readable storage medium of claim 9 , further comprising determining user-relative locations for a plurality of user interface elements, the determining comprising:

performing visual analysis on an image captured by an imaging device an augmented reality or virtual reality headset of a user;

identifying objects present within the image based on the visual analysis;

determining locations of the identified objects within the image;

ranking at least some of the user interface elements according to likelihoods of the user interface elements being selected by the user into high-likelihood user interface elements and low-likelihood user interface elements;

identifying lower-priority regions and higher-priority regions within the image displayed on the display based on the performed visual analysis of the image displayed on the display of the augmented reality or the virtual reality headset of the user, wherein lower-priority regions and higher-priority regions are distinguished from one another based on one or more of a presence of text, a uniformity of backgrounds, or an existence of movement; and

assigning locations of the user interface elements to be within the lower-priority regions.

12 . The computer-implemented method of claim 1 , further comprising:

receiving second sensor data associated with a second movement of a wearable device of the user;

determining that the second movement corresponds to a second location;

identifying, using the second determined location, a second user interface element among the rendered user interface elements;

determining that the second movement is part of the user removing a headset; and

determining, based on the determination that the second movement is part of the user removing the headset, that the user is not selecting the second user interface element.

13 . The non-transitory computer-readable storage medium of claim 9 , the actions further comprising:

determining, using the sensor data, a type of gesture corresponding to the movement; and identifying the called function at least in part based on the type of gesture.

14 . The non-transitory computer-readable storage medium of claim 9 , the actions further comprising:

determining second user-relative locations for a second plurality of user interface elements;

rendering the second plurality of user interface elements at the second user-relative locations, the rendering comprising:

generating synthetic speech corresponding to textual labels of the second plurality of user interface elements;

applying transfer functions to the generated synthetic speech to produce transformed speech that appears to emanate from the second user-relative locations; and

aurally outputting the transformed speech.

15 . The non-transitory computer-readable storage medium of claim 9 , the actions further comprising:

receiving second sensor data associated with a second movement of the user;

determining, using the second sensor data, a second location associated with the second movement; and

identifying, using the second location, a selected one of a second plurality of user interface elements.

16 . The non-transitory computer-readable storage medium of claim 9 , wherein a headset includes a physical interaction site, the actions further comprising:

receiving second sensor data associated with a second movement of the wearable device;

determining, using the second sensor data, a second location corresponding to the second movement;

determine that the determined second location corresponds to a location of the physical interaction site; and

calling a function registered in association with the physical interaction site.

17 . A client device comprising:

a computer processor; and

a non-transitory computer-readable storage medium storing instructions that when executed by the computer processor perform actions comprising:

transitioning from a first power state to a second power state by dimming or shutting off a display of a client device of a user, wherein the second power state is a lower power state than the first power state;

rendering two or more user interface elements at two or more user-relative locations by, in response to transitioning to the second power state, using spatial sound to cause audio content to appear to originate from the two or more user-relative locations;

receiving sensor data associated with a movement of at least one of the client device or a wearable device of the user;

determining, using the sensor data, a location corresponding to the movement;

identifying, using the determined location, a user interface element among the two or more user interface elements; and

calling a function registered in association with the user interface element,

wherein the two or more user interface elements represent options to proceed with a current action and to go back, the two or more user-relative locations are to left and right sides of the user, the two or more user interface elements are rendered using text-to-speech on respective labels thereof and application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations, and the sensor data indicates turning of a head of the user toward one of the two or more user-relative locations.

18 . The client device of claim 17 , wherein the application of the spatial sound to make speech for the respective labels appear to emanate from the two or more user-relative locations includes application of head-related transfer functions.

19 . The client device of claim 17 , further comprising determining user-relative locations for a plurality of user interface elements, the determining comprising:

performing visual analysis on an image captured by an imaging device of an augmented reality or virtual reality headset of a user;

identifying objects present within the image based on the visual analysis;

determining locations of the identified objects within the image;

ranking at least some of the user interface elements according to likelihoods of the user interface elements being selected by the user into high-likelihood user interface elements and low-likelihood user interface elements;

identifying lower-priority regions and higher-priority regions within the image displayed on the display based on the performed visual analysis of the image displayed on the display of the augmented reality or the virtual reality headset of the user, wherein lower-priority regions and higher-priority regions are distinguished from one another based on one or more of a presence of text, a uniformity of backgrounds, or an existence of movement; and

assigning locations of the user interface elements to be within the lower-priority regions.

20 . The client device of claim 17 , the actions further comprising:

receiving second sensor data associated with a second movement of a wearable device of the user;

determining that the second movement corresponds to a second location;

identifying, using the second determined location, a second user interface element among the rendered user interface elements;

determining that the second movement is part of the user removing a headset; and

determining, based on the determination that the second movement is part of the user removing the headset, that the user is not selecting the second user interface element.